Audio processing method and device, electronic equipment and storage medium
By distinguishing between voice and mechanical sounds, environmental audio processing is improved, thus solving the problem of misadjusting voice sounds due to household appliance noise in existing technologies, thereby enhancing the accuracy of audio processing and user experience.
Patent Information
- Application Number
- CN202110793296.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-14
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2041-07-14
AI Technical Summary
Existing noise reduction technologies cannot effectively distinguish between noise from household appliances and voice sources, leading to incorrect adjustment of voice sources and reduced user experience.
By acquiring ambient sounds, distinguishing between speech and mechanical sounds, processing the audio only when mechanical sounds are present in the ambient sounds, removing echoes, and differentiating based on sound energy fluctuations and correlation, the speech portion is enhanced.
It improves the accuracy of audio processing, reduces erroneous adjustments, enhances the user's auditory experience, and ensures that voice sounds are not mistaken for interference sources.
Smart Images

Figure CN115620735B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of communication technology, and in particular, to an audio processing method and device, electronic equipment and storage medium. BACKGROUND
[0002] Audio-visual entertainment devices represented by smart TVs, tablet computers and personal computers often use loudspeakers as carriers to perform audio-visual playing tasks. In audio-visual programs, more than half of the programs use voice as the main sound source. In daily home environment, audio-visual programs played through loudspeakers are inevitably disturbed by environmental noise, and the most important source of disturbance is household appliance noise, and voice sound sources are easily masked by household appliance noise, so that users cannot effectively obtain program sound information in a noisy environment. In current noise reduction technology, all sounds in the surrounding environment are identified as noise, and the real disturbance source represented by household appliance noise cannot be distinguished. This noise reduction scheme without distinguishing scenes will lead to misadjustment of voice sound sources, and reduce user experience. SUMMARY
[0003] The present disclosure provides an audio processing method, device, electronic equipment and storage medium.
[0004] According to a first aspect of the present disclosure, an audio processing method is provided, the method comprising:
[0005] obtaining sound in an environment to obtain a first environmental sound, the first environmental sound comprising at least a first audio echo, wherein the first audio echo refers to a sound signal collected by a microphone after a first audio is played by a loudspeaker;
[0006] removing the first audio echo in the first environmental sound according to a first audio reference sound to obtain a second environmental sound, wherein the first audio reference sound refers to audio source data of the first audio when the first audio is not played by the loudspeaker;
[0007] if the second environmental sound does not contain voice sound and contains mechanical sound, processing a second audio to be played to obtain a third audio;
[0008] sending the third audio to the loudspeaker to play the third audio by the loudspeaker.
[0009] In some embodiments, whether the second environmental sound contains voice sound or mechanical sound is determined according to the following steps:
[0010] determining whether the second environmental sound contains the voice sound;
[0011] if the second environmental sound does not contain the voice sound, determining whether the second environmental sound contains mechanical sound.
[0012] In some embodiments, whether the second ambient sound contains speech sound is determined according to the following steps:
[0013] sound energy of a plurality of audio frames in the second ambient sound within a preset time period is determined;
[0014] whether the second ambient sound contains the speech sound is determined according to fluctuation of sound energy of a plurality of audio frames in the second ambient sound within the preset time period.
[0015] In some embodiments, the determination of sound energy of the second ambient sound within a preset time period comprises:
[0016] a maximum frame energy and a minimum frame energy within the preset time period are determined; wherein the maximum frame energy is an audio frame with the maximum sound energy among a plurality of audio frames included in the preset time period; and the minimum frame energy is an audio frame with the minimum sound energy among the plurality of audio frames included in the preset time period:
[0017] the determination of whether the second ambient sound contains the speech sound according to fluctuation of sound energy of the second ambient sound within the preset time period comprises:
[0018] a ratio of the maximum frame energy and the minimum frame energy is determined;
[0019] if the ratio is greater than or equal to a first threshold value, it is determined that the second ambient sound contains the speech sound.
[0020] In some embodiments, the determination of whether the second ambient sound contains the speech sound according to fluctuation of sound energy of the second ambient sound within the preset time period comprises:
[0021] if the minimum frame energy among the plurality of audio frames included in the preset time period is less than a second threshold value, it is determined that the second ambient sound contains the speech sound.
[0022] In some embodiments, whether the second ambient sound contains mechanical sound is determined according to the following steps:
[0023] correlation of audio frames within a first sub-time period and a second sub-time period is determined according to the second ambient sound within the first sub-time period and the second sub-time period; wherein a maximum value of the first sub-time period is less than a maximum value of the second sub-time period;
[0024] if the correlation is greater than a correlation threshold value, it is determined that the second ambient sound contains the mechanical sound.
[0025] In some embodiments, the first sub-time period and the second sub-time period partially overlap and partially offset in time domain.
[0026] In some embodiments, whether the second ambient sound comprises the mechanical sound is determined according to the following steps:
[0027] respectively determining a mean value and a standard deviation of sound energy ratio values of a plurality of audio frames in a preset time period, wherein the sound energy ratio value of the audio frame is a ratio of sound energy of an audio frame of the second ambient sound in a first frequency domain to sound energy of the audio frame of the second ambient sound in a second frequency domain, wherein the first frequency domain is a frequency domain remaining after a preset frequency domain is removed from the second frequency domain, and a minimum value of frequencies of the first frequency domain is greater than a maximum value of frequencies of the preset frequency domain;
[0028] if the mean value is greater than a mean value threshold and the standard deviation is less than a difference value threshold, it is determined that the second ambient sound contains the mechanical sound.
[0029] In some embodiments, if the second ambient sound does not contain speech sound and contains mechanical sound, the second audio to be played is processed to obtain a third audio, comprising:
[0030] obtaining the second audio to be played;
[0031] obtaining prior information of the second ambient sound;
[0032] performing enhancement processing on a speech part contained in the second audio according to the prior information of the second ambient sound to obtain the third audio.
[0033] In some embodiments, the prior information of the second ambient sound comprises at least one of the following:
[0034] a frequency domain signal of the second ambient sound in the first frequency domain;
[0035] a frequency domain signal of the first audio reference sound in the second frequency domain;
[0036] a frequency domain signal of the second ambient sound in the second frequency domain;
[0037] sound energy of a plurality of audio frames of the second ambient sound in the first frequency domain;
[0038] sound energy of a plurality of audio frames of the second ambient sound in the second frequency domain;
[0039] sound energy of a plurality of audio frames of the first audio reference sound in the second frequency domain
[0040] wherein the first frequency domain is a frequency domain remaining after a preset frequency domain is removed from the second frequency domain, and a minimum value of frequencies of the first frequency domain is greater than a maximum value of frequencies of the preset frequency domain.
[0041] In some embodiments, the method further comprises:
[0042] determining sound energy of a plurality of audio frames in the second ambient sound in a unit time;
[0043] determining sound energy of the first audio reference sound in the unit time;
[0044] determining whether the second ambient sound contains the speech sound or the mechanical sound if the sound energy of the plurality of audio frames in the second ambient sound and the sound energy of the first audio reference sound satisfy a preset relationship.
[0045] According to an embodiment of the first aspect of the present disclosure, an audio processing apparatus is provided, and the apparatus comprises:
[0046] an acquisition module configured to acquire sound in an environment to obtain a first ambient sound, the first ambient sound comprising at least a first audio echo, wherein the first audio echo refers to a sound signal collected by a microphone after a first audio is played by a loudspeaker;
[0047] a first processing module configured to remove the first audio echo in the first ambient sound according to a first audio reference sound to obtain a second ambient sound, wherein the first audio reference sound refers to audio source data of the first audio when the first audio is not played by the loudspeaker;
[0048] a second processing module configured to process a second audio to be played to obtain a third audio when the second ambient sound does not contain a speech sound and contains a mechanical sound;
[0049] a third processing module configured to send the third audio to the loudspeaker to play the third audio by the loudspeaker.
[0050] In some embodiments, the apparatus further comprises:
[0051] a first determination module configured to determine whether the second ambient sound contains the speech sound
[0052] if the second ambient sound does not contain the speech sound, determining whether the second ambient sound contains a mechanical sound.
[0053] In some embodiments, the apparatus further comprises:
[0054] a second determination module configured to determine sound energy of a plurality of audio frames in the second ambient sound in a preset time period;
[0055] a third determination module configured to determine whether the second ambient sound contains the speech sound according to fluctuation of the sound energy of the second ambient sound in the preset time period.
[0056] In some embodiments, the second determining module is further configured to:
[0057] determine a maximum frame energy and a minimum frame energy in the preset time period; the maximum frame energy is an audio frame with the maximum sound energy among a plurality of audio frames included in the preset time period; the minimum frame energy is an audio frame with the minimum sound energy among the plurality of audio frames included in the preset time period;
[0058] the third determining module is further configured to:
[0059] determine a ratio of the maximum frame energy and the minimum frame energy;
[0060] if the ratio is greater than or equal to a first threshold value, determine that the second environmental sound contains the speech sound.
[0061] In some embodiments, the third determining module is further configured to:
[0062] if the minimum frame energy among the plurality of audio frames included in the preset time period is less than a second threshold value, determine that the second environmental sound contains the speech sound.
[0063] In some embodiments, the apparatus further comprises:
[0064] a fourth determining module configured to determine a correlation degree of audio frames in a first sub-time period and a second sub-time period according to the second environmental sound collected in the first sub-time period and the second sub-time period; a maximum value of the first sub-time period is less than a maximum value of the second sub-time period, and the first sub-time period and the second sub-time period are both located in the preset time period;
[0065] if the correlation degree is greater than a correlation degree threshold value, determine that the second environmental sound contains the mechanical sound.
[0066] In some embodiments, the apparatus further comprises:
[0067] a fifth determining module configured to determine a mean value and a standard deviation of sound energy ratios of a plurality of audio frames in the preset time period respectively; the sound energy ratio of the audio frame is a ratio of a sound energy of an audio frame of the second environmental sound in a first frequency domain to a sound energy of the audio frame of the second environmental sound in a second frequency domain; the first frequency domain is a frequency domain remaining after a preset frequency domain is removed from the second frequency domain, and a minimum value of frequencies of the first frequency domain is greater than a maximum value of frequencies of the preset frequency domain;
[0068] if the mean value is greater than a mean value threshold value and the standard deviation is less than a difference value threshold value, determine that the second environmental sound contains the mechanical sound.
[0069] In some embodiments, the second processing module is further configured to:
[0070] obtain second audio to be played;
[0071] obtain prior information of the second environmental sound;
[0072] perform enhancement processing on a speech part contained in the second audio according to the prior information of the second environmental sound, to obtain the third audio.
[0073] In some embodiments, the apparatus further includes:
[0074] a sixth determining module configured to determine sound energy of a plurality of audio frames in the second environmental sound in a unit time;
[0075] determine sound energy of the first audio reference sound in the unit time;
[0076] if the sound energy of the plurality of audio frames in the second environmental sound and the sound energy of the first audio reference sound satisfy a preset relationship, determine whether the second environmental sound contains the speech sound or the mechanical sound.
[0077] According to a third aspect of the present disclosure, an electronic device is provided, including:
[0078] a processor;
[0079] a memory for storing processor-executable instructions;
[0080] The processor is configured to implement the method steps of the first aspect.
[0081] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, which stores a computer program, when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the method steps of the first aspect.
[0082] The technical solutions provided by the embodiments of the present disclosure can include the following beneficial effects:
[0083] As can be seen from the above embodiments, the present disclosure adjusts the second audio data only when the second environmental sound does not contain speech sound and contains mechanical sound, so that the electronic device can distinguish different sounds in the second environmental sound, and distinguish the mechanical sound as the interference source, rather than regarding all sounds in the second environmental sound as the interference source, for example, the speech sound in the second environmental sound is not regarded as the interference source. This way of processing sound effectively reduces the misadjustment of the second audio, improves the accuracy of processing the second audio, and improves the user's auditory experience.
[0084] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0085] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and serve to explain the principles of the present disclosure, in which:
[0086] Figure 1 is one of flowchart diagrams of an audio processing method according to an exemplary embodiment;
[0087] Figure 2 is another of flowchart diagrams of an audio processing method according to an exemplary embodiment;
[0088] Figure 3 is one of structural schematic diagrams of an audio processing apparatus according to an exemplary embodiment;
[0089] Figure 4 is another of structural schematic diagrams of an audio processing apparatus according to an exemplary embodiment;
[0090] Figure 5 is a constituent structural block diagram of an apparatus for audio processing according to an exemplary embodiment. DETAILED DESCRIPTION
[0091] The exemplary embodiments will be described in detail herein with reference to the attached drawings; figures, wherein like numerals indicate like elements, and wherein the various elements in the figures are not necessarily to scale, unless specifically stated, the use of the same numbers across different figures is intended to denote like elements. The following detailed description is not meant to limit the present disclosure to all of the embodiments described herein, but rather the present disclosure is intended to be a device consistent with some aspects of the present disclosure as described in the appended claims.
[0092] In a first aspect of the present disclosure, an audio processing method is provided, as shown in Figure 1 The method comprises:
[0093] In step S110, a sound in an environment is acquired to obtain a first ambient sound, the first ambient sound comprising at least a first audio echo, wherein the first audio echo refers to a sound signal collected by a microphone after a first audio is played by a loudspeaker;
[0094] In step S120, a first audio reference sound is used to remove the first audio echo in the first ambient sound to obtain a second ambient sound, wherein the first audio reference sound refers to audio source data of the first audio when the first audio is not played by the loudspeaker.
[0095] Step S130, if the second ambient sound does not contain speech sound and contains mechanical sound, processing the second audio to be played to obtain third audio.
[0096] Step S130, sending the third audio to a loudspeaker to play the third audio by the loudspeaker.
[0097] In steps S110 and S120, the electronic device includes but is not limited to: a television, a mobile phone, a tablet computer, a notebook computer, a sound box or a wearable device, and other devices with audio playing function.
[0098] The first ambient sound includes mechanical sound and non-mechanical sound. The mechanical sound includes the sound emitted by household appliances or other devices when they are in operation. For example: the mechanical sound can be the sound emitted by household appliances such as hair dryers, egg beaters, washing machines, and cleaning robots when they are in operation. Or, the mechanical sound is the sound of a car horn, the sound of cutting vegetables and meat, the sound of a power drill during decoration, or other noises that affect the user's listening to audio. Non-mechanical sound refers to other sounds other than mechanical sound, for example: non-mechanical sound includes but is not limited to: speech sound, first audio echo, or accidental sound caused by object falling, etc. Among them, the speech sound can be human speech or the sound of pets such as cats and dogs, etc.
[0099] The second ambient sound obtained after removing the first audio echo from the first ambient sound reduces the interference of the first audio echo, making it easier to distinguish speech sound and mechanical sound from the second ambient sound in the subsequent process, and also helps to further improve the accuracy of processing the second audio.
[0100] Without limitation, the first audio echo in the first ambient sound can be removed by Fourier transform (Short-Time Fourier Transform, STFT).
[0101] In step S130, by determining the category of the sound in the second ambient sound, when the second ambient sound does not contain speech sound and contains mechanical sound, the second audio to be played is processed to obtain third audio. This way not only makes the electronic device adaptively adjust the intensity of the interference source (i.e. noise) in the second ambient sound, but also has a distinguishing effect on different sounds in the second ambient sound, distinguishing mechanical sound as the interference source, rather than considering all sounds in the second ambient sound as the interference source, effectively reducing the misadjustment of the second audio.
[0102] For example: if the speech sound is not considered as the interference source, the user's voice will not be mistaken as the interference source, and thus will not cause the processing of the second audio. In this scenario, the user can have a normal conversation, reducing the energy competition between the sound played by the loudspeaker and the user's speech sound, improving the accuracy of processing the second audio, and improving the user's auditory experience.
[0103] Different sounds have different characteristics, and by analyzing the characteristics of different sounds, the weight of different sounds in the second environmental sound can be distinguished, and it can be determined whether the environmental sound includes speech sound or mechanical sound.
[0104] For example, compared with mechanical sound, speech sound has sparseness (also known as intermittency), and the sound energy fluctuates greatly in the time domain. If the second environmental sound has this characteristic, it can be determined that the second environmental sound contains speech sound. If the second environmental sound does not have this characteristic, it can be determined that the second environmental sound does not contain speech sound. Similarly, mechanical sound has continuity, and the sound energy fluctuates less in the time domain. If the second environmental sound has this characteristic, it can be determined that the second environmental sound contains mechanical sound. If the second environmental sound does not have this characteristic, it can be determined that the second environmental sound does not contain mechanical sound.
[0105] In some embodiments, the second audio and the first audio are different frames of the same audio, and the first audio is a previous frame and the second audio is a subsequent frame.
[0106] In some embodiments, the processing of the second audio to be played includes:
[0107] The volume of the second audio is increased.
[0108] In some embodiments, the processing of the second audio to be played further includes:
[0109] The volume of the second audio is increased by a first adjustment amount. The first adjustment volume can be five volume degrees or ten volume degrees. In this way, in the case that the second environmental sound contains mechanical sound, increasing the volume by a fixed first adjustment amount can make the played third audio more clearly recognized by the user, and at the same time, the electronic device does not need to calculate the volume adjustment amount, reducing the computing resources of the electronic device.
[0110] In other embodiments, the increasing of the volume of the second audio includes:
[0111] According to the volume of the mechanical sound, a second adjustment amount of the second audio is determined to obtain the volume of the third audio.
[0112] It can be understood that the second adjustment amount is proportional to the volume of the mechanical sound. In this way, when the volume of the mechanical sound increases, the adjustment amount is also increased to ensure that the audio played by the electronic device in the mechanical sound scene can be recognized by the user, and the adjustment effect of the second audio is improved.
[0113] In addition to adjusting the volume, the playback speed of the second audio or the third audio played by the electronic device can also be adjusted.
[0114] In some embodiments, the processing the second audio to be played can include:
[0115] slowing down the second audio to obtain the third audio.
[0116] In this way, in a scenario where the second ambient sound does not include the speech sound and includes the mechanical sound, by slowing down the playing speed of the electronic device, the electronic device can be slowed down in the playing in the scenario including the mechanical sound, and the amount of interference of the playing of the audio being watched or listened by the user due to the existence of the mechanical sound can be reduced.
[0117] In other optional embodiments, the method further includes:
[0118] if the second ambient sound includes the speech sound or does not include the mechanical sound, the second audio is not processed;
[0119] sending the second audio to a loudspeaker to play the second audio by the loudspeaker.
[0120] In actual application, if the second ambient sound includes the speech sound or the short-time fluctuation sound, the second audio can not be processed. The short-time fluctuation sound includes but is not limited to the sound produced by a falling object.
[0121] In the embodiments of the present disclosure, the speech sound such as the speaking sound of a person is not taken as the interference source, so that the sound played by the electronic device and the speech sound do not form a competitive relationship and do not hinder the normal communication of the user. For example, if the second ambient sound is not distinguished, the speech sound will also be considered as the interference source in the process of playing the second audio by the electronic device, and if the speech sound is produced in the communication of the user at this time, the processing of the second audio will also be triggered, which affects the communication of the user. In the embodiments of the present disclosure, the speech sound is not taken as the interference source, and the speech sound will not trigger the adjustment of the second audio, which further improves the user experience.
[0122] In other optional embodiments, whether the second ambient sound includes the speech sound or the mechanical sound is determined according to the following steps:
[0123] determining whether the second ambient sound includes the speech sound;
[0124] if the second ambient sound does not include the speech sound, determining whether the second ambient sound includes the mechanical sound.
[0125] In actual application, whether the speech sound is included is determined first, and then whether the mechanical sound is included is determined. Compared with determining the speech sound and the mechanical sound at the same time, the processing efficiency is ensured while the amount of calculation is saved, which is beneficial to use on the electronic device with small amount of calculation resources.
[0126] In some other optional embodiments, whether the second environmental sound contains speech sound is determined according to the following steps:
[0127] The sound energy of the second environmental sound in a preset time period is determined.
[0128] Whether the second environmental sound contains the speech sound is determined according to the fluctuation of the sound energy of the second environmental sound in the preset time period.
[0129] If the distribution of the sound energy of the second environmental sound in the preset time is uneven, it is determined that the second environmental sound contains speech sound with sound energy fluctuation characteristics. Since the occurrence of sound energy fluctuation and the like is a specific manifestation of the sparseness of speech sound, at this time, the second environmental sound can continue to be monitored, and the second audio does not need to be processed.
[0130] The preset time period can be 1 second, 2 seconds, 3 seconds, 4 seconds or 5 seconds, but is not limited thereto.
[0131] In some other optional embodiments, the determination of the sound energy of the second environmental sound in the preset time period includes:
[0132] The maximum frame energy and the minimum frame energy in the preset time period are determined. The maximum frame energy is the audio frame with the maximum sound energy among the audio frames included in the preset time period. The minimum frame energy is the audio frame with the minimum sound energy among the audio frames included in the preset time period.
[0133] The determination of whether the second environmental sound contains the speech sound according to the fluctuation of the sound energy of the second environmental sound in the preset time period includes:
[0134] The ratio of the maximum frame energy and the minimum frame energy is determined.
[0135] If the ratio is greater than or equal to a first threshold value, it is determined that the second environmental sound contains the speech sound.
[0136] In the embodiments of the present disclosure, the fluctuation of the sound energy of the second environmental sound in the sound collected in the preset time period is determined according to the maximum frame energy and the minimum frame energy in the preset time period. This calculation process does not require the delay and calculation complexity brought by a large deep learning model, has low calculation process complexity, good real-time performance, high feasibility, does not require cloud support, and can even be performed offline.
[0137] Without limitation, the first preset condition includes that the ratio of the maximum frame energy and the minimum frame energy is greater than or equal to a first threshold value.
[0138] The first threshold value is a preset value, also known as a silence threshold value.
[0139] In a specific example, the preset time period is represented by [μ, η], where μ represents the first frame of audio frames in the preset time period, and η represents the last frame of audio frames in the preset time period. A first threshold value is represented by δ. A sound energy is represented by E. Corresponding to the maximum frame energy, Corresponding to the minimum frame energy. If the ratio satisfies the following relationship:
[0140]
[0141] It can be determined that the second environmental sound includes speech sound.
[0142] In order to distinguish different application scenarios, the first threshold value can be further divided into δ1-δ2, that is, δ can be δ1, δ2, or any value between δ1 and δ2. Wherein, δ2>δ1. If the second environmental sound is mainly speech sound, at this time, it can be considered that there is no interference source in the second environmental sound, and the sparseness of speech sound is more obvious, and the ratio satisfies the following relationship:
[0143]
[0144] If there are strong mechanical sound and speech sound in the second environmental sound, the ratio satisfies the following relationship:
[0145]
[0146] At this time, the sparseness of speech sound is not obvious, and the energy gap between the speech sound and the mechanical sound is small. In this scenario, it can be determined that the second environmental sound does not include speech sound, and if it is further confirmed that the second environmental sound contains mechanical sound, the second audio is processed. Or, it can be determined that the second environmental sound contains speech sound, and the second audio is not processed, and the second audio is played.
[0147] If there is no speech sound in the second environmental sound, or the speech sound is small, and the second environmental sound is mainly mechanical sound, because the influence of mechanical sound on speech sound is large, it is even impossible to determine the sparseness of speech sound. At this time, the ratio satisfies the following relationship:
[0148]
[0149] In this scenario, it is determined that the second environmental sound does not contain speech sound.
[0150] In other some optional embodiments, the determining whether the second environmental sound contains the speech sound according to the fluctuation of the sound energy of the second environmental sound in the preset time period comprises:
[0151] If the minimum frame energy in the plurality of audio frames included in the preset time period is less than the second threshold value, it is determined that the mechanical noise is not included in the environmental sound.
[0152] The second threshold value can be a preset threshold value, and the second threshold value is also referred to as a silence threshold value.
[0153] If the energy of an audio frame in the preset time period is less than the second threshold value, it is considered that the feature of intermittent silence occurs, which does not conform to the characteristics of mechanical sound, and at this time, it can be considered that the second environmental sound contains speech sound but does not contain mechanical sound. The second audio is not processed, and the electronic device continues to be monitored.
[0154] In some embodiments, the determining whether the second environmental sound contains the speech sound according to fluctuations of sound energy of the second environmental sound in the preset time period comprises:
[0155] determining a maximum frame energy and a minimum frame energy in the preset time period; wherein the maximum frame energy is an audio frame with the maximum sound energy in a plurality of audio frames included in the preset time period, and the minimum frame energy is an audio frame with the minimum sound energy in the plurality of audio frames included in the preset time period:
[0156] If the minimum frame energy is less than the second threshold value, it is preliminarily determined that the environmental sound contains the speech sound.
[0157] If the ratio of the maximum frame energy to the minimum frame energy is greater than or equal to the first threshold value, it is determined that the second environmental sound contains the speech sound.
[0158] The ratio of the maximum frame energy to the minimum frame energy is determined in relation to the first threshold value, and the relationship between the minimum frame energy and the second threshold value is also determined. If the ratio of the maximum frame energy to the minimum frame energy is greater than the first threshold value, and the minimum frame energy is less than the second threshold value, it is determined that the second environmental sound contains the speech sound. In this way, the detection accuracy of the speech sound in the second environmental sound is further improved.
[0159] In other optional embodiments, whether the second environmental sound contains the mechanical sound is determined according to the following steps:
[0160] According to the second environmental sound collected in the first sub time period and the second sub time period, the correlation degree of audio frames in the first sub time period and the second sub time period is determined; wherein the maximum value of the first sub time period is less than the maximum value of the second sub time period, and the first sub time period and the second sub time period are both located in the preset time period.
[0161] If the correlation degree is greater than a correlation degree threshold value, it is determined that the second environmental sound contains the mechanical sound.
[0162] Confirming the correlation of the environmental sound information in different time periods can exclude short-time jitter noise and improve the accuracy of the judgment of mechanical sound. The short-time jitter noise is strong in occurrence and can be ignored.
[0163] When the second environmental sound correlation in different time periods is greater than the correlation threshold, it indicates that the second environmental sound includes mechanical sound with good continuity. If the correlation is less than the correlation threshold, it indicates that the second environmental sound does not include mechanical sound with good continuity. In this case, the electronic device can continue to be monitored, and the second audio does not need to be processed.
[0164] In some other optional embodiments, the first sub-time period and the second sub-time period partially overlap and partially offset in the time domain.
[0165] In comparison, when the first sub-time period and the second sub-time period partially overlap in the time domain, that is, the first sub-time period and the second sub-time period include certain overlapping frames, the overlapping frames can effectively reduce the error of the calculation process and help to increase the calculation accuracy.
[0166] In a specific example, the preset time period is still denoted by [μ,η]. The correlation threshold is denoted by ε. A plurality of frames in the interval [μ,η] are taken to form data blocks x1 and x2, which satisfy the following relationship:
[0167]
[0168]
[0169]
[0170] wherein is a preset parameter, and the parameter setting is different for different electronic devices, may represent the first sub-time period, may represent the second sub-time period, wherein and include certain overlapping frames. In this example, the preset time period is exemplarily determined to be 2 seconds, is the frame signal corresponding to the 0.5th to 1.5th second in the 2-second preset time period, and x1 is taken as is the frame signal corresponding to the 1st to 2nd second in the 2-second preset time period. The threshold judgment relationship of x1 and x2 is:
[0171] r(x1,x2)>ε (11)
[0172] wherein r(x1,x2) represents the Pearson correlation coefficient of x1 and x2.
[0173] In some optional embodiments, it is determined whether the second ambient sound comprises mechanical sound according to the following steps:
[0174] respectively determine a mean value and a standard deviation of sound energy ratio values of a plurality of audio frames in the preset time period; wherein the sound energy ratio value of the audio frame is a ratio value of sound energy of the audio frame of the second ambient sound in a first frequency domain and sound energy of the audio frame of the second ambient sound in a second frequency domain; wherein the first frequency domain is a frequency domain remaining after removing a preset frequency domain from the second frequency domain, and a minimum value of the first frequency domain is greater than a maximum value of the preset frequency domain;
[0175] If the mean value is greater than a mean value threshold and the standard deviation is less than a difference value threshold, it is determined that the second ambient sound contains the mechanical sound.
[0176] Generally, there will be noise residues after eliminating the echo generated by the loudspeaker, i.e., the collected audio information in the preset frequency domain is audio information corresponding to the noise residues, which has low practical value and is removed to reduce the calculation resources and improve the calculation efficiency.
[0177] The audio frame in the first frequency domain is generally an audio frame in the second frequency domain after removing the noise residues in the second ambient sound, wherein the frequency domain of the noise residues is the preset frequency domain. The audio information corresponding to the noise residues often lacks practical value. The preset frequency domain can refer to audio information corresponding to a frequency band below 2.5 kHz, but is not limited thereto.
[0178] The mean value of the sound energy ratio value of the audio frame is used to remove some noise residues concentrated in low frequencies, and the standard deviation of the sound energy ratio value of the audio frame is used to remove irregular dithering sound sources similar to the mechanical sound (not belonging to the mechanical sound).
[0179] In some embodiments, the determination of whether the second ambient sound comprises mechanical sound comprises:
[0180] determining a mean value and a standard deviation of sound energy ratio values of a plurality of audio frames in the preset time period according to the second ambient sound in the latter half of the preset time period;
[0181] If the mean value is greater than a mean value threshold and the standard deviation is less than a difference value threshold, it is determined that the ambient sound contains the mechanical sound.
[0182] Only calculating the second ambient sound in the latter half of the preset time period can further improve the accuracy of the calculation. Because the audio frames in the former half of the preset time period are more likely to have irregular dithering sound, which can not be considered as mechanical sound, such as the moment when the electrical appliance is just turned on, discarding the audio frame information in the former half of the time can better suppress accidental errors.
[0183] For example, the preset time period is represented by [μ, η], and the second half of the preset time period is:
[0184] In some embodiments, the second half of the preset time period is the second sub-time period.
[0185] In a specific example, mean(·) represents the mean value, and std(·) represents the standard deviation, represents the sound energy of the audio frame of the second environmental sound collected in the first frequency domain, E sig (i) represents the sound energy of the audio frame of the second environmental sound collected in the second frequency domain. In In the time period, it is determined whether the mean value and the standard deviation of the plurality of audio frame sound energy ratio values satisfy the following relationship:
[0186]
[0187]
[0188] If it is satisfied, it indicates that the sound energy distribution of the second environmental sound is uniform, the sound energy fluctuation is small, and the second environmental sound includes mechanical sound characteristics; if the above double constraints are not satisfied, the electronic device is continuously detected, and the second audio is not processed.
[0189] In other some optional embodiments, if the second environmental sound does not contain speech sound and contains mechanical sound, the second audio to be played is processed to obtain a third audio, including:
[0190] Obtaining the second audio to be played;
[0191] Obtaining prior information of the second environmental sound;
[0192] According to the prior information of the second environmental sound, the speech part contained in the second audio is enhanced to obtain the third audio.
[0193] According to the second environmental sound, the second audio is arranged, so that the second audio is only enhanced according to the mechanical sound, and the processing accuracy of the second audio is improved.
[0194] In other some optional embodiments, the prior information of the second environmental sound includes at least one of the following:
[0195] The frequency domain signal of the second environmental sound in the first frequency domain (denoted as ref(i, k) in the disclosure);
[0196] The frequency domain signal of the first audio reference sound in the second frequency domain (denoted as ref(i, k) in the disclosure);
[0197] a frequency domain signal of the second ambient sound in the second frequency domain (denoted by sig(i, k) in the present disclosure);
[0198] a sound energy of a plurality of audio frames of the second ambient sound in the first frequency domain (denoted by E in the present disclosure);
[0199] a sound energy of a plurality of audio frames of the second ambient sound in the second frequency domain (denoted by E sig in the present disclosure);
[0200] a sound energy of a plurality of audio frames of the first audio reference sound in the second frequency domain (denoted by E ref in the present disclosure);
[0201] The first frequency domain is a frequency domain remaining after removing a preset frequency domain from the second frequency domain, and a minimum value of the first frequency domain is greater than a maximum value of the preset frequency domain.
[0202] According to the prior information, the second audio can be enhanced by using an enhancement algorithm to obtain a third audio. The enhancement algorithm includes but is not limited to a dynamic range compression algorithm.
[0203] In some other optional embodiments, the method further comprises:
[0204] determining a sound energy of a plurality of audio frames of the second ambient sound in a unit time;
[0205] determining a sound energy of the first audio reference sound in the unit time;
[0206] if the sound energy of the plurality of audio frames of the second ambient sound and the sound energy of the first audio reference sound satisfy a preset relationship, determining whether the second ambient sound contains the speech sound or the mechanical sound.
[0207] The unit time can be 1 second, 2 seconds, 3 seconds, 4 seconds, etc. But it is not limited to this.
[0208] When the sound energy of the plurality of audio frames of the second ambient sound and the sound energy of the first audio reference sound satisfy the preset relationship, it indicates that there is a sound with a large volume in the second ambient sound, and it may be necessary to arrange the second audio.
[0209] If the sound energy of the plurality of audio frames of the second environmental sound in a unit time and the sound energy of the first audio reference sound do not satisfy the preset relationship, it indicates that the volume of the second environmental sound is small, and even if the second environmental sound does not contain speech sound and contains mechanical sound, the second audio does not need to be arranged. Therefore, under the premise that the sound energy of the plurality of audio frames of the second environmental sound in a unit time and the sound energy of the first audio reference sound satisfy the preset relationship, it can be further determined whether the second environmental sound contains speech sound or mechanical sound.
[0210] Non-limitingly, the preset relationship includes the following formula:
[0211]
[0212] In the above formula, a1, a2, b1, b2 are preset parameters, which are different for different electronic devices; d is a preset constant term, for example, d = 0.001; and The sequence numbers of the first audio frame and the last audio frame in a unit time are represented as E ref (i) is the sound energy of the first audio reference sound.
[0213] In some embodiments, as shown in Figure 2 The microphone 10 and the speaker 20 are corresponding hardware of the electronic device, the streaming media data is the source content (i.e. the first audio data), and the noise detection step involves an algorithm process, wherein the noise detection step S200 contains three algorithm steps, namely the noise energy detection step S230, the speech sound judgment step S240 and the mechanical sound judgment step S250. The specific working steps of the noise detection process include:
[0214] Step S210, microphone acquisition. This step includes step S110, that is, acquiring the sound in the environment through the microphone to obtain the first environmental sound. Specifically, when the electronic device is working, the microphone will collect the sound in the environment in real time;
[0215] Step S220, echo cancellation. This step includes: removing the first audio echo in the first environmental sound according to the first audio reference sound to obtain the second environmental sound.
[0216] Specifically, the first environment sound and the first audio reference sound collected by the microphone are preprocessed by frame division, windowing, etc., and then the loudspeaker echo is eliminated after eliminating the first audio echo. Since the echo cancellation technology needs to be processed in the frequency domain through short-time Fourier transform (STFT), in order to improve the operation efficiency, the STFT frequency domain data after echo cancellation processing can be directly delivered to the subsequent noise energy detection step, and the STFT frequency domain data after echo cancellation processing can be delivered frame by frame. Since the echo cancellation technology cannot remove all echoes, noise residues inevitably exist in the STFT data, and the noise residues are often concentrated in the frequency domain part below 2.5 kHz, and the corresponding spectrum line lacks practical value. In order to improve the algorithm efficiency, only the information corresponding to the spectrum line of the frequency band above 2.5 kHz in each audio frame in the collected audio information is calculated to obtain the second environment sound in the first frequency domain. In the first frequency domain, the sound energy of the plurality of audio frames in the second environment sound is confirmed according to the frequency domain signal of the second environment sound, and the sound energy of the first audio reference sound is confirmed according to the first frequency domain (i.e., without removing the information corresponding to the spectrum line of the frequency band above 2.5 kHz);
[0217] Step S230, noise energy detection. This step includes:
[0218] determining the sound energy of the plurality of audio frames in the second environment sound in a unit time;
[0219] determining the sound energy of the first audio reference sound in the unit time;
[0220] if the sound energy of the plurality of audio frames in the second environment sound and the sound energy of the first audio reference sound satisfy a preset relationship, determining whether the second environment sound contains the voice sound or the mechanical sound
[0221] Specifically, the data of the audio frames are accumulated to 1 second (a preset parameter, which can be changed). According to the sound energy of the first audio reference sound, dynamic threshold detection is performed on the sound energy of the second environment sound. If the sound energy of the second environment sound exceeds the threshold, it is determined that there is a strong sound source in the second environment sound, and the subsequent voice sound judgment step is performed. If the dynamic threshold is not reached, the electronic device continues to be detected to achieve the effect of monitoring, and the second audio is not processed. If the sound energy of the second environment sound in the unit time and the sound energy of the first audio reference sound satisfy the preset relationship, it is determined that the sound energy of the second environment sound exceeds the threshold, otherwise, it is determined that the sound energy of the second environment sound does not reach the dynamic threshold.
[0222] Step S240, voice sound decision. Specifically, when the accumulated data of the audio frames continuously collected in the time domain reaches 2 seconds, first, voice pause detection is performed, that is, when the sound energy of a certain audio frame in 2 seconds is less than a second threshold value, it is considered that the audio information in this period meets the sparsity of voice, and is determined as voice sound and the process is terminated to return to step S230 to continue listening; if the sound energy of all audio frames is greater than the second threshold value, voice energy jitter detection is performed, that is, the ratio of the maximum frame energy and the minimum frame energy in 2 seconds is compared with a first threshold value, if it is greater than the first threshold value, it is considered that it meets the characteristic of large sound energy fluctuation in the time domain, and is determined as voice sound and the process is terminated to return to step S230 to continue listening, if it is less than the first threshold value, subsequent steps are continued.
[0223] Step S250, mechanical sound decision. Specifically, the correlation of the audio frames corresponding to the spectral lines above 2.5 kHz in the 0.5-1.5 second and 1-2 second periods in the 2-second period is calculated, if the correlation is less than a correlation threshold value, it is determined as short-time jitter noise, and the second environmental sound does not include mechanical sound, at this time the process is terminated to return to step S230 to continue listening. If the correlation is greater than the correlation threshold value, it is determined as mechanical sound with strong correlation, and subsequent calculation is continued. The ratio of the sound energy of each audio frame corresponding to the spectral lines above 2.5 kHz to the sound energy of each audio frame without removing the spectral lines below 2.5 kHz is calculated, the mean and standard deviation of the sound energy ratio of each audio frame in 2 seconds are calculated, when the mean is greater than a mean threshold value and the standard deviation is less than a difference threshold value, it is determined as mechanical sound, and the corresponding prior data is provided for the adjustment of the first audio data in the subsequent steps. If the mean is less than the mean threshold value and the standard deviation is greater than the difference threshold value, it is considered that the environmental sound includes accidental high-correlation noise, and does not include mechanical sound, the process is terminated to return to step S230 to continue listening.
[0224] Step S260: voice perception enhancement, that is, processing the second audio to obtain the third audio. Depending on the algorithm used, in addition to the streaming data, prior information such as microphone acquisition data may also be required, and this part can adjust the output of the second audio data according to the actual technical route.
[0225] In some embodiments, the audio processing method comprises:
[0226] Echo cancellation and noise removal, wherein the echo cancellation comprises: removing the first audio echo and noise residue. Specifically, removing the first audio echo comprises: removing the first audio echo in the first environmental sound according to the first audio reference sound to obtain the second environmental sound in the second frequency domain. Removing the noise residue comprises: removing the second environmental sound in the preset frequency domain from the second environmental sound in the second frequency domain to obtain the second environmental sound in the first frequency domain.
[0227] The collected audio information of the microphone of the electronic device is denoted as mic(n), and the loudspeaker reference sound information is denoted as ref(n), where n represents the nth sampling point in a certain audio frame, the frame length of the audio frame is N, the common audio signal sampling rate is 8000 Hz, 16000 Hz, 44100 Hz or 48000 Hz, etc., and the disclosure takes 16000 Hz which is widely used as an example, and the frame signals are obtained by overlapping and windowing mic(n) and ref(n); wherein the window function is denoted as win(n), and any one of the commonly used Hanning window, Hamming window, etc. can be selected; the overlap ratio can be selected according to the actual situation, and the common overlap ratio is 25%, 50% or 75%, etc., and the example takes 50% as an example, that is, the time domain offset of each audio frame is N / 2; denoted as i, the audio frame number, the audio frame signal expression of the microphone collected audio information and the loudspeaker reference sound after overlapping and frame preprocessing can be obtained as follows:
[0228]
[0229]
[0230] Taking N=256 as an example, the echo cancellation process is expressed by the following relationship (3):
[0231] sig(i,k)=Γ(mic(i,n),ref(i,n),K),0≤k<K (3)
[0232] Where Γ represents the echo cancellation process, sig(i,k) represents the frequency domain signal after echo cancellation processing, k represents the kth frequency point in the ith frame signal, and the total number of frequency points in the audio frame signal is K, that is, the STFT length is K, and K must satisfy the constraint K≥N; the example takes K=256 as an example.
[0233] Due to the interference of noise residues after echo cancellation, the low-frequency part of the audio frame signal data is often severely distorted, and the subsequent sig(i,k) only uses the frequency points above 2.5 kHz, and the example selects K=256, that is, only the data of 40≤k<128 is used (considering the symmetry of Fourier spectrum), and is denoted as , wherein This paper takes r=40. The sound energy corresponding to sig(i,k) is denoted as E sig (i), The sound energy corresponding to ref(i,k) is denoted as The sound energy E ref (i) corresponding to the loudspeaker reference sound ref(i,k) (using the frequency domain expression) can be represented as:
[0234]
[0235]
[0236]
[0237] Noise energy detection. First, when the sound energy of the environmental sound information and the sound energy E of the reference sound information of the microphone ref (i) when the dynamic threshold relationship of formula (7) is met, i.e. the sound energy of the environmental sound information is above the dynamic threshold, continue the subsequent calculation.
[0238]
[0239] In the above formula, α1, α2, β1, β2 are preset parameters, which are different for different electronic devices; d is a preset constant term, usually d = 0.001; θ and The sequence numbers of the first and last audio frames in the current interval are represented by θ and η, and the unit time interval can be represented as The length of the unit time interval can be adjusted according to the device condition. In the example condition, the length of the unit time interval is 1 second, i.e. When the above formula is not met, the process of formula (1) to formula (7) is repeated to achieve the monitoring effect, which can also be called the monitoring mode.
[0240] Determine whether the second environmental sound contains mechanical sound. The calculation interval of the voice sound decision, i.e. the preset time interval, can be inconsistent with the unit time interval. The sequence numbers of the first and last audio frames in the preset time interval are represented by μ and η, and the preset time interval can be represented as [μ, η]. When the interval used by the voice sound decision once is greater than the previous step noise energy detection, the calculation is performed after accumulating several frames of data. In the example in this paper, the length of the preset time interval is 2 seconds, i.e. η = μ + 248. In this stage, first, perform silence detection, and the judgment relationship is:
[0241]
[0242] Where γ is a preset silence threshold, i.e. a second threshold. If the energy of a certain frame in the current calculation interval is less than γ, it is considered that the characteristics of intermittent silence of voice have appeared, the current sound source is determined to be voice, which belongs to a pseudo noise source and does not meet the characteristics of mechanical sound, and returns to the monitoring mode. The preset value of γ is different for different electronic devices. When min(·) is not less than γ, continue to the next step of maximum energy difference judgment, and the maximum energy difference judgment relationship is:
[0243]
[0244] wherein δ is a preset silence threshold, i.e., the first threshold value. If the ratio of the maximum frame energy and the minimum frame energy in the current interval is greater than δ, it is considered that the sound energy jitter characteristic in the time domain occurs, it is determined that the second environmental sound contains speech sound, belongs to the pseudo noise source, does not conform to the characteristics of the mechanical sound, and the second environmental sound does not include the mechanical sound, and the monitoring mode is returned. The preset value of δ is different for different electronic devices. When the ratio does not conform to the relationship of formula (9), it is indicated that the second environmental sound does not contain speech sound, and the subsequent steps are continued.
[0245] Next, the correlation measure is performed, and a plurality of frames in the interval μ to η are taken to form data blocks x1 and x2, and x1 and x2 satisfy the following relationship:
[0246]
[0247] wherein is a preset parameter, and the parameter setting is different for different electronic devices, and include a certain number of overlapping frames, and the overlapping frames can effectively reduce the mathematical calculation error caused by the STFT block effect. In this example, the frame signals corresponding to 0.5-1.5 seconds in a 2-second calculation interval are taken as x1, and the frame signals corresponding to 1-2 seconds in the 2-second calculation interval are taken as x2. The threshold value determination relationship of x1 and x2 is:
[0248] r(x1,x2) > ε (11)
[0249] wherein r(x1,x2) represents the Pearson correlation coefficient of x1 and x2; when the correlation coefficient (also referred to as the correlation degree) is greater than a preset correlation degree threshold value ε, it is considered that the mechanical sound has the characteristic of continuous correlation, and the subsequent steps are continued, and if the correlation degree is not greater than the correlation degree threshold value ε, it is required to return to the monitoring mode. When the condition of formula (11) is met, whether the mean value mean(·) and the standard deviation std(·) of the high-frequency energy proportion (i.e., the sound energy ratio of the audio frame) in each audio frame in the latter half of the preset time period [μ, η] satisfy the following threshold value determination relationship:
[0250]
[0251] wherein the mean value threshold φ and the standard deviation threshold The preset parameters are different for different electronic devices. If the double constraints of the mean value and the standard deviation are satisfied, it indicates that the audio spectrum energy distribution is uniform, the sound energy fluctuation is small, and the sound energy has typical mechanical sound characteristics. If the double constraints are not satisfied, the mean value of the sound energy ratio of the audio frame in the value monitoring mode is used to remove other sounds with low frequency concentrated in some frequency bands, and the standard deviation of the sound energy ratio of the audio frame is used to remove irregular jitter sound sources similar to mechanical sound. The formula (12) only calculates the audio frame in the second half of the interval [μ,η] because the probability of irregular jitter in the first half of the audio frame is higher, for example, the moment when the electrical appliance is just turned on. Discarding the audio frame in the first half can better reduce accidental errors.
[0252] The second audio is adjusted to obtain third audio. After determining that the second environmental sound contains mechanical sound, sig(i,k), ref(i,k) and E sig (i) of the perceptual enhancement algorithm are provided. E ref (i) and a series of prior information.
[0253] According to a second aspect of the embodiments of the present disclosure, as Figure 3 shown in the figure, a speech processing apparatus is provided, and the apparatus 300 includes:
[0254] An acquisition module 310 is configured to acquire sound in an environment to obtain first environmental sound, the first environmental sound at least including first audio echo, wherein the first audio echo refers to a sound signal collected by a microphone after first audio is played by a loudspeaker.
[0255] A first processing module 320 is configured to remove the first audio echo in the first environmental sound according to a first audio reference sound to obtain second environmental sound, wherein the first audio reference sound refers to audio source data of the first audio when the first audio is not played by the loudspeaker.
[0256] A second processing module 330 is configured to process second audio to be played to obtain third audio when the second environmental sound does not contain speech sound and contains mechanical sound.
[0257] A third processing module 340 is configured to send the third audio to a loudspeaker to play the third audio by the loudspeaker.
[0258] In some embodiments, the apparatus further includes:
[0259] A first determination module is configured to determine whether the second environmental sound contains the speech sound
[0260] If the second ambient sound does not include the speech sound, it is determined whether the second ambient sound includes a mechanical sound.
[0261] In some embodiments, the apparatus further includes:
[0262] a second determining module configured to determine sound energy of a plurality of audio frames in the second ambient sound within a preset time period;
[0263] a third determining module configured to determine, according to fluctuation of the sound energy of the second ambient sound within the preset time period, whether the second ambient sound includes the speech sound.
[0264] In some embodiments, the second determining module is further configured to:
[0265] determine a maximum frame energy and a minimum frame energy within the preset time period; wherein the maximum frame energy is an audio frame with the maximum sound energy among a plurality of audio frames included in the preset time period, and the minimum frame energy is an audio frame with the minimum sound energy among the plurality of audio frames included in the preset time period;
[0266] the third determining module is further configured to:
[0267] determine a ratio of the maximum frame energy and the minimum frame energy;
[0268] if the ratio is greater than or equal to a first threshold value, it is determined that the second ambient sound includes the speech sound.
[0269] In some embodiments, the third determining module is further configured to:
[0270] if the minimum frame energy among the plurality of audio frames included in the preset time period is less than a second threshold value, it is determined that the second ambient sound includes the speech sound.
[0271] In some embodiments, the apparatus further includes:
[0272] a fourth determining module configured to determine, according to the second ambient sound collected within a first sub-time period and a second sub-time period, a correlation degree of audio frames within the first sub-time period and the second sub-time period; wherein a maximum value of the first sub-time period is less than a maximum value of the second sub-time period, and the first sub-time period and the second sub-time period are both located within the preset time period;
[0273] the third determining module is configured to, if the correlation degree is greater than a correlation degree threshold value, determine that the second ambient sound includes the mechanical sound.
[0274] In some embodiments, the apparatus further includes:
[0275] The fifth determining module is used to determine the mean and standard deviation of the sound energy ratios of multiple audio frames within the preset time period; wherein, the sound energy ratio of the audio frames is: the ratio of the sound energy of the audio frame of the second ambient sound in the first frequency domain to the sound energy of the audio frame of the second ambient sound in the second frequency domain; wherein, the first frequency domain is: the frequency domain remaining after removing the preset frequency domain from the second frequency domain, and the minimum value of the frequency in the first frequency domain is greater than the maximum value of the frequency in the preset frequency domain;
[0276] If the mean is greater than the mean threshold and the standard deviation is less than the difference threshold, it is determined that the second ambient sound contains the mechanical sound.
[0277] In some embodiments, the second processing module is further configured to:
[0278] Get the second audio file to be played;
[0279] Obtain prior information about the second ambient sound;
[0280] Based on the prior information of the second ambient sound, the speech portion contained in the second audio is enhanced to obtain the third audio.
[0281] In some embodiments, the apparatus further includes:
[0282] The sixth determining module is used to determine the sound energy of multiple audio frames in the second ambient sound within a unit time period;
[0283] Determine the sound energy of the first audio reference tone within the unit time period;
[0284] If the sound energy of multiple audio frames of the second ambient sound satisfies a preset relationship with the sound energy of the first audio reference sound, it is determined whether the second ambient sound contains the speech sound or the mechanical sound.
[0285] like Figure 4 As shown, the electronic device is also called a multimedia terminal. The second or third audio is the audio content played by the electronic device. The mechanical sound specifically refers to the sound of household appliances. The microphone of the electronic device 11 collects the sound of the environment in real time, and the noise detection module 12 detects noise in real time. If the sound of household appliances 15, which is an interference source, is detected, the voice perception enhancement module 13 is notified to activate the enhancement function, process the second audio data, and provide the necessary prior information to the enhancement module. The enhancement algorithm in the voice perception enhancement module 13 includes, but is not limited to, dynamic range compression algorithms. After adjusting the first audio data to obtain the third audio, the user 14 will clearly feel that the played audio is clearer when the third audio is played. If the second audio is not adjusted, the noise of household appliances 11 will make it difficult for the user to hear the audio content played by the electronic device.
[0286] According to a third aspect of embodiments of the present disclosure, an electronic device is provided, comprising:
[0287] a processor;
[0288] a memory for storing processor-executable instructions;
[0289] wherein the processor is configured to implement the method steps of the first aspect of embodiments.
[0290] According to a third aspect of embodiments of the present disclosure, a computer-readable storage medium is provided, having stored thereon a computer program, which, when executed by a processor of an electronic device, causes the electronic device to perform the method steps of the first aspect of embodiments.
[0291] In exemplary embodiments, a plurality of modules, etc. in the audio processing apparatus can be implemented by one or more Central Processing Units (CPU), Graphics Processing Units (GPU), baseband processors (BP), Application-Specific Integrated Circuits (ASIC), DSP, Programmable Logic Devices (PLD), Complex Programmable Logic Devices (CPLD), Field-Programmable Gate Arrays (FPGA), general-purpose processors, controllers, microcontrollers (MCU), microprocessors, or other electronic elements, for executing the aforementioned methods.
[0292] Figure 5 is a block diagram of an apparatus 800 for audio processing according to an exemplary embodiment. The apparatus 800 can be, for example, a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, and the like.
[0293] Referring to Figure 5 , the apparatus 800 can include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0294] The processing component 802 generally controls the overall operations of the device 800, such as the operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 can include one or more processors 820 to execute instructions delivered from the memory 804 to complete all or part of the steps of the methods described above. In addition, the processing component 802 can include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 can include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0295] The memory 804 is configured to store various types of data to support the operations of the device 800. Examples of these data include instructions for any application or method operating on the device 800, contact data, phonebook data, messages, pictures, videos, and the like. The memory 804 can be implemented by any type of volatile or non-volatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0296] The power component 806 provides power to the various components of the device 800. The power component 806 can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 800.
[0297] The multimedia component 808 includes a screen providing an output interface between the device 800 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, swiping, and gestures on the touch panel. The touch sensors can not only sense the boundary of a touch or swipe action, but also detect duration and pressure associated with the touch or swipe action. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. The front and / or rear camera can receive external multimedia data when the device 800 is in an operation mode, such as a shooting mode or a video mode. Each of the front and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0298] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive an external audio signal when the device 800 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.
[0299] The I / O interface 812 provides an interface between the processing component 802 and peripheral interface modules, which can be a keypad, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0300] The sensor component 814 includes one or more sensors for providing status assessments of various aspects of the device 800. For example, the sensor component 814 can detect an open / closed position of the device 800, relative positioning of components, such as a display and a keypad of the device 800, a change of position of the device 800 or a component of the device 800, presence or absence of user contact with the device 800, changes in orientation or acceleration / deceleration / velocity of the device 800, and temperature changes of the device 800, among a plethora of other examples. The sensor component 814 can include proximity sensor configured to detect presence of an object in a proximity without any physical touch. The sensor component 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0301] The communication component 816 is configured to facilitate wired or wireless communication between the device 800 and other devices. The device 800 can access a wireless network based on a communication standard, such as WiFi, 4G, or 5G, or a combination thereof. In an example embodiment, the communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component 816 also includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) techniques, infrared data association (IrDA) techniques, ultra-wideband (UWB) techniques, Bluetooth (BT) techniques, and other techniques.
[0302] In exemplary embodiments, the apparatus 800 can be implemented using one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors or other electronic devices, to perform the above methods.
[0303] In exemplary embodiments, a non-transitory computer readable storage medium including instructions, such as the memory 804 including instructions, is also provided, which can be executed by the processor 820 of the apparatus 800 to complete the above methods. For example, the non-transitory computer readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disc, and an optical data storage device, etc.
[0304] Other embodiments of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the present application cover any and all variations of the application that come within the scope of the claims and their equivalents. It is intended that the specification and examples be considered exemplary only, with the true scope and spirit of the application being indicated by the following claims.
[0305] The methods disclosed in several method embodiments of the present disclosure can be combined, under the condition of no conflict, to obtain new method embodiments.
[0306] The features disclosed in several method or device embodiments of the present disclosure can be combined, under the condition of no conflict, to obtain new method embodiments or device embodiments.
[0307] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the present disclosure cover any and all variations of the application that come within the scope of the claims and their equivalents. It is intended that the specification and examples be considered exemplary only, with the true scope and spirit of the application being indicated by the following claims.
[0308] It should be understood that the present disclosure is not limited to the precise structures as set forth above and shown in the attached drawings and that various modifications and changes can be made without departing from the scope thereof, the scope being indicated by the appended claims.
Claims
1. An audio processing method, wherein, The method includes: Acquire sounds in the environment to obtain a first ambient sound, the first ambient sound including at least: a first audio echo, wherein the first audio echo refers to the sound signal collected by the microphone after the first audio is played by the speaker; The first audio echo in the first ambient sound is removed based on the first audio reference tone to obtain the second ambient sound, wherein the first audio reference tone refers to the audio source data of the first audio when it is not played by the speaker; If the second ambient sound does not contain speech but contains mechanical sounds, then the second audio to be played is processed to obtain the third audio. The third audio is sent to the speaker so that the speaker can play the third audio.
2. The method according to claim 1, wherein, Determine whether the second ambient sound contains speech or mechanical sound by following these steps; Determine whether the second ambient sound contains the speech sound; If the second ambient sound does not contain the spoken sound, then determine whether the second ambient sound contains mechanical sound.
3. The method according to claim 1, wherein, Determine whether the second ambient sound contains speech by following these steps: Determine the sound energy of multiple audio frames in the second ambient sound within a preset time period; Based on the fluctuation of sound energy in multiple audio frames of the second ambient sound within the preset time period, it is determined whether the second ambient sound contains the speech.
4. The method according to claim 3, wherein, The determination of the sound energy of the second ambient sound within a preset time period includes: Determine the maximum and minimum frame energy within the preset time period; wherein, the maximum frame energy is the audio frame with the highest sound energy among the multiple audio frames included in the preset time period; and the minimum frame energy is the audio frame with the lowest sound energy among the multiple audio frames included in the preset time period. The step of determining whether the second ambient sound contains the speech sound based on the fluctuation of the sound energy of the second ambient sound within the preset time period includes: Determine the ratio of the maximum frame energy to the minimum frame energy; If the ratio is greater than or equal to the first threshold, it is determined that the second ambient sound contains the speech.
5. The method according to claim 3 or 4, wherein, The step of determining whether the second ambient sound contains the speech sound based on the fluctuation of the sound energy of the second ambient sound within the preset time period includes: If the minimum frame energy among the multiple audio frames included in the preset time period is less than the second threshold, it is determined that the second ambient sound contains the speech.
6. The method according to claim 3, wherein, Determine whether the second ambient sound includes mechanical noise using the following steps: Based on the second ambient sound in the first and second sub-time periods, the correlation of audio frames in the first and second sub-time periods is determined; wherein the maximum value in the first sub-time period is less than the maximum value in the second sub-time period. If the correlation is greater than the correlation threshold, it is determined that the second ambient sound contains the mechanical sound.
7. The method according to claim 6, wherein, The first sub-time period and the second sub-time period partially overlap and partially offset each other in the time domain.
8. The method according to claim 3 or 6, wherein, Determine whether the second ambient sound includes mechanical noise using the following steps: The mean and standard deviation of the sound energy ratios of multiple audio frames within a preset time period are determined respectively; wherein, the sound energy ratio of the audio frames is: the ratio of the sound energy of the audio frame of the second ambient sound in the first frequency domain to the sound energy of the audio frame of the second ambient sound in the second frequency domain; wherein, the first frequency domain is: the frequency domain remaining after removing the preset frequency domain from the second frequency domain, and the minimum value of the frequency in the first frequency domain is greater than the maximum value of the frequency in the preset frequency domain; If the mean is greater than the mean threshold and the standard deviation is less than the difference threshold, it is determined that the second ambient sound contains the mechanical sound.
9. The method according to claim 1, wherein, If the second ambient sound does not contain speech but contains mechanical sounds, then the second audio to be played is processed to obtain a third audio, including: Obtain the second audio file to be played; Obtain prior information about the second ambient sound; Based on the prior information of the second ambient sound, the speech portion contained in the second audio is enhanced to obtain the third audio.
10. The method according to claim 9, wherein, The prior information of the second ambient sound includes at least one of the following: The frequency domain signal of the second ambient sound in the first frequency domain; The frequency domain signal of the first audio reference tone in the second frequency domain; The frequency domain signal of the second ambient sound in the second frequency domain; The sound energy of multiple audio frames of the second ambient sound in the first frequency domain; The sound energy of multiple audio frames of the second ambient sound in the second frequency domain; The sound energy of multiple audio frames of the first audio reference tone in the second frequency domain; Wherein, the first frequency domain is the frequency domain remaining after removing the preset frequency domain from the second frequency domain, and the minimum value of the first frequency domain frequency is greater than the maximum value of the preset frequency domain frequency.
11. The method according to claim 1, wherein, The method further includes: Determine the sound energy of multiple audio frames in the second ambient sound within a unit of time; Determine the sound energy of the first audio reference tone within the unit time period; If the sound energy of multiple audio frames of the second ambient sound satisfies a preset relationship with the sound energy of the first audio reference sound, it is determined whether the second ambient sound contains the speech sound or the mechanical sound.
12. An audio processing apparatus, wherein, The device includes: The acquisition module is used to acquire sounds in the environment to obtain a first ambient sound, the first ambient sound including at least: a first audio echo, wherein the first audio echo refers to the sound signal collected by the microphone after the first audio is played by the speaker; The first processing module is used to remove the first audio echo from the first ambient sound based on the first audio reference tone to obtain the second ambient sound, wherein the first audio reference tone refers to the audio source data of the first audio when it is not played by the speaker. The second processing module is used to process the second audio to be played when the second ambient sound does not contain speech but contains mechanical sound, to obtain the third audio. The third processing module sends the third audio to the speaker so that the speaker can play the third audio.
13. The apparatus according to claim 12, wherein, The device further includes: The first determining module is used to determine whether the second ambient sound contains the voice sound; If the second ambient sound does not contain the spoken sound, then determine whether the second ambient sound contains mechanical sound.
14. The apparatus according to claim 12, wherein, The device further includes: The second determining module is used to determine the sound energy of multiple audio frames in the second ambient sound within a preset time period; The third determining module is used to determine whether the second ambient sound contains the speech sound based on the fluctuation of the sound energy of the second ambient sound within the preset time period.
15. The apparatus according to claim 14, wherein, The second determining module is further configured to: Determine the maximum frame energy and minimum frame energy within the preset time period; wherein, the maximum frame energy is the audio frame with the highest sound energy among the multiple audio frames included in the preset time period; and the minimum frame energy is the audio frame with the lowest sound energy among the multiple audio frames included in the preset time period. The third determining module is further configured to: Determine the ratio of the maximum frame energy to the minimum frame energy; If the ratio is greater than or equal to the first threshold, it is determined that the second ambient sound contains the speech.
16. The apparatus according to claim 14 or 15, wherein, The third determining module is further configured to: If the minimum frame energy among the multiple audio frames included in the preset time period is less than the second threshold, it is determined that the second ambient sound contains the speech.
17. The apparatus according to claim 15, wherein, The device further includes: The fourth determining module is used to determine the correlation of audio frames in the first sub-time period and the second sub-time period based on the second ambient sound collected in the first sub-time period and the second sub-time period; wherein the maximum value of the first sub-time period is less than the maximum value of the second sub-time period, and both the first sub-time period and the second sub-time period are located within the preset time period; If the correlation is greater than the correlation threshold, it is determined that the second ambient sound contains the mechanical sound.
18. The apparatus according to claim 17, wherein, The device further includes: The fifth determining module is used to determine the mean and standard deviation of the sound energy ratios of multiple audio frames within the preset time period; wherein, the sound energy ratio of the audio frames is: the ratio of the sound energy of the audio frame of the second ambient sound in the first frequency domain to the sound energy of the audio frame of the second ambient sound in the second frequency domain; wherein, the first frequency domain is: the frequency domain remaining after removing the preset frequency domain from the second frequency domain, and the minimum value of the frequency in the first frequency domain is greater than the maximum value of the frequency in the preset frequency domain; If the mean is greater than the mean threshold and the standard deviation is less than the difference threshold, it is determined that the second ambient sound contains the mechanical sound.
19. The apparatus according to claim 12, wherein, The second processing module is further configured to: Get the second audio file to be played; Obtain prior information about the second ambient sound; Based on the prior information of the second ambient sound, the speech portion contained in the second audio is enhanced to obtain the third audio.
20. The apparatus according to claim 12, wherein, The device further includes: The sixth determining module is used to determine the sound energy of multiple audio frames in the second ambient sound within a unit time period; Determine the sound energy of the first audio reference tone within the unit time period; If the sound energy of multiple audio frames of the second ambient sound satisfies a preset relationship with the sound energy of the first audio reference sound, it is determined whether the second ambient sound contains the speech sound or the mechanical sound.
21. An electronic device, wherein, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute the steps of the method described in any one of claims 1 to 11 when implemented.
22. A computer-readable storage medium having a computer program stored thereon, wherein, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the steps of the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Voice enhancement processing method and device
CN104575509A
Audio suppression method, apparatus, medium, and apparatus for application
CN109166589A