Audio signal processing method and apparatus, storage medium, and vehicle
The proposed audio signal processing method addresses the limitations of current methods by using sound sensors to process noise signals and adjust audio sources in vehicles, leading to improved noise masking and user experience.
Patent Information
- Application Number
- JP2024568577
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-05-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-05-17
AI Technical Summary
Current audio signal processing methods in vehicles rely on non-acoustic measurement values, requiring extensive calibration and struggling to accurately adjust audio signals in changing environments, leading to a degraded auditory experience.
An audio signal processing method that utilizes sound sensors to collect and process audio signals, determining noise signals by processing human voice, harmonic, and burst sound information, and adjusting the original audio source based on these noise signals to improve noise masking and user experience.
This method accurately estimates current noise levels, reduces dependence on non-acoustic state information, and enhances the noise masking effect, resulting in an improved auditory experience for users.
Smart Images

Figure 2025517761000001_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to an audio signal processing method and apparatus, a storage medium, and a vehicle.
Background Art
[0002] In audio playback scenarios such as music playback, voice calls, navigation prompts, or human-machine interaction, the noise level affects people's audio experience. To obtain a better audio experience, the audio signal can be processed through volume adjustment or the like to reduce the energy of the noise perceived by people and reduce the noise interference received by people. However, for example, in a vehicle driving scenario, when the audio volume is manually adjusted, people's attention becomes distracted. This causes a safety risk and affects the driving experience.
[0003] In current solutions, the audio signal is usually processed by using non-acoustic measurement values (such as vehicle speed) to reduce noise interference. However, in this case, the relationship between the non-acoustic measurement value and the noise needs to be calibrated by relying on a large number of experiments. When the external environment changes, it is difficult to accurately determine the noise and adjust the reproduced audio signal. This causes a degraded user's auditory experience.
Summary of the Invention
[0004] In view of this, an audio signal processing method and apparatus, a storage medium, and a vehicle are proposed.
[0005] According to a first aspect, an embodiment of this application provides an audio signal processing method. The method includes: obtaining a first audio signal collected by a sound sensor; processing the first audio signal to determine a first noise signal in the first audio signal; adjusting a second audio signal based on the first noise signal and the second audio signal to obtain a third audio signal, where the second audio signal is an original audio source of a playback device, and the adjustment includes amplitude adjustment; and playing the third audio signal by using the playback device.
[0006] According to an embodiment of this application, the noise signal is determined by obtaining and processing an audio signal collected by a sound sensor. The audio signal collected by the sound sensor is fully utilized, and the dependence on non-acoustic state information is avoided. Furthermore, the current noise level can be accurately estimated, so that the estimated noise is more similar to the actual noise. The original audio source of the playback device is adjusted based on the noise signal and the original audio source to obtain an adjusted audio signal for playback by the playback device. The original audio source may be adjusted by using the noise signal, thereby achieving the effect of adapting to the noise environment. As a result, the adjusted audio signal has a better noise masking effect and the user's auditory experience is improved.
[0007] According to a first aspect, in a first possible implementation manner of the audio signal processing method, the step of processing the first audio signal to determine a first noise signal in the first audio signal includes processing one or more of human voice information, harmonic information, and burst sound information included in the first audio signal to determine the first noise signal in the first audio signal.
[0008] According to the embodiments of this application, one or more of the human voice information, harmonic information, and burst sound information included in the acquired audio signal collected by the sound sensor are processed to determine a noise signal. Therefore, the current noise level can be accurately estimated. As a result, the estimated noise is more similar to the actual noise, the adjusted audio signal has a better noise masking effect, the user's auditory experience is better, this adjustment method can be used in multiple scenarios, is more flexible, and supports rapid deployment.
[0009] According to a first aspect, in a first possible implementation manner of the audio signal processing method, the step of adjusting the second audio signal based on the first noise signal and the second audio signal to obtain a third audio signal is the step of determining a second noise signal based on the first noise signal and the transmission information, where the second noise signal is an estimated noise signal perceived by the user, and the step of adjusting the second audio signal based on the second noise signal and the second audio signal to obtain a third audio signal.
[0010] According to the embodiments of this application, the first noise signal is processed, and as a result, a second noise signal more similar to the noise actually perceived by the user can be obtained, and the obtained third audio signal can better mask the noise, thereby improving the user's auditory experience.
[0011] According to a first possible implementation manner of the first aspect, in a second possible implementation manner of the audio signal processing method, the transmission information includes the transmission information from the sound sensor to the user's human ear and / or the transmission information in the human ear.
[0012] According to the embodiments of this application, the noise transmission path may be more realistically simulated. As a result, the determined second noise signal is more similar to the noise actually perceived by the user.
[0013] According to the first or second possible implementation manner of the first aspect, in the third possible implementation manner of the audio signal processing method, the step of adjusting the second audio signal based on the second noise signal and the second audio signal to obtain a third audio signal includes a step of determining a gain curve based on the second noise signal and the second audio signal, and a step of adjusting the second audio signal based on the gain curve to obtain a third audio signal.
[0014] According to the embodiments of this application, the second audio signal can be adjusted by using the gain curve to obtain a third audio signal, thereby realizing the effect of masking noise by using the third audio signal and ensuring the user's hearing.
[0015] According to the first, second or third possible implementation manner of the first aspect, in the fourth possible implementation manner of the audio signal processing method, the step of adjusting the second audio signal based on the second noise signal and the second audio signal to obtain a third audio signal includes a step of determining a gain value based on the second noise signal and the second audio signal, and a step of adjusting the second audio signal based on the gain value to obtain a third audio signal.
[0016] According to the embodiments of this application, the second audio signal is adjusted by replacing the gain curve with a gain value. As a result, the obtained third audio signal has no sense of modulation, and the user's auditory experience is better.
[0017] According to the first, second, third, or fourth possible implementation manner of the first aspect, in the fifth possible implementation manner of the audio signal processing method, the step of adjusting the second audio signal based on the second noise signal and the second audio signal to obtain a third audio signal includes: determining a noise masking threshold of the second audio signal based on the second audio signal and psychoacoustic information, where the masking threshold indicates a volume threshold of noise masked by the second audio signal at each frequency, and noise with a volume lower than the volume threshold at each frequency is masked by the second audio signal; and adjusting the second audio signal based on the second noise signal and the masking threshold to obtain a third audio signal.
[0018] According to the embodiments of this application, the noise masking threshold of the second audio signal is determined by using psychoacoustic information. As a result, the third audio signal can be adjusted in a more targeted manner. In this way, a better noise masking effect can be obtained, and the user's hearing can be ensured.
[0019] According to the first aspect, or the first, second, third, fourth, or fifth possible implementation manner of the first aspect, in the sixth possible implementation manner of the audio signal processing method, the step of processing one or more of the human voice information, harmonic information, and burst sound information included in the first audio signal to determine a first noise signal in the first audio signal includes processing echo information and one or more of the human voice information, harmonic information, and burst sound information included in the first audio signal to determine a first noise signal in the first audio signal.
[0020] According to the embodiments of this application, echo information is processed, as a result, the first audio signal can be processed in a more targeted manner, and various scenarios are considered. Therefore, the noise signal can be more accurately separated from the first audio signal, and as a result, noise estimation is stabilized. In this way, after the second audio signal is adjusted, the noise signal can be better masked, and the user experience is improved.
[0021] According to the first aspect, or the first, second, third, fourth, fifth or sixth possible implementation manners of the first aspect, in the seventh possible implementation manner of the audio signal processing method, when processing one or more of the human voice information, harmonic information and burst sound information included in the first audio signal to determine the first noise signal in the first audio signal, the step of determining that the first noise signal is the first noise signal of the previous frame when it is determined that the first audio signal includes human voice information and / or harmonic information is included.
[0022] According to the embodiments of this application, the first noise signal may be directly obtained by using the first noise signal of the previous frame as the determined noise signal, and no processing for removing other information is required. This reduces the workload of the adjustment process and lowers the cost.
[0023] According to the first aspect, or the first, second, third, fourth, fifth, sixth or seventh possible implementation manners of the first aspect, in the eighth possible implementation manner of the audio signal processing method, the first audio signal includes the collected first audio signals of the current N frames, the second audio signal includes the second audio signals to be adjusted of the current N frames, the third audio signal includes the third audio signals of the current N frames, and N is a positive integer.
[0024] According to the embodiments of this application, the number of frames of the audio signal used is not limited. As a result, the amount of calculation in the adjustment process can be flexibly adjusted based on the actual situation, thereby facilitating the deployment in different scenarios.
[0025] According to a second aspect, an embodiment of this application provides an audio signal processing apparatus. The apparatus includes an acquisition module configured to acquire a first audio signal collected by a sound sensor, a first determination module configured to process one or more of human voice information, harmonic information, and burst sound information included in the first audio signal to determine a first noise signal in the first audio signal, a second determination module configured to adjust the second audio signal based on the first noise signal and the second audio signal to obtain a third audio signal, where the second audio signal is the original audio source of the playback device, and the adjustment includes amplitude adjustment, and a playback module configured to play the third audio signal by using the playback device.
[0026] According to a first possible implementation manner of the second aspect, in the second determination module of the audio signal processing apparatus, the second determination module is configured to determine a second noise signal based on the first noise signal and transmission information, where the second noise signal is an estimated noise signal perceived by the user, and is configured to adjust the second audio signal based on the second noise signal and the second audio signal to obtain a third audio signal.
[0027] According to the first possible implementation manner of the second aspect, in a second possible implementation manner of the audio signal processing apparatus, the transmission information includes transmission information from the sound sensor to the user's human ear and / or transmission information in the human ear.
[0028] According to the first or second possible implementation manner of the second aspect, in the third possible implementation manner of the audio signal processing apparatus, obtaining the third audio signal by adjusting the second audio signal based on the second noise signal and the second audio signal includes determining a gain curve based on the second noise signal and the second audio signal, and adjusting the second audio signal based on the gain curve to obtain the third audio signal.
[0029] According to the first, second or third possible implementation manner of the second aspect, in the fourth possible implementation manner of the audio signal processing apparatus, obtaining the third audio signal by adjusting the second audio signal based on the second noise signal and the second audio signal includes determining a gain value based on the second noise signal and the second audio signal, and adjusting the second audio signal based on the gain value to obtain the third audio signal.
[0030] According to the first, second, third or fourth possible implementation manner of the second aspect, in the fifth possible implementation manner of the audio signal device method, obtaining the third audio signal by adjusting the second audio signal based on the second noise signal and the second audio signal includes determining a noise masking threshold of the second audio signal based on the second audio signal and psychoacoustic information, where the masking threshold indicates a volume threshold of the noise masked by the second audio signal at each frequency, and the noise whose volume is lower than the volume threshold at each frequency is masked by the second audio signal, and adjusting the second audio signal based on the second noise signal and the masking threshold to obtain the third audio signal.
[0031] According to the second aspect, or the first, second, third, fourth or fifth possible implementation manner of the second aspect, in the sixth possible implementation manner of the audio signal processing apparatus, one determination module is configured to process echo information and one or more of human voice information, harmonic information and burst sound information included in the first audio signal to determine the first noise signal in the first audio signal.
[0032] According to the second aspect, or the first, second, third, fourth, fifth, or sixth possible implementation manner of the second aspect, in the seventh possible implementation manner of the audio signal processing apparatus, the first determination module is configured to determine that the first noise signal is the first noise signal of the previous frame when it is determined that the first audio signal includes human voice information and / or harmonic information.
[0033] According to the second aspect, or the first, second, third, fourth, fifth, sixth, or seventh possible implementation manner of the second aspect, in the eighth possible implementation manner of the audio signal processing apparatus, the first audio signal includes the collected first audio signals of the current N frames, the second audio signal includes the second audio signals to be adjusted of the current N frames, the third audio signal includes the third audio signals of the current N frames, and N is a positive integer.
[0034] According to the third aspect, an embodiment of this application provides an audio signal processing apparatus including a processor and a memory. The memory is configured to store a program, and the processor is configured to execute the program stored in the memory, so that the apparatus implements an audio signal processing method according to the first aspect or one or more of the plurality of possible implementation manners of the first aspect.
[0035] According to the fourth aspect, an embodiment of this application provides a terminal device. The terminal device may execute an audio signal processing method according to the first aspect or one or more of the plurality of possible implementation manners of the first aspect.
[0036] According to the fifth aspect, an embodiment of this application provides a computer-readable storage medium. The computer-readable storage medium stores program instructions, and when the program instructions are executed by a computer, the computer can implement an audio signal processing method according to the first aspect or one or more of the plurality of possible implementation manners of the first aspect.
[0037] According to a sixth aspect, an embodiment of this application provides a computer program product including program instructions. When the program instructions are executed by a computer, the computer can implement an audio signal processing method according to the first aspect or one or more of the plurality of possible implementation manners of the first aspect.
[0038] According to a seventh aspect, an embodiment of this application provides a vehicle. The vehicle includes a processor configured to execute an audio signal processing method according to the first aspect or one or more of the plurality of possible implementation manners of the first aspect.
[0039] These aspects and other aspects of this application will become clearer and more comprehensive in the following description of the (multiple) embodiments.
Brief Description of the Drawings
[0040] The accompanying drawings included in this specification and constituting a part of this specification, together with this specification, show exemplary embodiments, features, and aspects of this application and are intended to explain the principle of this application.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Embodiments for Carrying Out the Invention
[0041] Hereinafter, with reference to the accompanying drawings, various exemplary embodiments, features, and aspects of this application will be described in detail. The same reference numerals in the accompanying drawings indicate elements having the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the accompanying drawings are not necessarily drawn to scale unless otherwise specified.
[0042] The specific term "example" as used here means "used as an example, embodiment, or illustration." Any embodiment described as "exemplary" is not necessarily described as being superior or better than other embodiments.
[0043] Furthermore, for a better explanation of this application, in the following specific embodiments, a number of specific details are provided. Those skilled in the art should understand that this application can also be realized without some of these specific details. In some examples, methods, means, elements, and circuits well known to those skilled in the art are not described in detail, so that the subject matter of this application is emphasized.
[0044] For example, in a vehicle driving scenario, there are a large number of audio usage scenarios such as music playback, voice calls, navigation prompts, and human-machine interactions. The magnitude of the environmental noise affects people's auditory experience in the audio usage scenario. To obtain a better auditory experience, a method such as adjusting the volume may be used to adapt to the noise environment, reduce the energy of the noise perceived by people, and reduce the noise interference to people. However, frequent manual adjustment of the audio volume distracts people's attention, causes safety risks, and affects the driving experience. In current solutions, the audio signal is usually processed by using non-acoustic measurement values (e.g., vehicle speed) to reduce noise interference. However, in this case, the relationship between the non-acoustic measurement value and the noise needs to be calibrated by relying on a large number of experiments, and it is difficult to accurately determine the noise and adjust the reproduced audio signal when the external environment changes. Alternatively, the collected acoustic signal is used to fuzzily estimate the noise. However, in this case, the noise estimation is not accurate. Therefore, the user has a degraded auditory experience.
[0045] To solve the above technical problems, this application provides an audio signal processing method. In the audio signal processing method according to the embodiments of this application, an audio signal collected by a sound sensor is obtained, and the audio signal is processed to determine a noise signal in the audio signal. Therefore, the current noise level can be accurately estimated by using acoustic measurement values, the dependence on non-acoustic state information is avoided, and the estimated noise is similar to the actual noise. The original audio source of the playback device is adjusted based on the noise signal and the original audio source to obtain an adjusted audio signal for playback by the playback device. In this way, the original audio source can be adjusted by using the noise signal, thereby achieving the effect of adapting to the noise environment. As a result, the adjusted audio signal has a better noise masking effect and the user's auditory experience is better. Further, in the above process, based on the audio signal collected by the sound sensor, this adjustment method can be used in multiple scenarios, is more flexible, and supports rapid deployment.
[0046] (a) in FIG. 1 and (b) in FIG. 1 are schematic diagrams of application scenarios according to the embodiments of this application. As shown in (a) in FIG. 1 and (b) in FIG. 1, in a possible application scenario, the audio signal processing method according to the embodiments of this application may be applied to a scenario where noise masking is performed on a vehicle. The reproduced audio signal is adjusted to reduce the noise perception of the driver in the vehicle. The audio signal processing system according to the embodiments of this application may be arranged in a vehicle and includes a sound sensor, a processor, and a playback device.
[0047] The sound sensor (which may be the one shown in (a) in FIG. 1, for example, a microphone) may be disposed at any position within the vehicle, for example, near the driver within the vehicle, and is configured to collect an audio signal within the vehicle (which may be referred to as a first audio signal in (a) and (b) in FIG. 1) and determine ambient noise perceived by a user within the vehicle.
[0048] A processor, for example, a system on chip (SoC) or a digital signal processing (DSP) chip, may be incorporated into a head unit (or an audio system) within the vehicle as an in-vehicle computing unit. The processor may determine a noise signal corresponding to the ambient noise perceived by a user within the vehicle based on the audio signal collected by the sound sensor. The processor may further adjust the original audio source of the playback device (which may be referred to as a second audio signal with reference to (a) and (b) in FIG. 1) based on the determined noise signal and the original audio source, and determine an adjusted audio signal (which may be referred to as a third audio signal with reference to (a) and (b) in FIG. 1). Alternatively, the processor may be externally disposed in a cloud server. The server and the vehicle may communicate in a wireless connection manner, for example, by using mobile communication technologies such as 2G / 3G / 4G / 5G, and communicate in a wireless communication manner such as Wi-Fi, Bluetooth, frequency modulation (FM), wireless data transmission, or satellite communication. Through the communication between the vehicle and the server, the server may collect the audio signal collected by the sound sensor for calculation and return the calculation result to the corresponding vehicle.
[0049] The playback device (refer to (b) in FIG. 1) may be arranged in a vehicle, may include a speaker or the like, and may be configured to play an audio signal adjusted by a processor. FIG. 2 is a schematic diagram of adjusting an audio signal according to an embodiment of this application. As shown in FIG. 2, for example, in a scenario where music is played in a vehicle, when the noise outside the vehicle becomes large (for example, when the vehicle is passing through a congested road section) and the playback device still plays unadjusted music, as the noise increases, the noise perceived by the driver also increases, and it is certain that the driver's auditory experience will be affected. However, according to the embodiment of this application, the audio signal collected by the sound sensor is used to adjust the music played in this case, and the playback device plays the adjusted audio signal. For example, due to a change in volume (as shown in the drawing, the volume increases), the music heard by the driver can mask the noise perceived by the driver. In other words, since the music played is adjusted, the adjusted music can affect the noise-auditory effect in the driver's ears, and the noise perception of the driver in the vehicle can be reduced. In this way, since the driver cannot hear the noise, the auditory experience of the driver obtained when music is played in the vehicle is improved.
[0050] Although (a) in FIG. 1 and (b) in FIG. 1 show only one sound sensor, one processor, and one playback device, it should be understood that the audio signal processing system may alternatively include other numbers of sound sensors, processors, and playback devices.
[0051] The audio signal processing method in the embodiment of this application may also be applied to other scenarios where noise masking needs to be performed, other than the in-vehicle scenarios shown in (a) in FIG. 1, (b) in FIG. 1, and FIG. 2. For example, it may be applied to usage scenarios corresponding to electronic devices having an audio interaction function and a microphone, such as a mobile phone or a smart home. It should be noted that this is not limited in this application.
[0052] In the following, in order to describe in detail an audio signal processing method according to an embodiment of this application based on the above audio signal processing system, an in-vehicle scenario is used as an example.
[0053] FIG. 3 is a flowchart of an audio signal processing method according to an embodiment of this application. The method may be applied to the above audio signal processing system. As shown in FIG. 3, the method may include the following steps.
[0054] Step S301: Obtain a first audio signal collected by a sound sensor.
[0055] Regarding the sound sensor, refer to (a) in FIG. 1. The first audio signal may be a signal of one or more frames. In the case of a plurality of frames, the first audio signal may be a signal of N consecutive frames (the value of N may be preset), or a signal of N spaced-apart frames (for example, the N frames include signals determined at an interval of one frame). The environmental ambient acoustic information collected by the sound sensor is included. For example, it may include human voice information, harmonic information, burst sound information, echo information, and noise information.
[0056] The human voice information may include the voice of the driver inside the vehicle collected by the sound sensor. The harmonic information may include long vowels in the voice of the driver inside the vehicle, the sound of the speaker, etc. collected by the sound sensor. The burst sound information may include short-time burst sounds collected by the sound sensor, for example, when the door is opened and closed. The echo information may include the sound reproduced by the playback device collected by the sound sensor. The playback device may play audio such as music, navigation broadcasts, or other voice broadcasts, for example. The noise information may include ambient noise inside and outside the vehicle.
[0057] The ambient noise intensity perceived by a user inside a vehicle may be estimated by removing information other than noise information from a first audio signal, and thereby adjusting the audio signal in a more targeted manner so as to better mask the noise. For details of the process, refer to the following.
[0058] Step S302: Process the first audio signal to determine a first noise signal within the first audio signal.
[0059] The processing includes removing corresponding information from the first audio signal. To determine the first noise signal, for example, human voice information, harmonic information, burst sound information, echo information, etc. included in the first audio signal may be removed. FIG. 4 is a flowchart for processing a first audio signal according to an embodiment of this application. As shown in FIG. 4, the process of processing the first audio signal may include the following steps.
[0060] Step S401: Process the human voice information within the first audio signal.
[0061] The human voice information may correspond to audio information generated by the voice of the driver or a person outside the vehicle. A method such as voice activity detection (VAD) may be used for the processing. For example, whether the first audio signal includes human voice information may be determined by using the VAD method. If it is determined that the first audio signal includes human voice information, the human voice information included in the current first audio signal may be removed based on the first audio signal of the previous frame (for example, the first 3 to 5 frames) in a manner such as smooth interpolation, or the human voice information may be processed in other ways. This is not limited in this application.
[0062] Step S402: Process the harmonic information within the first audio signal.
[0063] The harmonic information may correspond to audio information generated by long vowels in sounds such as voice and speaker sound. A method such as long vowel detection (LVD) may be used for processing. For example, the LVD method may include collecting statistics regarding the energy peak of the first audio signal in the frequency domain, thereby determining whether the first audio signal contains harmonic information.
[0064] When it is determined that the first audio signal contains human voice information and / or harmonic information, the first noise signal may be determined as the first noise signal of the previous frame.
[0065] For example, in a per-frame processing scenario, the first noise signal of the previous frame may be directly used as the first noise signal of the current frame. Therefore, the first noise signal may be directly obtained, and no processing for removing other information is required. This reduces the workload of the adjustment process and lowers the cost.
[0066] Step S403: Process the burst sound information in the first audio signal.
[0067] The burst sound information may correspond to audio information generated by short-time sounds generated when the vehicle door is opened or closed, etc. A method such as minimum statistics (MS) may be used for processing. For example, the burst sound information included in the first audio signal may be estimated by using the MS method, and the estimated burst sound information is removed. Alternatively, the burst sound information included in the first audio signal may be estimated by using a method other than MS, thereby removing the burst sound information included in the first audio signal. This is not limited in this application.
[0068] The sequence for processing the above information is not limited in this application. For example, after the human voice information and harmonic information are processed, the burst sound information may be processed. Therefore, when the burst sound information is processed, the remaining human voice or harmonic information may be further removed. After the above one or more types of information included in the first audio signal are removed, the first noise signal may be determined.
[0069] Through the above process, one or more of the human voice information, harmonic information, and burst sound information included in the acquired audio signal collected by the sound sensor are processed, so that the current noise level can be estimated more accurately, and the estimated noise is more similar to the actual noise. In this way, the noise masking effect of the audio signal to be adjusted later may also be better, and the user's auditory experience may be better. Furthermore, this adjustment method may be used in multiple scenarios, is more flexible, and supports rapid deployment.
[0070] Optionally, the process of processing the first audio signal may further include the following steps.
[0071] Step S404: Process the echo information in the first audio signal.
[0072] The echo information may correspond to the audio information generated by the audio reproduced by the playback device. Therefore, the echo information in the first audio signal may be first removed by using a method such as a frequency domain adaptive filter (FDAF). This method may be a linear suppression method. For example, alternatively, any line echo cancellation (LEC) method other than FDAF may also be used.
[0073] When echo information is removed by a linear suppression method, residual echo information may exist, and the estimated noise value is not sufficiently accurate. This causes subsequent misadjustment and chain reactions of the second audio signal. Therefore, based on linear suppression, the residual echo information may be removed by using a residual echo suppression (RES) method. In this process, the first audio signals of several frames (for example, 3 to 5 frames) before the first audio signal of the current frame may be used.
[0074] In the process of removing echo information by using the FDAF and RES methods, spectral holes may be generated, that is, there may be an over-cancellation phenomenon on several frequencies of the first audio signal in the frequency domain. When the spectrum of the noise signal is generally considered to be smooth, a frequency smoothing (FS) method may be further used to compensate for the spectral holes.
[0075] In this way, the first audio signal may be processed in a more targeted manner, and various scenarios are considered. Therefore, the noise signal can be more accurately separated from the first audio signal, and as a result, the noise estimation becomes stable. In this way, after the second audio signal is adjusted, the noise signal can be better masked, and the user experience is better.
[0076] The order of processing echo information and processing human voice information, harmonic information, and burst sound information is not limited in this application. That is, it should be noted that the execution order of steps S401 to S404 is not limited. For example, the echo information may be processed first, and then one or more of the human voice information, harmonic information, and burst sound information are processed.
[0077] After the echo information and one or more of the human voice information, harmonic information, and burst sound information are removed, the acquired signal may be considered as a first noise signal, that is, the estimated ambient noise perceived by the user in the vehicle, and may be adjusted based on the first noise signal and the original audio source of the playback device to determine an adjusted audio signal for playback. Thus, the adjusted audio signal can mask the ambient noise perceived by the user in the vehicle, thereby achieving a noise masking effect. For the detailed process, refer to FIG. 3 again below.
[0078] Step S303: Adjust the second audio signal based on the first noise signal and the second audio signal to obtain a third audio signal.
[0079] The second audio signal is the original audio source of the playback device, and the adjustment may include amplitude adjustment. For the playback device, refer to (b) in FIG. 1. The original audio source may be music, navigation sound, voice call sound, etc. This is not limited in this application. The third audio signal obtained by performing amplitude adjustment on the second audio signal can mask the ambient noise, that is, can reduce the user's perception of the ambient noise, thereby achieving a noise masking effect.
[0080] Since there is a difference between the first noise signal and the noise actually perceived by the user's human ear, in order to enable the third audio signal obtained through adjustment to better mask the noise, the first noise signal may be processed to be more similar to the magnitude of the noise actually perceived by the user. Refer to the following.
[0081] Step S303 may include determining a second noise signal based on the first noise signal and the transmission information.
[0082] The second noise signal is an estimated noise signal perceived by the user.
[0083] For example, the second noise signal may be determined by weighting the first noise signal based on the transmission information. Alternatively, the second noise signal may be determined based on other methods by using the transmission information. This is not limited in this application.
[0084] The transmission information may include the transmission information from the sound sensor to the user's human ear and / or the transmission information in the human ear.
[0085] For example, the transmission information from the sound sensor to the user's human ear may indicate the noise transmission path from the sound sensor to the user's human ear (e.g., the area of the ear position), and may be determined by determining the relative position between the sound sensor and the user's human ear (or the area near the human ear). The transmission information in the human ear may indicate the noise transmission path in the user's semicircular canals (e.g., from the outer ear to the middle ear). The transmission information may be determined, for example, by using an attenuation function from the outer ear to the middle ear or an A-Weighted method, so that the calculation amount can be reduced and the performance can be improved. The second noise signal may alternatively be determined by other methods. This is not limited in this application.
[0086] According to the embodiments of this application, the noise transmission path may be simulated more realistically, and as a result, the determined second noise signal is more similar to the noise actually perceived by the user.
[0087] After the second noise signal is determined, the second audio signal may be adjusted based on the second noise signal and the second audio signal to obtain a third audio signal.
[0088] According to the embodiments of this application, the first noise signal is processed, and as a result, a second noise signal similar to the noise actually perceived by the user can be obtained. The obtained third audio signal can better mask the noise, thereby improving the user's auditory experience.
[0089] For example, the second audio signal may be multiplied by a gain based on the second noise signal and the second audio signal, thereby obtaining a third audio signal, and the gain may be a gain curve or a gain value. For details, refer to the following description.
[0090] Obtaining the third audio signal by adjusting the second audio signal based on the second noise signal and the second audio signal may include determining a gain curve based on the second noise signal and the second audio signal.
[0091] Determining a gain curve based on the second noise signal and the second audio signal may also be determining a gain curve based on the noise masking thresholds of the second noise signal and the second audio signal. The masking threshold may indicate the volume threshold of the noise masked by the second audio signal at each frequency. For the method of obtaining the masking threshold, refer to the following description. For example, the amplitude of the second noise signal corresponding to each frequency in the frequency domain may be subtracted from the volume threshold at the corresponding frequency in the masking threshold, thereby determining the gain curve. The gain curve may represent the amplitude gain corresponding to each frequency in the frequency domain.
[0092] After the gain curve is determined, the second audio signal may be adjusted based on the gain curve to obtain a third audio signal.
[0093] For example, the amplitude of the third audio signal corresponding to each frequency in the frequency domain may be determined by multiplying the value corresponding to each frequency on the gain curve by the amplitude of the second audio signal corresponding to the frequency in the frequency domain, thereby determining the third audio signal.
[0094] According to an embodiment of this application, the second audio signal can be adjusted by using the gain curve to obtain a third audio signal, thereby realizing the effect of masking noise by using the third audio signal and ensuring the user's hearing.
[0095] When the second audio signal is adjusted by using the gain curve, the adjusted audio signal may have an obvious sense of modulation and even distortion may occur. Therefore, the overall gain value may be used to replace the gain curve, thereby avoiding the excessive influence of some special values in the gain curve on the adjustment of the second audio signal and reducing the sense of modulation of the audio signal. For details, refer to the following description.
[0096] The second audio signal being adjusted based on the second noise signal and the second audio signal to obtain a third audio signal may include determining a gain value based on the second noise signal and the second audio signal.
[0097] Determining a gain value based on the second noise signal and the second audio signal may be to determine a gain curve based on the noise masking thresholds of the second noise signal and the second audio signal, and then determine the gain value based on the gain curve. The gain value may be, for example, the root mean square value or the weighted average value of all or some of the values on the gain curve. This is not limited in this application.
[0098] The gain value may be a single value. One gain value may be determined based on values corresponding to all frequencies on the gain curve, or may be determined based on values corresponding to some frequencies (e.g., 20 frequencies) on the gain curve. Alternatively, the gain value may be a plurality of values (e.g., 2 to 5 values), and the plurality of gain values may be individually determined based on, for example, the high-frequency, mid-frequency, and low-frequency portions of the gain curve.
[0099] After the gain value is determined, the second audio signal may be adjusted based on the gain value to obtain a third audio signal.
[0100] For example, the amplitude of the third audio signal corresponding to each frequency in the frequency domain may be determined by multiplying the gain value by the amplitude of the second audio signal at each frequency in the frequency domain, thereby determining the third audio signal. When the corresponding gain value is determined based on the high-frequency, intermediate-frequency, and low-frequency portions of the gain curve, the third audio signal multiplies the gain value corresponding to the high frequency by the amplitude of the second audio signal corresponding to the high-frequency portion in the frequency domain, multiplies the gain value corresponding to the intermediate frequency by the amplitude of the second audio signal corresponding to the intermediate-frequency portion in the frequency domain, and multiplies the gain value corresponding to the low frequency by the amplitude of the second audio signal corresponding to the low-frequency portion in the frequency domain, which may be determined.
[0101] According to an embodiment of this application, the second audio signal is adjusted by replacing the gain curve with the gain value, so that the obtained third audio signal has no modulation feeling and the user's auditory experience is better.
[0102] To simulate the noise masking of the audio signal, based on determining a second noise signal, a noise masking threshold of the second audio signal may be further obtained, which is the basis for calculating the gain value or the gain curve in the above description. For details, refer to the following description.
[0103] The second audio signal is adjusted based on the second noise signal and the second audio signal to obtain a third audio signal, which is It may include determining a noise masking threshold of the second audio signal based on the second audio signal and psychoacoustic information.
[0104] The masking threshold may indicate a volume threshold of noise masked by the second audio signal at each frequency, and noise with a volume lower than the volume threshold at each frequency may be masked by the second audio signal. For example, when the masking threshold of the second audio signal at 400 Hz is 30 dBspl, a noise signal below 30 dBspl may not be perceived by the user, thereby achieving the effect of masking the noise.
[0105] The psychoacoustic information may include information such as, for example, the user's threshold of hearing, loudness, pitch, and sound masking. The psychoacoustic information may be obtained based on, for example, a psychoacoustic model, such as a perceptual evaluation of audio Quality (PEAQ) model, a Johnston model, or a Terhardt model. This is not limited in this application. The volume threshold of noise that can be masked by the second audio signal at different frequencies may be determined based on the psychoacoustic information.
[0106] After the masking threshold is determined, the second audio signal may be adjusted based on the second noise signal and the masking threshold to obtain a third audio signal.
[0107] For example, refer to the above description. A gain value or a gain curve may be obtained by using the second noise signal and the masking threshold, and as a result, the second audio signal may be adjusted to obtain a third audio signal.
[0108] According to the embodiments of this application, the noise masking threshold of the second audio signal is determined by using psychoacoustic information. As a result, the third audio signal can be adjusted in a more targeted manner. In this way, a better noise masking effect can be obtained and the user's hearing can be ensured.
[0109] Since the human ear has different sensitivities to different loudness levels, in order to ensure the user's hearing, the third audio signal determined in step S303 may be further corrected in the frequency domain (for example, equal-loudness compensation is performed). For example, the third audio signal may be corrected based on loudness information, hearing threshold information, etc. in the psychoacoustic information. The loudness information may include an equal-loudness curve. That is, the compensation correction may be performed on the third audio signal at different frequencies based on the relationship between different pure tone sound pressure levels and frequencies obtained when the loudness perceived by the user's human ear within the audible frequency range is the same, thereby adapting to the sensitivity of the human ear to different loudness levels.
[0110] Step S304: Play the third audio signal by using a playback device.
[0111] According to the embodiments of this application, the noise signal is determined by acquiring and processing the audio signal collected by the sound sensor. The audio signal collected by the sound sensor is fully utilized, and the dependence on non-acoustic state information is avoided. Furthermore, the current noise level can be accurately estimated, and as a result, the estimated noise is more similar to the actual noise. The original audio source of the playback device is adjusted based on the noise signal and the original audio source to obtain an adjusted audio signal for playback by the playback device. The original audio source may be adjusted by using the noise signal, thereby achieving the effect of adapting to the noise environment. As a result, the adjusted audio signal has a better noise masking effect, and the user's auditory experience is improved.
[0112] In the process of processing the audio signal, the first audio signal may include the first audio signal collected from the current N frames, the second audio signal may include the second audio signal to be adjusted from the current N frames, and the third audio signal may include the third audio signal from the current N frames, where N is a positive integer.
[0113] The value of N may be preset. The signals of N frames may be N separated frames (for example, the N frames include signals determined at an interval of one frame), or N consecutive frames. This is not limited in this application.
[0114] For example, N may be 1. As a result, frame-by-frame processing may be performed to dynamically determine the third audio signal.
[0115] Therefore, the amount of calculation in the adjustment process may be flexibly adjusted based on the actual situation, thereby facilitating the deployment in different scenarios.
[0116] FIG. 5 is a diagram of the structure of an audio signal processing apparatus according to an embodiment of this application. As shown in FIG. 5, the apparatus includes an acquisition module 501 configured to acquire a first audio signal collected by a sound sensor, a first determination module 502 configured to process one or more of human voice information, harmonic information, and burst sound information included in the first audio signal to determine a first noise signal in the first audio signal, a second determination module 503 configured to adjust a second audio signal based on the first noise signal and the second audio signal to obtain a third audio signal, where the second audio signal is an original audio source of a playback device, and the adjustment includes amplitude adjustment, a playback module 504 configured to play the third audio signal by using the playback device and includes.
[0117] According to an embodiment of this application, an audio signal collected by a sound sensor is acquired, and one or more of human voice information, harmonic information, and burst sound information included in the audio signal are processed to determine a noise signal. Therefore, the current noise level can be accurately estimated, and as a result, the estimated noise is more similar to the actual noise. The original audio source of the playback device is adjusted based on the noise signal and the original audio source to obtain an adjusted audio signal for playback by the playback device. The original audio source may be adjusted by using the noise signal, thereby achieving the effect of adapting to the noise environment. As a result, the adjusted audio signal has a better noise masking effect and the user's auditory experience is better. Further, in the above process, non-acoustic measurement values are not used, the audio signal collected by the sound sensor is fully used, the dependence on non-acoustic state information is avoided, this adjustment method can be used in multiple scenarios, is more flexible, and supports rapid deployment.
[0118] Optionally, when the first decision module 502 determines that the first audio signal includes human voice information and / or harmonic information, the first decision module 502 may be configured to determine that the first noise signal is the first noise signal of the previous frame.
[0119] According to an embodiment of this application, the first noise signal may be directly obtained by using the first noise signal of the previous frame as the determined noise signal, and no process of removing other information is required. This reduces the workload of the adjustment process and lowers the cost.
[0120] Optionally, the first decision module 502 may be configured to process echo information and one or more of human voice information, harmonic information, and burst sound information included in the first audio signal to determine the first noise signal in the first audio signal.
[0121] According to an embodiment of this application, the echo information is processed, so that the first audio signal can be processed in a more targeted manner, and various scenarios are considered. Therefore, the noise signal can be more accurately separated from the first audio signal, and as a result, the noise estimation is stabilized. In this way, after the second audio signal is adjusted, the noise signal can be better masked, and the user experience is improved.
[0122] For example, the second decision module 503 may be configured to determine the second noise signal based on the first noise signal and the transmission information, where the second noise signal is an estimated noise signal perceived by the user, and may be configured to adjust the second audio signal based on the second noise signal and the second audio signal to obtain a third audio signal.
[0123] According to the embodiments of this application, the first noise signal is processed, and as a result, a second noise signal similar to the noise actually perceived by the user can be obtained, and the obtained third audio signal can better mask the noise, thereby improving the user's auditory experience.
[0124] The transmitted information may include the transmitted information from the sound sensor to the human ear of the user and / or the transmitted information in the human ear.
[0125] According to the embodiments of this application, the noise transmission path may be simulated more realistically, and as a result, the determined second noise signal is similar to the noise actually perceived by the user.
[0126] Optionally, obtaining the third audio signal by adjusting the second audio signal based on the second noise signal and the second audio signal may include determining a gain curve based on the second noise signal and the second audio signal, and adjusting the second audio signal based on the gain curve to obtain the third audio signal.
[0127] According to the embodiments of this application, the second audio signal can be adjusted by using the gain curve to obtain the third audio signal, thereby realizing the effect of masking the noise by using the third audio signal and ensuring the user's hearing.
[0128] Optionally, obtaining the third audio signal by adjusting the second audio signal based on the second noise signal and the second audio signal may include determining a gain value based on the second noise signal and the second audio signal, and adjusting the second audio signal based on the gain value to obtain the third audio signal.
[0129] According to the embodiments of this application, the second audio signal is adjusted by replacing the gain curve with a gain value, and as a result, the obtained third audio signal has no sense of modulation, and the user's auditory experience is better.
[0130] Optionally, obtaining the third audio signal by adjusting the second audio signal based on the second noise signal and the second audio signal may include determining a noise masking threshold of the second audio signal based on the second audio signal and psychoacoustic information, where the masking threshold indicates a volume threshold of the noise masked by the second audio signal at each frequency, and the noise whose volume is lower than the volume threshold at each frequency is masked by the second audio signal, and adjusting the second audio signal based on the second noise signal and the masking threshold to obtain the third audio signal.
[0131] According to the embodiments of this application, the noise masking threshold of the second audio signal is determined by using psychoacoustic information, and as a result, the third audio signal can be adjusted in a more targeted manner. In this way, a better noise masking effect can be obtained, and the user's hearing can be ensured.
[0132] The first audio signal may include the collected first audio signals of the current N frames, the second audio signal may include the second audio signals to be adjusted of the current N frames, the third audio signal may include the third audio signals of the current N frames, and N is a positive integer.
[0133] According to the embodiments of this application, the number of frames of the audio signal used is not limited, and as a result, the amount of calculation in the adjustment process can be flexibly adjusted based on the actual situation, thereby facilitating the deployment in different scenarios.
[0134] FIG. 6 is a diagram of the structure of an electronic device according to an embodiment of this application. The electronic device may be a terminal, such as a vehicle or a head unit, or may be a chip built into a terminal, and may implement the steps of the audio signal processing method shown in FIGS. 3 and 4, or may implement the functions of the modules of the audio signal processing apparatus shown in FIG. 5. As shown in FIG. 6, the electronic device 600 includes a processor 601 and an interface circuit 602 coupled to the processor. Although only one processor and one interface circuit are shown in FIG. 6, it should be understood that the electronic device 600 may include other numbers of processors and interface circuits.
[0135] The interface circuit 602 is configured to connect to other components of the terminal, such as a memory or other processors. The processor 601 is configured to perform signal interaction with other components by using the interface circuit 602. The interface circuit 602 may be an input / output interface of the processor 601.
[0136] The processor 601 may be a processor in an in-vehicle device such as a head unit, or may be a separately sold processing device.
[0137] For example, the processor 601 reads a computer program or instruction in a memory connected to the processor 601 by using the interface circuit 602, and decodes and executes the computer program or instruction. When the corresponding program or instruction is decoded and executed by the processor 601, the electronic device 600 may be capable of implementing the solution in the audio signal processing method provided in the embodiment of this application.
[0138] Optionally, these programs or instructions are stored in a memory outside the electronic device 600. When the above programs or instructions are decoded and executed by the processor 601, the memory temporarily stores some or all of the content of the above programs or instructions.
[0139] Optionally, these programs or instructions are stored in a memory within the electronic device 600. When the memory within the electronic device 600 stores a program or instruction, the electronic device 600 may be arranged as a terminal in the embodiments of this application.
[0140] Optionally, some of the content of these programs or instructions is stored in a memory outside the electronic device 600, and other content of these programs or instructions is stored in a memory within the electronic device 600.
[0141] FIG. 7 is a diagram of the structure of an electronic device according to an embodiment of this application. The electronic device may be a terminal, for example, a vehicle or a head unit, or may be a chip built into a terminal, and realizes the steps of the audio signal processing method shown in FIGS. 3 and 4, or realizes the functions of the modules of the audio signal processing device shown in FIG. 5. As shown in FIG. 7, the electronic device 700 includes a processor 701 and a memory 702 coupled to the processor. Although only one processor and one memory are shown in FIG. 7, it should be understood that the electronic device 700 may include other numbers of processors and memories.
[0142] The memory 702 is configured to store a computer program or computer instructions. When these computer programs or instructions are executed by the processor 701, the electronic device 700 may be enabled to realize the steps of the audio signal processing method in the embodiments of this application.
[0143] FIG. 8 is a diagram of the structure of an electronic device according to an embodiment of this application. As shown in FIG. 8, the electronic device 800 may be a terminal, for example, a vehicle or a head unit, or may be a chip built into the terminal, may implement the steps of the audio signal processing method shown in FIGS. 3 and 4, or may implement the functions of the modules of the audio signal processing apparatus shown in FIG. 5. The electronic device 800 includes at least one processor 1801, at least one memory 1802, and at least one communication interface 1803. Further, the electronic device may further include common components such as an antenna. Details are not described here.
[0144] Hereinafter, with reference to FIG. 8, the components in the electronic device 800 will be specifically described.
[0145] The processor 1801 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to control the execution of the above solution program. The processor 1801 may include one or more processing units. For example, the processor 1801 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, a neural-network processing unit (NPU), etc. Different processing units may be independent components or may be integrated into one or more processors.
[0146] The communication interface 1803 is configured to communicate with other electronic devices or communication networks, such as Ethernet, a radio access network (RAN), a core network, or a wireless local area network (WLAN).
[0147] The memory 1802 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, or a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Compact Disc Read-Only Memory (CD-ROM) or other compact disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store program code in the form of instructions or data structures and is accessible to a computer, but is not limited thereto. The memory may exist independently or be connected to the processor through a bus. The memory may alternatively be integrated with the processor.
[0148] The memory 1802 is configured to store application code for executing the above solutions, and the processor 1801 controls the execution. The processor 1801 is configured to execute the application code stored in the memory 1802.
[0149] In one example, referring to the audio signal processing apparatus shown in FIG. 5, the acquisition module 501 in FIG. 5 may be implemented by the communication interface 1803 in FIG. 8. The first determination module 502 and the second determination module 503 in FIG. 5 may be implemented by the processor 1801 in FIG. 8.
[0150] FIG. 9 is a diagram of the structure of an electronic device according to an embodiment of this application. The electronic device may be the above terminal, for example, a vehicle or a head unit, or may be a chip built into the terminal, and may execute the audio signal processing method shown in any one of FIGS. 3 and 4, or may implement the functions of the modules of the audio signal processing apparatus shown in FIG. 5. The electronic device 900 includes a sound sensor 901, a processing unit 902 coupled to the sound sensor 901, and a speaker 903 coupled to the processing unit 902. Although only one sound sensor, one speaker, and one processing unit are shown in FIG. 9, it should be understood that the electronic device 900 may include other numbers of sound sensors, speakers, and processing units.
[0151] The sound sensor 901 may include a capacitive microphone, a moving coil microphone, a laser microphone, etc. The sound sensor 901 is configured to collect the above first audio signal. The processing unit 902 may be configured to process the first audio signal to determine a noise signal in the first audio signal, and may further adjust the original audio source based on the noise signal and the original audio source to obtain an adjusted audio signal. The speaker 903 may be configured to play the adjusted audio signal. As a result, the noise masking effect of the played audio is better, thereby improving the user's auditory experience.
[0152] The electronic device in the embodiments of this application may be implemented by software, for example, a computer program or instructions. It should be understood that the corresponding computer program or corresponding instructions may be stored in the memory within the terminal. The processor reads the corresponding computer program or corresponding instructions in the memory and implements the above functions. Alternatively, the electronic device in the embodiments of this application may be implemented by hardware. The processing unit 902 is a processor.
[0153] In the above embodiments, the descriptions of the embodiments have their respective focuses. For parts not described in detail in the embodiments, refer to the relevant descriptions in other embodiments.
[0154] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include portable computer disks, hard disks, random-access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random-access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital video disc (DVD), memory stick, floppy disk, a mechanical coding device that stores instructions, such as a punched card or a groove-protrusion structure, and any suitable combination thereof.
[0155] The computer-readable program instructions or code described herein may be downloaded from a computer-readable storage medium to respective computing / processing devices, or may be downloaded from an external computer or an external storage device via a network such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium within each computing / processing device.
[0156] The computer program instructions used to execute the operations in this application may be in the form of assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Smalltalk and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user computer, partially on the user computer, as a stand-alone software package, partially on the user computer and partially on a remote computer, or entirely on a remote computer or server. When accompanied by a remote computer, the remote computer may be connected to the user computer on any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., connected over the Internet using an Internet service provider). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), is customized by using the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions to implement various aspects of this application.
[0157] Various aspects of this application are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowcharts and / or block diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, may be implemented by computer-readable program instructions.
[0158] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to generate a machine, such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create an apparatus for realizing the functions / operations specified in one or more blocks within the flowchart and / or block diagram. These computer-readable program instructions may alternatively be stored in a computer-readable storage medium. These instructions enable a computer, programmable data processing apparatus, and / or other devices to operate in a particular manner. Accordingly, the computer-readable medium storing the instructions includes an artifact containing instructions for realizing various aspects of the functions / operations specified in one or more blocks within the flowchart and / or block diagram.
[0159] These computer-readable program instructions may alternatively be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are executed on the computer, other programmable data processing apparatus, or other device to generate a computer-implemented process. Accordingly, the instructions executed on the computer, other programmable data processing apparatus, or other device realize the functions / operations specified in one or more blocks within the flowchart and / or block diagram.
[0160] The flowcharts and block diagrams in the accompanying drawings illustrate the system architecture, functions, and operations of possible implementations of apparatuses, systems, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of an instruction, and a module, a program segment, or a part of an instruction may include one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions marked in the blocks may also be performed in an order different from the order marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or depending on the functions involved, may be executed in the reverse order in some cases.
[0161] It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by hardware (e.g., a circuit or an ASIC (Application-Specific Integrated Circuit)) that performs the corresponding functions or operations, or may be implemented by a combination of hardware and software, such as firmware.
[0162] Regarding this application, although embodiments have been described with reference to them, in the process of implementing this application for which protection is claimed, those skilled in the art may understand and implement other variant forms of the disclosed embodiments by looking at the accompanying drawings, the disclosed content, and the appended claims. In the claims, "comprising" does not exclude other components or other steps, and "one" does not exclude multiple cases. A single processor or other unit may implement some of the functions listed in the claims. Although some means are described in different dependent claims, this does not mean that these means cannot be combined to produce better effects.
[0163] The embodiments of this application have been described above. The above description is by way of example and not exhaustive and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other skilled artisans to understand the embodiments disclosed herein.
Claims
1. An audio signal processing method, comprising: obtaining a first audio signal collected by a sound sensor; processing one or more of human voice information, harmonic information, and burst sound information included in the first audio signal to determine a first noise signal in the first audio signal; adjusting the second audio signal based on the first noise signal and a second audio signal to obtain a third audio signal, wherein the second audio signal is an original audio source of a playback device, and the adjustment includes amplitude adjustment; playing the third audio signal by using the playback device. A method comprising the above steps.
2. The step of adjusting the second audio signal based on the first noise signal and the second audio signal to obtain a third audio signal includes: determining a second noise signal based on the first noise signal and transmission information, wherein the second noise signal is an estimated noise signal perceived by a user; adjusting the second audio signal based on the second noise signal and the second audio signal to obtain the third audio signal. The method according to claim 1, comprising the above steps.
3. The method according to claim 2, wherein the transmission information includes transmission information from the sound sensor to the human ear of the user and / or transmission information in the human ear.
4. The step of adjusting the second audio signal based on the second noise signal and the second audio signal to obtain a third audio signal includes: determining a gain curve based on the second noise signal and the second audio signal; adjusting the second audio signal based on the gain curve to obtain the third audio signal. The method according to claim 2 or 3, comprising the above steps.
5. The step of adjusting the second audio signal based on the second noise signal and the second audio signal to obtain a third audio signal includes: determining a gain value based on the second noise signal and the second audio signal; adjusting the second audio signal based on the gain value to obtain the third audio signal. The method according to any one of claims 2 to 4, comprising
6. The step of adjusting the second audio signal based on the second noise signal and the second audio signal to obtain the third audio signal includes: Determining a noise masking threshold of the second audio signal based on the second audio signal and psychoacoustic information, the masking threshold indicating a volume threshold of noise masked by the second audio signal at each frequency, and noise having a volume lower than the volume threshold at each frequency being masked by the second audio signal; Adjusting the second audio signal based on the second noise signal and the masking threshold to obtain the third audio signal The method according to any one of claims 2 to 5, comprising
7. The step of processing one or more of human voice information, harmonic information, and burst sound information included in the first audio signal to determine a first noise signal in the first audio signal includes: The method according to any one of claims 1 to 6, comprising processing echo information and one or more of the human voice information, harmonic information, and burst sound information included in the first audio signal to determine the first noise signal in the first audio signal.
8. The step of processing one or more of human voice information, harmonic information, and burst sound information included in the first audio signal to determine a first noise signal in the first audio signal includes: The method according to any one of claims 1 to 7, comprising determining that the first noise signal is the first noise signal of the previous frame when it is determined that the first audio signal includes the human voice information and / or the harmonic information.
9. The first audio signal includes the collected first audio signals of the current N frames, the second audio signal includes the second audio signals to be adjusted of the current N frames, the third audio signal includes the third audio signals of the current N frames, and N is a positive integer. The method according to any one of claims 1 to 8.
10. An audio signal processing apparatus, comprising An acquisition module configured to acquire a first audio signal collected by a sound sensor; A first determination module configured to process one or more of human voice information, harmonic information, and burst sound information included in the first audio signal to determine a first noise signal in the first audio signal; A second determination module configured to adjust the second audio signal based on the first noise signal and a second audio signal to obtain a third audio signal, wherein the second audio signal is an original audio source of a playback device, and the adjustment includes amplitude adjustment; A playback module configured to play the third audio signal by using the playback device An audio signal processing apparatus including the above.
11. An audio signal processing apparatus including a processor and a memory, The memory is configured to store a program, The processor is configured to execute the program stored in the memory to enable the apparatus to implement the method according to any one of claims 1 to 9. An audio signal processing apparatus.
12. A computer-readable storage medium, The computer-readable storage medium stores program instructions, and when the program instructions are executed by a computer, the computer can implement the method according to any one of claims 1 to 9. A computer-readable storage medium.
13. A computer program product, The computer program product includes program instructions, and when the program instructions are executed by a computer, the computer can implement the method according to any one of claims 1 to 9. A computer program product.
14. A vehicle, the vehicle includes a processor, and the processor is configured to execute the method according to any one of claims 1 to 9. A vehicle.
Citation Information
Patent Citations
Method and device for adjusting output audio according to environmental noise, equipment and medium
CN112306448A
SOUND MODULATOR, METHOD AND COMPUTER PROGRAM
JP2007500466A
Signal processing method, signal processing device and hearing device
JP2021157134A
Dynamic Audibility Enhancement
US20120045069A1