Sound field expansion method, audio device and computer-readable storage medium
By obtaining the target transfer function of the near-ear open audio device to eliminate crosstalk and adjusting the sound intensity weight ratio in the initial reverberation audio, the problem of the human voice becoming muffled due to the expansion of the sound field is solved, and the effect of expanding the sound field while ensuring the human voice sound effect is achieved.
Patent Information
- Application Number
- CN202211195319.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-29
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2042-09-29
AI Technical Summary
Existing sound field expansion technology in near-ear open audio devices easily causes the human voice to become muffled, especially when using the head-related transfer function (HRTF) algorithm for sound field expansion, the human voice effect is poor.
By obtaining the target transfer function between the near-ear open audio device and the user's ears, crosstalk cancellation processing is performed, the actual sound intensity weight ratio between the human voice audio and the accompaniment audio in the initial reverberation audio is identified, and the sound intensity is adjusted according to the actual weight ratio to ensure that the human voice effect remains unchanged when the sound field expands.
While effectively expanding the sound field, it improves the human voice effect in near-ear open audio devices, enhances the user's listening experience, and avoids crosstalk problems.
Smart Images

Figure CN115604630B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio processing technology, and in particular to a sound field expansion method, an audio device, and a computer-readable storage medium. Background Art
[0002] Sound field expansion refers to the acoustic phenomenon in which the perceived sound field is wider than the actual speaker location. Sound field expansion is similar to a virtual speaker that can extend the sound position to a wider position than the actual speaker location. That is, the sound played by the sound source sounds to the human ear equivalent to the effect of the sound being emitted from a virtual speaker in a wider position.
[0003] In the field of audio processing technology, actual audio signals are mostly two-channel stereo signals. Sound field expansion technology builds on two-channel stereo without adding channels or speakers. By processing the signal, the listener perceives that the sound comes from multiple directions, creating a simulated stereo field. Currently, sound field expansion technology (also known as virtual surround sound technology) has become an indispensable technology. It is mainly used for far-field sound sources, such as in scenarios using speakers. With the increasing market shipments of near-ear open-back audio devices such as VR and AR in recent years, the demand for sound field expansion functions of near-ear open-back audio devices has gradually increased.
[0004] However, current sound field expansion (also known as virtual surround sound) is primarily achieved through the Head Related Transfer Function (HRTF) algorithm. This often results in the human voice sounding muffled. Therefore, it is crucial to ensure that the human voice remains clear while effectively expanding the sound field. Summary of the Invention
[0005] The main purpose of the present invention is to provide a sound field expansion method, audio equipment and computer-readable storage medium, aiming to solve the technical problem that the human voice part in the audio played by the near-ear open audio device after adding the sound field expansion function is poor.
[0006] To achieve the above object, the present invention provides a sound field expansion method, which includes the following steps:
[0007] Obtaining a target transfer function between the near-ear open audio device and the user's ears;
[0008] performing crosstalk cancellation processing on the input audio received by the near-ear open audio device according to the target transfer function to obtain initial reverberation audio;
[0009] identifying an actual sound intensity weight ratio between the human voice audio and the accompaniment audio in the initial reverberation audio, and adjusting the sound intensity of the human voice audio and / or the accompaniment audio in the initial reverberation audio according to the actual sound intensity weight ratio to obtain a target reverberation audio;
[0010] Play the target reverb audio.
[0011] Optionally, the step of adjusting the intensity of the vocal audio and / or the accompaniment audio in the initial reverberation audio according to the actual intensity weight ratio to obtain the target reverberation audio includes:
[0012] Obtaining a target sound intensity weight ratio between the vocal audio and the accompaniment audio;
[0013] According to the actual sound intensity weight ratio and the target sound intensity weight ratio, the sound intensity of the human voice audio and / or the accompaniment audio in the initial reverberation audio is adjusted to adjust the actual sound intensity weight ratio to the target sound intensity weight ratio to obtain the target reverberation audio.
[0014] Optionally, the step of obtaining a target sound intensity weight ratio between the vocal audio and the accompaniment audio includes:
[0015] Identifying the initial reverberation audio using a converged neural network model to obtain an audio type corresponding to the initial reverberation audio;
[0016] According to the audio type, the sound intensity weight ratio of the audio type mapping is obtained from the preset mapping data table, and the sound intensity weight ratio of the audio type mapping is used as the target sound intensity weight ratio between the human voice audio and the accompaniment audio.
[0017] Optionally, the step of adjusting the intensity of the vocal audio and / or the accompaniment audio in the initial reverberation audio includes:
[0018] Increasing the intensity of the human voice audio in the initial reverberation audio; and / or
[0019] The volume of the accompaniment audio in the initial reverberation audio is reduced.
[0020] Optionally, the step of obtaining a target transfer function between the near-ear open audio device and the user's ears includes:
[0021] Obtaining a preset artificial head transfer function and a free-field transfer function;
[0022] performing an inverse operation on the free-field transfer function to obtain an inverse free-field transfer function;
[0023] The artificial head transfer function is multiplied by the free-field inverse transfer function to obtain a target transfer function between the near-ear open audio device and the user's ears.
[0024] Optionally, the step of obtaining a preset artificial head transfer function and a free-field transfer function includes:
[0025] When the near-ear open-back audio device is worn on a preset artificial head and the near-ear open-back audio device outputs a sound signal, measuring the artificial head transfer function through a preset microphone in the ear canal of the artificial head; and
[0026] When the artificial head is removed and the near-ear open audio device outputs a sound signal, a free-field transfer function is measured by preset microphones placed at the left and right ear positions before the artificial head is removed.
[0027] Optionally, the step of performing crosstalk cancellation processing on the input audio received by the near-ear open audio device according to the target transfer function to obtain initial reverberation audio includes:
[0028] performing an inverse operation on the target transfer function to obtain a target inverse transfer function;
[0029] The input audio received by the near-ear open audio device is multiplied by the target inverse transfer function to obtain initial reverberation audio.
[0030] Optionally, the step of identifying an actual sound intensity weight ratio between the human voice audio and the accompaniment audio in the initial reverberation audio includes:
[0031] Dividing the initial reverberation audio into a plurality of frames, wherein each frame of the initial reverberation audio has an accompaniment audio and a vocal audio in a time-synchronized relationship;
[0032] Performing windowing processing on each frame of the initial reverberation audio, and converting the windowed initial reverberation audio from the time domain to the frequency domain by fast Fourier transform to obtain an initial reverberation spectrum;
[0033] Decomposing the initial reverberation spectrum to obtain an accompaniment spectrum and a vocal spectrum in the initial reverberation spectrum;
[0034] Determining an actual sound intensity weight ratio between the vocal audio and the accompaniment audio in the initial reverberation spectrum based on the accompaniment spectrum and the vocal spectrum;
[0035] The step of adjusting the sound intensity of the vocal audio and / or the accompaniment audio in the initial reverberation audio according to the actual sound intensity weight ratio and the target sound intensity weight ratio to obtain the target reverberation audio includes:
[0036] According to the actual sound intensity weight ratio and the target sound intensity weight ratio, the human voice spectrum in the initial reverberation spectrum is subjected to sound intensity amplification processing, and / or the accompaniment spectrum in the initial reverberation audio is subjected to sound intensity reduction processing to obtain a target reverberation spectrum;
[0037] The target reverberation spectrum is converted from the frequency domain to the time domain to obtain the target reverberation audio.
[0038] In addition, to achieve the above-mentioned purpose, the present invention also provides an audio device, which includes: a memory, a processor, and a sound field expansion program stored in the memory and runnable on the processor, and when the sound field expansion program is executed by the processor, the steps of the sound field expansion method described above are implemented.
[0039] In addition, to achieve the above objectives, the present invention further provides a computer-readable storage medium, on which a sound field expansion program is stored. When the sound field expansion program is executed by a processor, the steps of the sound field expansion method described above are implemented.
[0040] The present invention obtains a target transfer function between a near-ear open-ear audio device and the user's ears, and then performs crosstalk cancellation processing on the input audio received by the near-ear open-ear audio device according to the target transfer function to obtain initial reverberation audio, so that the ears of the user wearing the near-ear open-ear audio device receive a sound signal consistent with the input audio, eliminating the interference of the near-ear open-ear audio device itself on the sound signal. In the case where the speaker of the near-ear open-ear audio device cannot be placed in the human ear like headphones, the listening effect of the sound played by the near-ear open-ear audio device when it is transmitted to the user's ears is consistent with the listening effect when wearing headphones, effectively improving the listening experience of the user group of the near-ear open-ear audio device and avoiding the crosstalk problem. However, since the current sound field expansion function is mainly achieved through the Head Related Transfer Function (HRTF) algorithm, the use of HRTF sound field expansion often results in the effect of making the human voice become muffled. That is, the initial reverberation audio obtained after the sound field expansion often has a smaller actual sound intensity weight ratio between the human voice audio and the accompaniment audio. In other words, the weight of the sound intensity of the human voice audio in the initial reverberation audio is often smaller, while the weight of the sound intensity of the accompaniment audio in the initial reverberation audio is often larger. Therefore, the present invention dynamically identifies the actual sound intensity weight ratio between the human voice audio and the accompaniment audio in the initial reverberation audio, and determines whether the actual sound intensity weight ratio is within a preset standard sound intensity weight ratio range. If it exceeds the preset standard sound intensity weight ratio range, it means that the audio currently played by the near-ear open audio device when performing HRTF sound field expansion has the problem of human voice becoming hollow. Therefore, the present invention adjusts the sound intensity of the human voice audio and / or the accompaniment audio in the initial reverberation audio, increases the weight of the sound intensity of the human voice audio in the initial reverberation audio, obtains the target reverberation audio and plays it, thereby improving the problem of human voice becoming hollow when the near-ear open audio device performs HRTF sound field expansion. That is, the present invention utilizes the extraction of the accompaniment audio signal and the vocal signal of the song to be processed, and then adjusts the sound intensity of the accompaniment audio signal and / or the vocal signal of the initial reverberation audio according to the reverberation degree values of the extracted accompaniment audio signal and the vocal signal, so as to achieve the goal of ensuring the sound effect of the vocal while effectively expanding the sound field, and overcome the technical problem of poor sound effect of the vocal part in the audio played by the near-ear open audio device after adding the sound field expansion function. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 Schematic diagram of the flow of the first embodiment of the sound field expansion method of the present invention;
[0042] Figure 2 Schematic diagram of the flow of the second embodiment of the sound field expansion method of the present invention;
[0043] Figure 3A schematic diagram of an application scenario of an embodiment of a sound field expansion method of the present invention;
[0044] Figure 4 1. A schematic diagram of a flow chart for identifying the actual intensity weight ratio between human voice audio and accompaniment audio in one embodiment of the present invention;
[0045] Figure 5 The figure is a schematic diagram of the structure of an audio device involved in an embodiment of the present invention.
[0046] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0047] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0048] The main solution of the embodiment of the present invention is: a sound field expansion method, which includes the following steps:
[0049] Obtaining a target transfer function between the near-ear open audio device and the user's ears;
[0050] performing crosstalk cancellation processing on the input audio received by the near-ear open audio device according to the target transfer function to obtain initial reverberation audio;
[0051] identifying an actual sound intensity weight ratio between the human voice audio and the accompaniment audio in the initial reverberation audio, and adjusting the sound intensity of the human voice audio and / or the accompaniment audio in the initial reverberation audio according to the actual sound intensity weight ratio to obtain a target reverberation audio;
[0052] Play the target reverb audio.
[0053] Since sound field expansion technology (i.e., virtual surround sound technology) has become an indispensable technology, it is mainly used in far-field sound sources, such as scenarios using speakers. With the increasing market shipments of near-ear open-back audio devices such as VR and AR in recent years, the demand for sound field expansion functions of near-ear open-back audio devices has gradually increased. However, the current sound field expansion function (i.e., virtual surround sound function) is mainly achieved through the Head Related Transfer Function (HRTF) algorithm. When using HRTF sound field expansion, it often makes the human voice sound muffled.
[0054] The present invention dynamically identifies the actual sound intensity weight ratio between the human voice audio and the accompaniment audio in the initial reverberation audio, and determines whether the actual sound intensity weight ratio is within a preset standard sound intensity weight ratio range. If it exceeds the preset standard sound intensity weight ratio range, it means that the audio currently played by the near-ear open audio device when performing HRTF sound field expansion has the problem of human voice becoming hollow. Therefore, the present invention adjusts the sound intensity of the human voice audio and / or the accompaniment audio in the initial reverberation audio, increases the weight of the sound intensity of the human voice audio in the initial reverberation audio, obtains the target reverberation audio and plays it, thereby improving the problem of human voice becoming hollow when the near-ear open audio device performs HRTF sound field expansion. That is, the present invention utilizes the extraction of the accompaniment audio signal and the vocal signal of the song to be processed, and then adjusts the sound intensity of the accompaniment audio signal and / or the vocal signal of the initial reverberation audio according to the reverberation degree values of the extracted accompaniment audio signal and the vocal signal, so as to achieve the goal of ensuring the sound effect of the vocal while effectively expanding the sound field, and overcome the technical problem of poor sound effect of the vocal part in the audio played by the near-ear open audio device after adding the sound field expansion function.
[0055] The embodiment of the present invention provides a sound field expansion method, referring to Figure 1 , Figure 1 FIG. 4 is a flow chart of an embodiment of a sound field expansion method according to the present invention.
[0056] In this embodiment, the sound field expansion method includes:
[0057] Step S10, obtaining a target transfer function between the near-ear open audio device and the user's ears;
[0058] In this embodiment, the execution subject is a near-ear open audio device, which includes but is not limited to AR (Augmented Reality), VR (Virtual Reality), smart audio glasses, neck-hanging speakers, open headphones and other products. Compared with the speaker scene, the position of the speaker or loudspeaker of the near-ear open audio device is closer to the human ear, and the near-ear open device is generally an all-in-one machine. The distance between the playback device and the human ear is almost unadjustable, that is, under normal circumstances, the sound played by the near-ear open audio device heard by the human ear is a near-field sound effect. Figure 3 To understand, Figure 3 This is a schematic diagram of an application scenario provided by this embodiment. Assume that the relative position of the user's head and the device speaker is as follows: Figure 3As shown, it is obvious that when users use near-ear open-back audio devices, they will be affected by the environment and crosstalk problems will occur. For example, the transfer function of the near-ear open-back audio device itself affects the results, and the sound signal passes through the user's head contour and then enters the ear canal. As a result, the listening experience cannot achieve the best effect. Therefore, it is necessary to improve the user's listening experience by obtaining the target transfer function between the near-ear open-back audio device and the user's ears.
[0059] In this embodiment, the target transfer function between the near-ear open-back audio device and the user's ears, that is, the transfer function from the output sound source (i.e., the speaker or loudspeaker) of the near-ear open-back audio device to the user's ears, is used to reflect the changes in the input audio during the process of the input audio of the near-ear open-back audio device being transmitted to the user's ears.
[0060] Based on this, in a feasible embodiment, the above step S10 may include:
[0061] Step S11, obtaining a preset artificial head transfer function and a free-field transfer function;
[0062] It should be noted that in this embodiment, the target transfer function can be understood as the effect of the user's head contour on the sound signal transmission result. However, this target transfer function cannot be directly derived, but is instead calculated after deriving two different acoustic transfer functions based on two different sound transmission scenarios. The artificial head transfer function is the acoustic transfer function measured by a preset microphone in the ear canal of the artificial head when the near-ear open-back audio device is worn on a preset artificial head and the near-ear open-back audio device outputs a sound signal. This function includes the effect of the playback device and the human head contour on the sound transmission result. The free-field transfer function is the acoustic transfer function measured by preset microphones placed at the left and right ear positions before the artificial head is removed and the near-ear open-back audio device outputs a sound signal. This function includes the effect of the playback device on the sound transmission result.
[0063] It is easy to understand that when currently near-ear open audio devices are used to expand the sound field, the various transfer functions mentioned (such as the artificial head transfer function and the free-field transfer function) are all head-related transfer functions.
[0064] Furthermore, in a feasible embodiment, the step of obtaining the artificial head transfer function in the above step S11 may include:
[0065] Step S111, when the near-ear open-back audio device is worn on a preset artificial head and the near-ear open-back audio device outputs a sound signal, measuring an artificial head transfer function through a preset microphone in the ear canal of the artificial head; and
[0066] It should be noted that the preset artificial head is an auxiliary device constructed to simulate the user's head for assisting in measuring the acoustic transfer function. It can simulate the scenario in which the user receives sound signals emitted from the speakers of a near-ear open audio device. The preset artificial head is provided with left and right ears and ear canals, and a microphone for receiving sound signals can be pre-placed in the ear canal.
[0067] As an example, combining Figure 3 As can be seen from the application scenario shown, the near-ear open audio device is worn on the artificial head, and the two preset microphones in the ear canal of the artificial head are used to measure the acoustic transfer function from the sound source (i.e., the speaker or horn of the near-ear open audio device) to the two ears of the artificial head, and it is recorded as H1.
[0068] Step S112 , when the artificial head is removed and the near-ear open audio device outputs a sound signal, a free-field transfer function is measured by using preset microphones placed at the left and right ear positions before the artificial head is removed.
[0069] As an example, combining Figure 3 As can be seen from the application scenario shown, two microphones consistent with those in the ear canal of the artificial head in the above step S111 are first placed at the positions of the left and right ears of the artificial head, and then the artificial head is removed. The acoustic transfer function of the sound source when working in the free field is measured using two microphones that are not affected by the artificial head, and is recorded as H2.
[0070] After step S11, step S12 is performed: performing an inverse operation on the free-field transfer function to obtain a free-field inverse transfer function;
[0071] Step S13: multiplying the artificial head transfer function by the free-field inverse transfer function to obtain a target transfer function between the near-ear open audio device and the user's ears.
[0072] In this embodiment, the free-field transfer function H2, obtained in step S112 and including the effect of the playback device on the sound transfer result, is first inverted to obtain an inverse free-field transfer function, denoted as H2'. The artificial head transfer function H1, obtained in step S111 and including the effect of the playback device and the human head contour on the sound transfer result, is then multiplied by H2' to obtain the target transfer function H. It should be noted that H2', obtained after the inversion operation, eliminates the effect of the playback device on the sound transfer result. Multiplying it by H1 eliminates the portion of H1 affected by the playback device, retaining the effect of the human head contour on the sound transfer result as the target transfer function H.
[0073] Step S20, performing crosstalk cancellation processing on the input audio received by the near-ear open audio device according to the target transfer function to obtain initial reverberation audio;
[0074] It is understandable that since the speakers or horns of the near-ear open audio device are not ideal sound sources, and the speakers or horns of the near-ear open audio device cannot be directly placed in the user's ear canals as playback devices, crosstalk problems will inevitably occur during the transmission of the initial reverberation audio. In order to avoid this problem, the input audio is first subjected to crosstalk cancellation processing before the initial reverberation audio is generated, that is, the crosstalk problem that occurs during the transmission of the initial reverberation audio after playback is offset. That is, the initial reverberation audio is obtained after the input audio is subjected to crosstalk cancellation processing, which can offset the influence of the playback device itself during the sound transmission process and the influence of the user's head on the sound signal.
[0075] As an example, the above step S20 may include:
[0076] Step S21, performing an inverse operation on the target transfer function to obtain a target inverse transfer function;
[0077] Step S22: multiply the input audio received by the near-ear open audio device by the target inverse transfer function to obtain initial reverberation audio.
[0078] From the above steps, it can be seen that the target transfer function H represents the influence of the human head contour on the sound transmission result. It should be understood that the target inverse transfer function obtained after inverting H is equivalent to a unit matrix, which represents the elimination of the influence of the human head contour on the sound transmission result. The initial reverberation audio obtained after applying it to the input audio for processing can obviously offset the influence of the human head contour on the sound transmission result when the sound signal is transmitted, so that the audio received by the user's two ears can be consistent with the input audio.
[0079] As an example, combining Figure 3 As can be seen from the application scenario shown, when an input audio signal X is given from a near-ear open-type audio device, the input audio signal X is processed by the crosstalk cancellation algorithm module and then output by the SPK (speaker). The output signal is then transmitted to the human ear through the human head model. The basic idea behind implementing the crosstalk cancellation algorithm module is to first obtain the transfer function H from the sound produced by the SPK to the human ear, and then invert this transfer function through the crosstalk cancellation algorithm module. The combined effect of these two functions can achieve the effect of reducing crosstalk and eliminating crosstalk. If the inverse of H is denoted as C, then the initial reverberated audio signal Y = XCH is the audio signal after crosstalk elimination.
[0080] Step S30, identifying an actual sound intensity weight ratio between the human voice audio and the accompaniment audio in the initial reverberation audio, and adjusting the sound intensity of the human voice audio and / or the accompaniment audio in the initial reverberation audio according to the actual sound intensity weight ratio to obtain a target reverberation audio;
[0081] In this embodiment, the actual sound intensity weight ratio refers to the weight ratio of the sound intensity of the human voice audio to the sound intensity of the accompaniment audio in the initial reverberation audio. It is easy to understand that the sound intensity, also known as volume or loudness, refers to the strength of the sound felt by the human ear. It is a subjective feeling of the person about the size of the sound. In other words, the sound intensity is the degree of loudness of the sound. It should be noted that since the current sound field expansion function (i.e., virtual surround sound function) is mainly implemented through the Head Related Transfer Function (HRTF) algorithm, when the HRTF sound field expansion is adopted, it often brings about the effect of the human voice becoming virtual. That is, the actual sound intensity weight ratio between the human voice audio and the accompaniment audio in the initial reverberation audio obtained after the sound field expansion is often small, that is, the sound intensity of the human voice audio in the initial reverberation audio is often small, while the sound intensity of the accompaniment audio in the initial reverberation audio is often large. Therefore, this embodiment dynamically identifies the actual sound intensity weight ratio between the human voice audio and the accompaniment audio in the initial reverberation audio through the human voice recognition module and the accompaniment sound recognition module (such as Figure 3 As shown), it is determined whether the actual sound intensity weight ratio is within the preset standard sound intensity weight ratio range. If it exceeds the preset standard sound intensity weight ratio range, it means that the audio currently played by the near-ear open audio device when the HRTF sound field expansion is performed has the problem of human voice becoming hollow. Therefore, this embodiment adjusts the sound intensity of the human voice audio and / or accompaniment audio in the initial reverberation audio, and increases the weight of the sound intensity of the human voice audio in the initial reverberation audio, thereby improving the problem of human voice becoming hollow when the near-ear open audio device is performing HRTF sound field expansion.
[0082] As an example, the step of adjusting the intensity of the vocal audio and / or the accompaniment audio in the initial reverberation audio includes:
[0083] Step S321: increasing the intensity of the human voice audio in the initial reverberation audio; and / or
[0084] Step S322: reducing the volume of the accompaniment audio in the initial reverberation audio.
[0085] After step S30, step S40 is performed: playing the target reverberation audio.
[0086] This embodiment obtains a target transfer function between a near-ear open-ear audio device and the user's ears, and then performs crosstalk cancellation on the input audio received by the near-ear open-ear audio device based on the target transfer function to obtain initial reverberation audio, so that the ears of the user wearing the near-ear open-ear audio device receive a sound signal consistent with the input audio. This embodiment calculates the head-related transfer function in different scenarios through sound source simulation, which can eliminate interference from the near-ear open-ear audio device itself on the sound signal. In situations where the speaker of the near-ear open-ear audio device cannot be placed in the human ear like headphones, the listening effect of the sound played by the near-ear open-ear audio device when transmitted to the user's ears is consistent with the listening effect when wearing headphones, effectively improving the listening experience of the near-ear open-ear audio device users and avoiding crosstalk issues. However, because the current sound field expansion function is mainly implemented through the Head-Related Transfer Function (HRTF) algorithm, the use of HRTF sound field expansion often results in the effect of making the human voice sound muffled. That is, the initial reverberation audio obtained after the sound field expansion often has a smaller actual sound intensity weight ratio between the human voice audio and the accompaniment audio. In other words, the weight of the sound intensity of the human voice audio in the initial reverberation audio is often smaller, while the weight of the sound intensity of the accompaniment audio in the initial reverberation audio is often larger. Therefore, this embodiment dynamically identifies the actual sound intensity weight ratio between the human voice audio and the accompaniment audio in the initial reverberation audio, and determines whether the actual sound intensity weight ratio is within a preset standard sound intensity weight ratio range. If it exceeds the preset standard sound intensity weight ratio range, it means that the audio currently played by the near-ear open audio device has the problem of human voice becoming muffled when performing HRTF sound field expansion. Therefore, this embodiment adjusts the sound intensity of the human voice audio and / or the accompaniment audio in the initial reverberation audio, increases the weight of the sound intensity of the human voice audio in the initial reverberation audio, obtains the target reverberation audio, and plays it, thereby improving the problem of human voice becoming muffled when the near-ear open audio device performs HRTF sound field expansion. That is, since the head transfer function is obtained based on the near field / far field / free field, and the sound field crosstalk elimination processing is performed through the head transfer function, the problem of weak human voice will arise. This embodiment utilizes the extraction of the accompaniment audio signal and human voice signal of the song to be processed, and then adjusts the sound intensity of the accompaniment audio signal and / or human voice signal of the initial reverberation audio according to the reverberation degree values of the extracted accompaniment audio signal and human voice signal, so as to achieve the goal of effectively expanding the sound field while ensuring the sound effect of the human voice, and overcome the technical problem of poor sound effect of the human voice part in the audio played by the near-ear open audio device after adding the sound field expansion function.
[0087] In one possible implementation, please refer to Figure 2The step of adjusting the intensity of the vocal audio and / or the accompaniment audio in the initial reverberation audio according to the actual intensity weight ratio to obtain the target reverberation audio includes:
[0088] Step S31, obtaining a target sound intensity weight ratio between the human voice audio and the accompaniment audio;
[0089] In one embodiment, the target sound intensity weight ratio may be obtained by experimental calibration before leaving the factory and pre-stored in the system of the near-ear open audio device. The near-ear open audio device with the sound field expansion function added may obtain the target sound intensity weight ratio from the system before leaving the factory. In another embodiment, the target sound intensity weight ratio may also be obtained after leaving the factory by the user inputting the target sound intensity weight ratio into the system of the near-ear open audio device based on the user's personal listening comfort experience and habits for audio. In yet another embodiment, the near-ear open audio device may obtain the theoretical sound intensity weight ratio corresponding to the same reverberation audio (i.e., the same audio as the initial reverberation audio, except that the sound field expansion processing is not performed) output by the near-ear open audio device when the sound field expansion function is not turned on, and use the theoretical sound intensity weight ratio as the target sound intensity weight ratio.
[0090] Step S32: Adjust the sound intensity of the vocal audio and / or the accompaniment audio in the initial reverberation audio according to the actual sound intensity weight ratio and the target sound intensity weight ratio, so as to adjust the actual sound intensity weight ratio to the target sound intensity weight ratio to obtain the target reverberation audio.
[0091] This embodiment obtains a target intensity weight ratio between the vocal audio and the accompaniment audio, and adjusts the intensity of the vocal audio and / or the accompaniment audio in the initial reverberation audio according to the actual intensity weight ratio and the target intensity weight ratio, so as to adjust the actual intensity weight ratio to the target intensity weight ratio, thereby more accurately adjusting the intensity of the accompaniment audio signal and / or the vocal signal of the initial reverberation audio, thereby effectively expanding the sound field while ensuring the sound effect of the vocals.
[0092] As an example, the step S31, obtaining the target intensity weight ratio between the vocal audio and the accompaniment audio, includes:
[0093] Step S311, identifying the initial reverberation audio through a converged neural network model to obtain an audio type corresponding to the initial reverberation audio;
[0094] Step S312: According to the audio type, obtain the sound intensity weight ratio of the audio type mapping from the preset mapping data table, and use the sound intensity weight ratio of the audio type mapping as the target sound intensity weight ratio between the human voice audio and the accompaniment audio.
[0095] In this embodiment, it will be understood by those skilled in the art that different audio types require different standard sound intensity weight ratios to be achieved, so that the sound intensity ratio of the accompaniment and the human voice is better, thereby improving the user's listening comfort experience. For example, the sound intensity ratio of the human voice in folk songs is often relatively higher, that is, the weight of the sound intensity of the human voice audio in folk songs is relatively large. However, ancient music often requires the sound intensity of the accompaniment to be relatively higher, that is, the weight of the sound intensity of the accompaniment audio in ancient music is relatively large. For example, rock music requires a relatively moderate sound intensity ratio between the accompaniment and the human voice (close to 1:1). In this embodiment, the neural network model can be trained in advance on audio samples of different audio types (such as rock music, folk music, ancient music, ethnic music, rap, etc.), and the prediction accuracy of the neural network model for the audio type can be manually verified. If the prediction accuracy of the audio sample obtained by testing a preset number of consecutive audio samples reaches a preset threshold (for example, 95%), it is determined that the neural network model has converged, and a converged neural network model is obtained.
[0096] This embodiment identifies the initial reverberation audio through a convergent neural network model to obtain the audio type corresponding to the initial reverberation audio, and based on the audio type, obtains the sound intensity weight ratio mapped to the audio type from a preset mapping data table, and uses the sound intensity weight ratio mapped to the audio type as the target sound intensity weight ratio between the human voice audio and the accompaniment audio, thereby improving the intelligence and accuracy of identifying the target sound intensity weight ratio of the initial reverberation audio.
[0097] Furthermore, in step S30, the step of identifying the actual sound intensity weight ratio between the human voice audio and the accompaniment audio in the initial reverberation audio includes:
[0098] Step S51, dividing the initial reverberation audio into a plurality of frames, wherein each frame of the initial reverberation audio has an accompaniment audio and a vocal audio in a time-synchronized relationship;
[0099] In this embodiment, each frame of the initial reverberant audio after framing may include a preset number of sound sampling points, and adjacent frames may have a preset number of overlapping sampling points. For example, in this embodiment, the time domain signal of the initial reverberant audio may be divided into N frames, each frame including 512 sound sampling points (at an audio sampling rate of 16 kHz), and adjacent frames may have 256 overlapping sampling points. This processing is intended to achieve a smooth transition between frames.
[0100] Step S52: performing windowing processing on each frame of the initial reverberation audio, and converting the windowed initial reverberation audio from the time domain to the frequency domain by means of a fast Fourier transform, to obtain an initial reverberation spectrum;
[0101] In this embodiment, the fast Fourier transform (FFT) method is a general term for efficient and fast computational methods for calculating discrete Fourier transforms (DFTs) using computers. The fast Fourier transform method can be used to convert the windowed initial reverberation audio from the time domain to the frequency domain, obtaining the amplitude and phase information of each frame of the initial reverberation audio, i.e., the initial reverberation spectrum.
[0102] Step S53, decomposing the initial reverberation spectrum to obtain an accompaniment spectrum and a vocal spectrum in the initial reverberation spectrum;
[0103] Step S54, determining an actual sound intensity weight ratio between the human voice audio and the accompaniment audio in the initial reverberation spectrum based on the accompaniment spectrum and the human voice spectrum;
[0104] In step S32, the step of adjusting the intensity of the vocal audio and / or the accompaniment audio in the initial reverberation audio according to the actual intensity weight ratio and the target intensity weight ratio to obtain the target reverberation audio includes:
[0105] Step S55: According to the actual sound intensity weight ratio and the target sound intensity weight ratio, the human voice spectrum in the initial reverberation spectrum is subjected to sound intensity amplification processing, and / or the accompaniment spectrum in the initial reverberation audio is subjected to sound intensity reduction processing to obtain a target reverberation spectrum;
[0106] Step S56: convert the target reverberation spectrum from the frequency domain to the time domain to obtain the target reverberation audio.
[0107] In this embodiment, the target reverberation spectrum may be converted from the frequency domain to the time domain by using an inverse Fourier transform method to obtain the target reverberation audio.
[0108] Among them, the human voice / accompaniment sound recognition algorithm logic is as follows Figure 4 As shown, it should be noted that, in the process of extracting the features of vocals / accompaniment sounds, the features used include but are not limited to: spectral entropy (Spectral Entropy), Linear Prediction Cepstrum Coefficients (LPCC) and Line Spectrum Pair (LSP), short-time energy, Mel-scale Frequency Cepstral Coefficients (MFCC), first-order difference Mel-scale Cepstrum Coefficients (first-order difference MFCC), loudness and glottal excitation pulse, etc.
[0109] In this embodiment, refer to Figure 4This embodiment performs frame segmentation, windowing, and fast Fourier transform processing on the initial reverberation audio, converting the initial reverberation audio from the time domain to the frequency domain to obtain an initial reverberation spectrum, and analyzing the frequency domain features of the initial reverberation spectrum to extract an accompaniment spectrum and a vocal spectrum. Based on the extracted accompaniment spectrum and vocal spectrum, an actual sound intensity weight ratio between the vocal audio and the accompaniment audio is determined, thereby accurately and effectively analyzing the actual sound intensity weight ratio of the initial reverberation audio. Then, based on the actual sound intensity weight ratio and the target sound intensity weight ratio, the vocal spectrum in the initial reverberation spectrum is intensity-intensified and / or the accompaniment spectrum in the initial reverberation audio is intensity-reduced to obtain a target reverberation spectrum. Finally, the target reverberation spectrum is converted from the frequency domain to the time domain to obtain a target reverberation audio. This allows for more accurate adjustment of the sound intensity of the accompaniment audio signal and / or the vocal signal of the initial reverberation audio, thereby effectively expanding the sound field while ensuring the sound effect of the vocals.
[0110] In addition, an embodiment of the present invention further provides an audio device, referring to Figure 5 , Figure 5 The figure is a schematic diagram of the structure of an audio device involved in an embodiment of the present invention.
[0111] like Figure 5 As shown, the audio device may include: a processor 1001, a communication bus 1002, a user interface 1003, a network interface 1004 and a memory 1005. Among them, the processor 1001 may be a central processing unit (CPU). The communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a wireless fidelity (WIreless-FIdelity, WI-FI) interface). The memory 1005 may be a high-speed random access memory (Random Access Memory, RAM) memory, or a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk storage. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0112] Those skilled in the art will understand that Figure 5 The structure shown in the figure does not constitute a limitation to the audio device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0113] like Figure 5As shown, the memory 1005 as a storage medium may include an operating system, a data storage module, a network communication module, a user interface module and a sound field expansion program.
[0114] exist Figure 5 In the audio device shown, the network interface 1004 is mainly used for data communication with other devices; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in this embodiment can be set in the audio device, and the audio device calls the sound field expansion program stored in the memory 1005 through the processor 1001 and performs the following operations:
[0115] Obtaining a target transfer function between the near-ear open audio device and the user's ears;
[0116] performing crosstalk cancellation processing on the input audio received by the near-ear open audio device according to the target transfer function to obtain initial reverberation audio;
[0117] identifying an actual sound intensity weight ratio between the human voice audio and the accompaniment audio in the initial reverberation audio, and adjusting the sound intensity of the human voice audio and / or the accompaniment audio in the initial reverberation audio according to the actual sound intensity weight ratio to obtain a target reverberation audio;
[0118] Play the target reverb audio.
[0119] Optionally, the processor 1001 may call the sound field expansion program stored in the memory 1005 and further perform the following operations:
[0120] Obtaining a target sound intensity weight ratio between the vocal audio and the accompaniment audio;
[0121] According to the actual sound intensity weight ratio and the target sound intensity weight ratio, the sound intensity of the human voice audio and / or the accompaniment audio in the initial reverberation audio is adjusted to adjust the actual sound intensity weight ratio to the target sound intensity weight ratio to obtain the target reverberation audio.
[0122] Optionally, the processor 1001 may call the sound field expansion program stored in the memory 1005 and further perform the following operations:
[0123] Identifying the initial reverberation audio using a converged neural network model to obtain an audio type corresponding to the initial reverberation audio;
[0124] According to the audio type, the sound intensity weight ratio of the audio type mapping is obtained from the preset mapping data table, and the sound intensity weight ratio of the audio type mapping is used as the target sound intensity weight ratio between the human voice audio and the accompaniment audio.
[0125] Optionally, the processor 1001 may call the sound field expansion program stored in the memory 1005 and further perform the following operations:
[0126] increasing the intensity of the human voice audio in the initial reverberation audio; and / or,
[0127] The volume of the accompaniment audio in the initial reverberation audio is reduced.
[0128] Optionally, the processor 1001 may call the sound field expansion program stored in the memory 1005 and further perform the following operations:
[0129] Obtaining a preset artificial head transfer function and a free-field transfer function;
[0130] performing an inverse operation on the free-field transfer function to obtain an inverse free-field transfer function;
[0131] The artificial head transfer function is multiplied by the free-field inverse transfer function to obtain a target transfer function between the near-ear open audio device and the user's ears.
[0132] Optionally, the processor 1001 may call the sound field expansion program stored in the memory 1005 and further perform the following operations:
[0133] When the near-ear open-back audio device is worn on a preset artificial head and the near-ear open-back audio device outputs a sound signal, measuring the artificial head transfer function through a preset microphone in the ear canal of the artificial head; and
[0134] When the artificial head is removed and the near-ear open audio device outputs a sound signal, a free-field transfer function is measured by preset microphones placed at the left and right ear positions before the artificial head is removed.
[0135] Optionally, the processor 1001 may call the sound field expansion program stored in the memory 1005 and further perform the following operations:
[0136] performing an inverse operation on the target transfer function to obtain a target inverse transfer function;
[0137] The input audio received by the near-ear open audio device is multiplied by the target inverse transfer function to obtain initial reverberation audio.
[0138] Optionally, the processor 1001 may call the sound field expansion program stored in the memory 1005 and further perform the following operations:
[0139] Decomposing the initial reverberation spectrum to obtain an accompaniment spectrum and a vocal spectrum in the initial reverberation spectrum;
[0140] Determining an actual sound intensity weight ratio between the vocal audio and the accompaniment audio in the initial reverberation spectrum based on the accompaniment spectrum and the vocal spectrum;
[0141] The step of adjusting the sound intensity of the vocal audio and / or the accompaniment audio in the initial reverberation audio according to the actual sound intensity weight ratio and the target sound intensity weight ratio to obtain the target reverberation audio includes:
[0142] According to the actual sound intensity weight ratio and the target sound intensity weight ratio, the human voice spectrum in the initial reverberation spectrum is subjected to sound intensity amplification processing, and / or the accompaniment spectrum in the initial reverberation audio is subjected to sound intensity reduction processing to obtain a target reverberation spectrum;
[0143] The target reverberation spectrum is converted from the frequency domain to the time domain to obtain the target reverberation audio.
[0144] In addition, an embodiment of the present invention also proposes a computer-readable storage medium, which is applied to a computer. The computer-readable storage medium can be a non-volatile computer-readable storage medium. A sound field expansion program is stored on the computer-readable storage medium. When the sound field expansion program is executed by a processor, the steps of the sound field expansion method of the present invention as described above are implemented.
[0145] The various embodiments of the audio device and the computer-readable storage medium of the present invention may refer to the various embodiments of the sound field expansion method of the present invention, and will not be described in detail here.
[0146] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.
[0147] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.
[0148] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0149] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A sound field expansion method, characterized in that: The sound field expansion method comprises the following steps: Obtaining a preset artificial head transfer function and a free-field transfer function, wherein the artificial head transfer function includes the influence of a near-ear open-back audio device on sound transfer results and the influence of a head contour on the sound transfer results, and the free-field transfer function includes the influence of a near-ear open-back audio device on sound transfer results; performing an inverse operation on the free-field transfer function to obtain an inverse free-field transfer function; Multiplying the artificial head transfer function by the free-field inverse transfer function to obtain a target transfer function between the near-ear open-back audio device and the user's ears, wherein the target transfer function represents the effect of the head contour on the sound transfer result; performing crosstalk cancellation processing on the input audio received by the near-ear open audio device according to the target transfer function to obtain initial reverberation audio; identifying an actual sound intensity weight ratio between the human voice audio and the accompaniment audio in the initial reverberation audio, and adjusting the sound intensity of the human voice audio and / or the accompaniment audio in the initial reverberation audio according to the actual sound intensity weight ratio to obtain a target reverberation audio; Play the target reverb audio.
2. The sound field expansion method according to claim 1, wherein: The step of adjusting the intensity of the vocal audio and / or the accompaniment audio in the initial reverberation audio according to the actual intensity weight ratio to obtain the target reverberation audio includes: Obtaining a target sound intensity weight ratio between the vocal audio and the accompaniment audio; According to the actual sound intensity weight ratio and the target sound intensity weight ratio, the sound intensity of the human voice audio and / or the accompaniment audio in the initial reverberation audio is adjusted to adjust the actual sound intensity weight ratio to the target sound intensity weight ratio to obtain the target reverberation audio.
3. The sound field expansion method according to claim 2, wherein: The step of obtaining a target sound intensity weight ratio between the human voice audio and the accompaniment audio comprises: Identifying the initial reverberation audio using a converged neural network model to obtain an audio type corresponding to the initial reverberation audio; According to the audio type, the sound intensity weight ratio of the audio type mapping is obtained from the preset mapping data table, and the sound intensity weight ratio of the audio type mapping is used as the target sound intensity weight ratio between the human voice audio and the accompaniment audio.
4. The sound field expansion method according to claim 2, wherein: The step of adjusting the intensity of the vocal audio and / or the accompaniment audio in the initial reverberation audio includes: Increasing the intensity of the human voice audio in the initial reverberation audio; and / or The volume of the accompaniment audio in the initial reverberation audio is reduced.
5. The sound field expansion method according to claim 1, wherein: The step of obtaining a preset artificial head transfer function and a free-field transfer function comprises: When the near-ear open-back audio device is worn on a preset artificial head and the near-ear open-back audio device outputs a sound signal, measuring the artificial head transfer function through a preset microphone in the ear canal of the artificial head; and When the artificial head is removed and the near-ear open audio device outputs a sound signal, a free-field transfer function is measured by preset microphones placed at the left and right ear positions before the artificial head is removed.
6. The sound field expansion method according to any one of claims 1 to 5, characterized in that: The step of performing crosstalk cancellation processing on the input audio received by the near-ear open audio device according to the target transfer function to obtain initial reverberation audio includes: performing an inverse operation on the target transfer function to obtain a target inverse transfer function; The input audio received by the near-ear open audio device is multiplied by the target inverse transfer function to obtain initial reverberation audio.
7. The sound field expansion method according to claim 2, wherein: The step of identifying the actual sound intensity weight ratio between the human voice audio and the accompaniment audio in the initial reverberation audio comprises: Dividing the initial reverberation audio into a plurality of frames, wherein each frame of the initial reverberation audio has an accompaniment audio and a vocal audio in a time-synchronized relationship; Performing windowing processing on each frame of the initial reverberation audio, and converting the windowed initial reverberation audio from the time domain to the frequency domain by fast Fourier transform to obtain an initial reverberation spectrum; Decomposing the initial reverberation spectrum to obtain an accompaniment spectrum and a vocal spectrum in the initial reverberation spectrum; Determining an actual sound intensity weight ratio between the vocal audio and the accompaniment audio in the initial reverberation spectrum based on the accompaniment spectrum and the vocal spectrum; The step of adjusting the sound intensity of the vocal audio and / or the accompaniment audio in the initial reverberation audio according to the actual sound intensity weight ratio and the target sound intensity weight ratio to obtain the target reverberation audio includes: According to the actual sound intensity weight ratio and the target sound intensity weight ratio, the human voice spectrum in the initial reverberation spectrum is subjected to sound intensity amplification processing, and / or the accompaniment spectrum in the initial reverberation audio is subjected to sound intensity reduction processing to obtain a target reverberation spectrum; The target reverberation spectrum is converted from the frequency domain to the time domain to obtain the target reverberation audio.
8. An audio device, characterized in that The audio device includes: a memory, a processor, and a sound field expansion program stored in the memory and executable on the processor. When the sound field expansion program is executed by the processor, the steps of the sound field expansion method according to any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a sound field expansion program, which, when executed by a processor, implements the steps of the sound field expansion method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Sound effect processing method and apparatus
CN105405448A
Audio separation method and device, electronic equipment and storage medium
CN110503976A
Audio signal processing method and device
CN114203163A