Sound processing method and device, electronic equipment, storage medium and vehicle
By determining the reference signal and target channel in a multi-zone audio playback scenario and using a preset model to filter echoes, the high algorithm complexity problem in existing technologies is solved, and efficient human voice signal filtering and sound source localization are achieved.
Patent Information
- Application Number
- CN202410544692.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-30
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies have high algorithm complexity and consume a lot of computing power when filtering human voice signals and locating sound sources in multi-zone audio playback scenarios, which affects processing performance.
By determining a reference signal, the mixed signals received by multiple microphones are obtained. A preset model is used to determine the target channel with the highest proportion of human voice signal in the mixed signal. Based on the reference signal, the media signal with echo in the target channel is filtered to obtain a clean human voice signal.
It simplifies the algorithm complexity, saves computing power, improves the processing performance of human voice signal screening and sound source localization, and ensures the accuracy of human voice recognition.
Smart Images

Figure CN120877752A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing, and in particular to a sound processing method, apparatus, electronic device, storage medium, and vehicle. Background Technology
[0002] To achieve a superior audio experience, multiple speakers and various playback settings are used during audio playback. By adjusting the volume and other sound parameters of different speakers, different sound effects are produced. In this scenario, when multiple microphones are used to record human voices, the sound picked up by each microphone includes not only the human voice but also the sound emitted by each speaker. The echo effect occurs when the microphones receive sound from each speaker.
[0003] Currently, in order to achieve the filtering of human voice signals and sound source localization in scenarios where audio is played in multiple audio zones, related technologies typically require setting up different independent functional modules. For example, a functional module using adaptive filters is used to filter human voice signals, and a functional module using blind separation algorithms is used to localize sound sources. This results in high algorithm complexity and high computational power consumption during the execution of related technologies, which in turn affects the processing performance of filtering human voice signals and localizing sound sources. Summary of the Invention
[0004] In view of this, this application provides a sound processing method, apparatus, electronic device, storage medium, and vehicle, the main purpose of which is to solve the technical problems of current related technologies that affect the processing performance of screening human voice signals and sound source localization.
[0005] To achieve the above objectives, the first aspect of this application discloses a sound processing method, which includes:
[0006] A reference signal is determined, wherein the reference signal is a signal source played by the speaker;
[0007] After the signal source is played by the speaker, a mixed signal received by multiple microphones is acquired. The multiple microphones correspond to different audio channels. The mixed signal includes human voice signals and media signals with echo corresponding to media audio.
[0008] The reference signal and the mixed signal are input into a preset model. The preset model is used to determine the target channel with the highest signal ratio of the human voice signal in the mixed signal. The echoing media signal in the target channel is filtered according to the reference signal to obtain a clean human voice signal in the target audio region corresponding to the target channel.
[0009] A second aspect of this application provides a sound processing apparatus, the apparatus comprising:
[0010] A determination module is used to determine a reference signal, wherein the reference signal is a signal source played by a speaker;
[0011] The acquisition module is used to acquire a mixed signal received by multiple microphones after the signal source is played by the speaker. The multiple microphones correspond to different audio channels. The mixed signal includes human voice signals and media signals with echo corresponding to media audio.
[0012] The filtering module is used to input the reference signal and the mixed signal into a preset model, use the preset model to determine the target channel with the highest signal proportion of the human voice signal in the mixed signal, and filter the echoing media signal in the target channel according to the reference signal to obtain a clean human voice signal in the target audio region corresponding to the target channel.
[0013] A third aspect of this application provides an electronic device, comprising:
[0014] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform any of the methods disclosed in the first aspect.
[0015] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0016] A fifth aspect of this application provides a vehicle in which the device as described in the second aspect or the electronic device as described in the third aspect is mounted.
[0017] In summary, based on the technical solution disclosed in this application, to address the problem of inaccurate voice recognition caused by echoes in the sound played by a speaker, this application discloses a solution. First, a reference signal is determined, which is the signal source played by the speaker. Then, after the signal source is played by the speaker, a mixed signal received by multiple microphones is acquired. Each microphone corresponds to a channel in a different audio range. The mixed signal contains both the voice signal and the echo-corresponding media signal. Finally, the reference signal and the mixed signal are input into a preset model. The preset model is used to determine the target channel with the highest proportion of the voice signal in the mixed signal. The echo-corresponding media signal in the target channel is then filtered based on the reference signal to obtain a clean voice signal within the target audio range corresponding to the target channel. In this application, a reference signal is recorded, and the target channel of the human voice signal is accurately located within the mixed signal. The mixed signal in the target channel is then filtered for echoing media signals, using the reference signal as a benchmark, to avoid interference with the human voice signal. Simultaneously, using the reference signal as a benchmark, the source of echoes caused by media signals in the mixed signal can be accurately determined, effectively extracting clean human voice signals. Furthermore, this application's technical solution trains a preset model that can accurately determine the target channel corresponding to the target audio region in an integrated manner, and further extract clean human voice signals from the target channel. This eliminates the need for multi-functional coordination, allowing for flexible acquisition of clean human voice signals in various scenarios in a relatively simple way, saving computational resources and further ensuring subsequent human voice recognition performance, thus improving the performance of human voice signal filtering and sound source localization processing.
[0018] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A flowchart of a sound processing method provided in an embodiment of this application is shown;
[0022] Figure 2 This illustration shows the input values and corresponding target output values of a model provided in an embodiment of this application;
[0023] Figure 3 A structural diagram of a sound processing device provided in an embodiment of this application is shown. Detailed Implementation
[0024] To better understand the above-mentioned objectives, features, and advantages of this application, the solution of this application will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0025] To address the technical problems that current related technologies affect the processing performance of screening human voice signals and localizing sound sources, this application provides the following embodiments to solve the above problems:
[0026] This embodiment provides a sound processing method. The executing entity can be a processor or a processing module, which can be located on a cloud server or a terminal, such as an in-vehicle infotainment system. Figure 1 The diagram shown is a flowchart of the method in this embodiment. The method in this embodiment may specifically include the following steps:
[0027] Step 101: Determine the reference signal, which is the signal source played by the speaker;
[0028] First, a reference signal is recorded. When a speaker plays sound, especially media, the signal it plays is generated by the sound control system and further transmitted to the speaker connected to the sound control system. The signal is then converted into a data format that the speaker can resolve and played back, allowing the user to hear the played sound signal. The signal generated by the sound control system and further applied to subsequent speaker playback serves as the reference signal. When the reference signal is transmitted to the speaker, it becomes the signal source for the speaker's playback. This reference signal includes at least media audio, such as music, voice announcements, or voice responses.
[0029] Step 102: After the signal source is played by the speaker, the mixed signal received by multiple microphones is acquired. The multiple microphones correspond to different audio channels. The mixed signal contains human voice signals and media signals with echo corresponding to media audio.
[0030] In the application scenarios described in this embodiment, besides supporting speaker playback, sound reception via microphone is also possible. In more complex scenarios, the microphone receives not only the required human voice but also the sound played by the speakers. Furthermore, when there are multiple speakers, the final sound received by the microphone also includes the echo of the speaker playback. The human voice and the speaker playback are mixed together and received by the microphone, generating a mixed signal. When a reference signal is used as a signal source and played by the speakers, the speakers generate corresponding media sounds. When the microphone receives media sounds played by multiple speakers and receives the corresponding media signals, these media signals appear as echo-containing media signals in the actual signal received by the microphone. The mixed signal includes the human voice signal corresponding to the human voice and the echo-containing media signal corresponding to the media sounds generated by the speaker playback signal source. For example, in this embodiment, media sounds refer to sound played through a playback device. Media sounds can be the display sounds of various software, such as the singing voice played by music software, the navigation prompts of navigation software, the response sounds of smart assistants, or the call sounds of voice calls, etc.
[0031] In one application scenario of this embodiment, the space is divided into multiple sound zones, each corresponding to a specific space. To ensure the acquisition of human voice signals in each space, a microphone is installed in each space to capture human voices within that space. The space corresponding to the microphone is then the sound zone corresponding to that microphone. This technical solution can be applied to the passenger compartment of an in-vehicle infotainment system.
[0032] Step 103: Input the reference signal and the mixed signal into the preset model, use the preset model to determine the target channel with the highest signal ratio of human voice signal in the mixed signal, and filter the media signal with echo in the target channel according to the reference signal to obtain a clean human voice signal in the target audio region corresponding to the target channel.
[0033] Since the reference signal acts as the signal source to generate media sound through the loudspeaker, the generated media sound corresponds to the echoing media signal in the mixed signal. Therefore, based on the extraction of the reference signal, when filtering the mixed signal, the media signal can be further determined from the mixed signal using the reference signal. At this point, the media signal is essentially noise. Simultaneously, due to the influence of multiple loudspeakers, the media signal may be received by the microphone as an echo signal. The reference signal can accurately determine the corresponding media signal and its echo signal. After determining the media signal, accurate filtering can be performed on the media signal and its echo signal, ultimately retaining a clean vocal signal.
[0034] In the actual process of filtering echoing media signals, the mixed signal is generated by acquiring sound signals from multiple microphones in their respective audio ranges. However, not all spaces corresponding to these audio ranges contain human voice signals. Therefore, before extracting human voice signals from the mixed signal, it is necessary to first identify the target channel, which contains the clearest human voice signal. The target channel has the highest signal proportion compared to other channels, meaning it contains the clearest human voice signal. Therefore, it can be understood that a user exists in the space corresponding to this target channel, and that user is emitting human voice signals.
[0035] For example, within the vehicle's infotainment system, to clearly play audio to occupants in various locations, the system can be equipped with multiple speakers. The played audio can be music, video audio, or media audio such as responses from a voice assistant. When an occupant speaks a voice command, the media audio continues playing. To accurately understand the user's voice command, it's crucial to first accurately capture the user's voice signal from the various media audio sources before extracting the voice command. At this point, the microphone in the infotainment system receives a mixed signal containing both the voice signal and media signals. To enhance the accuracy of voice command recognition, it's necessary to extract a clean voice signal from the mixed signal. Therefore, a reference signal corresponding to the media signal playing simultaneously with the voice signal is recorded, and the media signal is filtered based on this reference signal to ensure that the final obtained voice signal is a clean voice signal. In the process of acquiring clean human voice signals, the signal proportion of human voice signals in each mixed signal is calculated using a preset model. The channel with the highest signal proportion is selected as the target channel, and the human voice signal is extracted from the target channel. After filtering out media signals with echoes, a clean human voice signal representing the human voice in that audio range can be obtained.
[0036] In this application, a reference signal is recorded, and the target channel of the human voice signal is accurately located in the mixed signal. The mixed signal in the target channel is filtered with reference to the reference signal to avoid interference from the echoing media signal to the human voice signal. At the same time, the technical solution of this application trains a preset model, which can accurately determine the target channel corresponding to the target audio region in an integrated manner, and further extract clean human voice signals from the target channel. Without the need for multi-functional coordination, it can flexibly acquire clean human voice signals in various scenarios in a relatively simple way, and further ensure the subsequent human voice recognition effect, thereby improving the performance of human voice signal filtering and sound source localization processing.
[0037] In one possible embodiment, determining the target vocal channel with the highest signal proportion of the human voice signal in the mixed signal using the preset model includes:
[0038] Select the target human voice signal in the mixed signal. The target human voice signal exists in at least one set of channels in the mixed signal. The target human voice signal is the same human voice signal in at least one set of channels. Calculate the sound loudness value of the target human voice signal in each of the at least one set of channels using the signal ratio of the target human voice signal in the mixed signal. Select the channel with the maximum sound loudness value from the at least one set of channels as the target channel.
[0039] Simultaneously, considering the distribution of multiple microphones, the strongest voice information is found in the corresponding channel of the microphone that receives the voice of the nearest person. Therefore, a target channel can be selected based on the distribution of each microphone and the loudness value of the voice signal within each microphone. The voice signal in the target channel represents the voice signal obtained by the microphone that best receives the person's voice information. Furthermore, it can be understood as the voice information received by the microphone closest to the person. While the target channel is determined by the loudness value of the voice signal, in a mixed signal, the proportion of the voice signal within the mixed signal is determined by its signal percentage. The higher the loudness value of the voice signal, the higher its signal percentage in the mixed signal. Additionally, when recognizing voice signals, other characteristic signal parameters such as the amplitude value of the voice signal can also be used for identification. Meanwhile, in the target channel identification process of this embodiment, the signal proportion of each human voice signal in a channel is only compared with the signal proportion of the same human voice signal in another channel to determine the signal proportion of the human voice signal in different channels and to determine the magnitude relationship between the different signal proportions of the human voice signal. This avoids the problem that when two users speak at the same time, microphones in different audio ranges will identify two human voice signals with the same loudness value, thus making it impossible to determine the target channel. For example, in a car system, at least one microphone is set up for each seat to clearly receive the human voice of the occupant in that seat, and the distribution position of the microphones is fixed with the seat. In each channel, there is a mixed signal acquired by each microphone. Therefore, after filtering the media signal in the mixed signal, during the human voice signal splicing process, the human voice channel of the sound signal acquired by the microphone closest to the person's position is selected for human voice acquisition to obtain the final clean human voice signal.
[0040] For example, in the vehicle's infotainment system, multiple microphones are set up to receive the voices of occupants sitting in various positions. The microphone closest to the occupant, or the microphone specifically designed to receive the voice signal from a particular position, has the strongest voice volume. Furthermore, by combining the occupant's position with the position of the corresponding microphone, the system ultimately selects the most suitable microphone and the mixed signal it absorbs, and combines the segmented mixed signals to form a complete voice signal, which is then used as a clean voice signal.
[0041] In one possible embodiment, determining the reference signal includes:
[0042] Based on the number of speakers, acquire the audio signals of a first number of channels to be played by the speakers; select a target number of audio signals to be identified from the audio signals of the first number of channels, where the target number is greater than 1; determine the repeating audio signals among the target number of audio signals to be identified; use the signal ratio of the repeating audio signals in the target number of audio signals to be identified as the audio repeat value, and verify the relationship between the audio repeat value in the repeating audio signals and a first preset threshold; if the audio repeat value is greater than or equal to the first preset threshold, mix the multiple audio signals that generate the repeating audio signals to obtain the audio signals of a second number of channels, which serve as reference signals, where the first number is greater than the second number.
[0043] This embodiment further discloses a process for determining a reference signal. After receiving the reference signal, the loudspeaker can play audio based on the signal source represented by the reference signal. In a multi-speaker playback scenario, the sound source represented by the reference signal is transmitted through channels connected to each loudspeaker and played on the corresponding loudspeaker. However, in some application scenarios, the audio playback instructions of the reference signals to the loudspeakers are the same or highly similar, causing each loudspeaker to play the same or almost identical content. In the process of clean voice signal separation, if the reference signals from all channels are used to filter media signals separately, inputting the same reference signals into a preset model in multiple paths for target channel determination and voice signal stripping would undoubtedly be a waste of computing power. Therefore, this embodiment proposes selecting a target number of audio signals from a first number of channels, with each selected group of audio signals serving as an audio signal. The number of the selected target number of audio signals to be identified is at least greater than 1, and the specific value of the target number is an adjustable integer. The system acquires repeated audio signals among the audio signals to be identified and further calculates the proportion of the repeated audio signals in the audio signals to be identified that generated the repeated audio signals. This proportion is used as the audio repetition value. If the audio repetition value exceeds a first preset threshold, it can be determined that the media signals played between the channels that generated the repeated audio signals have a very high degree of overlap. Therefore, before inputting the model, the system can mix the target number of audio signals to be identified or delete one set of audio signals. In the process of continuously mixing and deleting multiple audio signals, the number of channels in the final audio signal is reduced. Inputting the reference signals from the second number of channels into the preset model can effectively reduce the computational power consumption of the model.
[0044] In one possible embodiment, the preset model building process includes:
[0045] The system plays sample media audio through a loudspeaker and sample human voices in locations where people can sit; it extracts the sample reference signal corresponding to the sample media audio; it acquires the sample media signal generated by the microphone receiving the sample media audio and the sample human voice signal generated by the microphone receiving the sample human voice; it mixes the sample reference signal, sample media signal, and sample human voice signal to generate a sample mixed signal set; it inputs the sample mixed signal set into an initial model corresponding to a preset model, trains the initial model until the time-frequency domain difference between the human voice signal output by the initial model and the signal of the sample human voice is less than a second threshold, and generates the preset model.
[0046] This embodiment further explains the model training process. The purpose of the final generated preset model is to separate the human voice signal from the mixed signal for further recognition. Therefore, during the model training process, sample media audio played from the speaker is first acquired. This sample media audio is generated from the sample reference signal played by the speaker as the signal source. In addition, during the playback of sample media audio, there is at least one speaker. When the microphone receives sample media audio played from multiple speakers, the sample media signal corresponding to the sample media audio received by the microphone is manifested as a media signal form with echo. Sample human voices are placed in the seating area so that the microphones can receive the sample human voice signals. This allows the microphones participating in the training to obtain a mixed signal of sample media signals and sample human voice signals. At the same time, the sample reference signal corresponding to the media sound is recorded as training data. The three types of training data, sample media signals, sample human voice signals, and sample reference signals, jointly participate in the training of the initial model. The training process continues until the output of the initial model converges. The output convergence can be expressed as the time-frequency domain difference between the human voice signal output by the initial model and the signal of the sample human voice is less than a second threshold. Finally, the initial model is trained to be a preset model, and the preset model is obtained.
[0047] Specifically, during the initial model training process, the initial model can first determine the channel where the sample human voice is located among multiple channels in the mixed sound, and determine this channel as the target channel. Then, human voice separation is performed in the target channel, and the separated human voice and the set sample human voice are subjected to time-frequency domain loss function calculation. The initial model is iterated based on the loss function value until the output result of the initial model converges, generating the preset model.
[0048] In one possible embodiment, playing sample human voices at a location where people can sit includes:
[0049] Set the number of sample human voice timbres, which is less than the number of available seating locations; confirm the distribution of sample human voices among the available seating locations; play the sample human voices of the specified number of timbres at each of the distributed locations.
[0050] During the training process of playing sample voices, voices with different timbres are set up in different locations or in different vocal registers. Multiple preset locations can be set for playing sample voices. Any location can be selected as a preset location, and the selected preset location can be adjusted based on the number of preset locations to ensure that the sample voice is played at that location. Since the timbre of a voice has a strong characteristic correspondence with a person, the timbre of the sample voices played in each vocal register should be different during model training. Furthermore, the playback of sample voices should simulate the actual distribution of people; therefore, the number of available seating locations should be greater than the number of timbre types in the sample voices, so that the model trained based on the sample voices can accurately imitate the process of people speaking in normal scenarios.
[0051] For example, training voices are played at various seats in the vehicle's infotainment system. These seats can be one or more, arranged arbitrarily, to simulate the speaking process of passengers in the cabin. Furthermore, the played training voices can be represented by sounds at different volumes. When selecting a training location, a corresponding identifier is added to the selected location. When playing sample voices at different locations, the timbre of the sample voices can be set. Since each voice corresponds to a different timbre, and each available seat can only accommodate one person, the number of timbre types in a given space is less than the number of available seats. Within these available seats, the distribution of passengers can be randomly selected according to a permutation and combination. The distribution of sample voices is then set according to the passenger distribution, and each played sample voice corresponds to a different timbre. This ensures that the correspondence between sound locations and the final clean voice is clearly defined during training.
[0052] In one possible embodiment, the sample reference signal, the sample media signal, and the sample human voice signal are mixed to generate a sample mixed signal set, including:
[0053] Obtain the ratio of different numbers of human voices contained in the sample human voice signal; obtain a random number of sample human voice signals from the sample human voices containing different ratios of human voices; mix the random number of sample human voice signals, the sample reference signal, and the sample media signal with a preset mixing ratio to generate a sample mixed signal set.
[0054] During the signal mixing process, a large number of sample media signals are collected, including different sound effects and different volume settings. At this time, the collected sample media signals are in the form of sample media signals with echo. The collected signals are divided into M+K channels, of which the first M channels are sample media signals received by the microphone, and the last K channels are sample reference signal channels. The signals are mixed to generate the first mixed training signal set.
[0055] A large number of sample human voice signals were collected, including different locations, different voice volumes, and different numbers of people. The corresponding vocal ranges were marked. The collected signals consisted of M+K channels, of which the first M channels were sample human voice signals received by the microphone, and the last K channels were sample reference signal channels (when the car stereo is not playing media audio, the last K channels should be all 0). The signals were then mixed to generate a second mixed training signal set.
[0056] The first mixed training signal set is downmixed, specifically M+K→M+K1.
[0057] The first and second mixed training signal sets after downmixing are mixed to generate a sample mixed signal set, which is used as the input to the initial model. The output of the model should be echo-free and able to distinguish the corresponding channels of human voice information. For example: There are four sound zones in the car: driver's seat, front passenger seat, second row left, and second row right. In the mixed signal, when someone speaks simultaneously in the driver's seat and the second row left position (represented by color intensity), the target signal should be as follows: only the driver's voice is heard in the driver's corresponding channel, only the second row left voice is heard in the second row left corresponding channel, and no signal is heard in other positions. Figure 2 As shown, Figure 2 This is a diagram illustrating the input values and corresponding target output values of a model.
[0058] In one possible embodiment, determining the reference signal includes:
[0059] Read the sound effect types corresponding to multiple signal sources; obtain the playback parameters set in multiple speakers under each sound effect type; determine the combination order of the channels of multiple speakers based on the allocation order of speakers in different sound zones under the playback parameters during the same playback sound time period; combine multiple signal sources according to the channel combination order of the speakers to obtain a reference signal.
[0060] In this embodiment, during the generation of a preset number of mixed signals, parameters related to the sound effect type are further incorporated, particularly for scenarios involving multiple speakers. In the current audio playback, multiple speakers can be used, and various sound effects, such as 3D sound effects and HiFi sound effects, can be played by adjusting the playback parameters between each speaker. Furthermore, in determining the reference signal, to improve the accuracy and efficiency of identifying clean human voice signals using the reference signal, this embodiment further considers the playback parameters corresponding to different sound effect types when playing media audio of various sound effect types. After determining the signal channels, the arrangement order of the multi-channel reference signals is further adjusted to achieve the division of the mixed signal. In some special sound effect scenarios, during the playback period, some speakers may have signals to be played, while others may not. In such sound effect scenarios, the channels can be adjusted so that the channels with signals to be played are prioritized for combination to generate the reference signal. In another audio effect scenario, the audio effect itself achieves an echo effect by controlling the playback time of the speakers. In order to ensure accurate echo filtering, the channel combination order of the speakers can be arranged according to the speaker working time corresponding to the echo effect to obtain a complete and orderly reference signal. This reference signal can be input into the preset model in an orderly combination order. On the basis of ensuring orderly data input, the model output result can be calculated quickly.
[0061] For example, when playing media information, the audio emitted by the speaker fluctuates in loudness. In this state, when extracting human voice information, the loudness fluctuations must be considered to filter the media information, thereby further ensuring the purity of the human voice signal after filtering. Therefore, this embodiment further proposes to further divide the mixed signal by combining the sound effect type of the reference signal.
[0062] In one possible embodiment, inputting a reference signal and a mixed signal to a preset model includes:
[0063] Based on the number of microphones, set the number of first channels; divide the mixed signal into a preset number of mixed signals based on the number of first channels; divide the reference signal into a preset number of reference signals based on the number of first channels; combine the preset number of reference signals to filter the media signals in the preset number of mixed signals to generate a preset number of human voice signals; combine the preset number of human voice signals to obtain a clean human voice signal.
[0064] In this embodiment, to achieve accurate signal filtering, a segmented filtering technique is proposed. Multiple microphones can be used to acquire sound from different locations. For example, in a vehicle infotainment system, microphones can be positioned at the driver's seat, passenger seat, or the left side of the first-row passenger compartment, each capturing sound from its corresponding location. Therefore, the sound signals extracted by different microphones differ. To accurately identify human voice signals, this embodiment proposes setting channels based on the number of microphones. The number of channels is the same as the number of first channels, used to transmit the sound signals received by the microphones. The number of channels can correspond one-to-one with the number of microphones, or it can be less than the number of microphones, adapting to the computing power required to execute the sound processing method. After setting the number of channels, the mixed signal is further divided, allowing filtering of the media signal to be performed in different channels, thus improving the efficiency of media signal filtering.
[0065] This embodiment provides a sound processing device, such as... Figure 3 The diagram shown is a structural diagram of the device in this embodiment, which may include:
[0066] Determining module 31 is used to determine a reference signal, wherein the reference signal is a signal source played by a speaker;
[0067] The acquisition module 32 is used to acquire a mixed signal received by multiple microphones after the signal source is played by the speaker. The multiple microphones correspond to different audio channels. The mixed signal includes human voice signals and media signals with echo corresponding to media audio.
[0068] The filtering module 33 is used to input the reference signal and the mixed signal into a preset model, use the preset model to determine the target channel with the highest signal proportion of the human voice signal in the mixed signal, and filter the echo-bearing media signal in the target channel according to the reference signal to obtain a clean human voice signal in the target audio region corresponding to the target channel.
[0069] In one possible embodiment, the filtering module 33 is specifically used for:
[0070] Select a target human voice signal from the mixed signal. The target human voice signal exists in at least one set of channels corresponding to the mixed signal. The target human voice signal is the same human voice signal in the at least one set of channels.
[0071] Using the signal proportion of the target human voice signal in the mixed signal, calculate the sound loudness value of the target human voice signal corresponding to each of the at least one set of channels;
[0072] The channel with the highest loudness value is selected from the at least one set of channels as the target channel.
[0073] In one possible embodiment, the determining module 31 is specifically used for:
[0074] Based on the number of speakers, obtain the audio signals of a first number of channels to be played by the speakers;
[0075] From the audio signals of the first number of channels, a target number of audio signals to be identified are selected, wherein the target number is greater than 1;
[0076] Determine the repeating audio signals among the target number of audio signals to be identified;
[0077] The proportion of the repeating audio signal in the target number of audio signals to be identified is used as the audio repetition value, and the relationship between the audio repetition value in the repeating audio signal and the first preset threshold is verified.
[0078] If the audio repetition value is greater than or equal to the first preset threshold, the multiple audio signals that generate the repetitive audio signal are mixed to obtain a second number of audio channels as the mixed signal, wherein the first number is greater than the second number.
[0079] In one possible embodiment, the determining module 31 is specifically used for:
[0080] Read the sound effect types corresponding to the multiple signal sources respectively;
[0081] Obtain the playback parameters set for each of the multiple speakers under the given sound effect type;
[0082] Based on the arrangement order of speakers in different sound zones under the playback parameters during the same sound playback time period, the combination order of the channels of the multiple speakers is determined;
[0083] The reference signal is obtained by combining multiple signal sources according to the channel combination order of the loudspeaker.
[0084] In one possible embodiment, the sound processing device further includes: a training module 34, configured to:
[0085] The speaker is used to play sample media audio, and sample human voices are played in a position where people can sit.
[0086] Extract the sample reference signal corresponding to the sample media audio;
[0087] Acquire the sample media signal generated by the microphone receiving the sample media audio, and acquire the sample human voice signal generated by the microphone receiving the sample human voice;
[0088] The sample reference signal, the sample media signal, and the sample human voice signal are mixed to generate a sample mixed signal set;
[0089] The sample mixed signal set is input into the initial model corresponding to the preset model, and the initial model is trained until the time-frequency domain difference between the human voice signal output by the initial model and the human voice signal of the sample is less than the second threshold, thereby generating the preset model.
[0090] In one possible embodiment, training module 34 is specifically used for:
[0091] Set the number of timbre types for the sample human voice, where the number of timbre types is less than the number of locations where a person can sit;
[0092] Confirm the distribution of the sampled human voices in locations where people can sit;
[0093] Play the number of sample human voices of the specified timbre types at the specified distribution locations.
[0094] In one possible embodiment, training module 34 is specifically used for:
[0095] Obtain the ratio of the number of different human voices contained in the sample human voice;
[0096] From the sample human voices containing different ratios of the number of human voices, a random number of sample human voice signals are obtained;
[0097] The random number of sample human voice signals, the sample reference signals, and the sample media signals are mixed at a preset mixing ratio to generate a sample mixed signal set.
[0098] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.
[0099] Based on the above, Figure 1 The method shown, and Figure 3To achieve the above objectives, this application also provides an electronic device, which can be configured on the end side of a vehicle (such as a new energy vehicle). This device includes at least one processor and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor. The processor executes a computer program to implement the above-described virtual device embodiments. Figure 1 The method shown.
[0100] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0101] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0102] Based on the above, Figure 1 The method illustrated in this application also provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the method corresponding to any embodiment. The storage medium may further include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device and supports the operation of the information processing program and other software and / or programs. The network communication module is used to realize communication between the components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0103] Based on the aforementioned electronic device, this application embodiment also provides a vehicle, which may specifically include: such as Figure 3 The device shown or the electronic equipment described above. The vehicle may specifically be a new energy vehicle or a traditional vehicle, etc.
[0104] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platform, or it can be implemented by hardware. By applying the solution of this embodiment, compared with the prior art, this embodiment first determines a reference signal, which is the signal source played by the speaker; then, after the signal source is played by the speaker, it acquires a mixed signal received by multiple microphones, each microphone corresponding to a channel of a different audio range, and the mixed signal contains human voice signal and media signal with echo corresponding to media audio; finally, it inputs the reference signal and the mixed signal into a preset model, uses the preset model to determine the target channel with the highest signal proportion of human voice signal in the mixed signal, and filters the media signal with echo in the target channel according to the reference signal to obtain a clean human voice signal in the target audio range corresponding to the target channel. In this embodiment, by recording a reference signal and accurately locating the target channel of the human voice signal within the mixed signal, the echoing media signal in the mixed signal within the target channel is filtered based on the reference signal to avoid interference from the echoing media signal to the human voice signal. Simultaneously, using the reference signal as a reference, the source of the echo caused by the media signal in the mixed signal can be accurately determined, effectively extracting the clean human voice signal. Furthermore, this embodiment trains a preset model that, based on accurately determining the target channel corresponding to the target audio region, can further extract the clean human voice signal from the target channel. Without requiring multi-functional coordination, it can flexibly acquire clean human voice signals in various scenarios in a relatively simple way, further ensuring the subsequent human voice recognition effect.
[0105] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0106] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A sound processing method, characterized in that, include: A reference signal is determined, wherein the reference signal is a signal source played by the speaker; After the signal source is played by the speaker, a mixed signal received by multiple microphones is acquired. The multiple microphones correspond to different audio channels. The mixed signal includes human voice signals and media signals with echo corresponding to media audio. The reference signal and the mixed signal are input into a preset model. The preset model is used to select the target channel with the highest signal ratio of the human voice signal in the mixed signal. The echoing media signal in the target channel is filtered according to the reference signal to obtain a clean human voice signal in the target audio region corresponding to the target channel.
2. The method according to claim 1, characterized in that, Determining the target vocal channel with the highest signal proportion of the human voice signal in the mixed signal using the preset model includes: Select a target human voice signal from the mixed signal. The target human voice signal exists in at least one set of channels corresponding to the mixed signal. The target human voice signal is the same human voice signal in the at least one set of channels. Using the signal proportion of the target human voice signal in the mixed signal, calculate the sound loudness value of the target human voice signal corresponding to each of the at least one set of channels; Select the channel with the maximum loudness value from the at least one set of channels as the target channel.
3. The method according to claim 1, characterized in that, The determination of the reference signal includes: Based on the number of speakers, obtain the audio signals of a first number of channels to be played by the speakers; From the audio signals of the first number of channels, a target number of audio signals to be identified are selected, wherein the target number is greater than 1; Determine the repeating audio signals among the target number of audio signals to be identified; The proportion of the repeating audio signal in the target number of audio signals to be identified is used as the audio repetition value, and the relationship between the audio repetition value in the repeating audio signal and the first preset threshold is verified. If the audio repetition value is greater than or equal to the first preset threshold, the multiple audio signals that generate the repetitive audio signal are mixed to obtain a second number of audio channels as the reference signal, wherein the first number is greater than the second number.
4. The method according to claim 1, characterized in that, The process of establishing the preset model includes: The speaker is used to play sample media audio, and sample human voices are played in a position where people can sit. Extract the sample reference signal corresponding to the sample media audio; Acquire the sample media signal generated by the sample media audio received by the microphone, and acquire the sample human voice signal generated by the sample human voice received by the microphone; The sample reference signal, the sample media signal, and the sample human voice signal are mixed to generate a sample mixed signal set; The sample mixed signal set is input into the initial model corresponding to the preset model, and the initial model is trained until the time-frequency domain difference between the human voice signal output by the initial model and the human voice signal of the sample is less than the second threshold, thereby generating the preset model.
5. The method according to claim 4, characterized in that, Playing sample human voices in locations where people can sit includes: Set the number of timbre types for the sample human voice, where the number of timbre types is less than the number of locations where a person can sit; Confirm the distribution of the sampled human voices in locations where people can sit; Play the number of sample human voices of the specified timbre types at the specified distribution locations.
6. The method according to claim 4, characterized in that, The process of mixing the sample reference signal, the sample media signal, and the sample human voice signal to generate a sample mixed signal set includes: Obtain the ratio of the number of different human voices contained in the sample human voice; From the sample human voices containing different ratios of the number of human voices, a random number of sample human voice signals are obtained; The random number of sample human voice signals, the sample reference signals, and the sample media signals are mixed at a preset mixing ratio to generate a sample mixed signal set.
7. The method according to claim 1, characterized in that, The determination of the reference signal includes: Read the sound effect types corresponding to multiple signal sources; Obtain the playback parameters set for each of the multiple speakers under the given sound effect type; Based on the arrangement order of speakers in different sound zones under the playback parameters during the same sound playback time period, the combination order of the channels of the multiple speakers is determined; The reference signal is obtained by combining multiple signal sources according to the channel combination order of the loudspeaker.
8. A sound processing device, characterized in that, include: A determination module is used to determine a reference signal, wherein the reference signal is a signal source played by a speaker; The acquisition module is used to acquire a mixed signal received by multiple microphones after the signal source is played by the speaker. The multiple microphones correspond to different audio channels. The mixed signal includes human voice signals and media signals with echo corresponding to media audio. The filtering module is used to input the reference signal and the mixed signal into a preset model, use the preset model to determine the target channel with the highest signal proportion of the human voice signal in the mixed signal, and filter the echoing media signal in the target channel according to the reference signal to obtain a clean human voice signal in the target audio region corresponding to the target channel.
9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1-7.
11. A vehicle, characterized in that, The vehicle is equipped with the device as described in claim 8, or the electronic device as described in claim 9.
Citation Information
Patent Citations
Voice recognition method, device and equipment and computer readable storage medium
CN110992974A
Speech recognition method, device and system, electronic device and storage medium
CN112102816A
Voice signal enhancement method and device
CN115472176A
Human voice positioning method, electronic equipment and storage medium
CN115713946A
Deep learning driven multi-channel filtering for speech enhancement
US20190172476A1