Human voice extraction method, device, product, vehicle-mounted KTV system without wheat and method

By segmenting and feature processing the mixed audio data, employing a dual-segment strategy to process the audio data into segments, and utilizing models such as bidirectional recurrent neural networks for feature extraction and overlay fusion, the problem of poor human voice extraction in closed, noisy scenarios is solved, achieving both lightweight design and improved accuracy.

CN117198317BActive Publication Date: 2026-07-21CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHONGQING CHANGAN AUTOMOBILE CO LTD
Filing Date
2023-09-19
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing voice extraction methods perform poorly in enclosed and noisy environments and consume significant computational resources, making it difficult to achieve lightweight voice extraction.

Method used

By segmenting and feature processing the mixed audio data, a two-segment strategy is adopted to process the audio data into segments, and human voice features are extracted and transformed separately. The bidirectional recurrent neural network and other models are used for feature extraction and superposition fusion, which reduces the demand for computing resources and improves the accuracy of human voice extraction.

Benefits of technology

With limited computing resources, the accuracy of voice extraction in closed scenarios has been improved, achieving both lightweight and enhanced precision in voice extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117198317B_ABST
    Figure CN117198317B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a human voice extraction method, device, product, a car-mounted KTV system without a microphone and a method, and relates to the technical field of audio processing. The method comprises the following steps: segmenting mixed audio data to obtain multiple pieces of audio data; performing feature processing on each piece of audio data to obtain a first processing result corresponding to each piece of audio data; segmenting each first processing result to obtain multiple processing sub-results corresponding to each first processing result; performing feature processing on each processing sub-result to obtain a second processing result corresponding to each processing sub-result; and superimposing and fusing multiple first processing results and multiple second processing results to obtain human voice audio data in the mixed audio data. Through the human voice extraction method, the technical effect of improving the human voice extraction accuracy in a closed scene on the basis of light weight is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, and in particular to a method, apparatus, product, microphone-free in-vehicle KTV system and method for extracting human voices. Background Technology

[0002] In recent years, with the development of artificial intelligence, applications based on the separation of human voice and ambient sound have penetrated into people's lives. Separating human voice from mixed audio has become a widely studied audio processing method.

[0003] However, current voice extraction methods perform poorly in enclosed and noisy environments and are computationally intensive. Therefore, how to improve the accuracy of voice extraction in enclosed environments while maintaining a lightweight approach with limited computational resources is a key technical problem that this invention aims to solve. Summary of the Invention

[0004] This invention provides a method, apparatus, product, and microphone-free in-vehicle KTV system and method for human voice extraction, which improves the accuracy of human voice extraction in closed scenarios while making human voice extraction more lightweight.

[0005] The first aspect of this invention provides a method for human voice extraction, the method comprising:

[0006] The mixed audio data is segmented to obtain multiple audio segments;

[0007] Each audio data segment is subjected to feature processing to obtain a first processing result corresponding to each audio data segment;

[0008] Each of the first processing results is segmented to obtain multiple sub-processing results corresponding to each of the first processing results;

[0009] Each sub-processing result is subjected to feature processing to obtain a second processing result corresponding to each sub-processing result.

[0010] The first processing results and the second processing results are superimposed and fused to obtain the human voice audio data in the mixed audio data.

[0011] Optionally, the step of performing feature processing on each audio data segment to obtain a first processing result corresponding to each audio data segment includes:

[0012] Human voice features are extracted from each audio data segment to obtain a three-dimensional human voice tensor for each audio data segment.

[0013] Perform feature transformation on the human voice three-dimensional tensor of each audio data segment to obtain a first three-dimensional tensor with the same shape corresponding to each audio data segment;

[0014] The first three-dimensional tensor is used as the first processing result.

[0015] Optionally, the step of performing feature processing on each processing sub-result to obtain a second processing result corresponding to each processing sub-result includes:

[0016] Human voice features are extracted from each of the multiple processing sub-results to obtain the human voice three-dimensional tensor of each processing sub-result.

[0017] The human voice three-dimensional tensor of each processing sub-result is subjected to feature transformation to obtain a second three-dimensional tensor with the same shape corresponding to each processing sub-result.

[0018] The second three-dimensional tensor is used as the second processing result.

[0019] Optionally, the method further includes:

[0020] Based on the segmented time step of each audio data segment, multiple first three-dimensional tensors are encapsulated to obtain a first concatenated three-dimensional tensor;

[0021] Based on the segmented time step of each processing sub-result, multiple second three-dimensional tensors of each first processing result are encapsulated to obtain a second spliced ​​three-dimensional tensor corresponding to each first processing result.

[0022] The step of superimposing and fusing multiple first processing results and multiple second processing results to obtain human voice audio data in the mixed audio data includes:

[0023] Based on the time dependency between the second stitched 3D tensor and the first stitched 3D tensor, the first stitched 3D tensor and multiple second stitched 3D tensors are superimposed and stitched together to obtain the human voice audio data.

[0024] Optionally, the segmentation of the mixed audio data to obtain multiple audio segments includes:

[0025] The mixed audio data is randomly divided into a preset number of segments to obtain multiple audio data segments with and / or without overlap.

[0026] The step of segmenting each of the first processing results to obtain multiple processing sub-results corresponding to each of the first processing results includes:

[0027] Each of the first processing results is randomly divided into the preset number of segments to obtain multiple processing sub-results corresponding to the first processing results that have overlaps and / or do not overlap.

[0028] A second aspect of the present invention provides a human voice extraction device, the device comprising:

[0029] The first segmentation module is used to segment the mixed audio data to obtain multiple audio segments;

[0030] The first processing module is used to perform feature processing on each audio data segment to obtain a first processing result corresponding to each audio data segment.

[0031] The second segmentation module is used to segment each of the first processing results to obtain multiple processing sub-results corresponding to each of the first processing results.

[0032] The second processing module is used to perform feature processing on each processing sub-result to obtain the second processing result corresponding to each processing sub-result.

[0033] The fusion module is used to superimpose and fuse multiple first processing results and multiple second processing results to obtain human voice audio data in the mixed audio data.

[0034] Optionally, the first processing module includes:

[0035] The first determining submodule is used to extract human voice features from each audio data segment to obtain the human voice three-dimensional tensor of each audio data segment.

[0036] The first conversion submodule is used to perform feature conversion on the human voice three-dimensional tensor of each audio data segment to obtain a first three-dimensional tensor with the same shape corresponding to each audio data segment.

[0037] The second determining submodule is used to take the first three-dimensional tensor as the first processing result.

[0038] Optionally, the second processing module includes:

[0039] The third determining submodule is used to extract human voice features from each of the multiple processing sub-results to obtain the human voice three-dimensional tensor of each processing sub-result.

[0040] The second conversion submodule is used to perform feature conversion on the human voice three-dimensional tensor of each processing sub-result to obtain a second three-dimensional tensor with the same shape corresponding to each processing sub-result.

[0041] The fourth determination submodule is used to take the second three-dimensional tensor as the second processing result.

[0042] Optionally, the device further includes:

[0043] The first encapsulation module is used to encapsulate multiple first three-dimensional tensors according to the segmented time step of each audio data segment to obtain a first spliced ​​three-dimensional tensor.

[0044] The second encapsulation module is used to encapsulate multiple second three-dimensional tensors of each first processing result according to the segmented time step of each processing sub-result, so as to obtain a second spliced ​​three-dimensional tensor corresponding to each first processing result.

[0045] The fusion module includes:

[0046] The fusion submodule is used to superimpose and splice the first spliced ​​three-dimensional tensor and multiple second spliced ​​three-dimensional tensors based on the time dependency relationship between the second spliced ​​three-dimensional tensor and the first spliced ​​three-dimensional tensor to obtain the human voice audio data.

[0047] Optionally, the first segmentation module includes:

[0048] The first segmentation submodule is used to randomly segment the mixed audio data a preset number of times to obtain the multiple audio data segments with and / or without overlap.

[0049] The second segmentation module includes:

[0050] The second segmentation submodule is used to randomly segment each of the first processing results by the preset number of segments to obtain multiple processing sub-results corresponding to the first processing results that have overlap and / or do not overlap.

[0051] A third aspect of the present invention provides an electronic device, the electronic device comprising: a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the computer program, when executed by the processor, implements the human voice extraction method as described in the first aspect of the present invention.

[0052] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the human voice extraction method of the first aspect of the present invention.

[0053] The fifth aspect of this invention provides a microphone-free in-vehicle karaoke system, the system comprising at least: an in-vehicle audio system; the in-vehicle audio system comprising at least: one or more microphones, an audio digital signal processor, a smart cockpit controller, an audio stream management module, an in-vehicle karaoke client, an amplifier module, and a speaker;

[0054] The MIC microphone is used to collect mixed audio data in the vehicle and send the mixed audio data in the vehicle to the audio digital signal processor;

[0055] The audio digital signal processor is used to collect the in-vehicle mixed audio data and send the in-vehicle mixed audio data to the smart cockpit main controller;

[0056] The intelligent cockpit main controller is used to extract human voice audio data from the mixed audio data in the vehicle through the human voice extraction method described in any of the above embodiments in the hardware abstraction layer, and send the human voice audio data to the audio stream management module;

[0057] The audio stream management module is used to mix the human voice audio data with the song accompaniment track sent by the in-vehicle karaoke client to obtain karaoke audio data and send it to the audio digital signal processor.

[0058] The audio digital signal processor is also used to output the karaoke audio data through the power amplifier module and the speaker.

[0059] Optionally, the vehicle audio system further includes: a signal processing module, a low-pass filter, and an analog-to-digital converter / A2B bus; the signal processing module, the low-pass filter, and the analog-to-digital converter / A2B bus are sequentially located between the MIC microphone and the audio digital signal processor.

[0060] The signal processing module is used to amplify and reduce noise in the in-vehicle mixed audio data after the MIC microphone acquires the mixed audio data, to obtain first in-vehicle mixed audio data and send it to the low-pass filter.

[0061] The low-pass filter is used to remove high-frequency signals from the first in-vehicle mixed audio data to obtain the second in-vehicle mixed audio data, and the second in-vehicle mixed audio data is sent to the audio digital signal processor through the analog-to-digital converter module / A2B bus.

[0062] Optionally, the audio digital signal processor is further configured to perform sound effect processing, earphone monitoring switch processing, echo cancellation, reverb adjustment, and delay processing on the karaoke audio data, and then output the processed karaoke audio data through the power amplifier module and the speaker.

[0063] Optionally, the system further includes: a cloud platform; the vehicle audio system further includes: a communication module; the cloud platform includes at least: an OTA module, a function subscription module, an audio source management module, a membership management module, a payment system, a commission system, and a push system;

[0064] The cloud platform is used to communicate with the in-vehicle karaoke client through the communication module, and realizes the functions of OTA update, function subscription, third-party audio source management, membership management, payment management, third-party commission, and information push through the OTA module, the function subscription module, the audio source management module, the membership management module, the payment system, the commission system, and the push system.

[0065] A sixth aspect of this invention provides a microphone-free in-vehicle karaoke method, the method being applied to a microphone-free in-vehicle karaoke system, the microphone-free in-vehicle karaoke system comprising at least: an in-vehicle audio system; the in-vehicle audio system comprising at least: one or more microphones, an audio digital signal processor, a smart cockpit controller, an audio stream management module, an in-vehicle karaoke client, an amplifier module, and a speaker; the method comprising:

[0066] The in-vehicle mixed audio data is collected by the MIC microphone and sent to the audio digital signal processor.

[0067] The in-vehicle mixed audio data is collected by the audio digital signal processor and sent to the smart cockpit main controller.

[0068] The intelligent cockpit main controller extracts human voice audio data from the mixed audio data in the vehicle through the human voice extraction method described in any of the above embodiments in the hardware abstraction layer, and sends the human voice audio data to the audio stream management module.

[0069] The audio stream management module mixes the human voice audio data with the song accompaniment track sent by the in-vehicle karaoke client to obtain karaoke audio data, which is then sent to the audio digital signal processor.

[0070] The audio digital signal processor transmits the karaoke audio data through the amplifier module and outputs it via the speaker.

[0071] Optionally, the in-vehicle audio system further includes: a signal processing module, a low-pass filter, and an analog-to-digital converter / A2B bus; the signal processing module, the low-pass filter, and the analog-to-digital converter / A2B bus are sequentially located between the MIC microphone and the audio digital signal processor; in this method, sending the in-vehicle mixed audio data to the audio digital signal processor includes:

[0072] After the MIC microphone acquires the mixed audio data in the vehicle, the signal processing module amplifies and reduces the noise of the mixed audio data in the vehicle to obtain the first mixed audio data in the vehicle and sends it to the low-pass filter.

[0073] The high-frequency signal in the first in-vehicle mixed audio data is removed by the low-pass filter to obtain the second in-vehicle mixed audio data, and the second in-vehicle mixed audio data is sent to the audio digital signal processor through the analog-to-digital converter module / A2B bus.

[0074] Optionally, the step of outputting the karaoke audio data through the amplifier module and the speaker via the audio digital signal processor includes:

[0075] The audio digital signal processor processes the karaoke audio data through sound effects processing, earphone monitoring, echo cancellation, reverb adjustment, and delay processing. The processed karaoke audio data is then output through the amplifier module and the speaker.

[0076] Optionally, the system further includes: a cloud platform; the in-vehicle audio system further includes: a communication module; the cloud platform includes at least: an OTA module, a function subscription module, an audio source management module, a membership management module, a payment system, a commission system, and a push system; the method further includes:

[0077] Through the cloud platform, the communication module communicates with the in-vehicle karaoke client, and the OTA module, function subscription module, audio source management module, membership management module, payment system, commission system, and push system respectively realize the functions of OTA update, function subscription, third-party audio source management, membership management, payment management, third-party commission, and information push.

[0078] The voice extraction method provided in this embodiment of the invention involves segmenting mixed audio data to obtain multiple audio segments; performing feature processing on each audio segment to obtain a first processing result corresponding to each audio segment; segmenting each first processing result to obtain multiple processing sub-results corresponding to each first processing result; performing feature processing on each processing sub-result to obtain a second processing result corresponding to each processing sub-result; and then superimposing and fusing the multiple first processing results and multiple second processing results to extract the voice audio data from the mixed audio data. In this embodiment, a dual-segment strategy is used to perform segmentation processing on the mixed audio data twice. The multiple first processing results after the first segmentation processing are then segmented and processed again, and the processing results of the two segmentation processing are superimposed and fused to obtain the final voice features. Thus, this embodiment can reduce the amount of data processed for each feature step by segmenting the mixed audio, thereby reducing computational resources and making the voice extraction method lightweight. It can also make up for the deficiencies of the first processing result of the first segmentation block processing in terms of voice features based on the second processing result of further segmentation block processing, thereby improving the accuracy of voice extraction in closed scenes with high environmental noise. It achieves the technical effect of improving the accuracy of voice extraction in closed scenes on the basis of lightweighting. Attached Figure Description

[0079] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0080] Figure 1 This is a flowchart illustrating a human voice extraction method according to an embodiment of the present invention;

[0081] Figure 2 This is a structural diagram of a human voice extraction algorithm network model shown in one embodiment of the present invention;

[0082] Figure 3 This is a structural block diagram of a human voice extraction device provided in an embodiment of the present invention;

[0083] Figure 4 This is a schematic diagram of an electronic device according to an embodiment of the present invention;

[0084] Figure 5 This is a structural block diagram of a microphone-free in-vehicle KTV system according to an embodiment of the present invention;

[0085] Figure 6This is a schematic diagram illustrating the installation position of a car microphone according to an embodiment of the present invention;

[0086] Figure 7 This is a flowchart illustrating the steps of a method for extracting human voices inside a vehicle, as shown in an embodiment of the present invention.

[0087] Figure 8 This is a hardware system block diagram of a microphone-free in-vehicle KTV system according to an embodiment of the present invention;

[0088] Figure 9 This is a software system block diagram of a microphone-free in-vehicle KTV system according to an embodiment of the present invention;

[0089] Figure 10 This is a flowchart illustrating a multi-zone microphone-free KTV method according to an embodiment of the present invention;

[0090] Figure 11 This is a flowchart of a microphone-free in-vehicle KTV method provided by an embodiment of the present invention. Detailed Implementation

[0091] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0092] refer to Figure 1 , Figure 1 This is a flowchart illustrating a human voice extraction method according to an embodiment of the present invention. Figure 1 As shown, the human voice extraction method in this embodiment may include the following steps:

[0093] Step S1: Segment the mixed audio data to obtain multiple audio segments.

[0094] In this embodiment, continuous mixed audio data can be acquired first. This embodiment does not limit the method for acquiring continuous mixed audio data. The mixed audio data in this embodiment refers to audio data including human voice audio and non-human voice audio. Non-human voice audio can be ambient sound, noise, or other non-human voice audio data; this embodiment does not limit this. After acquiring the mixed audio data, it needs to be segmented to obtain multiple audio segments. These multiple audio segments are the mixed audio data segments formed after segmentation, with each segment being a part of the mixed audio data.

[0095] Step S2: Perform feature processing on each audio data segment to obtain the first processing result corresponding to each audio data segment.

[0096] In this embodiment, after determining the segmented audio data, each segment needs to undergo feature processing to obtain a first processing result. This feature processing is the aforementioned "block processing." The purpose of this embodiment is to perform feature processing on each segment of audio data to process the human voice features of the segmented audio data blocks, obtaining a first processing result. This first processing result is the feature processing result based on the preliminary human voice feature processing of each segmented mixed audio data. Performing feature processing on each segment reduces the amount of data processed, thus achieving a lightweight human voice extraction method.

[0097] Step S3: Divide each of the first processing results into multiple processing sub-results corresponding to each of the first processing results.

[0098] In this embodiment, after obtaining the first processing results corresponding to multiple audio data segments, each first processing result needs to be further segmented to obtain multiple processing sub-results corresponding to each first processing result. Specifically, the multiple processing sub-results in this embodiment are multiple first processing results formed after segmenting the first processing results, with each sub-result being a part of the first processing result.

[0099] For example, if there are eight first processing results, each of the eight first processing results needs to be segmented separately to obtain multiple processing sub-results corresponding to each of the eight first processing results. The method for segmenting the first processing results can be the same as or different from the method for segmenting the mixed audio data; this embodiment does not impose any restrictions on this.

[0100] Step S4: Perform feature processing on each sub-result to obtain the second processing result corresponding to each sub-result.

[0101] In this embodiment, after obtaining the multiple processing sub-results corresponding to the first processing result, feature processing can be performed on each of the multiple processing sub-results to obtain the second processing result corresponding to each processing sub-result. The feature processing method for each processing sub-result in this embodiment can be the same as or similar to the feature processing method for each segment of audio data. The purpose of performing feature processing on each processing sub-result in this embodiment is to perform corresponding human voice feature processing on the segmented processing sub-result blocks again to obtain the second processing result. The second processing result in this embodiment is the feature processing result based on further human voice feature processing of each segmented processing sub-result. Performing feature processing on each processing sub-result can further reduce the amount of data processed, achieving a lightweight approach to human voice extraction.

[0102] It should be noted that the feature processing performed on each processing sub-result in this application can be performed immediately after each segment is obtained, or it can be performed simultaneously on each segment of the multiple processing sub-results after obtaining a first processing result, or it can be performed simultaneously on each segment of the multiple processing sub-results after obtaining multiple first processing results. This embodiment does not impose any restrictions on this. Furthermore, the segmentation of the first processing result can be performed sequentially for each first processing result, or it can be performed simultaneously on multiple first processing results. This embodiment also does not impose any restrictions on this.

[0103] Step S5: Superimpose and fuse multiple first processing results and multiple second processing results to obtain human voice audio data in the mixed audio data.

[0104] In this embodiment, after obtaining the second processing result corresponding to each processing sub-result, the multiple first processing results obtained from the first feature processing can be superimposed and fused with the multiple second processing results obtained from the second feature processing to extract the human voice audio data from the mixed audio data. The second processing result is a feature result obtained by further segmenting and "block processing" based on the first processing result. It can compensate for the accuracy of the human voice features in the first processing result. Therefore, by superimposing the two, the human voice audio data can be extracted from the mixed audio data, thus achieving human voice extraction.

[0105] It should be noted that the "feature processing" in this embodiment can be performed using any method capable of extracting human voice features. For example, "feature processing" can be performed using the Hourglass model, etc. This embodiment does not impose any restrictions on the specific method of feature processing. In this embodiment, after segmenting the mixed audio data, any "feature processing" method can be used to extract human voice features from each segment of audio data to obtain a first processing result. Then, the first processing result is segmented again, and each sub-result is further processed using the same "feature processing" method as arbitrarily selected above to extract human voice features, resulting in a second processing result. Finally, the first and second processing results are superimposed and fused to obtain the extracted human voice audio data.

[0106] In this embodiment, human voice extraction is performed on mixed audio data using a separation method of segmentation, block processing, and overlapping addition. Specifically, a two-segment strategy is used to sequentially perform segmentation and block processing on the mixed audio data twice, and the processing results of the two segmentation and block processing are superimposed and fused to obtain the final human voice features. In this way, this embodiment can reduce the amount of data for each feature processing by segmenting the data, achieving lightweight human voice extraction under limited computing resources. At the same time, it can compensate for the lack of accuracy of the human voice features in the first processing result by further segmenting and block processing the first processing result, thereby improving the accuracy of human voice extraction in closed scenes with high environmental noise. This achieves improved accuracy of human voice extraction in mixed audio in closed scenes while maintaining a lightweight approach.

[0107] In conjunction with the above embodiments, in one implementation, the present invention also provides a method for human voice extraction. In this method, step S2 may specifically include steps S21 to S23:

[0108] Step S21: Extract human voice features from each audio data segment to obtain the human voice three-dimensional tensor of each audio data segment.

[0109] In this embodiment, after obtaining multiple audio data segments, human voice features are extracted from each audio data segment to obtain a three-dimensional tensor of human voice corresponding to each audio data segment, which is a three-dimensional array representing human voice features. In this embodiment, human voice feature extraction can be performed using a pre-trained human voice extraction model, such as a human voice extraction model pre-trained using a bidirectional recurrent neural network (Bi-RNN).

[0110] Step S22: Perform feature transformation on the human voice three-dimensional tensor of each audio data segment to obtain a first three-dimensional tensor with the same shape corresponding to each audio data segment.

[0111] In this embodiment, after obtaining the human voice three-dimensional tensor of each audio data segment, feature transformation can be performed on the human voice three-dimensional tensor of each audio data segment to obtain a first three-dimensional tensor with the same shape after transformation. Here, the first three-dimensional tensor refers to a three-dimensional tensor with the same shape as the human voice three-dimensional tensor of each audio data segment. The same shape indicates that the two three-dimensional tensors have the same meaning in expressing features.

[0112] Step S23: Use the first three-dimensional tensor as the first processing result.

[0113] In this embodiment, after obtaining the first three-dimensional tensor corresponding to each audio data segment through feature extraction and feature transformation, the feature processing of each audio data segment is completed, and the first three-dimensional tensor corresponding to each audio data segment can be used as the first processing result corresponding to each audio data segment.

[0114] In conjunction with the above embodiments, in one implementation, the present invention also provides a method for human voice extraction. In this method, step S4 may specifically include steps S41 to S43:

[0115] Step S41: Extract human voice features from each of the multiple processing sub-results to obtain the human voice three-dimensional tensor of each processing sub-result.

[0116] In this embodiment, after obtaining multiple sub-results of the first processing result, i.e., after dividing the first three-dimensional tensor into multiple sub-three-dimensional tensors, it is possible to further extract human voice features from each sub-result (i.e., each sub-three-dimensional tensor), thereby obtaining the corresponding human voice three-dimensional tensor for each sub-result, which is a more refined three-dimensional array representing human voice features. In this embodiment, the human voice features of each sub-result can be extracted using a pre-trained human voice extraction model, such as a human voice extraction model pre-trained using a bidirectional recurrent neural network (Bi-RNN).

[0117] Step S42: Perform feature transformation on the human voice three-dimensional tensor of each processing sub-result to obtain a second three-dimensional tensor with the same shape corresponding to each processing sub-result.

[0118] In this embodiment, after obtaining the human voice three-dimensional tensor of each processing sub-result, feature transformation can be performed on the human voice three-dimensional tensor of each processing sub-result to obtain a second three-dimensional tensor with the same shape after transformation. The second three-dimensional tensor refers to a three-dimensional tensor with the same shape as the human voice three-dimensional tensor of each processing sub-result; the same shape indicates that the two three-dimensional tensors express the same features.

[0119] Step S43: Use the second three-dimensional tensor as the second processing result.

[0120] In this embodiment, after obtaining the second three-dimensional tensor corresponding to each processing sub-result through feature extraction and feature transformation, the feature processing of each processing sub-result is completed, and the second three-dimensional tensor corresponding to each processing sub-result can be used as the second processing result corresponding to each processing sub-result.

[0121] like Figure 2 As shown, Figure 2 This is a structural diagram of a network model for a human voice extraction algorithm according to an embodiment of the present invention. Figure 2 In this process, a continuous audio stream sequence (i.e., mixed audio data) is input, then the audio stream sequence is segmented to obtain multiple segments of audio data. Each segment is then processed in blocks. The structure for block processing includes, in sequence, a bidirectional recurrent neural network, a fully connected linear layer, and layer normalization, resulting in a first 3D tensor (i.e., the first processing result) after feature extraction and transformation. The model used for block processing can be trained on a benchmark dataset using methods such as convolutional encoders, splitting, and transposed convolutional decoders. The multiple first processing results (i.e., the first 3D tensors) are then further segmented to obtain multiple sub-processing results. Each sub-processing result is then processed in blocks, again using the same structure: a bidirectional recurrent neural network, a fully connected linear layer, and layer normalization, resulting in a second 3D tensor (i.e., the second processing result). Finally, the results of the second block processing (multiple second 3D tensors) are overlapped and added to the results of the first block processing (multiple first 3D tensors), resulting in the output: a sequence of human voice audio streams.

[0122] In conjunction with the above embodiments, in one implementation, the present invention also provides a method for human voice extraction. This method further includes steps S22.5 and S42.5, and step S5 specifically includes step S51:

[0123] Step S22.5: Based on the segmented time step of each audio data segment, encapsulate multiple first three-dimensional tensors to obtain a first concatenated three-dimensional tensor.

[0124] In this embodiment, after obtaining the first three-dimensional tensor corresponding to each audio data segment, a first sequence of each audio data segment can be defined based on the segment time step of each audio data segment during the segmentation of the mixed audio data. Then, according to the order of the first sequence, the multiple first three-dimensional tensors corresponding to the multiple audio data segments are concatenated and encapsulated to obtain a first concatenated three-dimensional tensor. Here, the first concatenated three-dimensional tensor refers to a three-dimensional tensor obtained by concatenating multiple first three-dimensional tensors.

[0125] Step S42.5: Based on the segmented time step of each processing sub-result, encapsulate the multiple second three-dimensional tensors of each first processing result to obtain a second spliced ​​three-dimensional tensor corresponding to each first processing result.

[0126] In this embodiment, after obtaining the second three-dimensional tensor corresponding to each processing sub-result, a second sequence of each processing sub-result can be defined based on the segmentation time step of each processing sub-result when the first processing result is segmented. Then, according to the order of the second sequence, multiple second three-dimensional tensors corresponding to the multiple processing sub-results in each first processing result are concatenated and encapsulated to obtain a second concatenated three-dimensional tensor corresponding to each first processing result. Here, the second concatenated three-dimensional tensor refers to a three-dimensional tensor obtained by concatenating multiple second three-dimensional tensors, and each first processing result corresponds to one second concatenated three-dimensional tensor.

[0127] Step S51: Based on the time dependency between the second stitched 3D tensor and the first stitched 3D tensor, the first stitched 3D tensor and multiple second stitched 3D tensors are superimposed and stitched together to obtain the human voice audio data.

[0128] In this embodiment, a time dependency exists between the obtained second stitched 3D tensor and the first stitched 3D tensor. This time dependency characterizes the time step relationship between the second stitched 3D tensor and the first stitched 3D tensor, that is, it characterizes the dependency relationship between the local (within the segmented "block") and the global (between the segmented "blocks") aspects. In this embodiment, during overlay and fusion, the first stitched 3D tensor and multiple second stitched 3D tensors can be overlaid and stitched based on the time dependency between the second stitched 3D tensor and the first stitched 3D tensor. This overlay and stitching is then converted into a sequential output to obtain the extracted human voice audio data.

[0129] In conjunction with the above embodiments, in one implementation, the present invention also provides a method for human voice extraction. In this method, step S1 may specifically include step S11, and step S3 may specifically include step S31:

[0130] Step S11: The mixed audio data is randomly divided into a preset number of segments to obtain multiple audio data segments with and / or without overlap.

[0131] In this embodiment, when segmenting the mixed audio data, the mixed audio data can be randomly segmented a preset number of times to obtain multiple audio segments with and / or without overlap. The preset number is a pre-defined number of segments, such as 48, 36, 24, etc., and the specific value of the preset number is not limited. This embodiment performs random segmentation, and the resulting multiple audio segments can be non-overlapping, overlapping, or partially overlapping; this embodiment does not impose any restrictions on this.

[0132] Step S31: Divide each of the first processing results into the preset number of random segments to obtain multiple processing sub-results corresponding to the first processing results that have overlap and / or do not overlap.

[0133] In this embodiment, when segmenting the first processing result, similar to segmenting mixed audio data, the first processing result can be randomly segmented a preset number of times to obtain multiple processing sub-results with overlapping and / or non-overlapping segments. The preset number of segments for the first processing result can be the same as or different from the preset number for segmenting mixed audio data; this is not limited. In this embodiment, when randomly segmenting each first processing result, the resulting multiple processing sub-results can be non-overlapping, overlapping, or partially overlapping; this embodiment does not impose any limitations on this.

[0134] In this embodiment, a separation method based on temporal segmentation, block processing, and overlapping is used to iteratively apply local (intra-block) and global (inter-block) modeling in an alternating manner. The output is converted back to a sequence output using an overlapping method to obtain the extracted human voice audio data. Thus, accurate human voice extraction and recognition in closed and noisy scenarios are achieved on a lightweight basis.

[0135] In conjunction with the above embodiments, in one implementation, the present invention also provides a method for extracting human voices. In this method, human voice audio can be extracted from mixed audio using the following formula (1):

[0136] Formula (1):

[0137]

[0138] In formula (1), V b Output: The extracted human voice audio; h b : Defined mapping function 1; T b : Three-dimensional tensor; D b : Feature dimension, where Db =G(f b [T b [:,:,i]), i=1,...,S)[:,:,i]+m,i=1,...,S;G:Weights of the fully connected layer;f b : Defined mapping function 2; i, j, s are coefficients 1, 2, 3 respectively; ":" is the separator; S: Intra-block input length; m: Bias of fully connected layer; ε: Numerical stability coefficient; Z: Rescaling factor 1; R: Rescaling factor 2; K: Inter-block input length; N: Feature dimension; ⊙: Hadamard product.

[0139] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0140] Based on the same inventive concept, one embodiment of the present invention provides a human voice extraction device 300. (See reference...) Figure 3 , Figure 3 This is a structural block diagram of a human voice extraction device provided in an embodiment of the present invention. Figure 3 As shown, the device 300 includes:

[0141] The first segmentation module 301 is used to segment the mixed audio data to obtain multiple audio segments;

[0142] The first processing module 302 is used to perform feature processing on each audio data segment to obtain a first processing result corresponding to each audio data segment.

[0143] The second segmentation module 303 is used to segment each of the first processing results to obtain multiple processing sub-results corresponding to each of the first processing results.

[0144] The second processing module 304 is used to perform feature processing on each processing sub-result to obtain the second processing result corresponding to each processing sub-result.

[0145] The fusion module 305 is used to superimpose and fuse multiple first processing results and multiple second processing results to obtain human voice audio data in the mixed audio data.

[0146] Optionally, the first processing module 302 includes:

[0147] The first determining submodule is used to extract human voice features from each audio data segment to obtain the human voice three-dimensional tensor of each audio data segment.

[0148] The first conversion submodule is used to perform feature conversion on the human voice three-dimensional tensor of each audio data segment to obtain a first three-dimensional tensor with the same shape corresponding to each audio data segment.

[0149] The second determining submodule is used to take the first three-dimensional tensor as the first processing result.

[0150] Optionally, the second processing module 304 includes:

[0151] The third determining submodule is used to extract human voice features from each of the multiple processing sub-results to obtain the human voice three-dimensional tensor of each processing sub-result.

[0152] The second conversion submodule is used to perform feature conversion on the human voice three-dimensional tensor of each processing sub-result to obtain a second three-dimensional tensor with the same shape corresponding to each processing sub-result.

[0153] The fourth determination submodule is used to take the second three-dimensional tensor as the second processing result.

[0154] Optionally, the device 300 further includes:

[0155] The first encapsulation module is used to encapsulate multiple first three-dimensional tensors according to the segmented time step of each audio data segment to obtain a first spliced ​​three-dimensional tensor.

[0156] The second encapsulation module is used to encapsulate multiple second three-dimensional tensors of each first processing result according to the segmented time step of each processing sub-result, so as to obtain a second spliced ​​three-dimensional tensor corresponding to each first processing result.

[0157] The fusion module 305 includes:

[0158] The fusion submodule is used to superimpose and splice the first spliced ​​three-dimensional tensor and multiple second spliced ​​three-dimensional tensors based on the time dependency relationship between the second spliced ​​three-dimensional tensor and the first spliced ​​three-dimensional tensor to obtain the human voice audio data.

[0159] Optionally, the first segmentation module 301 includes:

[0160] The first segmentation submodule is used to randomly segment the mixed audio data a preset number of times to obtain the multiple audio data segments with and / or without overlap.

[0161] The second segmentation module 303 includes:

[0162] The second segmentation submodule is used to randomly segment each of the first processing results by the preset number of segments to obtain multiple processing sub-results corresponding to the first processing results that have overlap and / or do not overlap.

[0163] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the human voice extraction method as described in any of the above embodiments of the present invention.

[0164] Based on the same inventive concept, another embodiment of the present invention provides an electronic device 400, such as... Figure 4 As shown. Figure 4 This is a schematic diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a memory 402, a processor 401, and a computer program stored in the memory and executable on the processor. When executed by the processor, the program implements the steps of the human voice extraction method described in any of the above embodiments of the present invention.

[0165] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0166] In conjunction with the above embodiments, in one implementation, to address the problems of current in-vehicle karaoke systems often requiring car owners to purchase separate microphones, resulting in high costs, and cumbersome system compatibility testing and poor user experience when multiple people are singing simultaneously, this invention also provides a microphone-free in-vehicle karaoke system. For example... Figure 5 As shown, Figure 5 This is a structural block diagram of a microphone-free in-vehicle KTV system according to an embodiment of the present invention. In this embodiment, the microphone-free in-vehicle KTV system 500 includes at least: an in-vehicle audio system 508; wherein, the in-vehicle audio system 508 includes at least: one or more MIC microphones 501 (shared microphone for voice recognition), an audio digital signal processor 502 (Audio DSP), an intelligent cockpit main controller 503 (AVN / HU / IVISOC), an audio stream management module 504, an in-vehicle karaoke client (in-vehicle karaoke application APP) 505, a power amplifier module 506, and a speaker 507.

[0167] The MIC microphone 501 is used to collect mixed audio data in the vehicle and send the mixed audio data in the vehicle to the audio digital signal processor 502.

[0168] In this embodiment, one or more microphones can be installed in the car's smart cockpit. When multiple microphones are installed, up to six independent karaoke audio signal acquisition zones can be supported. In one embodiment, the installation location of the multi-zone microphones is shown below. Figure 6 ,like Figure 6 As shown, Figure 6 This is a schematic diagram illustrating the installation position of a car microphone according to an embodiment of the present invention. Each microphone is installed above the seat at the edge of the roof, supporting up to six audio zones. Each microphone operates independently without interference. Furthermore, this embodiment of the microphone-free in-car karaoke system eliminates the need for a handheld microphone, while still being compatible with handheld microphones.

[0169] After a user sings karaoke in the car, the various microphones inside the vehicle can collect mixed audio data. This mixed audio data includes the accompaniment sound from the speakers, the user's voice, and other background noise (such as tire noise, road noise, etc.). After collecting the mixed audio data, the microphones can send it to the audio digital signal processor.

[0170] The audio digital signal processor 502 is used to collect the in-vehicle mixed audio data and send the in-vehicle mixed audio data to the smart cockpit main controller 503.

[0171] In this embodiment, the audio digital signal processor (ADSP) collects in-vehicle mixed audio data transmitted from the microphone (MIC) and sends the collected in-vehicle mixed audio data to the hardware abstraction layer (HAL) of the smart cockpit main controller. In one embodiment, the ADSP also performs digital noise reduction and a series of data processing on the collected in-vehicle mixed audio data, and then sends the processed in-vehicle mixed audio data to the hardware abstraction layer (HAL) of the smart cockpit main controller via the IIS bus.

[0172] The intelligent cockpit main controller 503 is used to extract human voice audio data from the mixed audio data in the vehicle using the human voice extraction method described in any of the above embodiments in the hardware abstraction layer, and send the human voice audio data to the audio stream management module 504.

[0173] In this embodiment, after the smart cockpit main controller receives the mixed in-vehicle audio data from the audio digital signal processor, it can extract the human voice audio data from the mixed in-vehicle audio data in the hardware abstraction layer of the smart cockpit main controller through the human voice extraction method described in any of the above embodiments, thereby obtaining the extracted human voice audio data, and sending the extracted human voice audio data to the audio stream management module.

[0174] Since singing karaoke in a car is a closed environment with high noise levels (such as tire noise, road noise, etc.), this embodiment uses the voice extraction method described in any of the aforementioned embodiments to extract voices from mixed audio data in the car. This method is not only more suitable for lightweight smart cockpit controllers, but also allows for more accurate extraction of voice audio data in the karaoke scenario.

[0175] In one method of human voice extraction, such as Figure 7 As shown, Figure 7 This is a flowchart illustrating the steps of a method for extracting human voices inside a vehicle, as shown in an embodiment of the present invention. Figure 7 The process begins with inputting microphone audio data for each sound zone. This data can be mixed audio data collected by microphones located in different sound zones within the vehicle. The input data is then segmented chronologically, with each segment defined by a time step. Next, each segment undergoes block processing to obtain a 3D tensor with extracted human voice features. This tensor is then transformed into another 3D tensor of the same shape, resulting in multiple transformed 3D tensors for each block. This process is repeated. The transformed 3D tensors (which can be understood as the feature-extracted and transformed microphone audio data for each sound zone) are then segmented chronologically, with each segment within each block defined by a time step. This process is repeated again, with each segment within each block undergoing block processing to obtain another 3D tensor with extracted human voice features. This tensor is then transformed into another 3D tensor of the same shape, resulting in multiple transformed 3D tensors for each segment within each block. Finally, the three-dimensional tensors obtained from the first block processing are concatenated in sequence to obtain a single three-dimensional tensor. Then, the three-dimensional tensors obtained from the second block processing are concatenated in sequence to obtain multiple three-dimensional tensors. The previously obtained three-dimensional tensor and the multiple three-dimensional tensors obtained afterward are converted back into a human voice audio sequence by applying an overlapping and addition method, thereby outputting the extracted human voice audio data.

[0176] The audio stream management module 504 is used to mix the human voice audio data with the song accompaniment track sent by the in-vehicle karaoke client 505 to obtain karaoke audio data and send it to the audio digital signal processor 502.

[0177] In this embodiment, the audio stream management module can receive human voice audio data sent by the smart cockpit main controller, and can also receive the song accompaniment track sent by the in-vehicle karaoke client. This song accompaniment track is the accompaniment track of the karaoke song currently selected by the user. Furthermore, the audio stream management module can perform relevant mixing processing on the received human voice audio data and the received song accompaniment track to obtain karaoke audio data, and then send the karaoke audio data to the audio digital signal processor.

[0178] The audio digital signal processor 502 is also used to output the karaoke audio data through the power amplifier module 506 and the speaker 507.

[0179] In this embodiment, after the audio digital signal processor receives the karaoke audio data, it can output the karaoke audio data through the power amplifier module and the speaker, thereby realizing karaoke entertainment in the car.

[0180] The microphone-free in-vehicle karaoke system based on this embodiment enables convenient and free karaoke singing in various audio zones within the in-vehicle smart cockpit, eliminating the need for purchasing microphones and system debugging. It boasts low operating costs, real-time sound effects, and allows multiple users to sing simultaneously across multiple audio zones, offering strong entertainment and interactivity. It achieves low investment costs while meeting user needs, directly utilizing the built-in voice microphone and sound system of the car's smart cockpit, resulting in low latency, excellent feedback noise suppression, and a wide range of application scenarios. Furthermore, the voice extraction method employed in this embodiment of the microphone-free in-vehicle karaoke system is more suitable for lightweight smart cockpit main controllers, achieving accurate voice extraction even in noisy, enclosed in-vehicle environments without requiring excessive computing resources.

[0181] In conjunction with the above embodiments, in one implementation, the present invention also provides a microphone-free in-vehicle KTV system. In this microphone-free in-vehicle KTV system, the in-vehicle audio system further includes: a signal processing module, a low-pass filter, and an analog-to-digital converter / A2B bus; and between the MIC microphone and the audio digital signal processor, the signal processing module, the low-pass filter, and the analog-to-digital converter / A2B bus (i.e., ADC / A2B) are sequentially located.

[0182] The signal processing module is used to amplify and reduce the noise of the in-vehicle mixed audio data after the MIC microphone acquires the mixed audio data, to obtain first in-vehicle mixed audio data and send it to the low-pass filter.

[0183] In this embodiment, after the MIC microphone collects the mixed audio data in the vehicle, the mixed audio data signal in the vehicle is amplified and noise-reduced by the signal processing module. For example, the signal processing module can amplify the mixed audio data in the vehicle through a gain amplifier and perform signal processing such as hardware noise reduction to obtain the first mixed audio data in the vehicle after signal processing, and then send the first mixed audio data in the vehicle to the low-pass filter.

[0184] The low-pass filter is used to remove high-frequency signals from the first in-vehicle mixed audio data to obtain the second in-vehicle mixed audio data, and the second in-vehicle mixed audio data is sent to the audio digital signal processor through the analog-to-digital converter module / A2B bus.

[0185] In this embodiment, after receiving the first in-vehicle mixed audio data, the low-pass filter can remove high-frequency signals from the first in-vehicle mixed audio data, thereby obtaining a second in-vehicle mixed audio data with high-frequency signals removed. This second in-vehicle mixed audio data is then sent to the audio digital signal processor via an analog-to-digital converter / A2B bus. In one embodiment, each audio zone is configured with independent audio signal acquisition, requiring a microphone sampling rate of 48kHz and an ADC (analog-to-digital converter) of 24bit. In another embodiment, the microphones of each audio zone can be connected to the intelligent cockpit system host via the analog-to-digital converter / A2B bus.

[0186] In conjunction with the above embodiments, in one implementation, the present invention also provides a microphone-free in-vehicle KTV system. In this microphone-free in-vehicle KTV system, the audio digital signal processor is further configured to perform sound effect processing, earphone monitoring switching processing, echo cancellation, reverb adjustment, and delay processing on the karaoke audio data, and then output the processed karaoke audio data through the power amplifier module and the speaker.

[0187] In this embodiment, after receiving the karaoke audio data sent by the audio stream management module, the audio digital signal processor can sequentially perform relevant sound effect processing, ear monitor switch processing, adaptive echo cancellation processing, audio reverb adjustment and sound equalization adjustment, and timing-based delay processing on the karaoke audio data. Specifically, the ear monitor switch processing determines whether the ear monitor switch is on and performs corresponding processing accordingly. The karaoke audio data, after undergoing these processing steps, is then output and played through the speaker via the power amplifier module, achieving a better karaoke sound quality.

[0188] In addition, in one embodiment, before performing relevant sound effect processing, the audio digital signal processor can also perform human voice noise reduction processing on the received karaoke audio data to achieve further noise reduction, and then perform sound effect processing, earphone switch processing, echo cancellation, reverb adjustment, and delay processing on the noise-reduced karaoke audio data in sequence.

[0189] In conjunction with the above embodiments, in one implementation, the present invention also provides a microphone-free in-vehicle KTV system. This microphone-free in-vehicle KTV system further includes: a cloud platform; the in-vehicle audio system further includes: a communication module; the cloud platform includes at least: an OTA module, a function subscription module, an audio source management module, a membership management module, a payment system, a commission system, and a push system.

[0190] The cloud platform is used to communicate with the in-vehicle karaoke client through the communication module, and realizes the functions of OTA update, function subscription, third-party audio source management, membership management, payment management, third-party commission, and information push through the OTA module, the function subscription module, the audio source management module, the membership management module, the payment system, the commission system, and the push system.

[0191] In this embodiment, the in-car karaoke client in the car audio system is connected to the cloud platform. The in-car karaoke client can communicate with the cloud platform via the communication module in the car audio system, allowing the cloud platform to update the client's functions. Specifically, the communication module in this embodiment can be a 4G / 5G / WIFI communication module, enabling communication between the cloud platform and the in-car karaoke client via 4G / 5G / WIFI signals.

[0192] In this embodiment, the cloud platform can implement OTA (Over-The-Air) updates for the in-vehicle karaoke client via the OTA module. Specifically, the OTA update function can support remote updates of the in-vehicle karaoke function by the cloud platform, supporting OTA updates for individual vehicles and vehicle models as a whole.

[0193] In this embodiment, the cloud platform can implement the function subscription function between the cloud platform and the in-vehicle karaoke client through the function subscription module. Specifically, the function subscription function in this embodiment supports users to freely configure and subscribe to the in-vehicle karaoke function on demand. For example: 1. Subscription period configuration: billing can be done daily, with the function turning off after 24 hours; billing can be done weekly, with the function turning off after 7*24 hours; billing can be done monthly / yearly or other durations; 2. Subscription price configuration: prices can be set according to different periods, and prices can be dynamically configured; 3. Limited-time free function configuration: limited-time free functions can be configured according to different vehicle groups, and the duration can be customized.

[0194] In this embodiment, the cloud platform can implement third-party audio source management functions through the audio source management module. Specifically, audio source management can support third-party CP / SP access and joint operation between OEMs and third parties. The cloud platform can implement membership management functions through the membership management module. Specifically, membership management has account and permission functions and information security. Furthermore, the cloud platform can also implement payment management functions, third-party commission sharing functions, and information push functions through the payment system, commission sharing system, and push system, respectively. Specifically, the payment system supports payment functions; the commission sharing system supports revenue sharing between OEMs and third-party CP / SPs; and the push system supports real-time push of operation-related activities and related information.

[0195] This microphone-free in-car karaoke system eliminates the need for car owners to purchase additional handheld microphones. It supports OTA (Over-The-Air) configuration updates and utilizes the built-in microphones in each audio zone of the car's smart cockpit for karaoke. Furthermore, it can be applied to other karaoke-related systems and terminals. This microphone-free in-car karaoke system allows for the addition of karaoke functionality subscriptions and OTA update subscriptions to the car's smart cockpit without increasing hardware costs, enhancing the profitability of the car manufacturer's ecosystem operations. Additionally, users do not need to purchase separate microphones, reducing user costs, and the system supports up to six independent karaoke zones, enhancing user experience and engagement.

[0196] like Figure 8 As shown, Figure 8 This is a hardware system block diagram of a microphone-free in-vehicle KTV system according to an embodiment of the present invention. Figure 8 In China, the hardware system of a microphoneless in-vehicle KTV system includes at least: multiple microphones, signal processing module, low-pass filter, analog-to-digital converter / A2B bus, audio digital signal processor, intelligent cockpit controller, communication module, cloud platform, digital-to-analog converter, audio amplifier module, and speakers.

[0197] After the microphones in each sound zone collect the mixed audio data in the vehicle, the signal processing module amplifies and reduces the noise of the mixed audio data signal. Then, the mixed audio data is sent to the audio digital signal processor via analog-to-digital conversion / A2B bus. After receiving the mixed audio data, the audio digital signal processor performs digital noise reduction and data processing on the mixed audio data again, and sends it to the AVN / HU / IVI SOC (intelligent cockpit main controller) via the IIS bus. Then, the intelligent cockpit main controller extracts human voice data from the mixed audio data in the HAL layer (hardware abstraction layer) using the human voice extraction method proposed in any of the aforementioned embodiments.

[0198] Meanwhile, the smart cockpit main controller runs the in-vehicle KTV APP, which can control the in-vehicle KTV APP to communicate with the cloud platform through the communication module. It can also send the extracted human voice audio data to Audio Flinger (audio stream management) for mixing with the song accompaniment track to obtain karaoke audio data. The karaoke audio data is then transmitted back to the audio digital signal processor, and after digital-to-analog conversion, it is output to the external speaker through the audio amplifier.

[0199] like Figure 9 As shown, Figure 9 This is a software system block diagram illustrating a microphone-free in-vehicle KTV system according to an embodiment of the present invention. Figure 9 In the process, after the MIC microphone collects the mixed audio signal in the vehicle, the signal processing module amplifies the audio signal gain and performs noise reduction. The high frequencies above 20kHz are filtered by a low-pass filter before reaching the audio digital signal processor. The audio digital signal processor can further reduce noise and process the mixed audio signal in the vehicle. Then, the processed mixed audio signal in the vehicle is transmitted to the HAL layer of the smart cockpit main controller. Specifically, it can be transmitted to the HAL layer through the audio input stream and audio input in the audio hardware device, and the human voice of the mixed audio in the vehicle is extracted in the Audio APP of the HAL layer audio splitting. Specifically, the human voice of the MIC is extracted through the human voice extraction method proposed in any of the above embodiments, and then the extracted human voice audio is transmitted to the audio stream management through the audio device input descriptor.

[0200] Furthermore, the cloud platform communicates with the KTV APP. When a user sings karaoke, the KTV APP sequentially transmits the song's accompaniment track to the audio stream management system via the audio track (application layer), system multimedia audio track, audio track (framework layer), and audio strategy service. In this way, the audio stream management system can mix the vocal audio and the song's accompaniment track to obtain karaoke audio data. It then transmits the karaoke audio data to the audio digital signal processor via the audio device output descriptor, audio output stream, and audio output stream. After receiving the karaoke audio data, the audio digital signal processor can make relevant adjustments to the karaoke audio data through frequency response and sound effects processing, and then output the adjusted karaoke audio data through the AMP (amplifier) ​​and speakers.

[0201] like Figure 10 As shown, Figure 10 This is a flowchart illustrating a multi-zone microphone-free KTV method according to an embodiment of the present invention. Figure 10In the process, the microphone-free in-car KTV system is powered on, and the in-car KTV APP is launched. Users select songs through the KTV APP. The system determines whether the song's audio source is local or cloud-based. If it's local, it retrieves the audio from the local source; if it's cloud-based, it retrieves the audio from the cloud and begins playback. The system continuously monitors whether the music has finished playing. If it has, it returns to the song selection process; otherwise, it activates the KTV function. Additionally, it checks whether the KTV function has been activated; if not, it returns to the music playback check.

[0202] If the KTV function is activated, according to relevant national traffic laws and regulations, the driver and passenger seats can use the KTV function when the vehicle is parked. While the vehicle is in motion, it must adhere to traffic rules. This means that after activating the KTV function, it needs to determine if the driver's vocal range meets the traffic rules (other vocal ranges do not require this check). If the traffic rules are not met, the KTV app will exit directly; if the traffic rules are met, the vocal and instrumental tracks will be separated to obtain the instrumental track.

[0203] Simultaneously, the system automatically detects microphone zones and is compatible with handheld microphones. If no handheld microphone is detected, the multi-zone microphone-free karaoke function is automatically activated, enabling in-vehicle karaoke via the vehicle's built-in voice microphone. During song playback, in-vehicle voice recognition commands are not supported. After the current song ends, the multi-zone microphone voice extraction function automatically exits, releasing the microphone zone from its occupancy, and then supports in-vehicle voice recognition commands. N microphone zones simultaneously acquire mixed audio data, and the voice extraction method provided in any of the above embodiments is used to extract voices from the mixed audio data, outputting voice audio data, and performing noise reduction and sound effect processing on the voice audio data. The processed vocal audio data and song accompaniment track are then mixed using an in-ear monitor switch, echo cancellation, reverb adjustment, and delay processing to obtain karaoke audio data. This karaoke audio data is then used to play the karaoke song. If the karaoke song is not finished playing, the process returns to the step of checking if the driver's seat meets the driving rules, and continues with mixing, vocal extraction, accompaniment acquisition, accompaniment and vocal mixing, and song playback until the karaoke song finishes playing and the KTV APP is exited.

[0204] The microphone-free in-vehicle KTV method and system proposed in this invention can be understood as consisting of an in-vehicle microphone installation location, a hardware system, a software system, an end-to-cloud integrated multi-zone microphone-free KTV method, and a multi-zone microphone voice extraction method. The in-vehicle microphone extracts the voice of the microphone in the sound zone in real time, performs low-latency processing, eliminates acoustic feedback, suppresses environmental noise, mixes and reverberates, and enhances the voice sound effects, thereby realizing a smart cockpit karaoke experience.

[0205] Based on the same inventive concept, one embodiment of the present invention provides a microphone-free in-vehicle karaoke method. (Reference) Figure 11 , Figure 11 This is a flowchart illustrating the steps of a microphone-free in-vehicle karaoke method according to an embodiment of the present invention. The method is applied to a microphone-free in-vehicle karaoke system, which includes at least: an in-vehicle audio system; the in-vehicle audio system includes at least: one or more microphones, an audio digital signal processor, a smart cockpit controller, an audio stream management module, an in-vehicle karaoke client, an amplifier module, and speakers; the method includes the following steps:

[0206] Step S11-1: Collect in-vehicle mixed audio data through the MIC microphone and send the in-vehicle mixed audio data to the audio digital signal processor;

[0207] Step S11-2: Collect the in-vehicle mixed audio data through the audio digital signal processor, and send the in-vehicle mixed audio data to the smart cockpit main controller;

[0208] Step S11-3: Through the intelligent cockpit main controller, the human voice audio data in the mixed audio data in the vehicle is extracted in the hardware abstraction layer using the human voice extraction method described in any of the above embodiments, and the human voice audio data is sent to the audio stream management module;

[0209] Step S11-4: The audio stream management module mixes the human voice audio data with the song accompaniment track sent by the in-vehicle karaoke client to obtain karaoke audio data and sends it to the audio digital signal processor.

[0210] Step S11-5: The karaoke audio data is output through the amplifier module and the speaker by the audio digital signal processor.

[0211] Optionally, the in-vehicle audio system further includes: a signal processing module, a low-pass filter, and an analog-to-digital converter / A2B bus; the signal processing module, the low-pass filter, and the analog-to-digital converter / A2B bus are sequentially located between the MIC microphone and the audio digital signal processor; in this method, the step S11-1 above, "sending the in-vehicle mixed audio data to the audio digital signal processor," may specifically include:

[0212] After the MIC microphone acquires the mixed audio data in the vehicle, the signal processing module amplifies and reduces the noise of the mixed audio data in the vehicle to obtain the first mixed audio data in the vehicle and sends it to the low-pass filter.

[0213] The high-frequency signal in the first in-vehicle mixed audio data is removed by the low-pass filter to obtain the second in-vehicle mixed audio data, and the second in-vehicle mixed audio data is sent to the audio digital signal processor through the analog-to-digital converter module / A2B bus.

[0214] Optionally, step S11-5 above may specifically include:

[0215] The audio digital signal processor processes the karaoke audio data through sound effects processing, earphone monitoring, echo cancellation, reverb adjustment, and delay processing. The processed karaoke audio data is then output through the amplifier module and the speaker.

[0216] Optionally, the system further includes: a cloud platform; the in-vehicle audio system further includes: a communication module; the cloud platform includes at least: an OTA module, a function subscription module, an audio source management module, a membership management module, a payment system, a commission system, and a push system; the method further includes:

[0217] Through the cloud platform, the communication module communicates with the in-vehicle karaoke client, and the OTA module, function subscription module, audio source management module, membership management module, payment system, commission system, and push system respectively realize the functions of OTA update, function subscription, third-party audio source management, membership management, payment management, third-party commission, and information push.

[0218] As the method embodiments are basically similar to the system embodiments, the description is relatively simple, and relevant parts can be found in the description of the system embodiments.

[0219] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0220] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0221] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0222] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0223] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0224] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0225] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0226] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0227] The present invention provides a detailed description of a human voice extraction method, apparatus, product, microphone-free in-vehicle KTV system and method. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for extracting human voice, characterized in that, The method includes: The mixed audio data is segmented to obtain multiple audio segments; Each audio data segment is subjected to feature processing to obtain a first processing result corresponding to each audio data segment; Each of the first processing results is segmented to obtain multiple sub-processing results corresponding to each of the first processing results; Each sub-processing result is subjected to feature processing to obtain a second processing result corresponding to each sub-processing result. The first processing results and the second processing results are superimposed and fused to obtain the human voice audio data in the mixed audio data; The step of performing feature processing on each audio data segment to obtain a first processing result corresponding to each audio data segment includes: Human voice features are extracted from each audio data segment to obtain a three-dimensional human voice tensor for each audio data segment. Perform feature transformation on the human voice three-dimensional tensor of each audio data segment to obtain a first three-dimensional tensor with the same shape corresponding to each audio data segment; The first three-dimensional tensor is used as the first processing result; The step of performing feature processing on each processing sub-result to obtain a second processing result corresponding to each processing sub-result includes: Human voice features are extracted from each of the multiple processing sub-results to obtain the human voice three-dimensional tensor of each processing sub-result. The human voice three-dimensional tensor of each processing sub-result is subjected to feature transformation to obtain a second three-dimensional tensor with the same shape corresponding to each processing sub-result. The second three-dimensional tensor is used as the second processing result.

2. The human voice extraction method according to claim 1, characterized in that, The method further includes: Based on the segmented time step of each audio data segment, multiple first three-dimensional tensors are encapsulated to obtain a first concatenated three-dimensional tensor; Based on the segmented time step of each processing sub-result, multiple second three-dimensional tensors of each first processing result are encapsulated to obtain a second spliced ​​three-dimensional tensor corresponding to each first processing result. The step of superimposing and fusing multiple first processing results and multiple second processing results to obtain human voice audio data in the mixed audio data includes: Based on the time dependency between the second stitched 3D tensor and the first stitched 3D tensor, the first stitched 3D tensor and multiple second stitched 3D tensors are superimposed and stitched together to obtain the human voice audio data.

3. The human voice extraction method according to claim 1, characterized in that, The process of segmenting the mixed audio data to obtain multiple audio segments includes: The mixed audio data is randomly divided into a preset number of segments to obtain multiple audio data segments with and / or without overlap. The step of segmenting each of the first processing results to obtain multiple processing sub-results corresponding to each of the first processing results includes: Each of the first processing results is randomly divided into the preset number of segments to obtain multiple processing sub-results corresponding to the first processing results that have overlaps and / or do not overlap.

4. A human voice extraction device, characterized in that, The device includes: The first segmentation module is used to segment the mixed audio data to obtain multiple audio segments; The first processing module is used to perform feature processing on each audio data segment to obtain a first processing result corresponding to each audio data segment. The second segmentation module is used to segment each of the first processing results to obtain multiple processing sub-results corresponding to each of the first processing results. The second processing module is used to perform feature processing on each processing sub-result to obtain the second processing result corresponding to each processing sub-result. The fusion module is used to superimpose and fuse multiple first processing results and multiple second processing results to obtain human voice audio data in the mixed audio data; The first processing module includes: The first determining submodule is used to extract human voice features from each audio data segment to obtain the human voice three-dimensional tensor of each audio data segment. The first conversion submodule is used to perform feature conversion on the human voice three-dimensional tensor of each audio data segment to obtain a first three-dimensional tensor with the same shape corresponding to each audio data segment. The second determining submodule is used to take the first three-dimensional tensor as the first processing result; The second processing module includes: The third determining submodule is used to extract human voice features from each of the multiple processing sub-results to obtain the human voice three-dimensional tensor of each processing sub-result. The second conversion submodule is used to perform feature conversion on the human voice three-dimensional tensor of each processing sub-result to obtain a second three-dimensional tensor with the same shape corresponding to each processing sub-result. The fourth determination submodule is used to take the second three-dimensional tensor as the second processing result.

5. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the human voice extraction method as described in any one of claims 1 to 3.

6. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the human voice extraction method as described in any one of claims 1 to 3.

7. A microphone-free in-vehicle KTV system, characterized in that, The system includes at least: a vehicle audio system; the vehicle audio system includes at least: one or more microphones, an audio digital signal processor, a smart cockpit controller, an audio stream management module, a vehicle karaoke client, an amplifier module, and speakers; The MIC microphone is used to collect mixed audio data in the vehicle and send the mixed audio data in the vehicle to the audio digital signal processor; The audio digital signal processor is used to collect the in-vehicle mixed audio data and send the in-vehicle mixed audio data to the smart cockpit main controller; The intelligent cockpit main controller is used to extract human voice audio data from the in-vehicle mixed audio data in the hardware abstraction layer using the human voice extraction method as described in any one of claims 1 to 3, and send the human voice audio data to the audio stream management module; The audio stream management module is used to mix the human voice audio data with the song accompaniment track sent by the in-vehicle karaoke client to obtain karaoke audio data and send it to the audio digital signal processor. The audio digital signal processor is also used to output the karaoke audio data through the power amplifier module and the speaker.

8. The microphone-free in-vehicle KTV system according to claim 7, characterized in that, The vehicle audio system further includes: a signal processing module, a low-pass filter, and an analog-to-digital converter / A2B bus; between the MIC microphone and the audio digital signal processor, the signal processing module, the low-pass filter, and the analog-to-digital converter / A2B bus are sequentially located. The signal processing module is used to amplify and reduce noise in the in-vehicle mixed audio data after the MIC microphone acquires the mixed audio data, to obtain first in-vehicle mixed audio data and send it to the low-pass filter. The low-pass filter is used to remove high-frequency signals from the first in-vehicle mixed audio data to obtain the second in-vehicle mixed audio data, and the second in-vehicle mixed audio data is sent to the audio digital signal processor through the analog-to-digital converter module / A2B bus.

9. The microphone-free in-vehicle KTV system according to claim 7, characterized in that, The audio digital signal processor is also used to process the karaoke audio data by performing sound effects processing, earphone monitoring, echo cancellation, reverb adjustment, and delay processing, and then output the processed karaoke audio data through the power amplifier module and the speaker.

10. The microphone-free in-vehicle KTV system according to claim 7, characterized in that, The system also includes: a cloud platform; the vehicle audio system also includes: a communication module; the cloud platform includes at least: an OTA module, a function subscription module, an audio source management module, a membership management module, a payment system, a commission system, and a push system; The cloud platform is used to communicate with the in-vehicle karaoke client through the communication module, and realizes the functions of OTA update, function subscription, third-party audio source management, membership management, payment management, third-party commission, and information push through the OTA module, the function subscription module, the audio source management module, the membership management module, the payment system, the commission system, and the push system.

11. A method for creating a microphone-free in-vehicle karaoke system, characterized in that, The method is applied to a microphone-free in-vehicle KTV system, which includes at least: an in-vehicle audio system; the in-vehicle audio system includes at least: one or more microphones, an audio digital signal processor, a smart cockpit controller, an audio stream management module, an in-vehicle karaoke client, an amplifier module, and speakers; the method includes: The in-vehicle mixed audio data is collected by the MIC microphone and sent to the audio digital signal processor. The in-vehicle mixed audio data is collected by the audio digital signal processor and sent to the smart cockpit main controller. The intelligent cockpit main controller extracts human voice audio data from the in-vehicle mixed audio data in the hardware abstraction layer using the human voice extraction method as described in any one of claims 1 to 3, and sends the human voice audio data to the audio stream management module. The audio stream management module mixes the human voice audio data with the song accompaniment track sent by the in-vehicle karaoke client to obtain karaoke audio data, which is then sent to the audio digital signal processor. The audio digital signal processor transmits the karaoke audio data through the amplifier module and outputs it via the speaker.