A sound collection and reconstruction method, device and vehicle-mounted singing system

By identifying the user in the karaoke system and performing noise reduction processing, personalized audio data is generated and played, solving the problem of noise impact in the vehicle environment and improving the sound quality and user experience.

CN118230700BActive Publication Date: 2025-10-03SHENZHEN WANSHENG CULTURE TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410220082.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-28
Publication Date
2025-10-03
Estimated Expiration
2044-02-28

AI Technical Summary

Technical Problem

In karaoke systems, especially in car environments, the user's singing voice is easily affected by noise, resulting in poor sound collection, affecting the sound quality and reducing the user experience.

Method used

By collecting the user's first audio data, identifying the user's identity using a pre-established timbre library, and performing noise reduction processing based on the accompaniment melody, the second audio data is generated and played, including adjusting the frequency and amplitude of the timbre data sample, superimposing and smoothing to improve the sound quality.

Benefits of technology

Effectively improve the quality and clarity of users' singing voices, provide personalized sound experience, reduce noise interference, and enhance the user experience in the car environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118230700B_ABST
    Figure CN118230700B_ABST
Patent Text Reader

Abstract

The present invention provides a sound collection and reconstruction method, device, and in-vehicle singing system. The method includes collecting first audio data of a user singing, and identifying the user's identity based on the first audio data and a pre-established timbre library. When the user's identity is identified, the first audio data is subjected to noise reduction processing based on the played accompaniment melody to obtain processed data, and second audio data is generated and played based on the processed data and the timbre data sample corresponding to the user in the timbre library. The present invention can effectively distinguish the sound characteristics of different users to achieve identity recognition. The user can complete identity authentication without performing additional operations, thereby generating audio content adapted to the user based on the user's identity characteristics, reducing interference caused by noise, and improving the user's experience. The present invention can be applied to a variety of scenarios such as in-vehicle environments and indoor environments, and has wide applicability and scalability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of audio processing, and in particular relates to a sound collection and reconstruction method, device and a car-mounted singing system. Background Art

[0002] A karaoke system usually includes a screen, audio equipment, and a microphone. Users can select songs from a song library and sing along with the lyrics displayed on the screen. This provides users with a way to interact with music, allowing them to fully enjoy the fun of singing and facilitating their entertainment and pastime.

[0003] However, when there is a lot of noise, the user's singing voice may be affected by the noise, causing the microphone to pick up poor sound, resulting in poor sound collection effect, which is very likely to affect the sound quality and reduce the user experience.

[0004] Currently, karaoke systems can also be used in in-vehicle environments. However, in in-vehicle scenarios, due to limited space and driving conditions, it is often inconvenient for users to use traditional handheld microphones to sing, which will also make the sound collection effect worse, further amplifying the problem of affected sound quality. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, the present invention proposes a sound collection and reconstruction method, which includes:

[0006] Collecting first audio data of a user performing a singing performance, and identifying the user based on the first audio data and a pre-established timbre library;

[0007] If the identity of the user is recognized, performing noise reduction processing on the first audio data based on the played accompaniment melody to obtain processed data;

[0008] Based on the processed data and the timbre data samples corresponding to the user in the timbre library, second audio data is generated and played.

[0009] Specifically, the first audio data includes accompaniment audio data and audio data of a user singing, and the “performing noise reduction processing on the first audio data based on the played accompaniment melody to obtain processed data” includes:

[0010] determining the frequency of the accompaniment audio data;

[0011] Noise reduction is performed by removing audio data in the first audio data whose real-time frequency difference from the accompaniment audio data reaches a preset threshold, so as to obtain the processed data.

[0012] Specifically, the timbre database records timbre data samples of multiple different users, and the collecting of the first audio data of the user further includes:

[0013] In response to a recording instruction, recording an audio clip of the user singing in a preset quiet environment;

[0014] The audio clip is subjected to data processing, features for identifying the user's timbre characteristics are extracted from the audio data, and the features are associated with the user's identity information, thereby establishing a timbre data sample corresponding to the user in the timbre library.

[0015] Furthermore, the first audio data includes the user's singing audio data,

[0016] The generating and playing second audio data based on the processed data and the timbre data sample corresponding to the user in the timbre library includes:

[0017] generating second audio data by using the remaining singing audio data in the processed data and the timbre data sample corresponding to the user in the timbre library;

[0018] The second audio data is played.

[0019] Specifically, the processed data includes the user's singing audio data after noise reduction processing; and "generating and playing the second audio data based on the processed data and the timbre data sample corresponding to the user in the timbre library" includes:

[0020] Adjusting the frequency and amplitude of the timbre data sample corresponding to the user in the timbre library to match the singing audio data;

[0021] Superimposing the adjusted timbre data sample and the singing audio data to obtain superimposed data;

[0022] The superimposed data is smoothed to generate and play the second audio data.

[0023] Preferably, the step of “superimposing the adjusted timbre data sample with the singing audio data to obtain superimposed data” includes:

[0024] By optimizing the network parameters based on the mixture density network model that maximizes the likelihood function, the adjusted timbre data samples and the singing audio data are mapped to the parameters of multiple Gaussian distributions, and then data sampling is performed according to the weights in the distribution based on the distribution parameters of the adjusted timbre data samples and the singing audio data to obtain the superimposed data.

[0025] Optionally, the smoothing comprises applying a low-pass filter to reduce high frequency noise and / or applying a perturbation.

[0026] The present invention also proposes a sound collection and reconstruction device, which includes:

[0027] A collection module, configured to collect first audio data of a user singing, and identify the user based on the first audio data and a pre-established timbre library;

[0028] a processing module, configured to, when the identity of the user is identified, perform noise reduction processing on the first audio data based on the played accompaniment melody to obtain processed data;

[0029] A reconstruction module is used to generate and play second audio data based on the processed data and the timbre data sample corresponding to the user in the timbre library.

[0030] Specifically, the acquisition module includes a multi-microphone array composed of multiple microphones.

[0031] The present invention also proposes a vehicle-mounted singing accompaniment system, which applies the sound collection and reconstruction method as described above.

[0032] The present invention has at least the following beneficial effects:

[0033] The method proposed in the present invention can process and identify the user's first audio data to obtain processed data, and generate and play second audio data, effectively improving the quality and clarity of the user's singing voice. It utilizes a pre-established timbre library to identify the user's identity, ensuring that the user's voice can be automatically identified and processed, thereby providing a personalized sound experience. The method can generate a playback effect suitable for the in-vehicle environment by collecting and processing audio data in a vehicle environment, providing users with a better user experience.

[0034] Furthermore, the present invention can process the accompaniment audio data by removing the audio data whose frequency does not match the accompaniment audio data, thereby improving the quality and clarity of the singing voice and the sound quality effect. The present invention can establish a timbre data sample corresponding to the user based on the user's timbre characteristics and identity information, more accurately identify the user, and ensure the accuracy and personalization of the timbre processing.

[0035] Specifically, this solution can effectively reduce the impact of noise and interference, improve sound quality, better preserve the user's original timbre characteristics, and improve the continuity and stability of the audio by adjusting the frequency and amplitude of the corresponding timbre data samples in the database, superimposing the user's current singing audio data and applying a low-pass filter for smoothing. Combined with deep learning algorithms, data mapping and sampling are performed through a hybrid density network model to ensure the accuracy and naturalness of timbre processing.

[0036] Therefore, the present invention provides a sound collection and reconstruction method, device and in-vehicle karaoke system. The present invention can effectively distinguish the sound characteristics between different users to achieve identity recognition. The user can complete identity authentication without performing additional operations, thereby generating audio content adapted to the user based on the user's identity characteristics, reducing interference caused by noise, and improving the user's experience. It can be applied to various scenarios such as in-vehicle environments and indoor environments, and has wide applicability and scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0038] Figure 1 A schematic diagram of the overall process of the sound collection and reconstruction method provided in Example 1;

[0039] Figure 2 A flow chart of a method for recording user timbre data samples for a timbre library;

[0040] Figure 3 A flowchart of a method for generating and playing second audio data;

[0041] Figure 4 This is a schematic diagram of the overall process of the sound collection and reconstruction device provided in Example 2.

[0042] Reference numerals

[0043] 10 - acquisition module; 20 - processing module; 30 - reconstruction module; 31 - adjustment unit; 32 - superposition unit; 33 - playback unit; 41 - recording module; 42 - storage module. DETAILED DESCRIPTION

[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0045] Hereinafter, various embodiments of the present invention will be described more fully. The present invention can have various embodiments, and modifications and variations can be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present invention to the specific embodiments disclosed herein, but rather that the present invention should be construed to encompass all modifications, equivalents, and / or alternatives falling within the spirit and scope of the various embodiments of the present invention.

[0046] Hereinafter, the terms "include" or "may include" used in various embodiments of the present invention indicate the presence of disclosed functions, operations, or elements, and do not limit the addition of one or more functions, operations, or elements. In addition, as used in various embodiments of the present invention, the terms "include", "have" and their cognates are intended only to indicate specific features, numbers, steps, operations, elements, components, or combinations of the foregoing, and should not be understood as excluding the presence of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing or the possibility of adding one or more features, numbers, steps, operations, elements, components, or combinations of the foregoing.

[0047] In various embodiments of the present invention, the expression "or" or "at least one of A or / and B" includes any or all combinations of the words listed simultaneously. For example, the expression "A or B" or "at least one of A or / and B" may include A, may include B, or may include both A and B.

[0048] The expressions (such as "first", "second", etc.) used in the various embodiments of the present invention may modify the various constituent elements in the various embodiments, but may not limit the corresponding constituent elements. For example, the above expressions do not limit the order and / or importance of the elements. The above expressions are only used to distinguish one element from other elements. For example, a first user device and a second user device indicate different user devices, although both are user devices. For example, without departing from the scope of the various embodiments of the present invention, a first element may be referred to as a second element, and similarly, a second element may also be referred to as a first element.

[0049] It should be noted that, in the present invention, unless otherwise expressly specified or defined, terms such as "mounted," "connected," and "fixed" should be understood broadly. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediary; and internal communication between two components. Those skilled in the art will understand the specific meanings of these terms in the present invention based on specific circumstances.

[0050] In the present invention, those skilled in the art need to understand that the terms indicating orientation or positional relationships herein are based on the orientation or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.

[0051] The terms used in various embodiments of the present invention are only used to describe the purpose of specific embodiments and are not intended to limit the various embodiments of the present invention. As used herein, the singular form is intended to also include the plural form, unless the context clearly indicates otherwise. Unless otherwise limited, all terms used here (including technical terms and scientific terms) have the same meaning as those of ordinary skill in the art generally understood by the various embodiments of the present invention. The terms (such as those defined in generally used dictionaries) will be interpreted as having the same meaning as the contextual meaning in the relevant technical field and will not be interpreted as having idealized meaning or too formal meaning, unless clearly defined in various embodiments of the present invention.

[0052] Example 1

[0053] This embodiment proposes a sound collection and reconstruction method, which is applied to a singing system in a car or indoor scene. The singing system can play music accompaniment and collect the user's voice and output it to provide the user with a karaoke-like effect in a private space. Figure 1 , the method comprising:

[0054] S100: collecting first audio data of a user singing, and identifying the user based on the first audio data and a pre-established timbre library.

[0055] In this embodiment, the singing accompaniment system includes a multi-microphone array composed of multiple microphones, and the user can collect audio data without holding the microphone. Step S100 uses the multi-microphone array to collect the first audio data of the user singing. The first audio data includes accompaniment audio data, the user's singing audio data and environmental noise. The singing accompaniment system can identify the user's identity based on the first audio data and the timbre library.

[0056] Specifically, this embodiment uses a convolutional neural network (CNN) to extract and classify features of the user's first audio data, thereby realizing automatic recognition of the user's identity. The convolutional neural network can be trained through a back-propagation algorithm to optimize network weights and biases, thereby improving classification accuracy.

[0057] S200: If the identity of the user is recognized, noise reduction processing is performed on the first audio data based on the played accompaniment melody to obtain processed data.

[0058] In a specific embodiment, step S200 determines the frequency of the accompaniment audio data and performs noise reduction by removing the audio data in the first audio data whose difference from the real-time frequency of the accompaniment audio data reaches a preset threshold, and finally obtains processed data, that is, selects the sound within the sound frequency that matches the accompaniment audio data to obtain the processed data. If the accompaniment system is set in a car scene, this method can filter out most of the ambient noise in the car and most of the human voices other than the user's singing voice.

[0059] S300: Generate and play second audio data based on the processed data and the timbre data sample corresponding to the user in the timbre library.

[0060] The singing accompaniment system in this embodiment includes an output device for playing the second audio data, and the output device may include but is not limited to speakers, headphones and other devices.

[0061] In step S300, the second audio data is generated by processing the remaining singing audio data in the data and the timbre data sample corresponding to the user in the timbre library.

[0062] The timbre library of this embodiment can record timbre data samples of multiple different users. If and only if the pre-established timbre library has recorded data corresponding to the user, step S200 can identify the user's identity and then continue to execute subsequent steps. Figure 2 The method for recording user timbre data samples in the timbre library includes:

[0063] S410: In response to a recording instruction, record an audio clip of the user's singing.

[0064] It should be noted that step S410 should be performed in a sufficiently quiet environment to minimize noise interference in the recorded audio segment.

[0065] S420: Based on the recorded audio clip, a timbre data sample corresponding to the user is created in the timbre library.

[0066] In this embodiment, step S420 requires data processing of the recorded audio clip, and then extracting features that can be used to identify the user's timbre characteristics from the audio data, associating the extracted features with the user's identity information, and storing them in the timbre library;

[0067] Furthermore, step S420 may perform standardization processing on the extracted feature vectors so that the timbre data of different users are comparable in the feature space; the standardization processing may include but is not limited to steps such as de-meaning, normalization and dimensionality reduction.

[0068] Preferably, in this embodiment, a Gaussian mixture model (GMM), a support vector machine (SVM) and / or a neural network are used to model the timbre data. By labeling the timbre data samples and using the labeled timbre data samples for model training, the recognition performance of the convolutional neural network (CNN) on unknown audio data is evaluated, thereby achieving high-precision recognition of the user's identity.

[0069] Preferably, step S300 is achieved by superimposing the timbre library on the singing audio data to reconstruct the sound, and the processed data includes the user's singing audio data after noise reduction processing, see Figure 3 , step S300 specifically includes:

[0070] S310: Adjust the frequency and amplitude of the timbre data sample corresponding to the user in the timbre library to match the singing audio data.

[0071] It should be noted that the singing audio data can be specifically divided into multiple data such as volume, audio melody and pitch characteristics. The timbre data sample adjusted in step S310 can be matched with each data in the singing audio data.

[0072] In this embodiment, step S310 can use a recurrent neural network model (RNN) to match the user's processed data with the data corresponding to the user in the timbre library. The recurrent neural network can process sequence data and is suitable for processing the timing characteristics of audio signals, thereby obtaining an adjusted timbre data sample.

[0073] Specifically, the recurrent neural network model used in step S310 may include but is not limited to a long short-term memory network recurrent neural network model (LSTM) and a gated recurrent unit recurrent neural network model (GRU).

[0074] S320: Superimposing the adjusted timbre data sample and the singing audio data to obtain superimposed data.

[0075] In this embodiment, step S320 can be implemented using a mixture density network model (MDN), which models the adjusted timbre data samples and singing audio data and generates an output distribution. Specifically, the mixture density network model can map the adjusted timbre data samples and singing audio data to parameters of multiple Gaussian distributions, and then sample the data according to the weights in the distribution to obtain superimposed data.

[0076] It should be noted that the mixture density network model is a neural network model used to model multimodal distributions. In superposition data generation, the mixture density network model can map the adjusted timbre data samples and processed data to the parameters of multiple Gaussian distributions. The parameters are used to describe the characteristics of each Gaussian distribution, which may include but are not limited to the mean, variance, and weight.

[0077] During the training process, the mixture density network model optimizes the network parameters by maximizing the likelihood function, so that the model can more accurately reflect the distribution of the input data. After the training is completed, the mixture density network model can sample the output data according to the weights in these distributions based on the distribution parameters of the adjusted timbre data samples and the singing audio data, thereby generating expected superposition data.

[0078] S330: Perform smoothing processing on the superimposed data to generate and play second audio data.

[0079] It should be noted that the smoothing process includes applying a low-pass filter to reduce high-frequency noise and / or applying perturbations, and through the smoothing process of the superimposed data in step S330, the accompaniment system can make the processed second audio data simulate the natural changes of the human voice.

[0080] Through the sound collection and reconstruction method proposed in this embodiment, the accompaniment system can be applied to the vehicle environment. By collecting and processing audio data, it can generate a playback effect suitable for the vehicle environment, providing users with a better user experience.

[0081] Example 2

[0082] This embodiment proposes a sound collection and reconstruction device for implementing the method proposed in Example 1, see Figure 4 , the device comprises:

[0083] The acquisition module 10 is used to collect first audio data of the user singing and identify the user based on the first audio data and a pre-established timbre library;

[0084] a processing module 20 for performing noise reduction processing on the first audio data based on the played accompaniment melody to obtain processed data when the user's identity is recognized;

[0085] The reconstruction module 30 is configured to generate and play second audio data based on the processed data and the timbre data samples corresponding to the user in the timbre library.

[0086] In this embodiment, the acquisition module 10 includes a multi-microphone array composed of multiple microphones. The acquisition module 10 can use the multi-microphone array to collect first audio data when the user sings. The first audio data includes accompaniment audio data, user's singing audio data and environmental noise. The acquisition module 10 can identify the user's identity based on the first audio data and the timbre library;

[0087] The reconstruction module 30 includes an output device for playing the second audio data, and the output device may include but is not limited to speakers, headphones and other devices; the second audio data is generated by processing the remaining singing audio data in the data and the timbre data samples corresponding to the user in the timbre library.

[0088] In a specific embodiment, the processing module 20 determines the frequency of the accompaniment audio data and performs noise reduction by removing the audio data in the first audio data whose difference from the real-time frequency of the accompaniment audio data reaches a preset threshold, and finally obtains the processed data, that is, selects the sound within the sound frequency that matches the accompaniment audio data to obtain the processed data.

[0089] Specifically, the timbre library in this embodiment can record timbre data samples of multiple different users. If and only if the pre-established timbre library has recorded data corresponding to the user, the processing module 20 can identify the user's identity and then continue to perform subsequent steps. The device further includes:

[0090] The recording module 41 is used to record the audio clip of the user's singing in response to the recording instruction;

[0091] The storage module 42 is configured to create a timbre data sample corresponding to the user in the timbre library based on the recorded audio clip.

[0092] It should be noted that when the recording module 41 performs the preset steps, it should be in a sufficiently quiet environment to minimize noise interference in the recorded audio clip. The storage module 42 needs to process the recorded audio clip, and then extract features that can be used to identify the user's timbre characteristics from the audio data, associate the extracted features with the user's identity information, and store them in the timbre library;

[0093] Furthermore, the storage module 42 may perform standardization processing on the extracted feature vectors so that the timbre data of different users are comparable in the feature space; the standardization processing may include but is not limited to steps such as de-meaning, normalization and dimensionality reduction.

[0094] Preferably, in this embodiment, the storage module 42 uses a Gaussian mixture model (GMM), a support vector machine (SVM) and / or a neural network to model the timbre data, and evaluates the recognition performance of the convolutional neural network (CNN) on unknown audio data by labeling the timbre data samples and using the labeled timbre data samples for model training, thereby achieving high-precision recognition of the user identity.

[0095] Preferably, the reconstruction module 30 implements the preset steps by superimposing the timbre library on the singing audio data to reconstruct the sound. The processed data includes the user's singing audio data after noise reduction processing. The reconstruction module 30 specifically includes:

[0096] An adjustment unit 31 is used to adjust the frequency and amplitude of the timbre data sample corresponding to the user in the timbre library to match the singing audio data;

[0097] a superposition unit 32 for superimposing the adjusted timbre data sample with the singing audio data to obtain superimposed data;

[0098] The playing unit 33 is configured to perform smoothing processing on the superimposed data to generate and play the second audio data.

[0099] It should be noted that the singing audio data can be specifically divided into multiple data such as volume, audio melody and pitch characteristics, and the timbre data samples adjusted by the adjustment unit 31 can be matched with the various data in the singing audio data.

[0100] In this embodiment, the adjustment unit 31 may include a recurrent neural network model (RNN), which can match the user's processed data with the data corresponding to the user in the timbre library. The recurrent neural network can process sequence data and is suitable for processing the timing characteristics of audio signals, thereby obtaining adjusted timbre data samples.

[0101] Specifically, the recurrent neural network model may include but is not limited to a long short-term memory network recurrent neural network model (LSTM) and a gated recurrent unit recurrent neural network model (GRU).

[0102] In this embodiment, the superposition unit 32 may include a mixture density network model (MDN), which models the adjusted timbre data samples and singing audio data and generates an output distribution. Specifically, the mixture density network model can map the adjusted timbre data samples and singing audio data to parameters of multiple Gaussian distributions, and then sample the data according to the weights in the distribution to obtain superimposed data.

[0103] Smoothing processing may include applying a low-pass filter to reduce high-frequency noise and / or applying perturbations, etc. Through the smoothing processing of the superimposed data by the playback unit 33, the accompaniment system can make the processed second audio data simulate the natural changes of the human voice.

[0104] Example 3

[0105] This embodiment proposes an in-vehicle karaoke system, which applies the sound collection and reconstruction method proposed in Example 1. The in-vehicle karaoke system can be set in a vehicle or other vehicle.

[0106] In summary, the present invention provides a sound collection and reconstruction method, device and in-vehicle karaoke system. The present invention can effectively distinguish the sound characteristics between different users to achieve identity recognition. The user can complete identity authentication without performing additional operations, thereby generating audio content adapted to the user based on the user's identity characteristics, reducing interference caused by noise, and improving the user's experience. It can be applied to various scenarios such as in-vehicle environments and indoor environments, and has wide applicability and scalability.

[0107] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A sound collection and reconstruction method, characterized in that: The method comprises: Collecting first audio data of a user performing a singing performance, and identifying the user based on the first audio data and a pre-established timbre library; If the identity of the user is recognized, performing noise reduction processing on the first audio data based on the played accompaniment melody to obtain processed data; the processed data includes the user's singing audio data after the noise reduction processing; generating and playing second audio data based on the processed data and a timbre data sample corresponding to the user in the timbre library; The generating and playing second audio data based on the processed data and the timbre data sample corresponding to the user in the timbre library includes: Adjusting the frequency and amplitude of the timbre data sample corresponding to the user in the timbre library to match the singing audio data; Superimposing the adjusted timbre data sample and the singing audio data to obtain superimposed data; performing smoothing processing on the superimposed data to generate and play the second audio data; The step of superimposing the adjusted timbre data sample with the singing audio data to obtain superimposed data comprises: By optimizing the network parameters based on the mixture density network model that maximizes the likelihood function, the adjusted timbre data samples and the singing audio data are mapped to the parameters of multiple Gaussian distributions, and then data sampling is performed according to the weights in the distribution based on the distribution parameters of the adjusted timbre data samples and the singing audio data to obtain the superimposed data.

2. The sound collection and reconstruction method according to claim 1, characterized in that: The first audio data includes accompaniment audio data and audio data sung by a user, and the performing of noise reduction processing on the first audio data based on the played accompaniment melody to obtain processed data includes: determining the frequency of the accompaniment audio data; Noise reduction is performed by removing audio data in the first audio data whose real-time frequency difference from the accompaniment audio data reaches a preset threshold, so as to obtain the processed data.

3. The sound collection and reconstruction method according to claim 1, characterized in that: The timbre library records timbre data samples of multiple different users, and the collecting of the first audio data of the user also includes: In response to a recording instruction, recording an audio clip of the user singing in a preset quiet environment; The audio clip is subjected to data processing, features for identifying the user's timbre characteristics are extracted from the audio data, and the features are associated with the user's identity information, thereby establishing a timbre data sample corresponding to the user in the timbre library.

4. The sound collection and reconstruction method according to claim 1, characterized in that: The smoothing process may include applying a low-pass filter to reduce high frequency noise and / or applying a perturbation.

5. A sound collection and reconstruction device, characterized in that: The device comprises: A collection module, configured to collect first audio data of a user singing, and identify the user based on the first audio data and a pre-established timbre library; a processing module configured to, upon identifying the identity of the user, perform noise reduction processing on the first audio data based on the played accompaniment melody to obtain processed data; the processed data including the user's singing audio data after the noise reduction processing; a reconstruction module, configured to generate and play second audio data based on the processed data and a timbre data sample corresponding to the user in the timbre library; The reconstruction module includes: an adjusting unit, configured to adjust the frequency and amplitude of the timbre data sample corresponding to the user in the timbre library to match the singing audio data; a superposition unit, configured to superimpose the adjusted timbre data sample with the singing audio data to obtain superimposed data; A playing unit, configured to perform smoothing on the superimposed data to generate and play the second audio data, wherein the adjusting timbre data sample is superimposed on the singing audio data to obtain the superimposed data, comprising: By optimizing the network parameters based on the mixture density network model that maximizes the likelihood function, the adjusted timbre data samples and the singing audio data are mapped to the parameters of multiple Gaussian distributions, and then data sampling is performed according to the weights in the distribution based on the distribution parameters of the adjusted timbre data samples and the singing audio data to obtain the superimposed data.

6. The sound collection and reconstruction device according to claim 5, characterized in that: The acquisition module includes a multi-microphone array composed of multiple microphones.

7. A car karaoke system, characterized in that: The method according to any one of claims 1 to 4 is applied.

Citation Information

Patent Citations

  • Audio processing method and device, terminal and storage medium

    CN110956971A

  • Speech synthesis method, device and equipment and storage medium

    CN112382270A

  • Singing speech synthesis method and synthesis device, and computer storage medium

    CN112767914A

  • Audio signal noise reduction method, electronic equipment and storage medium

    CN115083440A