Audio processing method and device, storage medium and electronic equipment
By matching audio gain and voiceprint features to the mixed audio signal, the problem of insufficient accuracy of vocal separation and optimization in the prior art is solved, and high-quality audio separation and optimization effects are achieved.
Patent Information
- Application Number
- CN202510208998.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to accurately separate and optimize the vocals of a specific object from the mixed audio signal, resulting in the vocals that may be blurred in the final mix, have poor sound quality, and have low degree of integration with other audio elements.
By acquiring the first audio data to be processed, an audio gain operation is performed to improve the clarity and expressiveness of the audio signal, and then a voiceprint feature matching is performed on the optimized audio signal to separate the target audio signal from the mixed audio signal of at least two objects.
It realizes the precise separation and optimization of audio signals of specific objects from complex chorus audio, improves the audio performance of the audio signal, ensures the natural fusion and clear expression of the target audio signal in the final mix, and improves the overall quality and user experience of the audio.
Smart Images

Figure CN119993189A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and in particular to an audio processing method and device, a storage medium, and an electronic device. Background Art
[0002] Existing technologies have difficulty accurately separating and optimizing the vocals of a specific object (such as a singer) from mixed audio signals, resulting in the vocals being unclear and of poor quality in the final mix, and not being well integrated with other audio elements (such as accompaniment). This not only affects the overall expressiveness of the audio, but also reduces the user experience.
[0003] Therefore, there is a technical problem in the related art that the accuracy of extracting an audio signal from a mixed audio signal is low. Summary of the invention
[0004] The embodiments of the present application provide an audio processing method and device, a storage medium, and an electronic device to at least solve the technical problem in the related art of low accuracy in extracting an audio signal from a mixed audio signal.
[0005] According to one aspect of an embodiment of the present application, an audio processing method is provided, comprising: obtaining first audio data to be processed, wherein the first audio data comprises a first audio signal of a first audio attribute, and the first audio signal is an audio signal mixed with at least two objects; performing an audio gain operation on the first audio signal to obtain a second audio signal of a second audio attribute, wherein the audio attribute of the audio signal is used to indicate the degree of audio performance of the audio signal in the audio data, and the second audio attribute is higher than the first audio attribute; performing voiceprint feature matching on the second audio signal to obtain a target audio signal after the at least two objects are separated.
[0006] According to another aspect of an embodiment of the present application, an audio processing device is also provided, including: an acquisition unit, used to acquire first audio data to be processed, wherein the first audio data includes a first audio signal of a first audio attribute, and the first audio signal is an audio signal mixed with at least two objects; a gain unit, used to perform an audio gain operation on the first audio signal to obtain a second audio signal of a second audio attribute, wherein the audio attribute of the audio signal is used to indicate the degree of audio performance of the audio signal in the audio data, and the second audio attribute is higher than the first audio attribute; a matching unit, used to perform voiceprint feature matching on the second audio signal to obtain a target audio signal after the at least two objects are separated.
[0007] According to another aspect of the embodiment of the present application, a computer program product is provided, the computer program product comprising a computer program / instruction, the computer instruction being stored in a computer-readable storage medium. A processor of a computer device reads the computer program / instruction from the computer-readable storage medium, and the processor executes the computer program / instruction, so that the computer device performs the above audio processing method.
[0008] According to another aspect of an embodiment of the present application, there is also provided an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the audio processing method through the computer program.
[0009] In this embodiment, the first audio data to be processed is obtained, wherein the audio data is an audio signal mixed by at least two objects, and an audio gain operation is performed to improve the clarity and expressiveness of the audio signal, so as to at least solve the problem of background noise and uneven signal strength during the mixing process. After the audio gain operation, the audio properties (such as volume, sound quality, and pitch) of the audio signal are optimized. Compared with the original audio properties, the optimized audio properties (second audio properties) can better highlight the target audio signal and improve the accuracy and efficiency of subsequent processing. Then, the optimized audio signal is matched with voiceprint features to separate the target audio signal from the mixed audio signal of at least two objects.
[0010] It can be understood that in this embodiment, the audio signal of a specific object can be accurately separated and optimized from the complex chorus audio, the audio performance of the audio signal can be improved, the natural fusion and clear expression of the target audio signal in the final mix can be ensured, the overall audio quality and user experience can be improved, and the technical effect of improving the accuracy of extracting audio from mixed audio can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0012] Figure 1 is a schematic diagram of a flow of an optional audio processing method according to an embodiment of the present application;
[0013] Figure 2 is a schematic diagram of an optional audio processing method according to an embodiment of the present application;
[0014] Figure 3 is a schematic diagram of an optional audio processing method according to an embodiment of the present application;
[0015] Figure 4is a schematic diagram of an optional audio processing device according to an embodiment of the present application;
[0016] Figure 5 A schematic structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0017] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0018] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0019] Optionally, as an optional implementation, as Figure 1 As shown, the audio processing method includes the following specific steps:
[0020] S102, obtaining first audio data to be processed, wherein the first audio data includes a first audio signal of a first audio attribute, and the first audio signal is an audio signal mixed by at least two objects;
[0021] S104, reading first motion effect description information of the video frame, wherein the first motion effect description information is used to indicate a first display effect of the motion effect element, and the first motion effect description information includes first texture information of the video frame;
[0022] S106: Perform voiceprint feature matching on the second audio signal to obtain a target audio signal after at least two objects are separated.
[0023] Optionally, in this embodiment, audio data refers to sound information captured by a recording device, which is stored in the form of digital signals and can be a signal from a single sound source or a mixed signal from multiple sound sources. In this embodiment, audio data specifically refers to a mixed audio signal containing at least two object sounds.
[0024] Optionally, in this embodiment, the audio attribute describes the characteristics of the audio signal, such as volume, pitch, sound quality, and spectrum, etc. In this embodiment, the audio attribute refers to the degree of representation of the audio signal in the audio data, including signal strength, clarity, and sound quality.
[0025] Optionally, in this embodiment, voiceprint features are key parameters in voiceprint recognition, which describe unique characteristics of an individual's voice, such as timbre, pitch, tone, etc. Voiceprint features are used to identify and separate audio signals of specific objects.
[0026] Optionally, in this embodiment, in the initial stage of audio processing, mixed audio data containing the target object and other object sounds are received from the user or read from a file. These data may come from chorus recordings, multi-person conversations, or other multi-sound source environments.
[0027] Optionally, in this embodiment, an audio gain operation is performed on the acquired audio data to improve the clarity and strength of the audio signal so that it is easier to identify and separate in subsequent processing. The gain operation may include volume increase, sound quality improvement, or spectrum adjustment, thereby creating an audio signal with better second audio properties.
[0028] Optionally, in this embodiment, after audio gain, the second audio signal is subjected to complex analysis and matching using the personalized voiceprint features of the target object to identify and separate the audio signal of the target object. This process can accurately extract the voice of the target object from the mixed audio signal, even in a noisy environment.
[0029] It is understandable that this embodiment can be used, but not limited to, to extract and optimize the audio signal of a specific object (such as a target human voice) from an audio signal containing multiple sound sources. First, a segment of audio data containing two or more sound sources is obtained from a user, and these data may come from a chorus recording, a conference recording, or other complex audio scenes. Subsequently, an audio gain operation is performed on the audio signal in the audio data to improve its audio performance, including increasing the volume, improving the sound quality, and adjusting the spectrum, so as to obtain a second audio signal with better audio properties. This operation makes the audio signal of the target object more prominent in subsequent processing, and improves the accuracy of recognition and separation. Finally, the optimized second audio signal can be matched with voiceprint features using a pre-established personalized voiceprint model, so as to separate the target audio signal from the mixed audio containing multiple objects. Through this series of steps, the target audio signal can be accurately extracted and optimized in a complex multi-sound source environment, which is particularly suitable for the extraction and secondary synthesis optimization of specific human voices in chorus audio, improving the clarity, sound quality and expressiveness of the audio signal, ensuring the natural fusion and clear expression of the target audio signal in the final mix, and improving the efficiency and quality of the overall audio processing.
[0030] Through the embodiment provided by the present application, the first audio data to be processed is obtained, wherein the audio data is an audio signal mixed by at least two objects, and an audio gain operation is performed to improve the clarity and expressiveness of the audio signal, so as to at least solve the problem of uneven signal strength during background noise and mixing. After the audio gain operation, the audio properties (such as volume, sound quality, and pitch) of the audio signal are optimized. Compared with the original audio properties, the optimized audio properties (second audio properties) can better highlight the target audio signal and improve the accuracy and efficiency of subsequent processing. Then, the optimized audio signal is matched with voiceprint features to separate the target audio signal from the mixed audio signal of at least two objects. In this embodiment, the audio signal of a specific object can be accurately separated and optimized from the complex chorus audio, the audio performance of the audio signal is improved, the natural fusion and clear expression of the target audio signal in the final mix are ensured, the overall quality of the audio and the user experience are improved, and the technical effect of improving the accuracy of extracting audio from the mixed audio is achieved.
[0031] As an optional solution, performing an audio gain operation on the first audio signal to obtain a second audio signal with a second audio attribute includes at least one of the following:
[0032] Performing a spectral gain operation on the first audio signal, wherein the spectral gain operation is used to enhance the spectral expression of the first audio signal;
[0033] Performing a volume gain operation on the first audio signal, wherein the spectral gain operation is used to enhance the volume performance of the first audio signal;
[0034] Performing a frequency gain operation on the first audio signal, wherein the spectrum gain operation is used to enhance the frequency expression of the first audio signal;
[0035] A sound quality gain operation is performed on the first audio signal, wherein the spectrum gain operation is used to enhance the sound quality performance of the first audio signal.
[0036] Optionally, in this embodiment, the spectral gain operation refers to enhancing specific frequency components of the signal by adjusting the spectral distribution of the audio signal, thereby improving the clarity and expressiveness of the audio signal. In this embodiment, the spectral gain operation is used to highlight the characteristics of the target audio signal from a noisy environment and improve its spectral expression.
[0037] Optionally, in this embodiment, the volume gain operation refers to increasing the volume of the audio signal as a whole without changing the spectral distribution of the audio signal. In this embodiment, the volume gain operation helps to increase the volume performance of the target audio signal in a multi-sound source environment, making it more prominent in the final synthesized audio.
[0038] Optionally, in this embodiment, the frequency gain operation refers to targetedly increasing the intensity of a specific frequency range in the audio signal to optimize the sound quality and clarity of the signal. In this embodiment, the frequency gain operation is used to enhance the frequency expression of the target audio signal to ensure its naturalness and expressiveness when synthesized with the accompaniment track.
[0039] Optionally, in this embodiment, the sound quality gain operation refers to improving the sound quality characteristics of the audio signal, such as timbre, pitch, etc., by filtering, equalization or other audio processing techniques. In this embodiment, the sound quality gain operation is used to improve the sound quality performance of the target audio signal in a multi-sound source environment, making it more pleasant and natural in the final audio.
[0040] Through the embodiments provided in this application, by implementing any one or more of the above audio gain operations, the audio performance of the target audio signal can be effectively improved, ensuring its clarity, naturalness and expressiveness in chorus audio and multi-person recording scenes, thereby achieving the purpose of optimizing audio synthesis quality and improving user experience. The implementation of these gain operations constitutes a key technical point in the audio signal preprocessing stage, laying a solid foundation for the subsequent accurate extraction of the target human voice and secondary synthesis optimization.
[0041] As an optional solution, before obtaining the first audio data to be processed, the method further includes:
[0042] Creating an audio voiceprint model, wherein the audio voiceprint model package is used to indicate audio voiceprint feature information of at least two objects;
[0043] After obtaining the first audio data to be processed, the method further includes:
[0044] Based on the audio voiceprint model, a first audio signal matching the audio voiceprint feature information is determined from the first audio data.
[0045] Optionally, in this embodiment, the audio voiceprint model includes voiceprint feature information of an individual voice, such as timbre, pitch, tone, etc. This model is used to separate and identify a target audio signal from a mixed audio signal.
[0046] Optionally, in this embodiment, the audio voiceprint feature information refers to data describing the sound features of a specific object, including but not limited to timbre, pitch, tone, etc., and is the basis for creating an audio voiceprint model.
[0047] Optionally, in this embodiment, creating an audio voiceprint model is the first step. The model contains audio voiceprint feature information of at least two objects, such as timbre, pitch, tone, etc., which is used for subsequent identification of target audio signals. To further illustrate, before a user participates in a chorus or other multi-source recording activity, the user will be guided to record a personal audio, and an individual audio voiceprint model will be created based on the audio. The innovation lies in that, through the establishment of a personalized voiceprint model, an individualized identification basis is provided for the subsequent processing and optimization of audio signals, thereby enhancing the accuracy of extracting the target audio signal.
[0048] Next, after obtaining the first audio data to be processed, the audio signal in the audio data will be matched based on the created audio voiceprint model to determine the first audio signal that matches the voiceprint feature information, that is, the target audio signal. This process is the key to solving the problem of accurate extraction of target audio signals in a multi-sound source environment. By matching the voiceprint model, the sound of the target object can be identified and separated from the mixed audio signal, ensuring the purity and clarity of the target audio signal even in complex scenes with background noise and multi-person chorus, providing an accurate audio signal source for subsequent audio gain operations and secondary synthesis optimization.
[0049] Through the embodiments provided in this application, it is possible to effectively identify and separate target audio signals in a multi-sound source environment, thereby improving the accuracy and efficiency of audio processing, and providing strong technical support for applications such as chorus recording, audio optimization in multi-person conversation scenarios, and target voice extraction.
[0050] As an optional solution, performing voiceprint feature matching on the second audio signal to obtain a target audio signal after at least two objects are separated includes:
[0051] Extracting a third audio signal corresponding to the preset text from the second audio signal, and performing audio separation on the third audio signal to obtain a plurality of separated audio signals, wherein one separated audio signal corresponds to one object;
[0052] The audio voiceprint feature information is used to perform voiceprint feature matching on multiple separated audio signals to obtain a target audio signal after at least two objects are separated.
[0053] Optionally, in this embodiment, the second audio signal refers to an audio signal after an audio gain operation, and its audio properties (such as volume, sound quality, pitch, etc.) are optimized and improved, so that the characteristics of the target audio signal are more prominent, which facilitates subsequent voiceprint feature matching and audio signal separation.
[0054] Optionally, in this embodiment, the preset text refers to the lyrics or speech content used to guide the system to identify and separate the target audio signal during the audio processing process. These texts are closely related to the target audio signal and are an important basis for audio signal separation and feature matching.
[0055] Optionally, in this embodiment, the third audio signal is a partial audio signal filtered out from the second audio signal according to a preset text, which contains sound information matching the preset text and is the processing object of subsequent audio separation and voiceprint feature matching.
[0056] Optionally, in this embodiment, audio separation refers to distinguishing and separating audio signals of different objects from a mixed audio signal. This process involves technologies such as multi-channel signal processing, spectrum analysis, and deep learning, aiming to create an independent audio signal for each object.
[0057] Optionally, in this embodiment, voiceprint feature matching refers to comparing and matching the separated audio signal with the voiceprint feature information in the audio voiceprint model to identify and separate the target audio signal. It quantitatively analyzes the feature parameters in the audio signal and compares them with the feature information in the model to find the audio signal with the highest matching degree, that is, the target audio signal.
[0058] Through the embodiments provided by this application, not only the personalized recognition capability of the audio voiceprint model is utilized, but also the guiding role of the preset text and the practicality of the audio separation technology are combined, thereby realizing efficient recognition and separation of the target audio signal based on the second audio signal. This technical solution effectively solves the problem of target human voice extraction and optimization in complex audio environments, and significantly improves the accuracy and effect of audio processing.
[0059] As an optional solution, create an audio voiceprint model, including:
[0060] Record initial audio for at least two subjects;
[0061] Performing frequency feature extraction, timbre feature extraction, and rhythm feature extraction on the initial audio respectively, to obtain frequency feature information, timbre feature information, and rhythm feature information of at least two objects;
[0062] The frequency feature information, the timbre feature information and the rhythm feature information are aggregated to obtain audio voiceprint feature information of at least two objects.
[0063] Optionally, in this embodiment, the initial audio refers to unprocessed raw audio data obtained from at least two objects in order to create an audio voiceprint model. Such data generally contains the sound features of the objects, such as frequency, timbre, and rhythm, etc., which are used for subsequent feature extraction and model creation.
[0064] Optionally, in this embodiment, frequency feature extraction refers to extracting characteristic parameters of sound frequency distribution from an audio signal. This process can reveal the pitch and frequency composition of the sound and provide basic data for voiceprint analysis.
[0065] Optionally, in this embodiment, timbre feature extraction refers to extracting timbre features of sound from audio signals, including texture, brightness, fullness, etc. of the sound, which can reflect the uniqueness and individuality of the sound.
[0066] Optionally, in this embodiment, rhythm feature extraction refers to analyzing rhythm characteristics in audio signals, including rhythm, speed and rhythm pattern of pronunciation, etc., which helps to identify differences in speaking or singing rhythms of different objects.
[0067] Optionally, in this embodiment, feature aggregation refers to integrating and summarizing frequency feature information, timbre feature information and rhythm feature information to form a comprehensive audio voiceprint feature information for creating an audio voiceprint model.
[0068] For example, an initial audio is recorded for each subject involved in the audio recording. This operation is performed in a relatively quiet environment to ensure that the recorded audio signal is of high quality and has clear sound features, providing a high-quality data source for subsequent feature extraction and model creation.
[0069] After the initial audio is recorded, frequency feature extraction, timbre feature extraction, and rhythm feature extraction are performed on each audio signal. This series of feature extraction operations aims to comprehensively analyze the sound characteristics in the audio signal, from the frequency distribution of pitch, the texture and timbre of the sound, to the rhythm and speed of pronunciation, etc., to construct a multi-dimensional feature representation of each object's sound. Frequency feature extraction can reveal the pitch changes of the sound, timbre feature extraction reflects the texture and personality of the sound, and rhythm feature extraction captures the rhythmic characteristics of speaking or singing. This information is crucial for subsequent voiceprint feature matching.
[0070] After feature extraction, the frequency feature information, timbre feature information and rhythm feature information are aggregated to form comprehensive audio voiceprint feature information. This process involves feature selection, weight allocation and fusion algorithms, with the goal of integrating information from multiple feature dimensions into a complete voiceprint model to ensure that the model can accurately reflect the sound characteristics of the object and provide an accurate basis for subsequent audio signal matching and separation. The audio voiceprint model after feature aggregation contains feature information on multiple aspects of the object's sound, such as frequency, timbre and rhythm, which can effectively assist the system in identifying and separating target audio signals in complex multi-sound source environments, improving the accuracy and efficiency of audio processing.
[0071] As an optional solution, feature aggregation is performed on the frequency feature information, the timbre feature information and the rhythm feature information to obtain audio voiceprint feature information of at least two objects, including:
[0072] Acquire environmental information where at least two objects are located, wherein the first audio data is mixed audio data generated under the environmental information;
[0073] Acquire a feature weight combination for matching the environmental information, wherein the feature weight combination includes a frequency feature weight corresponding to the frequency feature information, a timbre feature weight corresponding to the timbre feature information, and a rhythm feature weight corresponding to the rhythm feature information, and different environmental information corresponds to different feature weight combinations;
[0074] According to the frequency feature weight, the timbre feature weight and the rhythm feature weight, feature aggregation is performed on the frequency feature information, the timbre feature information and the rhythm feature information to obtain audio voiceprint feature information of at least two objects.
[0075] Optionally, in this embodiment, environmental information refers to specific environmental conditions during audio recording, including but not limited to noise type, background sound, recording device characteristics, etc., which can affect the frequency, timbre and rhythm characteristics of the audio signal.
[0076] Optionally, in this embodiment, the feature weight combination refers to a set of weight values assigned to the frequency feature information, timbre feature information and rhythm feature information respectively according to the environmental information. The characteristic performance of the sound may be different under different environmental information, so it is necessary to adjust the feature weight to adapt to the environment and improve the accuracy and robustness of the voiceprint model.
[0077] Optionally, in this embodiment, the mixed audio data refers to an audio signal recorded in a multi-sound source environment, which contains sound information of at least two objects, and the information may overlap with each other and needs to be separated.
[0078] Optionally, in this embodiment, environmental information of at least two objects is analyzed. Since environmental conditions (such as noise type, background sound, etc.) will affect the frequency, timbre and rhythm of the audio signal, environmental information is crucial for the extraction of voiceprint features.
[0079] Next, a feature weight combination matching the environmental information is obtained. This combination includes frequency feature weights, timbre feature weights and rhythm feature weights, which are used to reflect the relative importance of frequency, timbre and rhythm features for voiceprint recognition in a specific environment.
[0080] Finally, according to the frequency feature weight, timbre feature weight and rhythm feature weight specified in the feature weight combination, the frequency feature information, timbre feature information and rhythm feature information are aggregated to obtain comprehensive audio voiceprint feature information as the basis for creating an audio voiceprint model. This process ensures that the audio voiceprint model created under specific environmental conditions can more accurately reflect the sound characteristics of the object by weighted fusion of different feature information, thereby improving the accuracy of subsequent target audio signal recognition and separation.
[0081] Through the embodiments provided in this application, through this series of feature weight adjustments and feature aggregation based on environmental information, it is possible to create personalized and environmentally adaptable audio voiceprint models, providing strong technical support for target human voice extraction and optimized processing in complex audio environments.
[0082] As an optional solution, after performing voiceprint feature matching on the second audio signal to obtain the target audio signal after at least two objects are separated, the method further includes:
[0083] Extracting an accompaniment audio signal corresponding to a preset text from the first audio data;
[0084] The target audio signal and the accompaniment audio signal are used to perform audio synthesis to obtain target audio data synthesized by at least two objects, wherein audio parameters of the target audio signal and the accompaniment audio signal are kept matched.
[0085] Optionally, in this embodiment, the accompaniment audio signal is accurately extracted from the original first audio data based on a preset text (such as lyrics). The extraction process of the accompaniment audio signal ensures the purity and integrity of the accompaniment audio track, avoids interference from the vocal part, and provides high-quality background music for subsequent audio synthesis.
[0086] Then, the separated target audio signal is used to perform audio synthesis with the accompaniment audio signal to generate the final target audio data. The audio synthesis operation is not just a simple audio splicing, but a natural fusion of the target audio signal and the accompaniment audio signal by accurately matching audio parameters, including volume, pitch, rhythm and timing. This process involves the correction of the lyrics timeline, fine-tuning of volume and pitch, rhythm synchronization and correction of timing deviation to ensure the overall harmony and expressiveness of the synthesized audio.
[0087] It should be noted that the ultimate goal of audio synthesis is to generate high-quality target audio data, in which the combination of vocals and accompaniment tracks is both clear and natural, which can significantly improve the listening quality and artistic expression of the audio.
[0088] Through the above process, this embodiment not only achieves accurate separation of the target audio signal, but also achieves high-quality fusion of the target audio signal and the accompaniment audio track through extraction and audio synthesis of the accompaniment audio signal.
[0089] As an optional solution, the above audio processing method is applied to a scenario of voiceprint matching, precise voice extraction and dynamic mixing for user chorus scenarios. In conventional chorus recording scenarios, if the user does not wear headphones for recording, the recorded audio often contains accompaniment music, environmental noise, other voices and other irrelevant sounds, resulting in poor recording effects. Especially in noisy public places, such as shopping malls, streets, squares, etc., background noise will greatly interfere with the user's voice extraction and seriously affect the audio quality.
[0090] It should be noted that traditional chorus recording technology often relies on track separation. This method gradually analyzes and separates the audio signal, but its effect is limited in complex environments, especially in the case of strong accompaniment and background noise, it is difficult to accurately separate the target vocals.
[0091] In order to solve the above defects, this embodiment provides a voiceprint matching accurate voice extraction and dynamic mixing method based on the above audio processing method, such as Figure 2 As shown in the figure, the specific method includes: before the target user sings in chorus, it will detect whether there is a voiceprint model currently in use. If it exists, this step can be skipped. If it does not exist, the user is guided to record a short audio clip (a few seconds is enough) in a relatively quiet environment. The recorded audio is processed. A personalized voiceprint model is created. The model includes a digital representation of the target user's timbre, pitch, tone and other features as the basis for subsequent optimization and extraction.
[0092] When users participate in the chorus and record the chorus audio, the target voice is optimized in real time based on the previously created personalized voiceprint model. The focus of this stage is to optimize the target voice and improve the clarity, sound quality and expressiveness of the target voice in complex environments, ensuring that the target user's voice is clear and prominent in the chorus and does not conflict with the voices of other users.
[0093] Based on the personalized voiceprint characteristics of the target user, optimization strategies are implemented in real time: for example, the spectral characteristics of the target voice are identified and enhanced, background sounds that do not belong to the target user are further suppressed, and the expressiveness of the target voice is improved in a targeted manner; for example, the pitch range of the target voice is dynamically adjusted to ensure alignment with the pitch range of the accompaniment and other users, and the volume gain is adjusted at the same time to optimize the clarity and expressiveness of the target voice; for example, the frequency response and sound quality are dynamically adjusted to ensure that the timbre of the target voice is clearer and fuller, and better integrated with the other parts of the chorus track.
[0094] After the above-mentioned chorus audio recording is completed, the optimized audio is further used to accurately extract the target vocals.
[0095] Optionally, the chorus audio is further compared with the lyrics file, and the lyrics are preliminarily matched according to the timeline to identify the audio segments corresponding to the target lyrics. In this process, non-singing content (such as chat sounds, environmental noise, etc.) will be automatically removed because they do not match the lyrics. In this way, only the audio segments that match the lyrics content are retained, and those that do not match the lyrics are excluded.
[0096] This step not only filters out the singing part, but also reduces the amount of subsequent calculations, because only the audio segments corresponding to the lyrics will enter the subsequent analysis process.
[0097] If multiple people sing the same lyrics at the same time during the chorus, and these audios are mixed together, multi-channel signal processing technology is used to separate the audio of different users. In this case, multiple audio signals are still obtained, but these signals may not be directly aligned with the target user.
[0098] After multi-channel voice analysis, the personalized voiceprint model of the target user will be used to match the features in the audio segment to further confirm whether the audio segment belongs to the target user. By matching the audio features in the audio segment with the target user's voiceprint model, the vocal part belonging to the target user can be finally determined.
[0099] After the target vocal extraction is completed in the above steps, a secondary synthesis will be performed to ensure the high-quality fusion of the target vocal and the accompaniment. This step is not just a simple audio splicing, but a variety of optimization strategies to ensure that the synthesized audio is natural and harmonious.
[0100] Secondary synthesis and optimization strategies include the following:
[0101] Lyric Timeline Alignment with Audio: Ensure that the target vocal is aligned with the rhythm and timing of the backing track. By aligning the lyrics timeline with the audio content, any slight misalignment caused by the recording process can be further corrected.
[0102] Volume, pitch and rhythm optimization: During the synthesis process, fine-tuning is performed based on the overall volume, pitch and rhythm characteristics of the target vocal and accompaniment. This ensures that the target vocal is not drowned out by the accompaniment and that the pitch is consistent with the overall style of the chorus.
[0103] Spectrum matching and timbre correction: further optimize the spectral characteristics of the target vocal to ensure that it matches the spectrum of the accompaniment track. By adjusting the timbre, formant, etc. of the target vocal, the harmony and expressiveness of the overall audio can be improved.
[0104] The volume and rhythm are automatically adjusted, and the time deviation is corrected to finally generate high-quality chorus audio. The target user can choose to rate the final generated chorus audio. The system can accurately evaluate whether the matching degree of the voiceprint model is high enough through user feedback and continuously adjust the model.
[0105] It should be noted that the above personalized voiceprint model creation process is as follows Figure 3 As shown, specifically including:
[0106] Segmentation and preprocessing of sound signals: The user records the audio for the first time (record 3-5 seconds of audio in a quiet environment), and the recorded audio signal is segmented, and denoised and normalized to ensure data cleanliness and consistency.
[0107] Multi-level sound feature extraction: Frequency feature extraction, extracting the pitch, main frequency and harmonics of the user's voice to form the user's frequency spectrum. Tone feature extraction, capturing the timbre characteristics of the user's voice and generating a unique timbre vector for the user. Rhythm feature extraction, recording the user's speaking speed and vocal rhythm, analyzing the timing characteristics of time intervals and pauses.
[0108] Feature aggregation and dynamic weight assignment: Aggregate the extracted frequency, timbre, and rhythm features, and assign weights based on the user's vocal characteristics. Frequency feature weight: If the user's pitch changes frequently, a higher weight will be given to the frequency feature to ensure that in the case of multiple recordings or complex background noise, the pitch feature can be relied upon to distinguish the user's voice. Tone feature weight: For users with more unique timbres, the weight of the timbre feature will be increased, especially when the frequency feature is not obvious, the timbre can help the system accurately separate the human voice. Rhythm feature weight: If the user's vocal rhythm is more regular, the system will give a higher weight to the rhythm feature to ensure that in long recordings, the system can accurately match based on the rhythm features.
[0109] Dynamic weight adjustment mechanism: Real-time monitoring of sound characteristics during the recording process, and dynamic adjustment of feature weights based on the complexity of environmental noise and the dynamic changes of user voice. For example, in a noisy environment, timbre characteristics may be given a higher weight; in a multi-person chorus scene, frequency changes will become the dominant feature. The system continuously adjusts the feature weight ratio in the voiceprint model based on the quality of each recording, noise conditions, and user feedback (users can score the quality of the final chorus audio generated each time, and the system will continue to adjust based on the score) to ensure high accuracy in different environments.
[0110] Accurate vocal extraction and secondary synthesis in chorus recordings. Users are required to select the current environment category (such as indoors, shopping malls, streets, etc.) before singing in chorus. During the chorus, the system will monitor the environmental noise in real time and remove noise interference in real time according to the environment category pre-selected by the user. For example, in a street environment, the system will focus on removing traffic noise. In addition, although basic feature weights have been assigned in the process of creating the above voiceprint model, in actual recordings, the degree of interference from environmental noise may reduce the reliability of certain features. For example, in a very noisy environment, frequency features may be more recognizable than timbre features. Therefore, in this case, the weight of the feature will be dynamically increased according to the environmental features preset by the user to help extract the human voice more accurately. This is a real-time optimization for different environmental noises, not a change to the voiceprint model itself.
[0111] After the user completes the chorus recording, the system will generate a mixed audio containing the user's voice, accompaniment music and environmental noise. Using the created voiceprint model, the user's voice is extracted from the mixed audio containing accompaniment and background noise to ensure that the extracted voice is clear and interference-free. Different from traditional track separation, the accompaniment and noise are removed according to the matching degree of the voiceprint features, retaining the user's original voice.
[0112] After the vocals are extracted successfully, the system will no longer optimize the existing tracks, but will perform a new secondary synthesis of the extracted vocals and the accompaniment tracks, automatically adjust the volume and rhythm, correct the time deviation, and finally generate high-quality chorus audio. Users can choose to rate the final chorus audio. The system can accurately evaluate whether the matching degree of the voiceprint model is high enough through user feedback and continuously adjust the model.
[0113] This embodiment significantly simplifies the traditional audio track separation steps by introducing the voiceprint model, and improves the efficiency and accuracy of voice extraction. Combined with the dynamic weight allocation mechanism and secondary synthesis technology, it can quickly extract the user's voice in a complex recording environment and generate high-quality chorus audio. By optimizing the voice extraction process, the system not only reduces the processing steps, but also greatly improves the effect of audio processing.
[0114] It is understandable that in the specific implementation of this application, related data such as user information is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0115] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0116] According to another aspect of the embodiments of the present application, an audio processing device for implementing the above audio processing method is also provided. Figure 4 As shown, the device comprises:
[0117] An acquiring unit 402 is configured to acquire first audio data to be processed, wherein the first audio data includes a first audio signal of a first audio attribute, and the first audio signal is an audio signal mixed by at least two objects;
[0118] a gain unit 404, configured to perform an audio gain operation on the first audio signal to obtain a second audio signal with a second audio attribute, wherein the audio attribute of the audio signal is used to indicate the audio performance degree of the audio signal in the audio data, and the second audio attribute is higher than the first audio attribute;
[0119] The matching unit 406 is configured to perform voiceprint feature matching on the second audio signal to obtain a target audio signal after at least two objects are separated.
[0120] As an optional solution, the gain unit 404 includes at least one of the following:
[0121] A first gain module, configured to perform a spectral gain operation on the first audio signal, wherein the spectral gain operation is used to enhance the spectral expression of the first audio signal;
[0122] A second gain module, configured to perform a volume gain operation on the first audio signal, wherein the spectral gain operation is used to enhance the volume performance of the first audio signal;
[0123] A third gain module, configured to perform a frequency gain operation on the first audio signal, wherein the spectrum gain operation is used to enhance the frequency representation of the first audio signal;
[0124] The fourth gain module is used to perform a sound quality gain operation on the first audio signal, wherein the spectrum gain operation is used to enhance the sound quality performance of the first audio signal.
[0125] As an optional solution, the device also includes:
[0126] A creation module, used for creating an audio voiceprint model before obtaining the first audio data to be processed, wherein the audio voiceprint model package is used to indicate audio voiceprint feature information of at least two objects;
[0127] The device also includes:
[0128] The determination module is used to determine, after acquiring the first audio data to be processed, a first audio signal matching the audio voiceprint feature information from the first audio data based on the audio voiceprint model.
[0129] As an optional solution, the matching unit 404 includes:
[0130] A first extraction module is used to extract a third audio signal corresponding to a preset text from the second audio signal, and perform audio separation on the third audio signal to obtain a plurality of separated audio signals, wherein one separated audio signal corresponds to one object;
[0131] The matching module is used to use the audio voiceprint feature information to perform voiceprint feature matching on multiple separated audio signals to obtain target audio signals after at least two objects are separated.
[0132] As an alternative, create a module that includes:
[0133] a recording submodule, for recording initial audio for at least two subjects;
[0134] An extraction submodule, used to extract frequency features, timbre features, and rhythm features from the initial audio, respectively, to obtain frequency feature information, timbre feature information, and rhythm feature information of at least two objects;
[0135] The aggregation submodule is used to perform feature aggregation on the frequency feature information, the timbre feature information and the rhythm feature information to obtain the audio voiceprint feature information of at least two objects.
[0136] As an optional solution, aggregate submodules, including:
[0137] A first acquisition subunit is used to acquire environment information where at least two objects are located, wherein the first audio data is mixed audio data generated under the environment information;
[0138] A second acquisition subunit is used to acquire a feature weight combination for matching environmental information, wherein the feature weight combination includes a frequency feature weight corresponding to the frequency feature information, a timbre feature weight corresponding to the timbre feature information, and a rhythm feature weight corresponding to the rhythm feature information, and different environmental information corresponds to different feature weight combinations;
[0139] The aggregation subunit is used to perform feature aggregation on the frequency feature information, the timbre feature information and the rhythm feature information according to the frequency feature weight, the timbre feature weight and the rhythm feature weight to obtain the audio voiceprint feature information of at least two objects.
[0140] As an optional solution, the device also includes:
[0141] A second extraction module is used to extract an accompaniment audio signal corresponding to a preset text from the first audio data after performing voiceprint feature matching on the second audio signal to obtain a target audio signal after at least two objects are separated;
[0142] A synthesis module is used to perform audio synthesis using the target audio signal and the accompaniment audio signal after performing voiceprint feature matching on the second audio signal to obtain a target audio signal after at least two objects are separated, so as to obtain target audio data synthesized from at least two objects, wherein the audio parameters of the target audio signal and the accompaniment audio signal are kept matched.
[0143] According to another aspect of the embodiment of the present application, an electronic device for implementing the above audio processing method is also provided, further as follows Figure 5 As shown, the electronic device includes a memory 502 and a processor 504. The memory 502 stores a computer program, and the processor 504 is configured to execute the steps in any of the above method embodiments through the computer program.
[0144] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0145] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:
[0146] S1, obtaining first audio data to be processed, wherein the first audio data includes a first audio signal of a first audio attribute, and the first audio signal is an audio signal mixed by at least two objects;
[0147] S2, performing an audio gain operation on the first audio signal to obtain a second audio signal with a second audio attribute, wherein the audio attribute of the audio signal is used to indicate the audio performance degree of the audio signal in the audio data, and the second audio attribute is higher than the first audio attribute;
[0148] S3, performing voiceprint feature matching on the second audio signal to obtain a target audio signal after at least two objects are separated.
[0149] Alternatively, a person skilled in the art may understand that: Figure 5 The structure shown is for illustration only. Figure 5 The structure of the electronic device is not limited. Figure 5 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 5 Different configurations shown.
[0150] Among them, the memory 502 can be used to store software programs and modules, such as the program instructions / modules corresponding to the audio processing method and device in the embodiments of the present application. The processor 504 executes various functional applications and data processing by running the software programs and modules stored in the memory 502, that is, realizing the above-mentioned audio processing method. The memory 502 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 502 may further include a memory remotely located relative to the processor 504, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 502 can be specifically, but not limited to, used to store information such as first audio attributes and second audio attributes. As an example, such as Figure 5 As shown, the memory 502 may include but is not limited to the acquisition unit 402, the gain unit 404, and the matching unit 406 in the audio processing device. In addition, it may also include but is not limited to other module units in the audio processing device, which will not be repeated in this example.
[0151] Optionally, the transmission device 506 is used to receive or send data via a network. Specific examples of the network may include wired networks and wireless networks. In one example, the transmission device 506 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers via a network cable so as to communicate with the Internet or a local area network. In one example, the transmission device 506 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0152] In addition, the electronic device further includes: a display 508 for displaying information such as the first audio attribute and the second audio attribute; and a connection bus 510 for connecting various module components in the electronic device.
[0153] In other embodiments, the client or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes in the form of network communication. A peer-to-peer network may be formed between the nodes, and any form of computing device, such as a server, a client, or other electronic device, may become a node in the blockchain system by joining the peer-to-peer network.
[0154] According to one aspect of the present application, a computer program product is provided, the computer program product comprising a computer program / instruction, the computer program / instruction comprising a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit, various functions provided by the embodiments of the present application are executed.
[0155] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0156] It should be noted that the computer system of the electronic device is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0157] The computer system includes a central processing unit (CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) or the program loaded from the storage part to the random access memory (RAM). In the random access memory, various programs and data required for system operation are also stored. The central processing unit, the read-only memory and the random access memory are connected to each other through a bus. The input / output interface (I / O interface) is also connected to the bus.
[0158] The following components are connected to the input / output interface: an input part including a keyboard, a mouse, etc.; an output part including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage part including a hard disk, etc.; and a communication part including a network interface card such as a local area network card, a modem, etc. The communication part performs communication processing via a network such as the Internet. A drive is also connected to the input / output interface as needed. Removable media, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., are installed on the drive as needed so that the computer program read therefrom is installed into the storage part as needed.
[0159] In particular, according to an embodiment of the present application, the process described in each method flow chart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, and the computer program contains a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit, various functions defined in the system of the present application are executed.
[0160] According to one aspect of the present application, a computer-readable storage medium is provided, and a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above-mentioned various optional implementations.
[0161] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0162] S1, obtaining first audio data to be processed, wherein the first audio data includes a first audio signal of a first audio attribute, and the first audio signal is an audio signal mixed by at least two objects;
[0163] S2, performing an audio gain operation on the first audio signal to obtain a second audio signal with a second audio attribute, wherein the audio attribute of the audio signal is used to indicate the audio performance degree of the audio signal in the audio data, and the second audio attribute is higher than the first audio attribute;
[0164] S3, performing voiceprint feature matching on the second audio signal to obtain a target audio signal after at least two objects are separated.
[0165] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing hardware related to the electronic device through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0166] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0167] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers or network devices, etc.) to execute all or part of the steps of the methods of each embodiment of the present application.
[0168] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0169] In the several embodiments provided in the present application, it should be understood that the recorded client can be implemented in other ways. Among them, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0170] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0171] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0172] The above are only preferred implementations of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. An audio processing method, characterized in that: include: Acquire first audio data to be processed, wherein the first audio data includes a first audio signal of a first audio attribute, and the first audio signal is an audio signal mixed by at least two objects; Performing an audio gain operation on the first audio signal to obtain a second audio signal with a second audio attribute, wherein the audio attribute of the audio signal is used to indicate the audio performance degree of the audio signal in the audio data, and the second audio attribute is higher than the first audio attribute; Perform voiceprint feature matching on the second audio signal to obtain a target audio signal after the at least two objects are separated.
2. The method according to claim 1, characterized in that The step of performing an audio gain operation on the first audio signal to obtain a second audio signal with a second audio attribute includes at least one of the following: Performing a spectrum gain operation on the first audio signal, wherein the spectrum gain operation is used to enhance the spectrum expression of the first audio signal; Performing a volume gain operation on the first audio signal, wherein the spectral gain operation is used to enhance the volume performance of the first audio signal; Performing a frequency gain operation on the first audio signal, wherein the frequency gain operation is used to enhance the frequency representation of the first audio signal; A sound quality gain operation is performed on the first audio signal, wherein the spectrum gain operation is used to enhance the sound quality performance of the first audio signal.
3. The method according to claim 1, characterized in that Before obtaining the first audio data to be processed, the method further includes: Creating an audio voiceprint model, wherein the audio voiceprint model package is used to indicate audio voiceprint feature information of the at least two objects; After obtaining the first audio data to be processed, the method further includes: Based on the audio voiceprint model, the first audio signal matching the audio voiceprint feature information is determined from the first audio data.
4. The method according to claim 3, characterized in that: The performing voiceprint feature matching on the second audio signal to obtain the target audio signal after the at least two objects are separated includes: Extracting a third audio signal corresponding to the preset text from the second audio signal, and performing audio separation on the third audio signal to obtain a plurality of separated audio signals, wherein one separated audio signal corresponds to one object; The audio voiceprint feature information is used to perform voiceprint feature matching on the multiple separated audio signals to obtain the target audio signal after the at least two objects are separated.
5. The method according to claim 3, characterized in that: The step of creating an audio voiceprint model includes: recording initial audio for the at least two subjects; Performing frequency feature extraction, timbre feature extraction, and rhythm feature extraction on the initial audio respectively to obtain frequency feature information, timbre feature information, and rhythm feature information of the at least two objects; The frequency feature information, the timbre feature information and the rhythm feature information are aggregated to obtain the audio voiceprint feature information of the at least two objects.
6. The method according to claim 5, characterized in that The step of performing feature aggregation on the frequency feature information, the timbre feature information, and the rhythm feature information to obtain the audio voiceprint feature information of the at least two objects includes: Acquire environment information where the at least two objects are located, wherein the first audio data is mixed audio data generated under the environment information; Acquire a feature weight combination for matching the environmental information, wherein the feature weight combination includes a frequency feature weight corresponding to the frequency feature information, a timbre feature weight corresponding to the timbre feature information, and a rhythm feature weight corresponding to the rhythm feature information, and different environmental information corresponds to different feature weight combinations; According to the frequency feature weight, the timbre feature weight and the rhythm feature weight, feature aggregation is performed on the frequency feature information, the timbre feature information and the rhythm feature information to obtain the audio voiceprint feature information of the at least two objects.
7. The method according to any one of claims 1 to 6, characterized in that: After performing voiceprint feature matching on the second audio signal to obtain the target audio signal after the at least two objects are separated, the method further includes: Extracting an accompaniment audio signal corresponding to a preset text from the first audio data; The target audio signal and the accompaniment audio signal are used to perform audio synthesis to obtain target audio data synthesized by the at least two objects, wherein audio parameters of the target audio signal and the accompaniment audio signal are kept matched.
8. An audio processing device, characterized in that: include: An acquiring unit, configured to acquire first audio data to be processed, wherein the first audio data comprises a first audio signal of a first audio attribute, and the first audio signal is an audio signal mixed by at least two objects; a gain unit, configured to perform an audio gain operation on the first audio signal to obtain a second audio signal with a second audio attribute, wherein the audio attribute of the audio signal is used to indicate the audio performance degree of the audio signal in the audio data, and the second audio attribute is higher than the first audio attribute; The matching unit is used to perform voiceprint feature matching on the second audio signal to obtain a target audio signal after the at least two objects are separated.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of claims 1 to 7 when executed by an electronic device.
10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 7 through the computer program.