A multi-person chorus method, apparatus and storage device
Patent Information
- Application Number
- CN202310663603.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-06
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2043-06-06
AI Technical Summary
[0004]有鉴于此,本申请提供一种多人合唱方法、装置及存储介质,以解决现有技术中多人合唱音频对齐不准、合成效果不佳的技术问题
[0043]本发明提供的多人合唱方法、装置及存储介质,构建原唱音频曲谱和待处理音频中对应位置的发音信息之间的映射关系,以重建修音模板,能够利用曲谱和发音这一个信息的对应关系同时实现音速和音调的对齐,将所有的用户唱歌音频均与原唱音频曲谱相对应,提高了修音处理的效果,在保证音质质量,不过分修音的情况下,尽可能保证合唱音频的真实性,同时从速度和音调两个方面同时保证了合音音频的标准性,提高了合音的效果。
Smart Images

Figure CN116631362B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio and video data processing, and in particular to a method, apparatus and storage medium for multi-person chorus. Background Technology
[0002] Online multi-person duets represent an innovative application trend within the broader internet-based entertainment karaoke landscape. This involves multiple singers using an app to record the same section or consecutive sections of the same song online, collaboratively performing and creating a choral piece. Unlike traditional karaoke where individuals create and release their own work, multi-person duets allow multiple users to perform the same song together, generating short videos showcasing their shared vocals. Compared to regular karaoke apps, multi-person duets offer more engaging content, unique styles, and easier sharing. The participation method provides ample social opportunities, facilitating community development and expanding friendships with fellow music enthusiasts. Furthermore, multi-person duets offer flexible operational options, allowing for diverse audio and video formats and the creation of dynamic content that capitalizes on current events and trends.
[0003] The effectiveness of online multi-person duets relies on intelligent production technology that mixes multiple audio and video data. This technology is widely used in karaoke apps such as "Quanmin K Ge," "Changba," and "Huisen," representing a significant application scenario within the karaoke software category. Current online multi-person duet methods primarily employ two approaches to achieve the function of multiple people singing a song together. The first approach involves directly aligning the recorded audio from multiple users on the timeline without processing, merging the multiple audio tracks, and then performing basic mixing and adding accompaniment. However, insufficient alignment accuracy results in poor duet performance. The second approach provides users with pitch correction functions before alignment and merging. However, existing pitch correction technologies in the duet field only fine-tune pitch and cannot effectively adjust rhythm. The resulting duet performances using current technologies often lack the powerful impact and emotional resonance of multi-person duets. Summary of the Invention
[0004] In view of this, this application provides a method, apparatus and storage medium for multi-person chorus singing to solve the technical problems of inaccurate audio alignment and poor synthesis effect in the prior art.
[0005] The first aspect of the present invention provides a method for multi-person chorus singing, specifically comprising: forming an audio data pair by combining each of the multiple audio samples to be processed with its corresponding original audio sample, wherein the original audio sample is the original audio sample of the song sung by the singer in the multi-person chorus singing method;
[0006] For each audio data pair, the audio to be processed is tuned based on the original audio score and the pronunciation information in the audio to be processed, resulting in multiple audio to be mixed corresponding to the audio to be processed.
[0007] Among them, a mapping relationship is constructed between the original audio score and the corresponding pronunciation information in the audio to be processed, so as to reconstruct the pitch correction template, and the speed and pitch of the audio to be processed are processed based on the pitch correction template;
[0008] Multiple audio files to be mixed are mixed to obtain the initial chorus audio.
[0009] Preferably, before forming audio data pairs by combining each audio to be processed with its corresponding original audio, the process further includes:
[0010] In response to a user-triggered chorus generation request, the system retrieves singing videos of multiple users performing the same song via the chorus terminal. Each user's singing video represents a segment of the current chorus piece.
[0011] The positions of each user's singing content within the chorus are not entirely the same, and there is overlap in the singing content among the various users.
[0012] Extract the corresponding audio from the singing videos of multiple users to serve as multiple audio files to be processed.
[0013] Preferably, after forming audio data pairs by combining each audio to be processed with its corresponding original audio, the method further includes:
[0014] Noise reduction and echo cancellation are performed on each audio file to be processed;
[0015] Noise reduction is performed on each audio file, including expanding the number of sub-bands of the noise reduction module based on the frequency range of burst noise, and setting up hidden layers in multiple dimensions to construct an improved noise reduction module.
[0016] The improved noise reduction module filters out burst noise and continuous noise whose duration reaches a threshold.
[0017] Echo cancellation is performed on the background music of each audio file based on a neural network model.
[0018] Preferably, a mapping relationship is constructed between the original audio score and the corresponding pronunciation information in the audio to be processed to reconstruct the pitch correction template, specifically including:
[0019] For each audio data pair, a reference information set is constructed based on the original audio score.
[0020] Identify pronunciation information in the audio to be processed;
[0021] Matching and aligning the reference information set with the pronunciation information yields the mapping relationship between the original audio score and the pronunciation information at the corresponding positions in the audio to be processed, and the pitch correction template is reconstructed.
[0022] Preferably, the reference information set includes at least: the melody, chords, paragraph structure, style classification, and harmony information of the original audio score;
[0023] Identify pronunciation information in the audio to be processed, including using a pre-trained speech recognition model based on audio data to identify initials and finals in the audio to be processed, and obtain the location and tone of the initials and finals.
[0024] Preferably, the reconstruction pitch template specifically includes:
[0025] A reference information set is constructed based on the original audio score, and substitute notes are calculated at each point in the original audio score based on the reference information set to achieve score reconstruction.
[0026] Calculate the DTW distance between the alternative note at each point and the initials and finals in the audio to be processed, considering both pitch and temporal position data.
[0027] Using the mapping method that minimizes the sum of DTW distances in the entire audio file as the audio editing template, the audio editing template is reconstructed.
[0028] The audio editing template includes the positional mapping relationship and pitch mapping relationship between the original audio and the audio to be processed.
[0029] Preferably, after obtaining the initial chorus audio, the process also includes:
[0030] Extract the video corresponding to multiple audio files to be mixed.
[0031] Arrange the videos corresponding to the multiple audio tracks to be mixed according to a preset layout.
[0032] The mixing method based on the initial chorus audio plays video frames corresponding to multiple audios to be mixed, as well as the initial chorus audio.
[0033] Preferably, after adjusting the pitch and tempo of the audio to be processed based on the mapping relationship in the pitch correction template to obtain the initial audio to be mixed, the method further includes:
[0034] The initial audio to be mixed is input into the smoothing module to output a fine-grained audio sequence.
[0035] The audio to be processed is input into the noise reduction and filtering module to obtain the reference audio to be processed.
[0036] The small-grained audio sequence and the reference audio to be processed are input into the speed-down and pitch-down module to obtain the audio to be mixed.
[0037] A second aspect of the present invention exemplarily provides a multi-person chorus device, comprising:
[0038] The pairing module is used to form audio data pairs by pairing each audio to be processed with its corresponding original audio, wherein the original audio is the original audio of the song sung by the singer in the multi-person chorus method.
[0039] The audio correction module is used to perform audio correction on each audio data pair based on the original audio score and the pronunciation information in the audio to be processed, so as to obtain multiple audio to be mixed corresponding to the audio to be processed.
[0040] The pitch correction module is configured to construct a mapping relationship between the original audio score and the corresponding pronunciation information in the audio to be processed, so as to reconstruct the pitch correction template and process the speed and pitch of the audio to be processed based on the pitch correction template.
[0041] The mixing module is used to mix multiple audio files to obtain the initial chorus audio.
[0042] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that the program, when executed by a processor, implements the steps of any of the above-described multi-person chorus methods.
[0043] The multi-person chorus method, apparatus, and storage medium provided by this invention construct a mapping relationship between the original audio score and the corresponding pronunciation information in the audio to be processed, so as to reconstruct the pitch correction template. It can simultaneously achieve the alignment of tempo and pitch by utilizing the correspondence between the score and the pronunciation information, so that all user singing audio corresponds to the original audio score, thereby improving the effect of pitch correction. While ensuring sound quality and avoiding excessive pitch correction, it ensures the authenticity of the chorus audio as much as possible, and at the same time ensures the standardization of the chorus audio in terms of both tempo and pitch, thereby improving the effect of the chorus. Attached Figure Description
[0044] Figure 1 A flowchart of a multi-person chorus method provided in the prior art;
[0045] Figure 2 A flowchart of a pitch correction method provided in the prior art;
[0046] Figure 3 A flowchart illustrating a multi-person chorus method is shown in an exemplary embodiment of the present invention;
[0047] Figure 4 A flowchart illustrating the pitch correction process as shown in an exemplary embodiment of the present invention;
[0048] Figure 5The existing multi-person chorus display interface is provided in the current technology;
[0049] Figure 6 This is an exemplary embodiment of the present invention, showing a multi-person chorus display interface.
[0050] Figure 7 This is a schematic diagram of the structure of a multi-person chorus device shown in an exemplary embodiment of the present invention;
[0051] Figure 8 This is a schematic diagram of the structure of an electronic device as an exemplary embodiment of the present invention. Detailed Implementation
[0052] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0053] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0054] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0055] like Figure 1 As shown, existing multi-person chorus solutions mainly include two steps: recording the singing data of multiple users, optionally performing audio editing on the singing data; and mixing the multiple singing data. In the first step, when performing audio editing on the singing data, the existing technology, such as... Figure 2As shown, the existing pitch correction algorithm analyzes the audio pitch and then adjusts it, only adjusting the frequency information of the sound without manipulating the time, meaning the rhythm of the vocal audio remains unchanged. Specifically, the existing pitch correction algorithm analyzes the frequency curve of the audio to obtain the note endpoints. By comparing these note endpoints with the musical score, it generates the pitch adjustment information needed for pitch correction. After obtaining the pitch adjustment information, a frequency-sensitive sola algorithm is used to process the vocals, ensuring that each note endpoint in the audio is at the correct pitch, making the melody sound basically correct. The above process only focuses on pitch data, that is, only adjusting the pitch data. However, in choral singing, non-professional users are very prone to rushing, missing beats, and overall audio misalignment. Existing technology cannot accurately adjust the audio rhythm. In the second step, due to the limitations of mixing technology, only time point information can be obtained. Simply aligning according to the time points results in low-quality choral audio, limited adjustable content, insufficient interest, and difficulty in attracting user interest.
[0056] Choral singing is an application scenario with extremely high requirements for rhythm. To create a piece that doesn't evoke negative emotions, the singers need to have good vocal skills to ensure that each person's voice is harmoniously unified. This means that traditional methods often make users with average singing abilities sound particularly out of place in choral works, discouraging novice users from participating. Consequently, this results in a small user base and limited repertoire. Existing multi-person choral methods also produce choral songs of poor quality and performance, and users have limited options for additional operations on the mixed choral songs, reducing the enjoyment of the choral singing process.
[0057] To address the technical problems in the existing technologies mentioned above, such as poor sound correction and mixing in choral singing, resulting in low audio quality and a lack of enjoyment in the choral singing process, the first embodiment of this invention provides a method for multi-person choral singing, such as... Figure 3 As shown, the multi-person chorus method specifically includes:
[0058] Each audio file to be processed is paired with its corresponding original audio file to form an audio data pair.
[0059] For each audio data pair, the audio to be processed is processed based on the original audio score and the pronunciation information in the audio to be processed, resulting in multiple audio to be mixed corresponding to the audio to be processed. The original audio is the original audio of the song sung by the singer in the multi-person chorus method.
[0060] Among them, a mapping relationship is constructed between the original audio score and the corresponding pronunciation information in the audio to be processed, so as to reconstruct the pitch correction template, and the speed and pitch of the audio to be processed are processed based on the pitch correction template;
[0061] Multiple audio files to be mixed are mixed to obtain the initial chorus audio.
[0062] The multi-person chorus method provided by this invention can simultaneously align tempo and pitch by utilizing the correspondence between musical score and pronunciation information. It ensures that all user singing audio corresponds to the original audio score, thereby improving the effect of pitch correction. While ensuring sound quality and avoiding excessive pitch correction, it also ensures the authenticity of the chorus audio as much as possible. At the same time, it guarantees the standardization of the chorus audio in terms of both tempo and pitch, thus improving the effect of the chorus.
[0063] Specifically, before forming audio data pairs by combining each audio to be processed with its corresponding original audio, the process further includes: responding to a user-triggered chorus generation request, acquiring singing videos of multiple users performing the same song through a chorus terminal. Each user's singing video is a segment of the current chorus song, and the positions of each user's singing content within the song are not entirely identical, with some overlap. The corresponding audio is extracted from these videos to serve as multiple audio samples to be processed. In a chorus app, numerous users can access the app's chorus page through their terminals to select songs and record audio and video, enabling them to achieve a chorus effect with other users singing the same song on the platform. If multiple users are singing the same segment of the same song, the timbre attributes of the user who triggered the chorus generation request are analyzed. Based on the matching degree of the timbre attributes, singers corresponding to various singing segments in other segments are selected, thus acquiring singing videos of multiple users performing the same song through the chorus terminal.
[0064] After obtaining multiple audio samples to be processed, the multi-person chorus method includes: forming audio data pairs by pairing each audio sample with its corresponding original vocal audio. Specifically, the original vocal audio of the song is segmented according to paragraphs to obtain multiple original vocal audio segments A1, A2, ..., An; the original vocal audio segments corresponding to the multiple audio samples to be processed are matched to form audio data pairs S1.<A1,B1> S2<A2,B2> , ..., Sn<An,Bn> In each audio data pair, the text content of the original audio clip includes the text content of the audio to be processed. That is, the text content of the audio to be processed is within the text content of the original audio clip, and the text content of the original audio clip is longer than the text content of the audio to be processed. For example, the text content of the audio to be processed is "Where will you be?", and the text content of the original audio in the audio data pair is "Suddenly I miss you, where will you be?".
[0065] After forming audio data pairs between each audio file and its corresponding original audio file, the process further includes noise reduction and echo cancellation for each audio file. Noise reduction involves expanding the number of sub-bands in the noise reduction module based on the frequency range of sudden noise, setting multiple-dimensional hidden layers to construct an improved noise reduction module, and filtering sudden noise and persistent noise with a duration reaching a threshold based on the improved noise reduction module. For example, the noise reduction module is a neural network module. Specifically, for the frequency range of sudden noise in choral scenes, the neural network module is improved to construct an improved noise reduction module. Specifically, the number of sub-bands in the neural network module is expanded to 24, and hidden layers of 512 and 384 dimensions are set respectively, enabling data to support sampling rates higher than the industry average, providing a wider frequency range of operational space for choral data processing. On the other hand, echo cancellation is performed on the background music of each audio file based on a neural network model. This neural network model adopts a structure of a pre-speak-X and a post-neural network. The multi-person chorus method provided by the present invention preprocesses the acquired audio through the above-mentioned noise reduction and echo cancellation methods, eliminates the background music generated by the playback device, and obtains clearer vocals. In the case of external playback, it can obtain clear vocals in complex environmental scenarios, thereby improving the accuracy of subsequent harmony processing.
[0066] For each audio data pair, the audio to be processed is tuned based on the original audio score and the pronunciation information in the audio to be processed, resulting in multiple audio to be mixed corresponding to the audio to be processed.
[0067] like Figure 4The process involves constructing a mapping relationship between the original audio score and the corresponding pronunciation information in the audio to be processed, in order to reconstruct a pitch correction template. Based on this template, the audio to be processed is then adjusted for tempo and pitch. Specifically, constructing this mapping relationship between the original audio score and the corresponding pronunciation information in the audio to be processed to reconstruct the pitch correction template includes: for each audio data pair, constructing a reference information set based on the original audio score; identifying the pronunciation information in the audio to be processed; matching and aligning the reference information set with the pronunciation information to obtain the mapping relationship between the original audio score and the corresponding pronunciation information in the audio to be processed, thus reconstructing the pitch correction template. The reference information set includes at least: the melody, chords, paragraph structure, style classification, and harmonic information of the original audio score; and the identification of pronunciation information in the audio to be processed, including identifying the initials and finals in the audio based on a pre-trained speech recognition model using audio data, obtaining the location and pitch of the initials and finals, identifying the text content in the audio to obtain paragraph information, constructing a mapping relationship based on paragraphs, and using the pitch of the initials and finals to adjust the pitch of the audio to be processed. As an optional embodiment, the speech recognition model adopts an HMM structure based on lyrics and a post-processing module based on boundary reinforcement, and pre-trains the speech recognition model using audio data. Existing audio editing technologies only adjust the pitch of user audio, lacking a precise timing recognition system for singing voices, thus failing to accurately adjust audio rhythm. Pitch recognition systems typically have time information offsets exceeding 100ms, while in singing scenarios, users noticeably perceive offsets greater than 30ms. Without 30ms precision time information, the audio editing module struggles to adjust tempo at this time granularity, failing to meet the technical requirements of online multi-person choral singing scenarios. Compared to using general speech recognition models to identify audio segments for audio editing, the speech recognition model provided in this invention utilizes an Hidden Markov Model (HMM) structure combined with a boundary enhancement post-processing module to identify audio boundaries. It achieves a 93.2% accuracy rate for boundary recognition at the 30ms level. Based on accurate boundary recognition, it can obtain tempo and pitch at every fine granularity, thereby performing audio editing processing in each segment, improving the precision and accuracy of audio editing. The following design improvements enhance boundary recognition capabilities: extensive use of singing audio data for model training, use of a lyrics-based HMM structure, and the use of a post-processing module that operates on boundaries.
[0068] The reconstruction of the pitch correction template specifically includes: constructing a reference information set based on the original audio score; calculating substitute notes at each point in the original audio score based on the reference information set to achieve score reconstruction; calculating the DTW distance between the substitute notes at each point and the initials and finals in the audio to be processed, and using the mapping method that minimizes the sum of DTW distances in the entire audio to be processed as the pitch correction template to reconstruct the pitch correction template. The pitch correction template includes the positional mapping relationship and pitch mapping relationship between the original audio and the audio to be processed. The reference information set includes at least: the melody, chords, paragraph structure, style classification, and harmonic information of the original audio score; and calculating substitute notes with different recommendation levels at each point in the original audio score based on the melody and chord progression. Due to the uncertainty of the difference between the input data and the original template, adjusting the audio using a fixed adjustment range has poor adjustment effects, so it is difficult to assume a safe audio adjustment range. The multi-person chorus method provided by this invention requires adaptive template reconstruction based on the input data to obtain the most economical new template, providing different audio editing strategies with minimal adjustments. Through the template reconstruction algorithm, the adjustment range of any audio can be controlled within the algorithm's manageable range. Based on this, any audio can pass through the audio editing module, ensuring that the output sound quality degradation rate is lower than the acceptable range for streaming music products (0.005%). Furthermore, after obtaining the audio to be mixed, all audio to be mixed is placed in the same accompaniment, ensuring sufficient rhythmic harmony.
[0069] Based on the mapping relationship in the pitch correction template, after adjusting the pitch and tempo of the audio to be processed to obtain the initial audio to be mixed, the method further includes: inputting the initial audio to be mixed into a smoothing module to output a small-grained audio sequence; inputting the audio to be processed into a noise reduction filtering module to obtain a reference audio to be processed; and inputting the small-grained audio sequence and the reference audio to be processed into a speed-shifting and pitch-changing module to obtain the audio to be mixed. Preferably, the speed-shifting and pitch-changing module can also be a Psola module, a tdPsola module, a sola module, a wpsola module, or a vocoder; the smoothing module can be an intermediate value interpolation module or a granularity adjustment module, used to realize the granularity conversion between input and output signals. In the traditional sola acoustic module, there is an 8-half-tone pitch adjustment limit for pitch adjustment. Exceeding this range can easily lead to sound quality damage. In order to ensure that the sound quality is not interfered with during the pitch correction process, the pitch correction method provided by this invention uses a smoothing module to output a small-grained audio sequence to avoid reaching the pitch adjustment limit, ensuring the audio quality during the pitch correction process, and thus improving the effect of multi-person chorus. The granularity of information generated by speed and pitch changes is at the phoneme level, which can lead to abrupt changes and unnatural transitions when used directly. On the other hand, the smoothing module provided by this invention breaks down the processing unit to the waveform level, reducing the processing granularity of the smoothing module to below 5ms. At this fine granularity, the speed level information of the initial audio to be mixed can be naturally connected, allowing the original singer's voice and vocal characteristics to be preserved, thus improving the effect of multi-person chorus.
[0070] Compared to existing technologies that simply adjust pitch to achieve repair, the pitch correction method provided by this invention adopts different processing schemes for different audio. It performs uniform processing on all recorded audio to obtain phoneme-level information such as initials and finals for mapping, so as to adjust the speed and pitch of the audio simultaneously. This results in consistent performance in rhythm and pitch of the processed audio, improving the quality of the corrected audio. At the same time, with the assistance of peripheral modules such as the smoothing module, it ensures that the audio to be mixed is restored from the phoneme granularity to the audio level, resulting in a natural and realistic audio to be mixed after adjusting the speed and pitch, which improves the effect and realism of multi-person chorus.
[0071] The process of mixing multiple audio files to obtain an initial chorus audio specifically includes: identifying the segment structure of the song; providing reverb parameters corresponding to the segments based on the segment structure; predicting the recommended volume range of each audio file to be mixed under different business scenarios based on the segment structure; adjusting the volume values of multiple audio files to be mixed according to the recommended volume range; and mixing multiple audio files to be mixed with volume values based on the song's segment structure and reverb parameters to obtain the initial chorus audio.
[0072] Furthermore, after obtaining the initial chorus audio, the multi-person chorus method further includes: adjusting the initial chorus audio based on an intelligent mixing module to obtain the final multi-person chorus song. The intelligent mixing module can perform at least one or more of the following functions: highlighting key figures, spectrum repair, and stylization. Highlighting key figures includes obtaining the user requesting the chorus song, matching the corresponding audio to be mixed as the target audio to be mixed, adjusting the volume and / or pitch of the target audio to be mixed, obtaining the highlighted target audio to replace the target audio to be mixed, remixing the updated multiple audios to be mixed, and obtaining the final multi-person chorus song. Spectrum repair includes obtaining the spectrogram of the initial chorus audio, locating the missing spectral positions in the spectrogram, repairing the missing spectral positions in the spectrogram based on spectral interpolation of neighboring pixels, and inverting the spectrogram to obtain the final multi-person chorus song. The stylization process specifically includes: receiving the user's selected stylization request, determining the audio to be mixed that matches the selected user, outputting the stylized audio to be mixed corresponding to the selected audio based on a pre-trained model, replacing the audio to be mixed that matches the selected user with the stylized audio to be mixed, remixing the updated audio to be mixed, and obtaining the final multi-person chorus song. Existing technologies, due to limitations in mixing techniques, cannot use style-based and engaging mixing techniques, resulting in choral works that are flat and lack sufficient soundstage and extension. The multi-person chorus method provided by this invention, based on basic mixing, offers multiple different intelligent modules that can be used in combination to further optimize the quality and dissemination effect of multi-person chorus songs, increasing their interest and appeal.
[0073] As an optional embodiment, after obtaining the initial chorus audio, the method further includes: extracting videos corresponding to multiple audios to be mixed, arranging the videos corresponding to the multiple audios to be mixed according to a preset format, and playing the video frames corresponding to the multiple audios to be mixed and the initial chorus audio based on the mixing method of the initial chorus audio. For example... Figure 5 , 6 As shown, Figure 5 This is the display interface for multi-person choral songs in existing technologies. Figure 6The existing technology for displaying multi-person choral songs, as provided by this invention, only uses lyrics and animated avatars to show the choral performance. Different numbers of avatars appear at different points in the audio playback to indicate participation in the current segment. This approach is monotonous, lacks engagement and appeal, has poor dissemination, and results in low user engagement for first-time users, making it difficult for them to effectively showcase their individuality. The multi-person choral method provided by this invention produces works with greater appeal, a unique style, and easier dissemination. Participation in multi-person choral singing offers abundant social opportunities, facilitating community development and expanding friendships with fellow music enthusiasts. Furthermore, multi-person choral singing offers flexible operational options, allowing for the display of diverse audio and video formats and the creation of dynamic content that capitalizes on current events and trends.
[0074] A second embodiment of the present invention provides a multi-person chorus device, such as... Figure 7 As shown, the multi-person chorus device specifically includes:
[0075] The pairing module is used to form audio data pairs by pairing each audio to be processed with its corresponding original audio, wherein the original audio is the original audio of the song sung by the singer in the multi-person chorus method.
[0076] The audio correction module is used to perform audio correction on each audio data pair based on the original audio score and the pronunciation information in the audio to be processed, so as to obtain multiple audio to be mixed corresponding to the audio to be processed.
[0077] The pitch correction module is configured to construct a mapping relationship between the original audio score and the corresponding pronunciation information in the audio to be processed, so as to reconstruct the pitch correction template and process the speed and pitch of the audio to be processed based on the pitch correction template.
[0078] The mixing module is used to mix multiple audio files to obtain the initial chorus audio.
[0079] It is not difficult to see that this embodiment is a device embodiment corresponding to the first embodiment, and this embodiment can be implemented in conjunction with the first embodiment. The relevant technical details mentioned in the first embodiment are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the first embodiment.
[0080] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this invention, this embodiment does not introduce units that are not closely related to solving the technical problem proposed by this invention; however, this does not mean that other units are absent from this embodiment.
[0081] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 3 This provides a method for multi-person choral singing.
[0082] This instruction manual also provides Figure 8 The diagram shows a schematic structure of an electronic device. Figure 8 At the hardware level, the electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for the business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to achieve the above-mentioned functions. Figure 3 The method for multi-person chorus described above. Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. In other words, the execution subject of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.
[0083] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0084] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0085] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0086] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.
Claims
1. A method for multi-person choral singing, characterized in that, The method of multi-person chorus specifically includes: Each audio file to be processed is paired with its corresponding original audio file to form an audio data pair, wherein the original audio file is the original audio file of the song sung by the singers in the multi-person chorus method; after pairing each audio file to be processed with its corresponding original audio file to form an audio data pair, the method further includes: Noise reduction and echo cancellation are performed on each audio file to be processed; Noise reduction is performed on each audio file, including expanding the number of sub-bands of the noise reduction module based on the frequency range of burst noise, and setting up hidden layers in multiple dimensions to construct an improved noise reduction module. The improved noise reduction module filters out burst noise and continuous noise whose duration reaches a threshold. Echo cancellation is performed on the background music of each audio file based on a neural network model. For each audio data pair, the audio to be processed is tuned based on the original audio score and the pronunciation information in the audio to be processed, resulting in multiple audio to be mixed corresponding to the audio to be processed. Among them, a mapping relationship is constructed between the original audio score and the corresponding pronunciation information in the audio to be processed, so as to reconstruct the pitch correction template. Based on the reconstructed pitch correction template, the speed and pitch of the audio to be processed are processed. The process of constructing a mapping relationship between the original audio score and the corresponding pronunciation information in the audio to be processed in order to reconstruct the pitch correction template specifically includes: for each audio data pair, constructing a reference information set based on the original audio score; identifying the pronunciation information in the audio to be processed; matching and aligning the reference information set with the pronunciation information to obtain the mapping relationship between the original audio score and the corresponding pronunciation information in the audio to be processed, and reconstructing the pitch correction template. The process of identifying pronunciation information in the audio to be processed includes: using a pre-trained speech recognition model based on audio data to identify initials and finals in the audio to be processed, obtaining the location and pronunciation tone of the initials and finals, identifying text content in the audio to be processed to obtain paragraph information of the audio to be processed, constructing a mapping relationship with paragraphs as units, and using the pronunciation tone of the initials and finals to adjust the tone of the audio to be processed. The reconstruction to obtain the pitch correction template includes: constructing a reference information set based on the original audio score; calculating substitute notes at each point in the original audio score based on the reference information set to achieve score reconstruction; calculating the DTW distance between the substitute notes at each point and the initials and finals in the audio to be processed, and using the mapping method with the minimum sum of DTW distances in the entire audio to be processed as the pitch correction template to reconstruct the pitch correction template, wherein the pitch correction template includes the position mapping relationship and pitch mapping relationship between the original audio and the audio to be processed; Multiple audio files to be mixed are mixed to obtain the initial chorus audio.
2. The method for multi-person chorus singing according to claim 1, characterized in that, Before constructing audio data pairs by combining each audio file from the multiple audio files to be processed with its corresponding original audio file, the process also includes: In response to a user-triggered chorus generation request, the system retrieves singing videos of multiple users performing the same song via the chorus terminal. Each user's singing video represents a segment of the current chorus piece. The positions of each user's singing content within the chorus are not entirely the same, and there is overlap in the singing content among the various users. Extract the corresponding audio from the singing videos of multiple users to serve as multiple audio files to be processed.
3. The method for multi-person chorus singing according to claim 1, characterized in that, The reference information set includes at least the melody, chords, paragraph structure, style classification, and harmonic information of the original audio score.
4. The method for multi-person chorus singing according to claim 1, characterized in that, After obtaining the initial chorus audio, it also includes: Extract the video corresponding to multiple audio files to be mixed. Arrange the videos corresponding to the multiple audio tracks to be mixed according to a preset layout. The mixing method based on the initial chorus audio plays video frames corresponding to multiple audios to be mixed, as well as the initial chorus audio.
5. The method for multi-person chorus singing according to claim 1, characterized in that, Based on the mapping relationship in the audio editing template, after adjusting the pitch and tempo of the audio to be processed to obtain the initial audio to be mixed, the following steps are also included: The initial audio to be mixed is input into the smoothing module to output a fine-grained audio sequence. The audio to be processed is input into the noise reduction and filtering module to obtain the reference audio to be processed. The small-grained audio sequence and the reference audio to be processed are input into the speed and pitch shifting module to obtain the audio to be mixed.
6. A multi-person chorus device, characterized in that, The multi-person chorus device specifically includes: A pairing module is used to form audio data pairs by pairing each audio to be processed with its corresponding original audio, wherein the original audio is the original audio of the song sung by the singers in a multi-person chorus method; after forming audio data pairs by pairing each audio to be processed with its corresponding original audio, the module further includes: Noise reduction and echo cancellation are performed on each audio file to be processed; Noise reduction is performed on each audio file, including expanding the number of sub-bands of the noise reduction module based on the frequency range of burst noise, and setting up hidden layers in multiple dimensions to construct an improved noise reduction module. The improved noise reduction module filters out burst noise and continuous noise whose duration reaches a threshold. Echo cancellation is performed on the background music of each audio file based on a neural network model. The audio correction module is used to perform audio correction on each audio data pair based on the original audio score and the pronunciation information in the audio to be processed, so as to obtain multiple audio to be mixed corresponding to the audio to be processed. The pitch correction module is configured to construct a mapping relationship between the original audio score and the corresponding pronunciation information in the audio to be processed, so as to reconstruct the pitch correction template, and process the speed and pitch of the audio to be processed based on the reconstructed pitch correction template. The process of constructing a mapping relationship between the original audio score and the corresponding pronunciation information in the audio to be processed in order to reconstruct the pitch correction template specifically includes: for each audio data pair, constructing a reference information set based on the original audio score; identifying the pronunciation information in the audio to be processed; matching and aligning the reference information set with the pronunciation information to obtain the mapping relationship between the original audio score and the corresponding pronunciation information in the audio to be processed, and reconstructing the pitch correction template. The process of identifying pronunciation information in the audio to be processed includes: using a pre-trained speech recognition model based on audio data to identify initials and finals in the audio to be processed, obtaining the location and pronunciation tone of the initials and finals, identifying text content in the audio to be processed to obtain paragraph information of the audio to be processed, constructing a mapping relationship with paragraphs as units, and using the pronunciation tone of the initials and finals to adjust the tone of the audio to be processed. The reconstruction to obtain the pitch correction template includes: constructing a reference information set based on the original audio score; calculating substitute notes at each point in the original audio score based on the reference information set to achieve score reconstruction; calculating the DTW distance between the substitute notes at each point and the initials and finals in the audio to be processed, and using the mapping method with the minimum sum of DTW distances in the entire audio to be processed as the pitch correction template to reconstruct the pitch correction template, wherein the pitch correction template includes the position mapping relationship and pitch mapping relationship between the original audio and the audio to be processed; The mixing module is used to mix multiple audio files to obtain the initial chorus audio.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that... When the program is executed by the processor, it implements the steps of the multi-person chorus method according to any one of claims 1-5.
Citation Information
Patent Citations
A device for eliminating echo of mobile terminal
CN101262530A
Audio correction method, device and equipment, and storage medium
CN111383620A
Intelligent chorusing method and device
CN112489610A
Chorus processing method, server, terminal, system and storage medium
CN116206584A
Karaoke apparatus
US20100192753A1