Audio processing method and device, terminal equipment and computer readable storage medium
By evaluating and selecting the target audio data with the best audio quality, and using intelligent synthesis nodes to generate simulated sounds that match the user's timbre, the problem of poor sound presentation caused by differences in device and user sound quality in karaoke applications is solved, achieving better audio editing results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HISENSE VISUAL TECH CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-21
AI Technical Summary
Existing karaoke apps cannot adapt to the differences in sound quality between different devices and users, resulting in poor sound presentation.
By acquiring microphone audio, the audio quality of each processing node is evaluated, the target audio data with the best audio quality is selected, and a simulated sound matching the user's timbre is generated using an intelligent synthesis node. The audio data quality is then evaluated using multiple scoring rules.
It improves audio editing effects, adapts to the audio needs of different devices and users, and enhances the quality of audio data and singing performance.
Smart Images

Figure CN121905128A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of audio processing technology, and in particular relates to an audio processing method, an audio processing device, a terminal device, and a computer-readable storage medium. Background Technology
[0002] With the development of smart terminals, the application services they provide are also increasing. For example, some smart terminals can provide karaoke services, allowing users to enjoy karaoke anytime, anywhere. To ensure the singing effect, current karaoke applications usually modify the microphone audio to make the sound from the microphone pleasant and easy to hear.
[0003] In practical applications, due to differences in device sound quality (such as different microphones) and differences in user timbre, using the same sound correction scheme will not be able to adapt to the sound correction needs of different devices and different users, thus affecting the final sound presentation effect. Summary of the Invention
[0004] This application provides an audio processing method, an audio processing apparatus, a terminal device, and a computer-readable storage medium, which can effectively improve audio editing results.
[0005] In a first aspect, embodiments of this application provide an audio processing method, including: Acquire microphone audio; wherein, the microphone audio is the audio captured by the microphone used by the user when singing the target song; Acquire audio data corresponding to at least one processing node during the audio processing of the microphone audio; Calculate the quality score of the audio data corresponding to each of the processing nodes; Based on the quality score, target audio data is selected from the audio data corresponding to at least one processing node. Audio editing is performed based on the target audio data.
[0006] In this embodiment, the audio quality of audio data at each processing node during microphone audio processing can be automatically evaluated to select the target audio data with the best audio quality from multiple processing nodes, providing high-quality audio data for subsequent audio editing. For different devices and users, the above method can adaptively acquire high-quality audio data for audio editing, thus improving the audio editing effect.
[0007] In one possible implementation of the first aspect, the at least one processing node includes an intelligent synthesis node for generating a user's simulated voice. Accordingly, the process of acquiring the audio data corresponding to the intelligent synthesis node includes: Obtain the original audio of the target song; wherein, the original audio includes the accompaniment and the original vocals; The original vocal audio is subjected to sound separation processing to obtain the accompaniment and the original vocal audio; A simulated sound matching the user's timbre is generated based on the original vocals; wherein, the audio data corresponding to the intelligent synthesis node includes the accompaniment and the simulated sound.
[0008] In the above method, the simulated sound obtained through the processing of the intelligent synthesis node not only matches the user's timbre, but also retains the original rhythm and key matching of the song to the greatest extent, avoiding the problems of pitch and rhythm misalignment that occur when manually matching accompaniment, and effectively improving the user's singing effect.
[0009] In one possible implementation of the first aspect, generating an analog sound that matches the user's timbre based on the original vocal recording includes: Obtain a sound model that matches the user's voice timbre; wherein, the sound model is a model trained based on sample audio, and the sample audio matches the user's voice timbre; The original vocals are input into the sound model, and the simulated vocals are output.
[0010] In the above method, generating the user's simulated voice through a pre-trained sound model can improve the efficiency of simulated voice generation. In addition, since the sound model is trained based on sample audio that matches the user's voice, the simulated voice can not only highly reproduce the timbre of the user's voice, but also inherit the rhythm, emotion and technique of the original singer's voice, thereby improving the effect of the simulated voice.
[0011] In one possible implementation of the first aspect, calculating the quality score of the audio data corresponding to each processing node includes: The evaluation score of the audio data corresponding to the processing node is calculated based on multiple scoring rules. The quality score of the audio data corresponding to the processing node is calculated based on the evaluation score corresponding to each of the scoring rules.
[0012] In the above method, the quality of audio data is evaluated based on multiple different scoring rules, which is equivalent to evaluating audio quality from different angles / dimensions. Compared with the method of using a single scoring rule for quality evaluation, it can reduce the scoring error caused by the "subjectivity" of a single scoring rule and help improve the accuracy of quality scoring.
[0013] In one possible implementation of the first aspect, the plurality of scoring rules include a first rule, a second rule, and a third rule; The process of calculating the evaluation score of the audio data corresponding to the processing node based on multiple scoring rules includes: The evaluation score of the audio data corresponding to the processing node is calculated according to the first rule to obtain the first score; wherein, the first rule is used to evaluate the similarity between the audio data and the original singer's voice of the target song; The evaluation score of the audio data corresponding to the processing node is calculated according to the second rule to obtain the second score; wherein, the second rule is used to evaluate the sound quality of the audio data; The evaluation score of the audio data corresponding to the processing node is calculated according to the third rule to obtain the third score; wherein, the third rule is used to evaluate the pronunciation quality of human voice in the audio data.
[0014] The above method is equivalent to evaluating audio data from three perspectives: similarity to the original singer, sound quality, and pronunciation. This makes the quality evaluation of audio data more comprehensive, and the resulting quality score can more accurately reflect the quality of the audio data.
[0015] In one possible implementation of the first aspect, after acquiring the microphone audio, the method further includes: Detect user input selection information; Accordingly, the audio data corresponding to at least one processing node in the audio processing of acquiring the microphone audio includes: When the selection information indicates intelligent selection, audio data corresponding to at least one processing node in the audio processing of the microphone audio is obtained.
[0016] In one possible implementation of the first aspect, after detecting the user-input selection information, the method further includes: If the selection information indicates a non-intelligent selection, the processing node indicated by the selection information is obtained to obtain the target node; During the audio processing of the microphone audio, the audio data corresponding to the target node is obtained to obtain the target audio data; Audio editing is performed based on the target audio data.
[0017] In the above methods, users can choose audio data from any audio processing node, or the smart terminal can automatically identify the audio data from the node with the best quality through intelligent selection. These methods are more flexible, can meet different user needs, and help improve the user experience.
[0018] Secondly, embodiments of this application provide an audio processing apparatus, including: An audio acquisition unit is used to acquire microphone audio; wherein the microphone audio is the audio captured by the microphone used by the user when singing the target song; The data acquisition unit is used to acquire the audio data corresponding to at least one processing node in the audio processing of the microphone audio. A scoring calculation unit is used to calculate the quality score of the audio data corresponding to each processing node; A data selection unit is used to select target audio data from the audio data corresponding to at least one processing node based on the quality score. An audio editing unit is used to perform audio editing based on the target audio data.
[0019] Thirdly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the audio processing method as described in any one of the first aspects above.
[0020] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the audio processing method as described in any one of the first aspects above.
[0021] Fifthly, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute the audio processing method described in any one of the first aspects.
[0022] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a schematic flowchart of the audio processing method provided in the embodiments of this application; Figure 2 This is a schematic diagram of the processing flow of the intelligent synthesis node provided in the embodiments of this application; Figure 3 This is an example of the audio processing flow provided in the embodiments of this application; Figure 4 This is a flowchart illustrating an audio processing method provided in another embodiment of this application; Figure 5 This is a schematic diagram of the user interface provided in an embodiment of this application; Figure 6 This is a structural block diagram of the audio processing apparatus provided in the embodiments of this application; Figure 7 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application. Detailed Implementation
[0025] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0026] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0027] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0028] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0029] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0030] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.
[0031] With the development of smart terminals, the application services they provide are also increasing. For example, some smart terminals can provide karaoke services, allowing users to enjoy karaoke anytime, anywhere. To ensure the singing effect, current karaoke applications usually modify the microphone audio to make the sound from the microphone pleasant and easy to hear.
[0032] In practical applications, due to differences in device sound quality (such as different microphones) and differences in user timbre, using the same sound correction scheme will not be able to adapt to the sound correction needs of different devices and different users, thus affecting the final sound presentation effect.
[0033] Based on this, embodiments of this application provide an audio processing method. In these embodiments, the audio quality of audio data at each processing node during microphone audio processing can be automatically evaluated to select the target audio data with the best audio quality from the audio data of multiple processing nodes. This provides high-quality audio data for subsequent audio editing, thereby improving the audio editing effect.
[0034] See Figure 1 This is a schematic flowchart of an audio processing method provided in an embodiment of this application. It is intended as an example and not a limitation. The method may include the following steps: S101, acquire microphone audio.
[0035] Among them, microphone audio refers to the audio captured by the microphone used by the user when singing the target song.
[0036] In some application scenarios, a smart terminal is connected to a microphone, and a karaoke app is installed on the smart terminal. The user selects a song on the karaoke app, and the smart terminal plays the accompaniment to the selected song. The user sings into the microphone, and the microphone captures the audio while the user is singing. The smart terminal can acquire the microphone audio in real time or after the performance ends.
[0037] S102, acquire the audio data corresponding to at least one processing node in the audio processing of microphone audio.
[0038] In this embodiment, at least one audio processing step can be performed on the microphone audio before audio editing, such as a preprocessing node, an audio effect processing node, a volume processing node, etc. For example, the preprocessing node is used to reduce noise and prevent feedback in the audio. The audio effect processing node is used to add audio effects (such as KTV effects, concert effects, etc.) to the audio. The volume processing node is used to adjust the audio volume. It can be understood that the audio data corresponding to each processing node refers to the audio data after being processed by that processing node. For example, the audio data corresponding to the preprocessing node refers to the audio data after noise reduction and feedback prevention processing; the audio data corresponding to the audio effect processing node refers to the audio data after adding audio effects; and the audio data corresponding to the volume processing node refers to the audio data after adjusting the volume.
[0039] Optionally, multiple processing nodes can be processed serially. For example, microphone audio may pass through a preprocessing node, an audio effects processing node, and a volume processing node in sequence.
[0040] Optionally, the audio data corresponding to each processing node can be saved to a data pool for later selection.
[0041] In one embodiment, at least one processing node includes a smart synthesis node, which is used to generate the user's simulated voice. Accordingly, the process of acquiring the audio data corresponding to the smart synthesis node in S102 includes: Obtain the original audio of the target song; the original audio includes the instrumental and the original vocals. The original vocal audio is separated to obtain the accompaniment and the original vocal audio. A simulated sound matching the user's timbre is generated based on the original singer's voice; the audio data corresponding to the intelligent synthesis node includes accompaniment and simulated sound.
[0042] In the above method, the simulated sound obtained through the processing of the intelligent synthesis node not only matches the user's timbre, but also retains the original rhythm and key matching of the song to the greatest extent, avoiding the problems of pitch and rhythm misalignment that occur when manually matching accompaniment, and effectively improving the user's singing effect.
[0043] In one implementation, the process of generating an analog sound that matches the user's timbre based on the original vocals may include: performing sound separation processing on the microphone audio to obtain the accompaniment and the user's voice; and performing weighted calculations on the original vocals and the user's voice to obtain the analog sound.
[0044] In another implementation, the process of generating an analog sound that matches the user's timbre based on the original vocals may include: Obtain a sound model that matches the user's voice timbre; wherein, the sound model is a model trained based on sample audio that matches the user's voice timbre; The original vocals are input into the sound model, and the simulated sound is output.
[0045] For example, see Figure 2 This is a schematic diagram illustrating the processing flow of the intelligent synthesis node provided in an embodiment of this application. It is intended as an example and not a limitation. Figure 2 As shown, the sound model can be deployed on a cloud server, and the cloud server will execute the training and application processes of the sound model accordingly.
[0046] Specifically, during training, the karaoke app (i.e., the terminal side) collects sample audio and uploads it to a cloud server. The cloud server preprocesses the sample audio (e.g., noise reduction, audio segmentation), and then trains a sound model based on the preprocessed sample audio, resulting in a trained sound model. During application, the karaoke app (i.e., the terminal side) receives the user's song request, obtains the target song, and sends the target song's information (e.g., song name or identifier, number, etc.) to the cloud server. Correspondingly, the cloud server acquires the target song's resources, performs sound separation processing on these resources, and obtains the accompaniment and original vocals. The original vocals are then input into the trained sound model, outputting a simulated sound, which is then returned to the karaoke app along with the accompaniment. The karaoke app can save the simulated sound and accompaniment separately, or it can combine the simulated sound and accompaniment into a single song before saving it.
[0047] In some implementations, users can choose whether to perform intelligent synthesis. For example, a user in a karaoke app can choose whether to perform intelligent synthesis; the karaoke app sends the user's selection command to the cloud server; correspondingly, the cloud server receives the user's selection command; if the user selects intelligent synthesis, the cloud server executes the processing flow of the aforementioned intelligent synthesis node; if the user does not select intelligent synthesis, the cloud server can simply perform the step of sound separation processing on the original vocal audio and send the resulting accompaniment to the karaoke app.
[0048] It's understandable that "matching" here can refer to having the same timbre as the user's voice, a similar timbre, or belonging to the same timbre type (e.g., both being male, female, or child voices). For example, a cloud server can store multiple voice models corresponding to each user. During application, the cloud server selects the voice model corresponding to the user's voice in the microphone audio and trains the sample audio of that voice model to belong to the same user. As another example, the cloud server can classify users based on timbre and train a voice model corresponding to each user category based on sample audio from each category. During application, the cloud server selects the voice model that matches the user's voice in the microphone audio and trains the sample audio of that voice model to belong to the same user category.
[0049] In the above method, generating the user's simulated voice through a pre-trained sound model can improve the efficiency of simulated voice generation. In addition, since the sound model is trained based on sample audio that matches the user's voice, the simulated voice can not only highly reproduce the timbre of the user's voice, but also inherit the rhythm, emotion and technique of the original singer's voice, thereby improving the effect of the simulated voice.
[0050] S103, calculate the quality score of the audio data corresponding to each processing node.
[0051] In one embodiment, S103 may include: The evaluation score of the audio data corresponding to the processing node is calculated based on multiple scoring rules. The quality score of the audio data corresponding to each processing node is calculated based on the evaluation score corresponding to each scoring rule.
[0052] In the above method, the quality of audio data is evaluated based on multiple different scoring rules, which is equivalent to evaluating audio quality from different angles / dimensions. Compared with the method of using a single scoring rule for quality evaluation, it can reduce the scoring error caused by the "subjectivity" of a single scoring rule and help improve the accuracy of quality scoring.
[0053] In one implementation, multiple scoring rules include a first rule, a second rule, and a third rule. Accordingly, the evaluation score of the audio data corresponding to the processing node is calculated based on each of the multiple scoring rules, including: The evaluation score of the audio data corresponding to the processing node is calculated according to the first rule to obtain the first score; wherein, the first rule is used to evaluate the similarity between the audio data and the original singer's voice of the target song; The evaluation score of the audio data corresponding to the processing node is calculated according to the second rule to obtain the second score; wherein, the second rule is used to evaluate the sound quality of the audio data; The evaluation score of the audio data corresponding to the processing node is calculated according to the third rule to obtain the third score; the third rule is used to evaluate the pronunciation quality of human voice in the audio data.
[0054] For example, the first rule could be to calculate the similarity between the vocals in the audio data and the original vocals of the target song, obtaining a first score. The second rule's sound quality could include parameters such as vocal volume, clarity, smoothness, reverberation, noise / dropping, feedback, and popping sounds; the second rule could be to calculate the weighted sum of the values for each sound quality parameter, obtaining a second score. The third rule's pronunciation quality could include parameters such as pronunciation accuracy, articulation clarity, and lyric matching; the third rule could be to calculate the weighted sum of the values for each pronunciation quality parameter, obtaining a third score.
[0055] The above method is equivalent to evaluating audio data from three perspectives: similarity to the original singer, sound quality, and pronunciation. This makes the quality evaluation of audio data more comprehensive, and the resulting quality score can more accurately reflect the quality of the audio data.
[0056] Optionally, a weighted sum of the first, second, and third scores can be calculated to obtain the quality score of the audio data. The weight of each score can be set according to the importance of each rule. Alternatively, the maximum / minimum value among the first, second, and third scores can be used as the quality score of the audio data.
[0057] S104, select the target audio data from the audio data corresponding to at least one processing node based on the quality score.
[0058] Optionally, the audio data corresponding to the highest quality score can be used as the target audio data.
[0059] For example, see Figure 3 This is an example of an audio processing flow provided in an embodiment of this application. It is intended as an example and not a limitation. Figure 3As shown, the microphone audio passes through three nodes: a preprocessing node, a sound effects processing node, and a smart synthesis node. Accordingly, the evaluation scores of the microphone audio are calculated using the first, second, and third rules, respectively, yielding a first score G1, a second score H1, and a third score V1. The first score G1, second score H1, and third score V1 are then weighted and summed to obtain the microphone audio quality score N1. Similarly, the evaluation scores of the audio data corresponding to the preprocessing node are calculated using the first, second, and third rules, yielding a first score G2, a second score H2, and a third score V2. The first score G2, second score H2, and third score V2 are then weighted and summed to obtain the audio data quality score N2. Finally, the evaluation scores of the audio data corresponding to the sound effects processing node are calculated using the first, second, and third rules, yielding a first score G3, a second score H2, and a third score V2. H3 and the third score V3 are used to calculate the quality score N3 of the audio data corresponding to the audio effect processing node by weighted summing of the first score G3, the second score H3, and the third score V3. The evaluation scores of the audio data corresponding to the intelligent synthesis node are calculated according to the first rule, the second rule, and the third rule, respectively, to obtain the first score G4, the second score H4, and the third score V4 of the audio data corresponding to the intelligent synthesis node. The first score G4, the second score H4, and the third score V4 of the audio data corresponding to the intelligent synthesis node are weighted summed to obtain the quality score N4 of the audio data corresponding to the intelligent synthesis node. Then, the audio data corresponding to the maximum value among the microphone audio quality score N1, the audio data quality score N2 corresponding to the preprocessing node, the audio data quality score N3 corresponding to the audio effect processing node, and the audio data quality score N4 corresponding to the intelligent synthesis node is used as the target audio data.
[0060] S105, performs audio editing based on the target audio data.
[0061] In this embodiment, the audio quality of audio data at each processing node during microphone audio processing can be automatically evaluated to select the target audio data with the best audio quality from multiple processing nodes, providing high-quality audio data for subsequent audio editing. For different devices and users, the above method can adaptively acquire high-quality audio data for audio editing, thus improving the audio editing effect.
[0062] See Figure 4 This is a schematic flowchart of an audio processing method provided in another embodiment of this application. It is intended as an example and not a limitation. Figure 4 As shown, the audio processing method may include the following steps: S401, acquire microphone audio.
[0063] S402, detect the selection information entered by the user.
[0064] If the selected information indicates intelligent selection, execute S403-S405 and S408; if the selected information indicates non-intelligent selection, execute S406-S408.
[0065] For example, see Figure 5 This is a schematic diagram of the user interface provided in an embodiment of this application. It is intended as an example and not a limitation. Figure 5 As shown, users can select audio data by interacting with the controls in front of each option (such as clicking, long-pressing, etc.). Figure 5 As shown in (a) above, the control before "Smart Selection" is selected, indicating that the selection information represents smart selection. Figure 5 As shown in (b), the control in front of "Sound Effects Processing" is selected, indicating that the selection information is not intelligent selection.
[0066] Optionally, the user interface can also provide options for saving methods. For example... Figure 5 As shown, users can choose to "save to local" or "upload to cloud".
[0067] S403, when the selection information indicates intelligent selection, acquire the audio data corresponding to at least one processing node in the audio processing of the microphone audio.
[0068] S404, calculate the quality score of the audio data corresponding to each processing node.
[0069] S405, select the target audio data from the audio data corresponding to at least one processing node based on the quality score.
[0070] The above steps S401 and Figure 1 Step S101 in the embodiment is the same, and steps S403-S405 above are the same as those in the embodiment. Figure 1 Steps S102-S104 in the embodiments are the same, and can be found in the description of steps S101-S104 in the embodiments, which will not be repeated here.
[0071] S406, if the selection information indicates a non-intelligent selection, obtain the processing node indicated by the selection information to obtain the target node.
[0072] S407, acquire the audio data corresponding to the target node during the audio processing of the microphone audio, and obtain the target audio data.
[0073] continue Figure 5 Examples, such as Figure 5As shown in (b), the current selection is non-smart, and the control in front of "sound effect processing" is selected, indicating that the user has selected to use the audio data corresponding to the sound effect processing node. In this case, the smart terminal will use the sound effect processing node as the target node, and correspondingly, use the audio data corresponding to the sound effect processing node as the target audio data.
[0074] S408 performs audio editing based on the target audio data.
[0075] Figure 4 In the illustrated embodiment, users can independently select audio data from any audio processing node, or the smart terminal can automatically identify the audio data from the node with the highest quality through intelligent selection. This approach is more flexible, can meet different user needs, and improves the user experience.
[0076] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0077] Corresponding to the audio processing method described in the above embodiments, Figure 6 This is a structural block diagram of the audio processing apparatus provided in the embodiments of this application. For ease of explanation, only the parts related to the embodiments of this application are shown.
[0078] Reference Figure 6 The device 6 includes: The audio acquisition unit 61 is used to acquire microphone audio; wherein the microphone audio is the audio collected by the microphone used by the user when singing the target song.
[0079] The data acquisition unit 62 is used to acquire the audio data corresponding to at least one processing node in the audio processing of the microphone audio.
[0080] The scoring calculation unit 63 is used to calculate the quality score of the audio data corresponding to each processing node.
[0081] The data selection unit 64 is used to select target audio data from the audio data corresponding to at least one processing node based on the quality score.
[0082] The audio editing unit 65 is used to perform audio editing based on the target audio data.
[0083] Optionally, the at least one processing node includes a smart synthesis node for generating the user's simulated voice.
[0084] Accordingly, the data acquisition unit 62 is also used for: Obtain the original audio of the target song; wherein, the original audio includes the accompaniment and the original vocals; The original vocal audio is subjected to sound separation processing to obtain the accompaniment and the original vocal audio; A simulated sound matching the user's timbre is generated based on the original vocals; wherein, the audio data corresponding to the intelligent synthesis node includes the accompaniment and the simulated sound.
[0085] Optionally, the data acquisition unit 62 is also used for: Obtain a sound model that matches the user's voice timbre; wherein, the sound model is a model trained based on sample audio, and the sample audio matches the user's voice timbre; The original vocals are input into the sound model, and the simulated vocals are output.
[0086] Optionally, the scoring calculation unit 63 is also used for: The evaluation score of the audio data corresponding to the processing node is calculated based on multiple scoring rules. The quality score of the audio data corresponding to the processing node is calculated based on the evaluation score corresponding to each of the scoring rules.
[0087] Optionally, the plurality of scoring rules include a first rule, a second rule, and a third rule.
[0088] Correspondingly, the scoring calculation unit 63 is also used for: The evaluation score of the audio data corresponding to the processing node is calculated according to the first rule to obtain the first score; wherein, the first rule is used to evaluate the similarity between the audio data and the original singer's voice of the target song; The evaluation score of the audio data corresponding to the processing node is calculated according to the second rule to obtain the second score; wherein, the second rule is used to evaluate the sound quality of the audio data; The evaluation score of the audio data corresponding to the processing node is calculated according to the third rule to obtain the third score; wherein, the third rule is used to evaluate the pronunciation quality of human voice in the audio data.
[0089] Optionally, device 6 also includes: The detection unit 66 is used to detect the user's selection information after acquiring the microphone audio.
[0090] Correspondingly, the data acquisition unit 62 is also used to acquire audio data corresponding to at least one processing node in the audio processing of the microphone audio when the selection information indicates intelligent selection.
[0091] Optionally, the data acquisition unit 62 is also used for: When the selection information indicates a non-intelligent selection, the processing node indicated by the selection information is obtained to obtain the target node; the audio data corresponding to the target node during the audio processing of the microphone audio is obtained to obtain the target audio data.
[0092] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0093] in addition, Figure 6 The audio processing device shown can be a software unit, a hardware unit, or a combination of software and hardware built into an existing terminal device, or it can be integrated into the terminal device as a separate component, or it can exist as a standalone terminal device.
[0094] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0095] Figure 7 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application. For example... Figure 7 As shown, the terminal device 7 of this embodiment includes: at least one processor 70 ( Figure 7 (Only one is shown in the diagram) a processor, a memory 71, and a computer program 72 stored in the memory 71 and executable on the at least one processor 70, wherein the processor 70 executes the computer program 72 to implement the steps in any of the above-described audio processing method embodiments.
[0096] The terminal device may be a desktop computer, laptop, handheld computer, or cloud server, etc. This terminal device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 7This is merely an example of terminal device 7 and does not constitute a limitation on terminal device 7. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.
[0097] The processor 70 can be a Central Processing Unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0098] In some embodiments, the memory 71 may be an internal storage unit of the terminal device 7, such as a hard disk or memory of the terminal device 7. In other embodiments, the memory 71 may be an external storage device of the terminal device 7, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the terminal device 7. Furthermore, the memory 71 may include both internal and external storage units of the terminal device 7. The memory 71 is used to store the operating system, applications, boot loader, data, and other programs, such as the program code of the computer program. The memory 71 can also be used to temporarily store data that has been output or will be output.
[0099] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps in the above-described method embodiments.
[0100] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments.
[0101] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / terminal equipment, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0102] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0103] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0104] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0105] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0106] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. An audio processing method, characterized in that, include: Acquire microphone audio; wherein, the microphone audio is the audio captured by the microphone used by the user when singing the target song; Acquire audio data corresponding to at least one processing node during the audio processing of the microphone audio; Calculate the quality score of the audio data corresponding to each of the processing nodes; Based on the quality score, target audio data is selected from the audio data corresponding to at least one processing node. Audio editing is performed based on the target audio data.
2. The audio processing method as described in claim 1, characterized in that, The at least one processing node includes an intelligent synthesis node, which is used to generate the user's simulated voice. Accordingly, the process of acquiring the audio data corresponding to the intelligent synthesis node includes: Obtain the original audio of the target song; wherein, the original audio includes the accompaniment and the original vocals; The original vocal audio is subjected to sound separation processing to obtain the accompaniment and the original vocal audio; A simulated sound matching the user's timbre is generated based on the original vocals; wherein, the audio data corresponding to the intelligent synthesis node includes the accompaniment and the simulated sound.
3. The audio processing method as described in claim 2, characterized in that, The step of generating a simulated sound that matches the user's timbre based on the original vocals includes: Obtain a sound model that matches the user's voice timbre; wherein the sound model is a model trained based on sample audio, and the sample audio matches the user's voice timbre; The original vocals are input into the sound model, and the simulated vocals are output.
4. The audio processing method as described in claim 1, characterized in that, The calculation of the quality score of the audio data corresponding to each processing node includes: The evaluation score of the audio data corresponding to the processing node is calculated based on multiple scoring rules. The quality score of the audio data corresponding to the processing node is calculated based on the evaluation score corresponding to each of the scoring rules.
5. The audio processing method as described in claim 4, characterized in that, The multiple scoring rules include the first rule, the second rule, and the third rule; The process of calculating the evaluation score of the audio data corresponding to the processing node based on multiple scoring rules includes: The evaluation score of the audio data corresponding to the processing node is calculated according to the first rule to obtain the first score; wherein, the first rule is used to evaluate the similarity between the audio data and the original singer's voice of the target song; The evaluation score of the audio data corresponding to the processing node is calculated according to the second rule to obtain the second score; wherein, the second rule is used to evaluate the sound quality of the audio data; The evaluation score of the audio data corresponding to the processing node is calculated according to the third rule to obtain the third score; wherein, the third rule is used to evaluate the pronunciation quality of human voice in the audio data.
6. The audio processing method according to any one of claims 1 to 5, characterized in that, After acquiring the microphone audio, the method further includes: Detect user input selection information; Accordingly, the audio data corresponding to at least one processing node in the audio processing of acquiring the microphone audio includes: When the selection information indicates intelligent selection, audio data corresponding to at least one processing node in the audio processing of the microphone audio is obtained.
7. The audio processing method as described in claim 6, characterized in that, After detecting the user's selection input, the method further includes: If the selection information indicates a non-intelligent selection, the processing node indicated by the selection information is obtained to obtain the target node; During the audio processing of the microphone audio, the audio data corresponding to the target node is obtained to obtain the target audio data; Audio editing is performed based on the target audio data.
8. An audio processing apparatus, characterized in that, include: An audio acquisition unit is used to acquire microphone audio; wherein the microphone audio is the audio captured by the microphone used by the user when singing the target song; The data acquisition unit is used to acquire the audio data corresponding to at least one processing node in the audio processing of the microphone audio. A scoring calculation unit is used to calculate the quality score of the audio data corresponding to each processing node; A data selection unit is used to select target audio data from the audio data corresponding to at least one processing node based on the quality score. An audio editing unit is used to perform audio editing based on the target audio data.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.