Audio production method, device, terminal equipment and readable storage medium

By acquiring and synthesizing the user's voice feature parameters and generating singing audio that meets the user's voice features, the problem of users being unable to sing or having poor singing quality is solved, and the singing quality and user experience are improved.

CN114974184BActive Publication Date: 2025-09-19MIGU MUSIC CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210563372.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-09-19
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

In the existing way of singing songs, users may not be able to sing or the singing quality may be poor, resulting in a poor user experience.

Method used

By obtaining the voice characteristic parameters of the first user, including voice line, timbre, pitch and range, the singing audio is synthesized based on these parameters to generate singing audio that conforms to the user's voice characteristics.

Benefits of technology

The singing quality and user experience are improved, and the generated singing audio conforms to the user's own voice characteristics, which improves the singing effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114974184B_ABST
    Figure CN114974184B_ABST
Patent Text Reader

Abstract

The present invention discloses an audio production method, apparatus, terminal device and readable storage medium, the method comprising: obtaining a first user's singing audio for a target song; obtaining the first user's voice characteristic parameters, the voice characteristic parameters including at least one of a voice line parameter, a timbre parameter, a pitch parameter and a range parameter, and the voice characteristic parameters are obtained based on the audio data corresponding to the first user when reciting or singing the target content of the target song; synthesizing the singing audio according to the voice characteristic parameters to obtain the first user's singing audio for the target song. The method of the present invention synthesizes the singing audio according to the first user's voice characteristic parameters to obtain the first user's singing audio for the target song, so that the obtained first user's singing audio for the target song has better singing quality, thereby achieving the purpose of improving the first user's singing quality of the target song while improving the user's singing experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to an audio production method, apparatus, terminal equipment and readable storage medium. Background Art

[0002] With the development and progress of society, people's entertainment is becoming increasingly diverse. Among them, listening to music and singing are the most popular entertainment activities. Current music software can be used to record music, play music, and allow users to fine-tune the audio to achieve the audio effects they desire. It can also be used to provide a singing platform for songs. When users sing songs, they can usually randomly select songs to sing. However, this method of singing songs may lead to a poor user experience because the user may not be able to sing the songs or the final singing quality of the songs may be poor.

[0003] The above content is only used to assist in understanding the technical solution of the present invention and does not constitute an admission that the above content is related technology. Summary of the Invention

[0004] The embodiments of the present invention aim to solve the technical problem of poor user experience caused by the existing method of singing songs, which is that the user may not be able to sing the song or the quality of the song finally sung is poor, by providing an audio production method, apparatus, terminal device and readable storage medium.

[0005] An embodiment of the present invention provides an audio production method, the audio production method comprising:

[0006] Obtaining the first user's singing audio of the target song;

[0007] Acquiring voice characteristic parameters of the first user, the voice characteristic parameters including at least one of a voice line parameter, a timbre parameter, a pitch parameter, and a range parameter, and the voice characteristic parameters are obtained based on audio data corresponding to when the first user recites or sings the target content of the target song;

[0008] The singing audio is synthesized according to the sound feature parameters to obtain the singing audio of the first user for the target song.

[0009] Optionally, the step of obtaining the voice characteristic parameters of the first user includes:

[0010] Obtaining keywords corresponding to each song section of the target song, and outputting the keywords, wherein the target content of the target song includes the keywords;

[0011] Acquire audio data corresponding to the keyword by the first user;

[0012] The audio data is identified to obtain voice characteristic parameters of the first user.

[0013] Optionally, after the step of synthesizing the singing audio according to the sound feature parameters to obtain the singing audio of the target song by the first user, the method further includes:

[0014] Determining a chorus section in the target song that matches the sound feature parameters based on a matching degree between the audio feature parameters corresponding to each song section of the target song and the sound feature parameters;

[0015] Generating a first sub-audio corresponding to the chorus segment according to the singing audio;

[0016] Generating second sub-audio corresponding to other song sections other than the chorus section based on the singing audio associated with the second user;

[0017] A chorus song audio is generated based on the first sub audio and the second sub audio.

[0018] Optionally, the step of generating a first sub-audio corresponding to the chorus section according to the singing audio includes:

[0019] Acquire audio data corresponding to the vocals of the chorus section in the performance audio to generate a first sub-audio corresponding to the chorus section; or,

[0020] The audio data corresponding to other song sections except the chorus section is deleted from the performance audio to generate a first sub-audio corresponding to the chorus section.

[0021] Optionally, the audio production method further includes:

[0022] determining the second user who sings the target song with the first user;

[0023] When the second user has sung the target song, the step of generating second sub-audio corresponding to other song sections except the chorus section based on the singing audio associated with the second user is executed.

[0024] Optionally, after the step of determining the second user who sings the target song with the first user, the method further includes:

[0025] When the second user does not sing the target song, obtaining the voice feature parameters of the second user;

[0026] generating, based on the voice characteristic parameters of the second user and the singing audio corresponding to the target song, the singing audio of the second user for the target song;

[0027] The second user is associated with the singing audio.

[0028] Optionally, the singing audio associated with the second user is the audio recorded when the second user sings the target song.

[0029] In addition, to achieve the above-mentioned object, the present invention further provides an audio production device, the audio production device comprising:

[0030] A first acquisition module is used to acquire the singing audio of the first user for the target song;

[0031] a second acquisition module, configured to acquire voice characteristic parameters of the first user, the voice characteristic parameters including at least one of a voice line parameter, a timbre parameter, a pitch parameter, and a range parameter, and the voice characteristic parameters being acquired based on audio data corresponding to the first user reciting or singing the target content of the target song;

[0032] An audio synthesis module is used to synthesize the singing audio according to the sound feature parameters to obtain the singing audio of the first user for the target song.

[0033] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal device, which includes: a memory, a processor, and an audio production program stored on the memory and runnable on the processor, and the audio production program implements the steps of the above-mentioned audio production method when executed by the processor.

[0034] In addition, to achieve the above-mentioned purpose, the present invention also provides a readable storage medium, which stores an audio production program. When the audio production program is executed by a processor, the steps of the above-mentioned audio production method are implemented.

[0035] An audio production method, apparatus, terminal device and readable storage medium provided in an embodiment of the present invention obtain the singing audio of the first user for a target song by synthesizing the singing audio according to the voice characteristic parameters of the first user. The method can adapt to the voice characteristic parameters of the first user and adjust and synthesize the singing audio of the first user singing the target song with his own voice, so that the obtained singing audio of the first user for the target song conforms to the voice characteristics of the first user's own voice, improves the singing audio of the first user singing the target song with his own voice, and makes the final singing audio have better singing quality, thereby achieving the purpose of improving the singing quality of the first user singing the target song while improving the user's singing experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 Schematic diagram of the structure of the terminal device involved in each embodiment of the audio production method of the present invention;

[0037] Figure 2 1 is a flow chart of a first embodiment of an audio production method according to the present invention;

[0038] Figure 3 Schematic diagram of the process of obtaining sound characteristic parameters in the first embodiment of the audio production method of the present invention;

[0039] Figure 4 is a flow chart of a second embodiment of the audio production method of the present invention;

[0040] Figure 5 is a flow chart of a second embodiment of the audio production method of the present invention;

[0041] Figure 6 This is a schematic diagram of the module composition of the audio production device provided by the present invention. DETAILED DESCRIPTION

[0042] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0043] In the subsequent description, the suffixes such as "module", "component" or "unit" used to represent elements are only used to facilitate the description of the present invention and have no specific meaning. Therefore, "module", "component" or "unit" can be used interchangeably.

[0044] The present invention provides an audio production method, the audio production method comprising:

[0045] Obtaining the first user's singing audio of the target song;

[0046] Acquiring voice characteristic parameters of the first user, the voice characteristic parameters including at least one of a voice line parameter, a timbre parameter, a pitch parameter, and a range parameter, and the voice characteristic parameters are obtained based on audio data corresponding to when the first user recites or sings the target content of the target song;

[0047] The singing audio is synthesized according to the sound feature parameters to obtain the singing audio of the first user for the target song.

[0048] The audio production method of the present invention synthesizes the singing audio according to the voice characteristic parameters of the first user to obtain the singing audio of the first user for the target song. The audio production method can adapt to the voice characteristic parameters of the first user and adjust and synthesize the singing audio of the first user singing the target song with his own voice, so that the obtained singing audio of the first user for the target song conforms to the voice characteristics of the first user's own voice, improves the singing audio of the first user singing the target song with his own voice, and makes the final singing audio have better singing quality, thereby achieving the purpose of improving the singing quality of the first user singing the target song while improving the user's singing experience.

[0049] Please refer to Figure 1 , Figure 1 Schematic diagram of the structure of the terminal device involved in each embodiment of the audio production method of the present invention. The terminal device involved in the audio production method of the present invention may include terminal devices such as mobile phones, tablet computers, laptops, PDAs, and personal digital assistants (PDAs).

[0050] like Figure 1 As shown, the terminal device may include: a memory 101 and a processor 102. Those skilled in the art will understand that Figure 1 The illustrated block diagram of the terminal structure does not limit the terminal; the terminal may include more or fewer components than shown, or may combine certain components or arrange the components differently. Memory 101 stores an operating device and an audio production program. Processor 102 is the control center of the terminal device. Processor 102 executes the audio production program stored in memory 101 to implement the steps of each embodiment of the audio production method of the present invention.

[0051] Optionally, the terminal device may further include a display unit 103, which includes a display panel. The display panel may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc., for outputting an interface for displaying the user browsing.

[0052] Optionally, the terminal device may further include a communication unit, which establishes data communication with other terminal devices such as computers through a network protocol (the data communication may be IP communication or a Bluetooth channel) to achieve data transmission between other terminal devices.

[0053] The embodiment of the present invention provides an embodiment of the audio production method. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in an order different from that shown here.

[0054] Based on the structural block diagram of the terminal device, various embodiments of the audio production method of the present invention are proposed. In the first embodiment, the present invention provides an audio production method, please refer to Figure 2 , Figure 2 1 is a flow chart of a first embodiment of the audio production method of the present invention. In this embodiment, the audio production method includes the following steps:

[0055] Step S10, obtaining the first user's singing audio of the target song;

[0056] The singing audio of the first user for the target song refers to the audio corresponding to the first user singing the target song.

[0057] To obtain the first user's singing audio of the target song, the audio of the first user singing the target song may be pre-recorded and stored to obtain the first user's singing audio of the target song from the stored target storage area, or the first user's singing audio of the target song may be obtained in real time when the first user sings the target song. This embodiment does not limit this step.

[0058] Step S20: Acquire the voice characteristic parameters of the first user.

[0059] The sound characteristic parameters include at least one of a voice parameter, a timbre parameter, a pitch parameter, and a range parameter, and the sound characteristic parameters are obtained based on audio data corresponding to when the first user recites or sings the target content of the target song;

[0060] Sound characteristic parameters include voice parameters, timbre parameters, pitch parameters, and range parameters. The range parameter refers to the parameter range from the lowest to the highest pitch achievable by a human voice or instrument. In this embodiment, the sound characteristic parameters are obtained based on the audio data corresponding to the first user reciting or singing the target content of the target song.

[0061] To obtain the voice characteristic parameters of the first user, the audio data corresponding to the target content of the target song recited or sung by the first user can be obtained first, and then the obtained audio data can be analyzed and identified by audio analysis software to obtain the voice characteristic parameters of the first user.

[0062] Optionally, the audio data corresponding to the target content of the target song recited or sung by the first user can be obtained directly when the first user recites or sings the target content of the target song, or the audio data corresponding to the target content of the target song recited or sung by the first user can be recorded first, and then obtained indirectly by obtaining the recorded audio data. This embodiment does not limit this.

[0063] Optionally, the target content of the target song includes at least one of all or part of the lyrics corresponding to the target song, keywords in the lyrics corresponding to the target song, and all or part of the melody corresponding to the target song.

[0064] As an optional implementation, please refer to Figure 3 , Figure 3 This is a flow chart of obtaining sound characteristic parameters in the first embodiment of the audio production method of the present invention, step S20 includes:

[0065] Step S21, obtaining keywords corresponding to each song section of the target song, and outputting the keywords, wherein the target content of the target song includes the keywords;

[0066] Step S22: obtaining audio data corresponding to the keyword from the first user;

[0067] Step S23: Identify the audio data to obtain voice feature parameters of the first user.

[0068] It should be noted that target songs can be pre-labeled. Labels include, but are not limited to, song language, song emotion, and song style. Song language includes, but is not limited to, Mandarin, Cantonese, English, and French; song emotion includes, but is not limited to, joy, anger, sorrow, happiness, sadness, excitement, and indignation; and song style includes, but is not limited to, pop, folk, rock, and rap.

[0069] In addition, corresponding to each target song, paragraph labels can be pre-set for each song paragraph of each target song based on the emotional tone reflected by the lyrics, melody and vocals corresponding to each song paragraph in the target song, where the paragraph labels include but are not limited to emotional labels, style labels, pitch parameters, and range parameters.

[0070] Optionally, each song paragraph in the target song may be a lyrics paragraph corresponding to the target song, wherein the lyrics paragraph may be a sentence of lyrics, or may be at least two consecutive sentences of lyrics.

[0071] Optionally, the target song can be obtained from a music library associated with the music player software, from locally downloaded and stored songs, or from a chorus platform provided by the music software, by receiving a song selected by the first user's song selection instruction from the chorus platform. This embodiment does not limit this.

[0072] Optionally, the target song can be obtained by obtaining the first user's mental state and then pushing songs to the first user based on the mental state. For example, with the first user's authorization, dynamic information about the first user's use of social software can be obtained and parsed to obtain the first user's mental state. Dynamic information from social software can include Moments, Weibo, and chat messages within a specific time period, such as user chat messages within 24-48 hours. The dynamic information is parsed to determine the first user's mental state, such as life status and mood, such as job promotion, marriage and childbirth, love, or heartbreak. Songs can then be pushed to the first user based on the psychological state, ensuring that the pushed songs are more relevant to the first user's life status and mood.

[0073] It should be noted that keywords can be used to obtain the user's own voice feature parameters.

[0074] Obtain keywords corresponding to each song paragraph of the target song and output the keywords. The first user reads or sings according to the output keywords, and then obtains audio data corresponding to the keywords of the first user, identifies the audio data, and obtains the voice feature parameters of the first user.

[0075] Optionally, a method for determining the keywords corresponding to each song paragraph of the target song can be to obtain the lyrics corresponding to the highest pitch and / or lowest pitch in each song paragraph, determine the lyrics as keywords, and then obtain the user's own voice feature parameters through the keywords.

[0076] Step S30: synthesize the singing audio according to the sound feature parameters to obtain the singing audio of the first user for the target song.

[0077] The singing audio is sung in chorus according to the sound characteristic parameters to obtain the singing audio of the first user for the target song and dance. The audio editing software can be used to adjust various parameters based on the audio data of the singing audio, using the sound characteristic parameters of the first user, such as vocal line parameters, timbre parameters, pitch parameters and range parameters, to synthesize the singing audio according to the sound characteristic parameters to obtain the singing audio of the first user for the target song.

[0078] It is easy to understand that compared with the singing audio of the target song by the first user, the singing audio is synthesized according to the voice characteristic parameters of the first user to obtain the singing audio of the target song by the first user. The singing audio of the target song sung by the first user with his own voice can be adjusted and synthesized to adapt to the voice characteristic parameters of the first user, so that the obtained singing audio of the target song by the first user conforms to the voice characteristics of the first user's own voice, and the singing audio of the target song sung by the first user with his own voice is improved, so that the final singing audio has better singing quality, thereby achieving the purpose of improving the singing quality of the target song sung by the first user while improving the user's singing experience.

[0079] Optionally, in actual application, the user's mood is different, and correspondingly, the user's voice characteristic parameters are different. When the voice characteristic parameters are directly obtained based on the audio data corresponding to the target content of the first user reciting or singing the target song in real time, and the audio data is analyzed and identified by audio analysis software, the voice characteristic parameters can be used to provide real-time feedback on the first user's current psychological state, such as the user's emotional state and / or affective state. The singing audio is synthesized according to the first user's voice characteristic parameters to obtain the first user's singing audio for the target song, so that the finally obtained singing audio has the sound characteristics that match the first user's own voice, improves the singing audio of the target song sung by the first user with his own voice, and can also provide feedback on the first user's psychological state at the time, so that the singing audio has more personal characteristics of the user himself.

[0080] In the technical solution disclosed in this embodiment, the singing audio is synthesized according to the voice characteristic parameters of the first user to obtain the singing audio of the first user for the target song. The singing audio of the first user singing the target song with his own voice can be adjusted and synthesized to adapt to the voice characteristic parameters of the first user, so that the obtained singing audio of the first user for the target song conforms to the voice characteristics of the first user's own voice, and the singing audio of the first user singing the target song with his own voice is improved, so that the final singing audio has better singing quality, thereby achieving the purpose of improving the singing quality of the first user singing the target song while improving the user's singing experience.

[0081] Based on the above first embodiment, a second embodiment of the audio production method of the present invention is proposed. Please refer to Figure 4 , Figure 4 1 is a flow chart of a second embodiment of the audio production method of the present invention. In this embodiment, after step S30, the method further includes:

[0082] Step S40, determining a chorus section in the target song that matches the sound feature parameters based on the matching degree between the audio feature parameters corresponding to each song section of the target song and the sound feature parameters;

[0083] The audio feature parameters corresponding to each section of the target song refer to the audio feature parameters corresponding to each section of the target song determined based on the standard audio associated with the target song. Optionally, the standard audio associated with the target song may refer to the audio corresponding to the target song released by the original singer of the target song, or may also refer to the audio corresponding to the target song sung by the cover singer of the target song, which is not limited in this embodiment.

[0084] Corresponding to the sound feature parameters, the audio feature parameters corresponding to each song paragraph include but are not limited to at least one of vocal parameters, timbre parameters, pitch parameters and range parameters. The audio feature parameters corresponding to each song paragraph of the target song are compared with the sound feature parameters through at least one of the four dimensions of vocal line, timbre, pitch and range, so as to determine the matching degree between the audio feature parameters and the sound feature parameters according to the comparison results, and then determine the chorus paragraph in the target song that matches the sound feature parameters according to the matching degree, that is, determine the chorus paragraph from the target song that meets the sound characteristics of the first user, so that the audio data obtained when the first user sings the chorus paragraph in the target song has a higher singing quality.

[0085] The matching degree between the audio feature parameters and the sound feature parameters is determined based on the comparison results. For example, the audio feature parameters and sound feature parameters corresponding to each song paragraph of the target song can be compared from four dimensions of voice, timbre, pitch and range. The number of matching dimensions is determined based on the comparison results, and the matching degree is determined based on the number of dimensions. For example, whether the voice parameters in the audio feature parameters are the same or similar to the voice parameters in the sound feature parameters, if they are the same or similar, the number of dimensions is +1. For example, whether the numerical range corresponding to the range parameter in the sound feature parameters contains the numerical range corresponding to the range parameter in the audio feature parameters, if it does, the number of dimensions is +1. Assuming that among the four dimensions of voice, timbre, pitch and range, the comparison results of the two dimensions of voice and range are matched, it indicates that the number of dimensions is 2, and the matching degree determined based on the number of dimensions is 2.

[0086] The chorus section that matches the sound feature parameters is determined according to the matching degree, and when the matching degree is greater than or equal to a preset matching degree, the song section of the target song is determined as the chorus section that matches the sound feature parameters.

[0087] Step S50, generating a first sub-audio corresponding to the chorus segment based on the singing audio;

[0088] Step S60, generating second sub-audio corresponding to other song sections other than the chorus section based on the singing audio associated with the second user;

[0089] Step S70: Generate chorus song audio based on the first sub-audio and the second sub-audio.

[0090] The singing audio associated with the second user refers to the audio of the second user singing the entire or partial song section corresponding to the target song, wherein the second user is the user who sings the target song together with the first user.

[0091] Optionally, the number of second users may be one or at least two.

[0092] Optionally, after step S30, it includes: associating the first user with the singing audio, or associating the first user, the singing audio and the singing audio.

[0093] As an optional implementation, based on the singing audio, a first sub-audio corresponding to the chorus section is generated. After obtaining the singing audio associated with the first user, the audio data corresponding to the vocals of other song sections other than the chorus section in the singing audio can be deleted through audio editing software to generate the first sub-audio corresponding to the chorus section. Correspondingly, based on the singing audio associated with the second user, a second sub-audio corresponding to other song sections other than the chorus section is generated. After obtaining the singing audio associated with the second user, the audio data corresponding to the vocals of other song sections other than the chorus section in the singing audio can be obtained through audio editing software to generate the second sub-audio. Then, a chorus song audio is generated based on the first sub-audio and the second sub-audio to obtain a chorus song.

[0094] As an optional implementation, based on the singing audio, a first sub-audio corresponding to the chorus section is generated. After obtaining the singing audio associated with the first user, the audio data corresponding to the vocals of other song sections other than the chorus section in the singing audio can be deleted through audio editing software to generate the first sub-audio corresponding to the chorus section. Correspondingly, based on the singing audio associated with the second user, a second sub-audio corresponding to other song sections other than the chorus section is generated. After obtaining the singing audio associated with the second user, the audio data corresponding to the vocals of other song sections other than the chorus section in the singing audio can be obtained through audio editing software to generate the second sub-audio. Then, a chorus song audio is generated based on the first sub-audio and the second sub-audio to obtain a chorus song.

[0095] Optionally, the singing audio associated with the second user is audio recorded when the second user sings a chorus song.

[0096] Compared with the method in which all users who need to sing a chorus are online at the same time to complete the chorus, in this embodiment, the singing audio recorded and synthesized by the first user when singing the target song can be used as the singing audio associated with the first user, and / or the audio recorded in advance by the second user when singing the chorus song can be used as the singing audio associated with the second user. The first user and the second user who do not need to sing a chorus are online at the same time to realize the chorus method that is more flexible and not restricted by time and space, so as to achieve the purpose of the first user and the second user completing the chorus song.

[0097] The technical solution disclosed in this embodiment, based on obtaining the voice characteristic parameters of the first user, determines the chorus section in the target song that matches the voice characteristic parameters according to the matching degree between the audio characteristic parameters corresponding to each song section of the target song and the voice characteristic parameters, so as to determine the chorus section in the target song that matches the voice characteristic of the first user, so that the first user can easily sing the chorus section of the target song, thereby improving the user experience of the first user; generates a first sub-audio corresponding to the chorus section based on the singing audio of the first user, so as to obtain the first sub-audio corresponding to the chorus section that the first user sings that matches his own voice characteristics, so that the audio data obtained when the first user sings the chorus section of the target song has higher singing quality; generates a second sub-audio corresponding to other song sections outside the chorus section based on the singing audio associated with the second user, so as to obtain the second sub-audio corresponding to other song sections outside the chorus section sung by the second user, and then generates chorus song audio based on the first sub-audio and the second sub-audio, so as to successfully obtain the chorus song audio of the first user and the second user, thereby improving the chorus quality of the chorus song audio of the first user and the second user.

[0098] The third embodiment of the audio production method of the present invention is proposed based on any one of the above embodiments. Please refer to Figure 5 , Figure 5 FIG. 1 is a flow chart of a third embodiment of the audio production method of the present invention. In this embodiment, the audio production method further includes the following steps:

[0099] Step S80, determining the second user who sings the target song with the first user;

[0100] Step S90: When the second user has sung the target song, execute step S60.

[0101] To determine the second user who will sing the target song with the first user, a selection interface containing the users to be sung together can be output based on the chorus platform provided by the music software, so that when a selection instruction is received on the selection interface, the user to be sung together selected by the selection instruction is obtained as the second user who will sing the target song with the first user; the second user who will sing the target song with the first user can also be determined by inputting the name corresponding to the second user. This embodiment does not limit this step.

[0102] It is understandable that, in actual application, the second user may be a user who has never released a song, or a user who has never sung the target song.

[0103] To determine whether the second user has sung the target song, the songs sung by the second user can be searched in the music player software. If there is a target song matching the target song among the found songs, it indicates that the second user has sung the target song. If there is no target song matching the target song among the found songs, it indicates that the second user has not sung the target song.

[0104] When the second user has performed the target song, step S60 may be executed to generate second sub-audio corresponding to the song sections other than the chorus section based on the second user's associated singing audio. The specific implementation of this step can be found in the first embodiment and will not be repeated here. In this embodiment, the second user's associated singing audio is the acquired song audio corresponding to the target song performed by the second user.

[0105] As an optional implementation manner, after step S80, the method further includes:

[0106] When the second user does not sing the target song, obtaining the voice feature parameters of the second user;

[0107] generating, based on the voice feature parameters of the second user and the audio data corresponding to the target song, an audio recording of the second user singing the target song;

[0108] The second user is associated with the singing audio.

[0109] In actual application, the second user may be a user who has never released a song, or a user who has never sung the target song. That is, when the second user does not sing the target song, in order to achieve the purpose of the target song of the first user and the second user, the target song audio is finally obtained. According to the voice feature parameters of the second user and the audio data corresponding to the target song, the singing audio of the target song is generated to obtain the singing audio corresponding to the target song sung by the second user, and the singing audio of the second user and the target song are associated, and then step S60 is executed, that is, according to the singing audio associated with the second user, the second sub-audio corresponding to other song sections except the chorus section is generated, so as to achieve the purpose of the target song of the first user and the second user even if the second user does not sing the target song, improve the fun of the target song, and provide a new audio production method for users who like singing or users who like chorus.

[0110] It should be noted that, the specific implementation of the step of obtaining the voice characteristic parameters of the second user can refer to the specific implementation method of obtaining the voice characteristic parameters of the first user in the first embodiment, and the specific implementation of the step of generating the singing audio of the target song according to the voice characteristic parameters of the second user and the audio data corresponding to the target song can refer to the specific implementation method of generating the singing audio of the target song according to the voice characteristic parameters of the first user and the audio data corresponding to the target song in the first embodiment, and will not be described in detail in this embodiment.

[0111] In the technical solution disclosed in this embodiment, by selectively determining the second user of the target song of the first user, when the first user has sung the target song, step S60 can be directly executed to obtain the second sub-audio corresponding to other song sections excluding the chorus section sung by the second user; when the first user has not sung the target song, the singing audio of the target song can be generated according to the voice feature parameters of the second user and the audio data corresponding to the target song to obtain the singing audio of the target song sung by the second user, and associate the second user with the singing audio of the target song, so as to achieve the purpose of successfully singing the target song with the second user even if the second user has not sung the target song.

[0112] like Figure 6 As shown, Figure 6 This is a schematic diagram of the module composition of the audio production device provided by the present invention. The audio production device 100 includes:

[0113] A first acquisition module 110 is configured to acquire a first user's singing audio of a target song;

[0114] A second acquisition module 120 is configured to acquire voice characteristic parameters of the first user, the voice characteristic parameters including at least one of a voice line parameter, a timbre parameter, a pitch parameter, and a range parameter, and the voice characteristic parameters are acquired based on audio data corresponding to the first user reciting or singing the target content of the target song;

[0115] The audio synthesis module 130 is used to synthesize the singing audio according to the sound feature parameters to obtain the singing audio of the first user for the target song.

[0116] The specific implementation of the audio production device of the present invention is basically the same as the various embodiments of the above-mentioned audio production method, and will not be repeated here.

[0117] The present invention also proposes a terminal device, which includes: a memory, a processor, and an audio production program stored in the memory and runnable on the processor. When the audio production program is executed by the processor of the terminal device, the steps of the audio production method in any of the above embodiments are implemented.

[0118] The present invention also provides a readable storage medium, which stores an audio production program. When the audio production program is executed by a processor, the steps of the audio production method described in any of the above embodiments are implemented.

[0119] In the embodiments of the terminal device and the readable storage medium provided by the present invention, all the technical features of the above-mentioned embodiments of the audio production method are included. The expanded and explained contents of the specification are basically the same as those of the above-mentioned embodiments of the audio production method, and will not be repeated here.

[0120] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatus, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0121] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0122] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0123] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0124] It should be noted that in the claims, any reference signs placed between parentheses shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claim. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by one and the same item of hardware. The use of the words first, second, third etc. does not indicate any order. These words may be interpreted as names.

[0125] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0126] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. An audio production method, characterized in that: The audio production method comprises: Obtaining real-time singing audio of a first user for a target song; Obtaining keywords corresponding to each song section of the target song, and outputting the keywords, wherein the target content of the target song includes the keywords; Obtaining audio data corresponding to the keyword from the first user, identifying the audio data to obtain voice feature parameters of the first user, the voice feature parameters including at least one of a voice line parameter, a timbre parameter, a pitch parameter, and a range parameter, and the voice feature parameters are obtained based on audio data corresponding to the first user reciting or singing the target content of the target song; synthesizing the singing audio according to the sound characteristic parameters to obtain the singing audio of the first user for the target song; Determining a chorus section in the target song that matches the sound feature parameters based on a matching degree between the audio feature parameters corresponding to each song section of the target song and the sound feature parameters; Generating a first sub-audio corresponding to the chorus segment according to the singing audio; Generating second sub-audio corresponding to other song sections other than the chorus section based on the singing audio associated with the second user; A chorus song audio is generated based on the first sub audio and the second sub audio.

2. The method according to claim 1, wherein The step of generating the first sub-audio corresponding to the chorus section according to the singing audio includes: Acquire audio data corresponding to the vocals of the chorus section in the performance audio to generate a first sub-audio corresponding to the chorus section; or, The audio data corresponding to other song sections except the chorus section is deleted from the performance audio to generate a first sub-audio corresponding to the chorus section.

3. The method according to claim 1, wherein The audio production method further includes: determining the second user who sings the target song with the first user; When the second user has sung the target song, the step of generating second sub-audio corresponding to other song sections except the chorus section based on the singing audio associated with the second user is executed.

4. The method according to claim 3, wherein After the step of determining the second user who sings the target song with the first user, the method further includes: When the second user does not sing the target song, obtaining the voice feature parameters of the second user; generating, based on the voice characteristic parameters of the second user and the singing audio corresponding to the target song, the singing audio of the second user for the target song; The second user is associated with the singing audio.

5. The method according to claim 1, wherein The singing audio associated with the second user is the audio recorded when the second user sings the target song.

6. An audio production device, characterized in that The audio production device comprises: A first acquisition module is used to obtain real-time singing audio of the first user for the target song; a second acquisition module, configured to acquire keywords corresponding to respective song paragraphs of the target song and output the keywords, wherein the target content of the target song includes the keywords; acquire audio data corresponding to the keywords by the first user, identify the audio data to acquire voice feature parameters of the first user, wherein the voice feature parameters include at least one of a voice line parameter, a timbre parameter, a pitch parameter, and a range parameter, and the voice feature parameters are acquired based on the audio data corresponding to the first user when reciting or singing the target content of the target song; an audio synthesis module, configured to synthesize the singing audio according to the sound characteristic parameters to obtain the singing audio of the first user for the target song; The audio production device is also used to determine the chorus section in the target song that matches the sound feature parameters based on the matching degree between the audio feature parameters corresponding to each song section of the target song and the sound feature parameters; generate a first sub-audio corresponding to the chorus section based on the singing audio; generate a second sub-audio corresponding to other song sections outside the chorus section based on the singing audio associated with the second user; and generate chorus song audio based on the first sub-audio and the second sub-audio.

7. A terminal device, characterized in that: The terminal device includes: a memory, a processor, and an audio production program stored in the memory and executable on the processor. When the audio production program is executed by the processor, the steps of the audio production method according to any one of claims 1 to 5 are implemented.

8. A readable storage medium, characterized in that: The readable storage medium stores an audio production program, and when the audio production program is executed by a processor, the steps of the audio production method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Singing sound converter

    CN110782866A

  • Song processing method, device and equipment, and computer readable storage medium

    CN113703882A