Method, Server, Terminal, System and Storage Medium for Chorus Processing
The server processes the original dry sound audio to generate chorus accompaniment audio, which solves the problem of incoherent audio switching in chorus songs, realizes the chorus effect of solo voice without the first singer, and improves the user experience.
Patent Information
- Application Number
- CN202310270230.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-15
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-03-15
AI Technical Summary
In the prior art, chorus songs are prone to incoherent connection problems when switching between the original audio and the accompaniment audio, which affects the user experience.
The original dry sound audio is processed through the server to generate the dry sound audio to be synthesized to ensure that it does not have the solo voice of the first singer during playback, and synthesizes the chorus accompaniment audio with the original accompaniment audio. The terminal only needs to play the synthesized audio to achieve the unmanned vocal effect of the solo part of the first singer.
It effectively avoids the problem of frequent switching and incoherent connection between the original singer audio and the accompaniment audio, and improves the user's karaoke experience.
Smart Images

Figure CN116206584B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio technology, and in particular, to a method, a server, a terminal, a system, and a storage medium for chorus processing. Background Art
[0002] With the rise of mobile karaoke, various karaoke forms have emerged continuously. Among them, the form of star chorus has become increasingly popular among people.
[0003] In the current star chorus form, the server sends the original audio, accompaniment audio, and lyrics file of the chorus song to the terminal together. The terminal plays the original audio during the singing time of the first singer and the accompaniment audio during the singing time of the second singer, and the user can sing while the accompaniment audio is being played. In this way, the effect of singing with a star can be achieved.
[0004] In the above chorus method, the terminal needs to alternately play the original audio and the accompaniment audio according to the singing time of different singers. In this way, at the connection point where the accompaniment audio switches to the original audio or the original audio switches to the accompaniment audio, there may be a problem of discontinuous connection, which affects the user experience. Summary of the Invention
[0005] Embodiments of this application provide a method, a server, a terminal, a system, and a storage medium for chorus processing. This method can avoid the problem of discontinuous connection that occurs when switching between multiple audios in the prior art. The technical solutions are as follows:
[0006] In a first aspect, a method for chorus processing is provided. The method is applied to a server and includes:
[0007] Obtain the original accompaniment audio, original dry vocal audio, and lyrics information of the song, where the lyrics information is in the format of word-by-word timestamp lyrics and includes segmentation information of at least two lyrics segments. Each segmentation information includes the start time and end time of the corresponding lyrics segment;
[0008] Process the original dry vocal audio according to the segmentation information of the at least two lyrics segments to obtain a synthesized dry vocal audio. When the synthesized dry vocal audio is played, there is no solo voice of the first singer. The singers of the original dry vocal audio include the first singer and the second singer;
[0009] Generate the chorus accompaniment audio of the song according to the original accompaniment audio and the synthesized dry vocal audio;
[0010] Receive a chorus request for the song sent by the terminal;
[0011] Send the chorus accompaniment audio of the song to the terminal.
[0012] In a possible implementation, processing the original dry audio according to the segmentation information of the at least two lyric segments to obtain the dry audio to be synthesized includes:
[0013] Determining a first lyric segment sung solo by the second singer among the at least two lyric segments;
[0014] Taking the start time and end time of the first lyric segment as the singing start time and singing end time of the solo dry audio segment of the second singer respectively;
[0015] In the original dry audio, intercepting the solo dry audio segment of the second singer according to the singing start time and singing end time of the solo dry audio segment of the second singer, and performing mixed-stream splicing on the solo dry audio segment of the second singer to obtain the dry audio to be synthesized.
[0016] In a possible implementation, processing the original dry audio according to the segmentation information of the at least two lyric segments to obtain the dry audio to be synthesized includes:
[0017] Determining a first lyric segment sung solo by the second singer among the at least two lyric segments;
[0018] Subtracting a first duration from the start time of the lyric segment sung solo by the second singer as the singing start time of the solo dry audio segment of the second singer, and adding the first duration to the end time of the lyric segment sung solo by the second singer as the singing end time of the solo dry audio segment of the second singer;
[0019] In the original dry audio, intercepting the solo dry audio segment of the second singer according to the singing start time and singing end time of the solo dry audio segment of the second singer, and performing mixed-stream splicing on the solo dry audio segment of the second singer to obtain the dry audio to be synthesized.
[0020] In a possible implementation, processing the original dry audio according to the segmentation information of the at least two lyric segments to obtain the dry audio to be synthesized includes:
[0021] Determining a first lyric segment sung solo by the first singer among the at least two lyric segments;
[0022] Taking the start time and end time of the first lyric segment as the singing start time and singing end time of the solo dry audio segment of the first singer respectively;
[0023] In the original dry audio, according to the singing start time and singing end time of the solo dry audio frequency band of the first singer, adjust the volume of the solo dry audio frequency band of the first singer to 0 to obtain the dry audio to be synthesized.
[0024] In a possible implementation, before generating the chorus accompaniment audio of the song according to the original accompaniment audio and the dry audio to be synthesized, the method further includes:
[0025] Determine the lag duration of the singing start time of the original accompaniment audio compared to the singing start time of the original dry audio;
[0026] Determine the leading duration of the singing end time of the original accompaniment audio compared to the singing end time of the original dry audio;
[0027] If the lag duration is greater than the second duration, cut off a part of the audio from the beginning of the original accompaniment audio so that the lag duration of the singing start time of the cut original accompaniment audio compared to the singing start time of the original dry audio is equal to the second duration;
[0028] If the leading duration is greater than the second duration, cut off a part of the audio from the end of the original accompaniment audio so that the leading duration of the singing end time of the cut original accompaniment audio compared to the singing end time of the original dry audio is equal to the second duration;
[0029] The generating the chorus accompaniment audio of the song according to the original accompaniment audio and the dry audio to be synthesized includes:
[0030] Generate the chorus accompaniment audio of the song according to the cut original accompaniment audio and the dry audio to be synthesized.
[0031] In a possible implementation, the sending the chorus accompaniment audio of the song to the terminal includes:
[0032] Send the chorus accompaniment audio of the song and the lyric information to the terminal.
[0033] In a possible implementation, before sending the chorus accompaniment audio of the song to the terminal, the method further includes:
[0034] Obtain the music short film of the song;
[0035] Remove the subtitles and audio in the music short film to obtain the accompaniment video to be synthesized;
[0036] Segment the at least two lyric segments in the lyric information and render them word by word onto the to-be-synthesized accompaniment video, wherein lyrics sung by different singers displayed in the to-be-synthesized accompaniment video have different colors;
[0037] Synthesize the rendered to-be-synthesized accompaniment video and the chorus accompaniment audio to obtain the chorus accompaniment video of the song;
[0038] The sending the chorus accompaniment audio of the song to the terminal includes:
[0039] Send the chorus accompaniment video of the song to the terminal.
[0040] In a second aspect, a chorus processing method is provided. The method is applied to a terminal and includes:
[0041] Obtain a chorus instruction of a song;
[0042] Send a chorus request for the song to the server;
[0043] Receive the chorus accompaniment video of the first song sent by the server;
[0044] Play the chorus accompaniment video, where the chorus accompaniment video includes a first singer and a second singer, and there is no singing voice during the solo time corresponding to the first singer when the chorus accompaniment video is played;
[0045] During the playing of the chorus accompaniment video, collect the singing dry audio of the user;
[0046] Mix the chorus accompaniment video and the singing dry audio to obtain a chorus video.
[0047] In a third aspect, a chorus processing device is provided. The device is applied to a server and includes:
[0048] An acquisition module, configured to acquire the original accompaniment audio, the original dry audio, and the lyric information of a song, wherein the lyric information is in the format of word-by-word timestamp lyrics and includes the segmentation information of at least two lyric segments, and each segmentation information includes the start time and the end time of the corresponding lyric segment;
[0049] A generation module, configured to process the original dry audio according to the segmentation information of the at least two lyric segments to obtain a to-be-synthesized dry audio, wherein there is no solo voice of the first singer when the to-be-synthesized dry audio is played, and the singers of the original dry audio include the first singer and the second singer; generate the chorus accompaniment audio of the song according to the original accompaniment audio and the to-be-synthesized dry audio;
[0050] A receiving module, configured to receive the chorus request of the song sent by the terminal;
[0051] A sending module, configured to send the chorus accompaniment audio of the song to the terminal.
[0052] In a possible implementation, the generating module is configured to:
[0053] Determine a first lyric segment sung solo by the second singer from among the at least two lyric segments;
[0054] Use the start time and end time of the first lyric segment as the singing start time and singing end time of the solo dry voice frequency band of the second singer, respectively;
[0055] In the original dry audio, intercept the solo dry voice frequency band of the second singer according to the singing start time and singing end time of the solo dry voice frequency band of the second singer, and perform mixing and splicing on the solo dry voice frequency band of the second singer to obtain a dry audio to be synthesized.
[0056] In a possible implementation, the generating module is configured to:
[0057] Determine a first lyric segment sung solo by the second singer from among the at least two lyric segments;
[0058] Subtract a first duration from the start time of the lyric segment sung solo by the second singer as the singing start time of the solo dry voice frequency band of the second singer, and add the first duration to the end time of the lyric segment sung solo by the second singer as the singing end time of the solo dry voice frequency band of the second singer;
[0059] In the original dry audio, intercept the solo dry voice frequency band of the second singer according to the singing start time and singing end time of the solo dry voice frequency band of the second singer, and perform mixing and splicing on the solo dry voice frequency band of the second singer to obtain a dry audio to be synthesized.
[0060] In a possible implementation, the generating module is configured to:
[0061] Determine a first lyric segment sung solo by the first singer from among the at least two lyric segments;
[0062] Use the start time and end time of the first lyric segment as the singing start time and singing end time of the solo dry voice frequency band of the first singer, respectively;
[0063] In the original dry audio, according to the singing start time and singing end time of the solo dry audio frequency band of the first singer, adjust the volume of the solo dry audio frequency band of the first singer to 0 to obtain the dry audio to be synthesized.
[0064] In a possible implementation manner, the generating module is configured to:
[0065] Determine the lag duration of the singing start time of the original accompaniment audio compared to the singing start time of the original dry audio;
[0066] Determine the leading duration of the singing end time of the original accompaniment audio compared to the singing end time of the original dry audio;
[0067] If the lag duration is greater than the second duration, cut off a part of the audio from the beginning of the original accompaniment audio so that the lag duration of the singing start time of the cut original accompaniment audio compared to the singing start time of the original dry audio is equal to the second duration;
[0068] If the leading duration is greater than the second duration, cut off a part of the audio from the end of the original accompaniment audio so that the leading duration of the singing end time of the cut original accompaniment audio compared to the singing end time of the original dry audio is equal to the second duration;
[0069] Generate the chorus accompaniment audio of the song according to the cut original accompaniment audio and the dry audio to be synthesized.
[0070] In a possible implementation manner, the sending module is configured to:
[0071] Send the chorus accompaniment audio of the song and the lyric information to the terminal.
[0072] In a possible implementation manner, the generating module is further configured to:
[0073] Obtain the music short film of the song;
[0074] Remove the subtitles and audio in the music short film to obtain the accompaniment video to be synthesized;
[0075] Segment the at least two lyric segments in the lyric information and render them word by word on the accompaniment video to be synthesized, where the lyrics sung by different singers displayed in the accompaniment video to be synthesized have different colors;
[0076] Synthesize the rendered accompaniment video to be synthesized and the chorus accompaniment audio to obtain the chorus accompaniment video of the song;
[0077] The sending module is configured to:
[0078] Send the chorus accompaniment video of the song to the terminal.
[0079] In a fourth aspect, a chorus processing device is provided. The method is applied to a terminal, and the device includes:
[0080] An acquisition module, configured to acquire a chorus instruction of a song;
[0081] A sending module, configured to send a chorus request of the song to a server;
[0082] A receiving module, configured to receive the chorus accompaniment video of the song sent by the server;
[0083] A playing module, configured to play the chorus accompaniment video. The chorus accompaniment video includes a first singer and a second singer, and there is no singing voice during the solo time corresponding to the first singer when the chorus accompaniment video is played;
[0084] An acquisition module, configured to acquire the dry singing audio of a user during the playing of the chorus accompaniment video;
[0085] A mixing module, configured to mix the chorus accompaniment video and the dry singing audio to obtain a chorus video.
[0086] In a fifth aspect, a server is provided. The server includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the chorus processing method as described in the first aspect above.
[0087] In a sixth aspect, a terminal is provided. The terminal includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the chorus processing method as described in the second aspect above.
[0088] In a seventh aspect, a chorus processing system is provided. The system includes the server as described in the fifth aspect above and the terminal as described in the sixth aspect above.
[0089] In an eighth aspect, a computer-readable storage medium is provided. The storage medium stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the chorus processing method as described in the first aspect or the second aspect.
[0090] In a ninth aspect, a computer program product is provided. The storage medium stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the chorus processing method as described in the first aspect or the second aspect.
[0091] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:
[0092] In the embodiments of the present application, the server processes the original dry audio of the song to obtain the dry audio to be synthesized. There is no solo voice of the first singer in the dry audio to be synthesized, and the first singer is any one of the multiple singers of the song. On this basis, the server synthesizes the original accompaniment audio of the song and the dry audio to be synthesized into the chorus accompaniment audio of the song and sends it to the terminal. Since there is no solo voice of the first singer in the dry audio to be synthesized, there is also no solo voice of the first singer in the synthesized chorus accompaniment audio. In this way, the terminal only needs to play the chorus accompaniment audio normally to achieve the effect that there is only accompaniment without the singer's voice in the solo part of the first singer, effectively avoiding the problem of discontinuous connection that occurs when frequently switching between the original dry audio and the original accompaniment audio in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0094] Figure 1 is a flowchart of a method for chorus processing provided by an embodiment of the present application;
[0095] Figure 2 is a flowchart of a method for chorus processing provided by an embodiment of the present application;
[0096] Figure 3 is a flowchart of a method for chorus processing provided by an embodiment of the present application;
[0097] Figure 4 is a schematic diagram of the processing of the original dry audio provided by an embodiment of the present application;
[0098] Figure 5 is a schematic diagram of the processing of the original dry audio provided by an embodiment of the present application;
[0099] Figure 6 is a schematic diagram of the chorus accompaniment video provided by an embodiment of the present application;
[0100] Figure 7 is a flowchart of a method for chorus processing provided by an embodiment of the present application;
[0101] Figure 8 is a schematic structural diagram of a device for chorus processing provided by an embodiment of the present application;
[0102] Figure 9 It is a schematic structural diagram of a device for chorus processing provided by an embodiment of the present application;
[0103] Figure 10 It is a schematic structural diagram of a terminal provided by an embodiment of the present application;
[0104] Figure 11 It is a schematic structural diagram of a server provided by an embodiment of the present application. Detailed implementation manners
[0105] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0106] An embodiment of the present application provides a method for chorus processing, which can be jointly implemented by a server and a terminal. Among them, the terminal can be a mobile phone, a tablet computer, etc. A KTV application program can be installed in the terminal. The KTV application program can be an application program specifically designed for KTV (or recording singing), or other application programs with KTV (or recording singing) functions. The server can be the background server of the KTV application program.
[0107] The method for chorus processing provided by an embodiment of the present application can be applied to the scenario of KTV. The server is used to produce a chorus accompaniment audio for KTV and provide the chorus accompaniment audio to the terminal. The terminal is used to play the chorus accompaniment audio when the user is singing KTV and collect the user's singing audio. Thus, the chorus accompaniment audio and the user's singing audio are mixed to obtain the user's chorus file.
[0108] In the above scenario, the server processes the original dry audio of the song to obtain a dry audio to be synthesized, in which there is no solo voice of the first singer, and the first singer is any one of the multiple singers of the song. On this basis, the server synthesizes the original accompaniment audio of the song and the dry audio to be synthesized into a chorus accompaniment audio of the song and sends it to the terminal. Since there is no solo voice of the first singer in the dry audio to be synthesized, there is also no solo voice of the first singer in the synthesized chorus accompaniment audio. In this way, the terminal only needs to play the chorus accompaniment audio normally to achieve the effect that there is only accompaniment without the singer's voice in the solo part of the first singer, effectively avoiding the problem of discontinuous connection that occurs when frequently switching between the original dry audio and the original accompaniment audio in the prior art to achieve this effect.
[0109] The following will describe the method for chorus processing provided by an embodiment of the present application with reference to the accompanying drawings. Refer to Figure 1 , this method is executed by the server, and the processing flow of the method can include the following steps:
[0110] Step 101: Obtain the original accompaniment audio, original dry vocal audio, and lyric information of the song.
[0111] Among them, the lyric information is in the format of word-by-word timestamp lyrics and includes segmentation information of at least two lyric segments. Each segmentation information includes the start time and end time of the corresponding lyric segment.
[0112] Step 102: Process the original dry vocal audio according to the segmentation information of at least two lyric segments to obtain the dry vocal audio to be synthesized. Generate the chorus accompaniment audio of the song based on the original accompaniment audio and the dry vocal audio to be synthesized.
[0113] Among them, when the dry vocal audio to be synthesized is played, there is no solo voice of the first singer, and the first singer is any one of the multiple singers of the original dry vocal audio.
[0114] Step 103: Receive the chorus request for the above song sent by the terminal.
[0115] Step 104: Send the chorus accompaniment audio of the above song to the terminal.
[0116] See Figure 2 , the chorus processing method provided by the embodiment of the present application can be executed by the terminal. The processing flow of the method can include the following steps:
[0117] Step 201: Obtain the chorus instruction of the song.
[0118] Step 202: Send a chorus request for the song to the server.
[0119] Step 203: Receive the chorus accompaniment video of the song sent by the server.
[0120] Step 204: Play the chorus accompaniment video. The chorus accompaniment video includes a first singer and a second singer. When the chorus accompaniment video is played, there is no singing voice during the solo time corresponding to the first singer.
[0121] Step 205: Collect the user's dry vocal audio during the playback of the chorus accompaniment video.
[0122] Step 206: Mix the chorus accompaniment video and the dry vocal audio to obtain a chorus video.
[0123] The chorus processing method provided by the embodiment of the present application is executed by the interaction between the terminal and the server. In this method, the server generates the chorus accompaniment audio and provides it for the terminal to use. The timing for the server to generate the chorus accompaniment audio can include:
[0124] Timing 1: After the server receives the chorus request for the first song sent by the terminal, it generates the chorus accompaniment audio for the first song.
[0125] Timing 2: Generate the corresponding chorus accompaniment audio for the chorus songs in the song library in advance. After the terminal sends the chorus request for the first song, directly return the chorus accompaniment audio of the first song to the terminal.
[0126] See Figure 3 , in the case where the server generates the chorus accompaniment audio using the above Timing 1, the chorus processing method provided by the embodiments of the present application may include the following steps:
[0127] Step 301: The terminal obtains the chorus instruction for the first song.
[0128] Wherein, the first song may be a chorus song of a first singer and a second singer. The first song includes the solo part of the first singer, the solo part of the second singer, and may also include the chorus part of the first singer and the second singer.
[0129] In implementation, the terminal may be installed with a KTV application program, and the KTV application program may provide the function of recording and singing ordinary songs and the function of chorus songs. When the user wants to record and sing a chorus song, the user can select the first song to be recorded and sung in the chorus song list, trigger the corresponding chorus instruction for the first song. Further, the terminal can obtain the chorus instruction for the corresponding first song.
[0130] Step 302: The terminal sends a chorus request for the first song to the server.
[0131] In implementation, after obtaining the chorus instruction for the first song, the terminal may send a chorus request for the first song to the server. In the chorus request, the identifier of the first song may be carried.
[0132] Step 303: The server obtains the original accompaniment audio, the original dry vocal audio, and the lyrics information of the first song.
[0133] Wherein, the original dry vocal audio includes the solo dry vocal audio of the first singer and the solo dry vocal audio of the second singer. The lyrics information may be in the format of word-by-word timestamp lyrics, and the lyrics information includes the segmentation information of each lyrics segment of the first song. Each segmentation information includes the start time and the end time of the corresponding lyrics segment. In addition, the song information may also include the start time and the end time of each word in the lyrics.
[0134] In implementation, after the server receives the chorus request of the first song sent by the terminal, it obtains the identifier of the first song carried therein. Furthermore, the server obtains the original accompaniment audio and the original dry voice audio of the first song. There are various ways to obtain the original accompaniment audio and the original dry voice audio. Several of the acquisition methods are listed below:
[0135] Method 1: The server previously separates the vocals from the song audio of the chorus songs in the song library to obtain the original accompaniment audio and the original dry voice audio of each chorus song, and stores them in the storage device. After receiving the chorus request for the corresponding first song, the original accompaniment audio and the original dry voice audio of the first song can be directly obtained from the storage device.
[0136] Method 2: After the server receives the chorus request for the corresponding first song, it obtains the song audio of the first song and separates the vocals from the song audio to obtain the original accompaniment audio and the original dry voice audio of the first song.
[0137] Method 3: The original accompaniment audio and the original dry voice audio of the song are already stored in the server. This situation may be that some users record their own songs and upload the original accompaniment audio and the original dry voice audio (in this case, the original dry voice audio refers to the dry voice audio recorded by the user) after recording, and the server stores the original accompaniment audio and the original dry voice audio separately.
[0138] Step 304: The server processes the original dry voice audio according to the lyric information to obtain the dry voice audio to be synthesized, and generates the chorus accompaniment audio of the first song based on the original dry voice audio and the dry voice audio to be synthesized.
[0139] Among them, there is no singing voice of the first singer during the singing time corresponding to the first singer when the chorus accompaniment audio is played.
[0140] In implementation, this step 304 can have various implementation methods. Several of the implementation methods are listed below for description.
[0141] Implementation method 1:
[0142] Case 1: When the first song includes the solo part of the first singer and the solo part of the second singer, and does not include the chorus part of the first singer and the second singer, the original dry voice audio of the first song includes the solo dry voice audio of the first singer and the solo dry voice audio of the second singer, and does not include the chorus dry voice audio of the first singer and the second singer.
[0143] In this case, the processing of implementation method 1 is as follows:
[0144] In the original dry audio, obtain the solo dry audio segment of the second singer. According to the singing start time of the solo dry audio segment of the second singer, mix the solo dry audio segment of the second singer and the original accompaniment audio to obtain the chorus accompaniment audio of the first song.
[0145] Specifically, first obtain the segmentation information of the lyrics segments of the first song. Among them, the segmentation information includes the start time, end time, and singer identifier of each lyrics segment in the first song. The singer identifier can be used to identify the singer's name. The structure of the lyrics segmentation information can be represented as follows:
[0146] type SegLyricInfo struct (structure of the lyrics segmentation information) {
[0147] StartMs (start time)
[0148] EndMs (end time)
[0149] Role (singer identifier)
[0150] }
[0151] Then, according to the singer identifier in the segmentation information, determine the first lyrics segment corresponding to the second singer. Use the start time of the first lyrics segment as the first singing start time of the second singer, and use the end time of the first lyrics segment as the first singing end time of the second singer. Then, according to the first singing start time and the first singing end time, intercept the solo dry audio segment of the second singer in the original dry audio. Here, there are also multiple ways to perform the interception process. Several of them are listed below for explanation.
[0152] Method 1: Directly intercept the dry audio segment between the first singing start time and the first singing end time in the original dry audio as the solo dry audio segment of the second singer.
[0153] Method 2: Subtract the first preset duration from the first singing start time to obtain the updated singing start time, and subtract the first preset duration from the first singing end time to obtain the updated singing end time. Intercept the dry audio segment between the updated singing start time and the updated singing end time in the updated original dry audio as the solo dry audio segment of the second singer. Then, perform audio fade-in processing on the first preset duration at the start of the intercepted solo dry audio segment of the second singer, and perform audio fade-out processing on the first preset duration at the end. In this case, use the updated singing start time as the singing start time of the solo dry audio segment of the second singer, and use the updated singing end time as the singing end time of the solo dry audio segment of the second singer.
[0154] For example, if the start time of the first performance is 1 min 10 s (1 minute and 10 seconds), the end time of the first performance is 1 min 20 s, and the first preset time is 300 ms (milliseconds), then the dry audio segment between (1 min 10 s - 300 ms = 1 min 9 s 700 ms) and (1 min 20 s + 300 ms = 1 min 20 s 300 ms) in the original dry audio is intercepted as the solo dry audio segment of the second singer. 1 min 9 s 700 ms is used as the start time of this solo dry audio segment, and 1 min 20 s 300 ms is used as the end time of this solo dry audio segment.
[0155] After obtaining each solo dry audio segment of the second singer in the original dry audio, according to the start time of each solo dry audio segment of the second singer, each solo dry audio segment of the second singer is subjected to a mixing and splicing process to obtain the dry audio to be synthesized. Among them, the dry audio to be synthesized has the same duration as the original dry audio, and the difference is that the dry audio to be synthesized does not include the solo dry audio of the first singer. As Figure 4 shown, the dry audio to be synthesized includes each solo dry audio segment of the second singer and empty audio, and no longer includes the solo dry audio of the first singer.
[0156] Then, the dry audio to be synthesized and the original accompaniment audio are mixed to obtain the chorus accompaniment audio of the first song.
[0157] In addition, if the original accompaniment audio and the original dry audio of the first song are directly obtained from the storage in Method 3 in Step 303, the lag duration of the start time of the original dry audio relative to the start time of the original accompaniment audio and the lead duration of the end time of the original dry audio relative to the end time of the original accompaniment audio can also be determined. If the lag duration is greater than the second preset duration, start cropping from the beginning of the original accompaniment audio so that the start time of the cropped original accompaniment audio is ahead of the original dry audio by the second preset duration. Similarly, if the lead duration is greater than the second preset duration, start cropping from the end of the original accompaniment audio so that the end time of the cropped original accompaniment audio lags behind the original dry audio by the second preset duration. After the cropping is completed, the first third preset duration of the cropped original accompaniment audio can be subjected to audio fade-in processing, and the last third preset duration can be subjected to audio fade-out processing, where the third preset duration is less than the second preset duration. For example, the second preset duration is 3 s and the third preset duration is 2 s. Then, the processed original accompaniment audio and the dry audio to be synthesized are mixed to obtain the chorus accompaniment audio of the first song.
[0158] For example, the singing start time of the original dry audio is 1 minute and 10 seconds, the singing start time of the original accompaniment audio is 0 minutes and 0 seconds, the singing end time of the original dry audio is 2 minutes and 10 seconds, the singing start time of the original accompaniment audio is 3 minutes and 50 seconds, and the second preset duration is 3 seconds. First, it is determined that the lag duration of the singing start time of the original dry audio relative to the singing start time of the original accompaniment audio is 1 minute and 10 seconds - 0 minutes and 0 seconds = 1 minute and 10 seconds, which is greater than 3 seconds. Then, the part from 0 minutes and 0 seconds to 1 minute and 7 seconds of the original accompaniment audio is cropped. After cropping, the singing start time of the original accompaniment audio is 1 minute and 7 seconds, which is 3 seconds ahead of the singing start time of the original dry audio, which is 1 minute and 10 seconds. It is determined that the leading duration of the singing end time of the original dry audio relative to the singing start time of the original accompaniment audio is 3 minutes and 50 seconds - 2 minutes and 10 seconds = 1 minute and 40 seconds, which is greater than 3 seconds. Then, the part from 2 minutes and 13 seconds to 3 minutes and 50 seconds of the original accompaniment audio is cropped. After cropping, the singing end time of the original accompaniment audio is 2 minutes and 13 seconds, which is 3 seconds behind the singing end time of the original dry audio, which is 2 minutes and 10 seconds.
[0159] Case 2: When the first song includes the solo part of the first singer, the solo part of the second singer, and the combined part of the first singer and the second singer, the original dry audio of the first song includes the solo dry audio of the first singer, the solo dry audio of the second singer, and the chorus dry audio of the first singer and the second singer.
[0160] In this case, the processing of Implementation Method 1 is as follows:
[0161] In the original dry audio, obtain the solo dry audio segment of the second singer and the chorus dry audio segment. According to the singing start time of the solo dry audio segment of the second singer and the singing start time of the chorus dry audio, mix the solo dry audio segment of the second singer, the chorus dry audio segment, and the original accompaniment audio to obtain the chorus accompaniment audio of the first song.
[0162] In Case 2, the method for obtaining the chorus dry audio segment is the same as the method for obtaining the solo dry audio segment of the second singer in Case 1 above, and will not be elaborated here.
[0163] After obtaining the solo dry audio segment of the second singer and the chorus dry audio segment, perform a mixed stream splicing process on the solo dry audio segment of the second singer and the chorus dry audio segment to obtain the dry audio to be synthesized. For example Figure 5As shown, compared with the original dry audio, the dry audio to be synthesized includes solo dry audio segments, chorus dry audio segments, and empty audio of the second singer, and no longer includes the solo dry audio of the first singer. The empty audio replaces the solo dry audio of the first singer.
[0164] Implementation 2: In the original dry audio, adjust the volume of the solo dry audio of the first singer to 0 to obtain the chorus accompaniment audio of the first song.
[0165] First, obtain the segmentation information of the first song. According to the singer identifier in the segmentation information, determine the second lyric segment corresponding to the first singer. Use the start time of the second lyric segment as the second singing start time of the first singer, and use the end time of the second lyric segment as the second singing end time of the first singer. Then, according to the singing start time and singing end time of the solo dry audio segment of the first singer, adjust the volume of the solo dry audio segment of the first singer in the original dry audio to 0 to obtain the audio to be synthesized of the first song. Obtain the chorus accompaniment audio of the first song. Finally, mix the dry audio to be synthesized and the original accompaniment audio to obtain the chorus accompaniment audio of the first song.
[0166] Step 305: The server sends the chorus accompaniment audio of the first song to the terminal.
[0167] In implementation, after obtaining the chorus accompaniment audio of the first song, the server can directly send the chorus accompaniment audio of the first song to the terminal.
[0168] In a possible implementation manner, when the server sends the chorus accompaniment audio of the first song to the terminal, it can also send the lyric information of the first song.
[0169] In another possible implementation manner, the server can also generate a chorus accompaniment video of the first song according to the chorus accompaniment audio, lyric information, and MV (Music Video) of the first song, and send the chorus accompaniment video of the first song to the terminal. Specifically, the method for generating the chorus accompaniment video can be as follows:
[0170] First, obtain the MV of the first song, remove the subtitles and audio in the MV to obtain the accompaniment video to be synthesized. When obtaining the MV of the first song, the MV of the first song can be directly obtained from the pre-established MV library, or the MV of the first song can be obtained from the open resources on the network.
[0171] Then, according to the start time and end time of each lyric segment in the lyric information, each lyric segment in the lyric information is rendered word by word onto the accompaniment video to be synthesized. During rendering, the lyrics sung by different singers can be rendered in different ways to distinguish the lyrics sung by different singers. For example, the lyrics sung by different singers can be rendered in different ways such as different font colors, different font weights, different font sizes, etc. In addition, when rendering the lyrics, the name of the singer of the lyric segment can also be rendered in front of the first word of the corresponding lyric segment, and the rendering method of the singer's name can be different from that of the lyrics to distinguish the singer's name and the lyrics. Refer to Figure 6 , which shows a frame of the chorus accompaniment video. Among them, the lyrics sung by the first singer are represented by bold black font, and the lyrics sung by the second singer are represented in italics, and the font size of the lyrics sung by the second singer is larger than that of the lyrics sung by the first singer. It should be noted that Figure 6 This is only an example, and in actual implementation, various rendering methods that can distinguish the lyrics sung by different singers can be adopted, and the embodiments of the present application do not limit this.
[0172] Finally, the chorus accompaniment audio of the first song and the rendered accompaniment video to be synthesized are combined to obtain the chorus accompaniment video of the first song. Among them, the chorus accompaniment video can be in formats such as MP4 (Moving Picture Experts Group Audio Layer IV).
[0173] In a possible implementation, when rendering each word in the lyrics, the rendering method of this word during the singing time of this word and the rendering method of this word outside the singing time of this word can be different. For example, during the singing time of this word, this word can be highlighted, and outside the singing time of this word, this word is not highlighted. Among them, the highlighting method can be various methods such as adding a background color, bolding, and flashing.
[0174] In another possible implementation, when rendering each lyric segment in the lyrics, the rendering method of this lyric segment during the singing time of this lyric segment and the rendering method of this lyric segment outside the singing time of this lyric segment can be different. For example, during the singing time of this lyric segment, this lyric segment can be highlighted, and outside the singing time of this lyric segment, this lyric segment is not highlighted. Among them, the highlighting method can be various methods such as adding a background color, bolding, and flashing.
[0175] Step 306: The terminal plays the chorus accompaniment audio of the first song.
[0176] In implementation, for the case where in step 305, the server only sends the chorus accompaniment audio of the first song, the terminal can play the chorus accompaniment audio of the first song after the user clicks the start option.
[0177] For the case where in step 305, when the server sends the chorus accompaniment audio of the first song, it also sends the lyrics information of the first song, the terminal can play the chorus accompaniment audio of the first song after the user clicks the start option, and highlight the lyrics according to the lyrics information. Specifically, when the chorus accompaniment audio is played, according to the playing progress of the chorus accompaniment audio, highlight the lyrics whose start time is the current playing progress. For example, if the current playing progress of the chorus accompaniment audio is 1 min 30 s, then highlight the lyrics whose start time is 1 min 30 s. In addition, when each lyric segment is displayed, the name of the singer of this lyric segment can also be displayed at the same time.
[0178] For the case where in step 306, the server sends the chorus accompaniment video of the first song, the terminal can play the chorus accompaniment video of the first song after the user clicks the start option.
[0179] Step 307: During the playing of the chorus accompaniment audio, the terminal collects the user's dry singing audio.
[0180] In implementation, while the terminal is playing the chorus accompaniment audio or the synthesized accompaniment video, it collects the user's dry singing audio through an audio collection device. Among them, the audio collection device can be a built-in microphone, or can be a peripheral such as a wired headset or a wireless headset.
[0181] Step 308: The terminal mixes the chorus accompaniment audio and the dry singing audio to obtain a chorus file.
[0182] In implementation, after the synthesized accompaniment audio or the synthesized accompaniment video is played, the terminal mixes the chorus accompaniment audio and the collected dry singing audio to obtain a chorus audio.
[0183] In a possible implementation, in the case where the server sends a chorus accompaniment video to the terminal, in this step 308, the collected dry singing audio and the synthesized accompaniment video are mixed to obtain a synthesized video.
[0184] In the embodiments of the present application, the server processes the original dry audio of a song to obtain a dry audio to be synthesized, in which there is no solo voice of a first singer, and the first singer is any one of multiple singers of the song. On this basis, the server synthesizes the original accompaniment audio of the song and the dry audio to be synthesized into a chorus accompaniment audio of the song and sends it to the terminal. Since there is no solo voice of the first singer in the dry audio to be synthesized, there is also no solo voice of the first singer in the synthesized chorus accompaniment audio. In this way, the terminal only needs to play the chorus accompaniment audio normally to achieve the effect that there is only accompaniment without the singer's voice in the solo part of the first singer, effectively avoiding the problem of discontinuous connection that occurs when frequently switching between the original dry audio and the original accompaniment audio in the prior art.
[0185] In addition, the chorus accompaniment audio or the synthesized accompaniment video is synthesized by the server and provided to the terminal, and the terminal can directly play it without complex switching processing by the terminal, and the requirements for the terminal performance are relatively low.
[0186] See Figure 7 , in the case where the server generates the chorus accompaniment audio using the above-mentioned timing two, the chorus processing method provided by the embodiments of the present application may include the following steps:
[0187] Step 601, the server obtains the original accompaniment audio, the original dry audio and the lyric information of the first song.
[0188] In implementation, the server can pre-separate the human voice from the song audio of the chorus songs in the song library to obtain the original accompaniment audio and the original dry audio of each chorus song, and store them in the storage device. In addition, each song is pre-configured with lyric information, and correspondingly, the original accompaniment audio, the original dry audio and the lyric information of the song can be stored correspondingly. After receiving the chorus request corresponding to the first song, the original accompaniment audio, the original dry audio and the lyric information of the first song can be directly obtained from the storage device.
[0189] In a possible implementation, the original accompaniment audio, the original dry audio and the lyric information of the song have already been stored in the server. Correspondingly, in this step 601, the server can directly obtain the original accompaniment audio, the original dry audio and the lyric information of the song from the storage device.
[0190] Step 602, the server processes the original dry audio according to the lyric information to obtain a dry audio to be synthesized, and generates a chorus accompaniment audio of the first song according to the original dry audio and the dry audio to be synthesized.
[0191] Among them, there is no singing voice of the first singer during the singing time corresponding to the first singer when the chorus accompaniment audio is played.
[0192] The processing of this step 602 is the same as that of the above step 304, and will not be elaborated here.
[0193] Step 603: The terminal obtains the chorus instruction corresponding to the first song.
[0194] The processing of this step 603 is the same as that of the above step 301, and will not be elaborated here.
[0195] Step 604: The terminal sends a chorus request for the first song to the server.
[0196] The processing of this step 604 is the same as that of the above step 302, and will not be elaborated here.
[0197] Step 605: The server sends the chorus accompaniment audio of the first song to the terminal.
[0198] The processing of this step 605 is the same as that of the above step 305, and will not be elaborated here.
[0199] Step 606: The terminal plays the chorus accompaniment audio of the first song.
[0200] The processing of this step 606 is the same as that of the above step 306, and will not be elaborated here.
[0201] Step 607: During the playing of the chorus accompaniment audio, the terminal collects the user's dry singing audio.
[0202] The processing of this step 607 is the same as that of the above step 307, and will not be elaborated here.
[0203] Step 608: The terminal mixes the chorus accompaniment audio and the dry singing audio to obtain a chorus file.
[0204] The processing of this step 608 is the same as that of the above step 308, and will not be elaborated here.
[0205] In the embodiments of the present application, the server processes the original dry audio of the song to obtain the dry audio to be synthesized. There is no solo voice of the first singer in the dry audio to be synthesized, and the first singer is any one of the multiple singers of the song. On this basis, the server synthesizes the original accompaniment audio of the song and the dry audio to be synthesized into the chorus accompaniment audio of the song and sends it to the terminal. Since there is no solo voice of the first singer in the dry audio to be synthesized, there is also no solo voice of the first singer in the synthesized chorus accompaniment audio. In this way, the terminal only needs to play the chorus accompaniment audio normally to achieve the effect that there is only accompaniment without the singer's voice in the solo part of the first singer, effectively avoiding the problem of discontinuous connection that occurs when frequently switching between the original dry audio and the original accompaniment audio in the prior art.
[0206] Based on the same technical concept, the embodiments of the present application also provide a chorus processing device, which can be the server in the above embodiments, such as Figure 8 As shown, the device includes: an acquisition module 710, a generation module 720, a reception module 730, and a transmission module 740, where:
[0207] The acquisition module 710 is configured to acquire the original accompaniment audio, the original dry audio, and the lyric information of the song, where the lyric information is in the format of word-by-word timestamp lyrics and includes the segmentation information of at least two lyric segments, and each segmentation information includes the start time and the end time of the corresponding lyric segment;
[0208] The generation module 720 is configured to process the original dry audio according to the segmentation information of the at least two lyric segments to obtain the dry audio to be synthesized, where there is no solo voice of the first singer when the dry audio to be synthesized is played, and the singers of the original dry audio include the first singer and the second singer; generate the chorus accompaniment audio of the song according to the original accompaniment audio and the dry audio to be synthesized;
[0209] The reception module 730 is configured to receive the chorus request of the song sent by the terminal;
[0210] The transmission module 740 is configured to send the chorus accompaniment audio of the song to the terminal.
[0211] In a possible implementation manner, the generation module 720 is configured to:
[0212] Determine the first lyric segment sung solo by the second singer among the at least two lyric segments;
[0213] Use the start time and the end time of the first lyric segment as the start time and the end time of the solo dry audio segment sung by the second singer respectively;
[0214] In the original dry audio, according to the singing start time and singing end time of the solo dry audio frequency band of the second singer, intercept the solo dry audio frequency band of the second singer, and perform mixing and splicing on the solo dry audio frequency band of the second singer to obtain the dry audio to be synthesized.
[0215] In a possible implementation manner, the generating module 720 is configured to:
[0216] Determine a first lyric segment sung solo by the second singer from among the at least two lyric segments;
[0217] Subtract a first duration from the start time of the lyric segment sung solo by the second singer as the singing start time of the solo dry audio frequency band of the second singer, and add the first duration to the end time of the lyric segment sung solo by the second singer as the singing end time of the solo dry audio frequency band of the second singer;
[0218] In the original dry audio, according to the singing start time and singing end time of the solo dry audio frequency band of the second singer, intercept the solo dry audio frequency band of the second singer, and perform mixing and splicing on the solo dry audio frequency band of the second singer to obtain the dry audio to be synthesized.
[0219] In a possible implementation manner, the generating module 720 is configured to:
[0220] Determine a first lyric segment sung solo by the first singer from among the at least two lyric segments;
[0221] Respectively use the start time and end time of the first lyric segment as the singing start time and singing end time of the solo dry audio frequency band of the first singer;
[0222] In the original dry audio, according to the singing start time and singing end time of the solo dry audio frequency band of the first singer, adjust the volume of the solo dry audio frequency band of the first singer to 0 to obtain the dry audio to be synthesized.
[0223] In a possible implementation manner, the generating module 720 is configured to:
[0224] Determine the lag duration of the singing start time of the original accompaniment audio compared to the singing start time of the original dry audio;
[0225] Determine the leading duration of the singing end time of the original accompaniment audio compared to the singing end time of the original dry audio;
[0226] If the lag time is greater than the second time duration, a part of the audio is cropped from the beginning of the original accompaniment audio, so that the lag time of the singing start time of the cropped original accompaniment audio compared to the singing start time of the original dry audio is equal to the second time duration;
[0227] If the lead time is greater than the second time duration, a part of the audio is cropped from the end of the original accompaniment audio, so that the lead time of the singing end time of the cropped original accompaniment audio compared to the singing end time of the original dry audio is equal to the second time duration;
[0228] Generate the chorus accompaniment audio of the song according to the cropped original accompaniment audio and the to-be-synthesized dry audio.
[0229] In a possible implementation manner, the sending module 740 is configured to:
[0230] Send the chorus accompaniment audio of the song and the lyric information to the terminal.
[0231] In a possible implementation manner, the generating module 720 is further configured to:
[0232] Obtain the music short film of the song;
[0233] Remove the subtitles and audio in the music short film to obtain a to-be-synthesized accompaniment video;
[0234] Segment the at least two lyric segments in the lyric information and render them word by word onto the to-be-synthesized accompaniment video, where the lyrics sung by different singers displayed in the to-be-synthesized accompaniment video have different colors;
[0235] Synthesize the rendered to-be-synthesized accompaniment video and the chorus accompaniment audio to obtain the chorus accompaniment video of the song;
[0236] The sending module 740 is configured to:
[0237] Send the chorus accompaniment video of the song to the terminal.
[0238] In the embodiments of the present application, the server processes the original dry audio of a song to obtain the dry audio to be synthesized, in which there is no solo voice of a first singer, and the first singer is any one of multiple singers of the song. On this basis, the server synthesizes the original accompaniment audio of the song and the dry audio to be synthesized into the chorus accompaniment audio of the song and sends it to the terminal. Since there is no solo voice of the first singer in the dry audio to be synthesized, there is also no solo voice of the first singer in the synthesized chorus accompaniment audio. In this way, the terminal only needs to play the chorus accompaniment audio normally to achieve the effect that there is only accompaniment without the singer's voice in the solo part of the first singer, effectively avoiding the problem of discontinuous connection that occurs when frequently switching between the original dry audio and the original accompaniment audio in the prior art to achieve this effect.
[0239] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0240] It should be noted that: when the chorus processing device provided in the above embodiments performs chorus processing, only the above-mentioned division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the server is divided into different functional modules to complete all or part of the functions described above. In addition, the chorus processing device provided in the above embodiments and the method embodiments of chorus processing belong to the same concept, and the specific implementation process can be seen in the method embodiments, which will not be elaborated here.
[0241] Based on the same technical concept, the embodiments of the present application also provide a chorus processing device, which can be the terminal in the above embodiments, such as Figure 9 As shown, the device includes: an acquisition module 810, a sending module 820, a receiving module 830, a playing module 840, a collecting module 850, and a mixing module 860, where:
[0242] The acquisition module 810 is used to acquire a chorus instruction of a song;
[0243] The sending module 820 is used to send a chorus request of the song to the server;
[0244] The receiving module 830 is used to receive the chorus accompaniment video of the song sent by the server;
[0245] The playing module 840 is used to play the chorus accompaniment video, where the chorus accompaniment video includes a first singer and a second singer, and there is no singing voice during the solo time corresponding to the first singer when the chorus accompaniment video is played;
[0246] The acquisition module 850 is configured to acquire the user's dry singing audio during the playback of the chorus accompaniment video;
[0247] The mixing module 860 is configured to mix the chorus accompaniment video and the dry singing audio to obtain a chorus video.
[0248] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0249] It should be noted that: when the chorus processing device provided in the above embodiments performs chorus processing, only the above division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the terminal is divided into different functional modules to complete all or part of the functions described above. In addition, the chorus processing device provided in the above embodiments and the method embodiments of chorus processing belong to the same concept. For the specific implementation process, please refer to the method embodiments, which will not be repeated here.
[0250] Figure 10 The block diagram of an electronic device 900 provided by an exemplary embodiment of the present application is shown. The electronic device 900 may be a portable mobile terminal, such as: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player, a laptop computer or a desktop computer. The electronic device 900 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0251] Generally, the electronic device 900 includes: a processor 901 and a memory 902.
[0252] The processor 901 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 901 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 901 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 901 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 901 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0253] The memory 902 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 902 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 902 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 901 to implement the chorus processing method provided in the method embodiments of the present application.
[0254] In some embodiments, the electronic device 900 may further optionally include: a peripheral device interface 903 and at least one peripheral device. The processor 901, the memory 902, and the peripheral device interface 903 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 903 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 904, a display screen 905, a camera assembly 906, an audio circuit 907, a positioning assembly 908, and a power supply 909.
[0255] The peripheral device interface 903 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 901 and the memory 902. In some embodiments, the processor 901, the memory 902, and the peripheral device interface 903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 901, the memory 902, and the peripheral device interface 903 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0256] The radio frequency circuit 904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 904 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 904 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 904 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. The radio frequency circuit 904 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, each generation of mobile communication network (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 904 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0257] The display screen 905 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 905 is a touch display screen, the display screen 905 also has the ability to collect touch signals on or above the surface of the display screen 905. The touch signals can be input as control signals to the processor 901 for processing. At this time, the display screen 905 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there may be one display screen 905, which is provided on the front panel of the electronic device 900; in other embodiments, there may be at least two display screens 905, which are respectively provided on different surfaces of the electronic device 900 or are in a folding design; in other embodiments, the display screen 905 may be a flexible display screen, which is provided on the curved surface or folding surface of the electronic device 900. Even, the display screen 905 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 905 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0258] The camera module 906 is used to capture images or videos. Optionally, the camera module 906 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to implement the function of background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera module 906 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0259] The audio circuit 907 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 901 for processing, or input to the radio frequency circuit 904 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 900. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 901 or the radio frequency circuit 904 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 907 may also include a headphone jack.
[0260] The positioning component 908 is used to locate the current geographical location of the electronic device 900 to achieve navigation or LBS (Location Based Service). The positioning component 908 may be a positioning component based on the GPS (Global Positioning System) of the United States, the Beidou system of China, or the Galileo system of Russia.
[0261] The power supply 909 is used to supply power to each component in the electronic device 900. The power supply 909 may be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 909 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0262] In some embodiments, the electronic device 900 further includes one or more sensors 910. The one or more sensors 910 include but are not limited to: an acceleration sensor 911, a gyroscope sensor 912, a pressure sensor 913, a fingerprint sensor 914, an optical sensor 915, and a proximity sensor 916.
[0263] The acceleration sensor 911 can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established with the electronic device 900. For example, the acceleration sensor 911 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 901 can control the display screen 905 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 911. The acceleration sensor 911 can also be used for collecting game or user's motion data.
[0264] The gyroscope sensor 912 can detect the body orientation and rotation angle of the electronic device 900. The gyroscope sensor 912 can cooperate with the acceleration sensor 911 to collect the 3D actions of the user on the electronic device 900. Based on the data collected by the gyroscope sensor 912, the processor 901 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0265] The pressure sensor 913 can be disposed on the side frame of the electronic device 900 and / or the lower layer of the display screen 905. When the pressure sensor 913 is disposed on the side frame of the electronic device 900, it can detect the holding signal of the user on the electronic device 900, and the processor 901 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 913. When the pressure sensor 913 is disposed on the lower layer of the display screen 905, the processor 901 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 905. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0266] The fingerprint sensor 914 is used to collect the fingerprints of the user. The processor 901 can identify the user's identity according to the fingerprints collected by the fingerprint sensor 914, or the fingerprint sensor 914 can identify the user's identity according to the collected fingerprints. When the identified user identity is a trusted identity, the processor 901 authorizes the user to perform relevant sensitive operations, and the sensitive operations include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 914 can be disposed on the front, back, or side of the electronic device 900. When there are physical buttons or manufacturer logos on the electronic device 900, the fingerprint sensor 914 can be integrated with the physical buttons or manufacturer logos.
[0267] The optical sensor 915 is used to collect the ambient light intensity. In one embodiment, the processor 901 can control the display brightness of the display screen 905 according to the ambient light intensity collected by the optical sensor 915. Specifically, when the ambient light intensity is high, the display brightness of the display screen 905 is increased; when the ambient light intensity is low, the display brightness of the display screen 905 is decreased. In another embodiment, the processor 901 can also dynamically adjust the shooting parameters of the camera module 906 according to the ambient light intensity collected by the optical sensor 915.
[0268] The proximity sensor 916, also known as a distance sensor, is usually disposed on the front panel of the electronic device 900. The proximity sensor 916 is used to collect the distance between the user and the front of the electronic device 900. In one embodiment, when the proximity sensor 916 detects that the distance between the user and the front of the electronic device 900 is gradually decreasing, the processor 901 controls the display screen 905 to switch from the lit state to the off state; when the proximity sensor 916 detects that the distance between the user and the front of the electronic device 900 is gradually increasing, the processor 901 controls the display screen 905 to switch from the off state to the lit state.
[0269] Those skilled in the art can understand that Figure 10 the structure shown in does not constitute a limitation on the electronic device 900, and may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0270] Figure 11 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1000 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 1001 and one or more memories 1002. Among them, at least one instruction is stored in the memory 1002, and the at least one instruction is loaded and executed by the processor 1001 to implement the above-mentioned chorus processing method.
[0271] In an exemplary embodiment, a computer-readable storage medium is further provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by the processor to implement the chorus processing method in the above embodiment. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0272] In an exemplary embodiment, a computer program product is further provided. The computer program product includes at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the method such as chorus processing.
[0273] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0274] The foregoing are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0275] It should be noted that the information involved in the present application (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, videos, audios, lyrics, etc.), and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) are all authorized by users or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the chorus requests, chorus instructions, the user's singing dry audio, the original dry audio, the MV of the song, the lyrics of the song, etc. involved in the present application are all obtained under full authorization.
Claims
1. A method for chorus processing, characterized in that, The method is applied to a server, and the method includes: Obtain the original accompaniment audio, original dry vocal audio, and lyric information of a song. The lyric information is in the format of word-by-word timestamp lyrics and includes segmentation information of at least two lyric segments. Each segmentation information includes the start time and end time of the corresponding lyric segment. Process the original dry vocal audio according to the segmentation information of the at least two lyric segments to obtain a dry vocal audio to be synthesized. When the dry vocal audio to be synthesized is played, there is no solo voice of the first singer. The singers of the original dry vocal audio include the first singer and the second singer. Generate the chorus accompaniment audio of the song according to the original accompaniment audio and the dry vocal audio to be synthesized. Receive the chorus request of the song sent by the terminal. Send the chorus accompaniment audio of the song to the terminal.
2. The method according to claim 1, wherein The processing the original dry vocal audio according to the segmentation information of the at least two lyric segments to obtain a dry vocal audio to be synthesized includes: Determine a first lyric segment sung solo by the second singer among the at least two lyric segments. Respectively use the start time and end time of the first lyric segment as the singing start time and singing end time of the solo dry vocal frequency band of the second singer. In the original dry vocal audio, intercept the solo dry vocal frequency band of the second singer according to the singing start time and singing end time of the solo dry vocal frequency band of the second singer, and perform mixing and splicing on the solo dry vocal frequency band of the second singer to obtain a dry vocal audio to be synthesized.
3. The method according to claim 1, wherein The processing the original dry vocal audio according to the segmentation information of the at least two lyric segments to obtain a dry vocal audio to be synthesized includes: Determine a first lyric segment sung solo by the second singer among the at least two lyric segments. Subtract a first duration from the start time of the lyric segment sung solo by the second singer as the singing start time of the solo dry vocal frequency band of the second singer, and add the first duration to the end time of the lyric segment sung solo by the second singer as the singing end time of the solo dry vocal frequency band of the second singer. In the original dry vocal audio, intercept the solo dry vocal frequency band of the second singer according to the singing start time and singing end time of the solo dry vocal frequency band of the second singer, and perform mixing and splicing on the solo dry vocal frequency band of the second singer to obtain a dry vocal audio to be synthesized.
4. The method according to claim 1, wherein The processing the original dry vocal audio according to the segmentation information of the at least two lyric segments to obtain a dry vocal audio to be synthesized includes: Determine a first lyric segment sung solo by the first singer among the at least two lyric segments. Respectively use the start time and end time of the first lyric segment as the singing start time and singing end time of the solo dry vocal frequency band of the first singer. In the original dry vocal audio, adjust the volume of the solo dry vocal frequency band of the first singer to 0 according to the singing start time and singing end time of the solo dry vocal frequency band of the first singer to obtain a dry vocal audio to be synthesized.
5. The method according to any one of claims 1 to 4, characterized in that, Before generating the chorus accompaniment audio of the song based on the original accompaniment audio and the dry audio to be synthesized, the method further includes: Determine the lag duration of the singing start time of the original accompaniment audio compared to the singing start time of the original dry audio; Determine the leading duration of the singing end time of the original accompaniment audio compared to the singing end time of the original dry audio; If the lag duration is greater than the second duration, trim a part of the audio from the beginning of the original accompaniment audio so that the lag duration of the singing start time of the trimmed original accompaniment audio compared to the singing start time of the original dry audio is equal to the second duration; If the leading duration is greater than the second duration, trim a part of the audio from the end of the original accompaniment audio so that the leading duration of the singing end time of the trimmed original accompaniment audio compared to the singing end time of the original dry audio is equal to the second duration; Generating the chorus accompaniment audio of the song according to the original accompaniment audio and the dry audio to be synthesized includes: Generating the chorus accompaniment audio of the song according to the trimmed original accompaniment audio and the dry audio to be synthesized.
6. The method according to any one of claims 1 to 4, characterized in that, Sending the chorus accompaniment audio of the song to the terminal includes: Sending the chorus accompaniment audio of the song and the lyric information to the terminal.
7. The method according to any one of claims 1-4, characterized in that, Before sending the chorus accompaniment audio of the song to the terminal, the method further includes: Obtain the music short film of the song; Remove the subtitles and audio from the music short film to obtain the accompaniment video to be synthesized; Segment the at least two lyric segments in the lyric information and render them word by word onto the accompaniment video to be synthesized, wherein the lyrics sung by different singers displayed in the accompaniment video to be synthesized have different colors; Synthesize the rendered accompaniment video to be synthesized and the chorus accompaniment audio to obtain the chorus accompaniment video of the song; Sending the chorus accompaniment audio of the song to the terminal includes: Sending the chorus accompaniment video of the song to the terminal.
8. A server, characterized in that, The server includes a processor and a memory, and at least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the chorus processing method according to any one of claims 1 to 7.
9. A system for chorus processing, characterized in that, The system includes the server and the terminal according to claim 8.
10. A computer-readable storage medium, characterized in that, At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the chorus processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Voice synthesis method and system
CN108269560A
Audio calibration method and device and storage medium
CN111785238A