Music generation method, device, electronic device and storage medium

By splitting the target pronunciation into word voice clips and corresponding to the strong beat points in the background music, the problem of the inability to accurately generate rap music in the prior art is solved, and efficient and accurate rap music generation is achieved.

CN115171629BActive Publication Date: 2025-06-06BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110285915.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-17
Publication Date
2025-06-06
Estimated Expiration
2041-03-17

AI Technical Summary

Technical Problem

The prior art cannot accurately and efficiently generate rap music.

Method used

By obtaining the target voice and background music, splitting the target voice into word voice clips, determining the strong beat points in the background music, establishing the correspondence between the word voice clip and the strong beat interval, and fusing the word voice clips with the background music to generate the target rap music.

Benefits of technology

It realizes automatic recognition of paragraphs that can be rapped in the background music and automatic finding of strong beat points, ensuring that the word voice clips are accurately stuck on strong beat points, and improving the accuracy and efficiency of rap music generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115171629B_ABST
    Figure CN115171629B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a music generation method, device, electronic device and storage medium, the method comprising: obtaining target speech and background music; segmenting the speech segments corresponding to each word from the target speech to obtain a sequence of word speech segments; determining the strong beat points in the pure music section of the background music, and the adjacent strong beat nodes constitute a strong beat interval; establishing a correspondence between the word speech segments and the strong beat interval in the word speech segment sequence; fusing the word speech segment sequence and the background music according to the correspondence to obtain the target rap music. The present disclosure improves the accuracy and efficiency of rap music generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a music generation method, device, electronic device and storage medium. Background Art

[0002] Rap music is a musical style, also known as rap, which refers to a special form of singing that speaks rhythmically. It is characterized by quickly telling a series of rhyming sentences against the background of mechanical rhythmic sounds.

[0003] In the related art, speech synthesis technology can be used to perform speech synthesis on text, but it is not possible to accurately and efficiently generate rap music. Therefore, a reliable or effective technical solution is needed to accurately and efficiently generate rap music. Summary of the invention

[0004] The present disclosure provides a music generation method, device, electronic device and storage medium to at least solve the problem in the related art that rap music cannot be generated accurately and efficiently. The technical solution of the present disclosure is as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a music generation method, comprising:

[0006] Get the target voice and background music;

[0007] Segmenting the speech segment corresponding to each word from the target speech to obtain a sequence of word speech segments;

[0008] Determine the strong beat points in the pure music section of the background music, and the adjacent strong beat nodes constitute a strong beat interval;

[0009] Establishing a correspondence between the word speech fragments in the word speech fragment sequence and the strong beat intervals;

[0010] The word speech segment sequence and the background music are fused according to the corresponding relationship to obtain the target rap music.

[0011] In an exemplary embodiment, the establishing of the correspondence between the word speech fragments in the word speech fragment sequence and the strong beat intervals includes:

[0012] Determining the duration of each word speech segment in the word speech segment sequence;

[0013] Determining the duration of the strong beat interval;

[0014] According to the duration and the interval duration, the word speech segments in the word speech segment sequence are sequentially matched to the strong beat intervals;

[0015] Each strong beat interval corresponds to at least one word speech segment, and the sum of the duration of the at least one word speech segment is smaller than the interval duration of the strong beat interval.

[0016] In an exemplary implementation, the acquiring the target speech includes:

[0017] Get the text to be synthesized;

[0018] Inputting the text to be synthesized into a speech synthesis model to obtain an output synthesized audio;

[0019] The synthesized audio is used as the target speech.

[0020] In an exemplary implementation, the acquiring the target speech includes:

[0021] Get the user's voice input;

[0022] The user voice is used as the target voice.

[0023] In an exemplary embodiment, the segmenting the speech segment corresponding to each word from the target speech includes:

[0024] Performing word segmentation processing on the text to be synthesized to obtain a word sequence corresponding to the text to be synthesized;

[0025] Determine the start time and end time of each word in the word sequence in the synthesized audio;

[0026] The synthesized audio is segmented according to the start time and the end time of each word in the synthesized audio to obtain a speech segment corresponding to each word.

[0027] In an exemplary embodiment, the segmenting the speech segment corresponding to each word from the target speech includes:

[0028] Performing speech recognition on the user's speech to obtain a recognition text corresponding to the user's speech;

[0029] Performing word segmentation processing on the recognized text to obtain a word sequence corresponding to the recognized text;

[0030] Determine the start time and end time of each word in the word sequence in the user's speech;

[0031] The user voice is segmented according to the starting time and the ending time of each word in the user voice to obtain a voice segment corresponding to each word.

[0032] In an exemplary embodiment, determining a strong beat point in the pure music section of the background music includes:

[0033] Performing audio event detection on the background music to determine pure music segments in the background music;

[0034] Performing beat detection on the background music to determine the strong beat points of the background music;

[0035] The markers are located at the strong beat points of the pure music passage.

[0036] According to a second aspect of an embodiment of the present disclosure, there is provided a music generating device, comprising:

[0037] A first acquisition unit is configured to acquire target voice and background music;

[0038] A segmentation unit is configured to segment the speech segment corresponding to each word from the target speech to obtain a sequence of word speech segments;

[0039] A strong beat point determination unit is configured to determine the strong beat points in the pure music section in the background music, and the adjacent strong beat nodes constitute a strong beat interval;

[0040] A correspondence establishing unit, configured to establish a correspondence between the word speech fragments in the word speech fragment sequence and the strong beat intervals;

[0041] The fusion unit is configured to fuse the word speech fragment sequence with the background music according to the corresponding relationship to obtain the target rap music.

[0042] In an exemplary implementation, the correspondence establishing unit includes:

[0043] A first duration determination unit is configured to determine the duration of each word speech segment in the word speech segment sequence;

[0044] A second duration determining unit is configured to determine the duration of the strong beat interval;

[0045] A corresponding unit is configured to sequentially correspond the word speech segments in the word speech segment sequence to the strong beat intervals according to the duration and the interval duration;

[0046] Each strong beat interval corresponds to at least one word speech segment, and the sum of the duration of the at least one word speech segment is smaller than the interval duration of the strong beat interval.

[0047] In an exemplary embodiment, the first acquiring unit includes:

[0048] A text acquisition unit, configured to acquire the text to be synthesized;

[0049] The audio synthesis unit is configured to input the text to be synthesized into a speech synthesis model to obtain an output synthesized audio; and use the synthesized audio as the target speech.

[0050] In an exemplary embodiment, the first acquiring unit includes:

[0051] The user voice acquisition unit is configured to acquire the input user voice and use the user voice as the target voice.

[0052] In an exemplary embodiment, the segmentation unit includes:

[0053] A first word segmentation unit is configured to perform word segmentation processing on the text to be synthesized to obtain a word sequence corresponding to the text to be synthesized;

[0054] A first time determination unit is configured to determine the start time and the end time of each word in the word sequence in the synthesized audio;

[0055] The first segmentation unit is configured to segment the synthesized audio according to the start time and end time of each word in the synthesized audio to obtain a speech segment corresponding to each word.

[0056] In an exemplary embodiment, the segmentation unit includes:

[0057] A recognition unit, configured to perform speech recognition on the user's speech to obtain a recognition text corresponding to the user's speech;

[0058] A second word segmentation unit is configured to perform word segmentation processing on the recognized text to obtain a word sequence corresponding to the recognized text;

[0059] A second time determination unit is configured to determine the start time and the end time of each word in the word sequence in the user's voice;

[0060] The second segmentation subunit is configured to segment the user voice according to the start time and the end time of each word in the user voice to obtain a voice segment corresponding to each word.

[0061] In an exemplary embodiment, the strong beat point determination unit includes:

[0062] a pure music segment determination unit, configured to perform audio event detection on the background music to determine the pure music segment in the background music;

[0063] A beat detection unit, configured to perform beat detection on the background music and determine a strong beat point of the background music;

[0064] The strong beat point marking unit is configured to mark the strong beat points located in the pure music section.

[0065] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:

[0066] processor;

[0067] a memory for storing instructions executable by the processor;

[0068] Wherein, the processor is configured to execute the instructions to implement the music generation method of the first aspect mentioned above.

[0069] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the music generation method of the first aspect described above.

[0070] According to a fifth aspect of an embodiment of the present disclosure, there is provided a computer program product, comprising a computer program / instructions, which, when executed by a processor, implement the music generation method of the first aspect.

[0071] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:

[0072] By acquiring the target speech and background music, the speech fragments corresponding to each word are segmented from the target speech to obtain a sequence of word speech fragments, and the strong beat points in the pure sound paragraphs in the background music are determined. Adjacent strong beat points constitute a strong beat area, and then a correspondence between the word speech fragments and the strong beat intervals in the word speech fragment sequence is established. According to the correspondence, the word speech fragments and the background music are fused to obtain the target rap music, thereby realizing automatic recognition of the paragraphs in the background music that can be rapped, and automatically finding the strong beat points to associate the word speech fragments, so that the word speech fragments can be accurately stuck on the strong beat points, thereby improving the accuracy and efficiency of rap music generation.

[0073] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0075] Figure 1 is a schematic diagram of an application environment of a music generation method according to an exemplary embodiment;

[0076] Figure 2 The present invention is a flowchart of a music generating method according to an exemplary embodiment.

[0077] Figure 3 is a flow chart showing a method for determining a strong beat point in a pure music section in background music according to an exemplary embodiment;

[0078] Figure 4 This is an example of marking a strong beat point in a pure music passage according to an exemplary embodiment;

[0079] Figure 5 A flowchart of establishing a correspondence between a word speech segment and a strong beat interval in a word speech segment sequence according to an exemplary embodiment;

[0080] Figure 6 This is an example of establishing a correspondence between a word speech segment and a strong beat interval in a word speech segment sequence according to an exemplary embodiment;

[0081] Figure 7 is a flow chart of another music generating method according to an exemplary embodiment;

[0082] Figure 8 is a flow chart of another music generating method according to an exemplary embodiment;

[0083] Fig. 9 is a block diagram of a music generating device according to an exemplary embodiment;

[0084] Fig.10 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0085] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0086] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0087] See also Figure 1 , which is a schematic diagram of an application environment of a music generation method according to an exemplary embodiment, the application environment may include a terminal 110 and a server 120, and the terminal 110 and the server 120 may be connected via a wired network or a wireless network.

[0088] The terminal 110 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto. The terminal 110 may be installed with client software such as an application (Application, referred to as App) that provides a music processing function. The application may be an application that specifically provides music processing, or may be other applications that have a music processing function, such as a live broadcast application with a music processing function, etc. The user of the terminal 110 may log in to the application using pre-registered user information, which may include an account number and a password.

[0089] The server 120 may be a server that provides background services for the application in the terminal 110 , or may be another server that is connected and communicated with the background server of the application. It may be a single server or a server cluster composed of multiple servers.

[0090] Exemplarily, the terminal 110 can execute the rap music generation method provided by the embodiment of the present disclosure through an application with a music processing function running thereon. It is understandable that the terminal 110 can also upload the target voice and background music to the server 120 through the above application, and the server 120 executes the rap music generation method provided by the embodiment of the present disclosure and returns the result to the terminal 110, which is not specifically limited in the embodiment of the present disclosure.

[0091] It can be seen that the rap music generation method of the embodiment of the present disclosure can be executed by an electronic device, which can be a terminal or a server. It can be executed by the terminal or the server alone, or by the terminal and the server working together.

[0092] Figure 2is a flowchart of a music generation method according to an exemplary embodiment. Figure 2 As shown, the music generation method is used for Figure 1 Taking the terminal shown in the figure as an example, the following steps are included:

[0093] In step S21, the target voice and background music are acquired.

[0094] Among them, the target voice refers to the voice containing text content used to generate rap music. The target voice can be the user voice obtained by recording the voice of the user reading text through a sound acquisition device. The text read by the user can be but not limited to lyrics text; the target voice can also be synthesized audio obtained based on the text to be synthesized using speech synthesis technology. The text to be merged can be the text selected and input by the user, and can be but not limited to lyrics text.

[0095] The background music may be obtained by responding to a user's selection operation on a music file stored in the terminal and obtaining the music file selected by the user as the background music; or by responding to a user's search operation on the Internet and obtaining the music file selected by the user from the search results as the background music.

[0096] In step S22, the speech segment corresponding to each word is segmented from the target speech to obtain a sequence of word speech segments.

[0097] The order of the word speech fragments in the word speech fragment sequence is consistent with the order in which the corresponding words appear in the target speech.

[0098] In step S23, the strong beat points in the pure music section of the background music are determined, and the adjacent strong beat nodes constitute a strong beat interval.

[0099] In the embodiments of the present disclosure, the acquired background music may be music without vocals, that is, the background music may be pure music or music containing vocal parts. When the background music contains vocal parts, since the vocal parts are not suitable for generating rap music, the strong beat points in the pure music section of the background music may be determined, and then the rap music may be generated based on the strong beat points in the pure music section, thereby improving the accuracy and quality of the generated rap music.

[0100] Based on this, in an exemplary embodiment, the above step S23 may include: Figure 3 The following steps:

[0101] In step S231, audio event detection is performed on the background music to determine pure music segments in the background music.

[0102] A pure music passage refers to a music passage in the background music that does not contain lyrics. We can determine whether the music passage is a pure music passage based on whether it contains vocals.

[0103] Exemplarily, when performing audio event detection on background music, human voice detection can be performed on the background music, and music segments containing human voices can be determined, and then pure music segments that do not contain human voices can be determined based on the music segments containing human voices. In a specific implementation, a human voice detection network model can be pre-trained, and the background music is input into the human voice detection network model to obtain an output music segment identifier containing human voices. The music segment identifier can be the position information of the music segment containing human voices in the background music, and then the pure music segment in the background music can be determined by the difference between the background music and the position information. Among them, the human voice detection network model can be a deep learning model.

[0104] In step S232, the beat of the background music is detected to determine the strong beat points of the background music.

[0105] The strong beat point indicates the starting moment of the strong beat in the background music. In the disclosed embodiment, in order to improve the accuracy of determining the strong beat points in the pure music segment, the beat detection can be performed on the whole background music to determine all the strong beat points in the background music, and then the strong beat points in the pure music segment can be determined based on all the strong beat points in the background music.

[0106] Exemplarily, candidate beat points may be determined according to the energy intensity value of the background music, and then a strong beat point may be selected from the candidate beat points according to the time interval between the frames where two adjacent candidate beat points are located.

[0107] Exemplarily, the background music may be filtered, and the filtered background music may be Fourier transformed to obtain a corresponding spectrum, and then the energy change value of each detection point may be determined based on the spectrum, and the strong beat points in the detection points may be determined based on the energy change values.

[0108] In step S233, the strong beat points of the pure music segment are marked.

[0109] The marking method may be to record the starting time of each strong beat point in the pure music segment, and mark the position of the strong beat point on the time axis of the background music according to the starting time of each strong beat point in the pure music segment, and adjacent strong beat points constitute a strong beat interval. Figure 4 An example of marking strong beat points in a pure music segment is shown, where B1, B2, and B3 correspond to strong beat points in a pure music segment respectively, [B1, B2] constitute a strong beat interval, and [B2, B3] constitute another strong beat interval.

[0110] The disclosed embodiment improves the generation efficiency of rap music while also improving the accuracy of the generated rap music by automatically detecting pure music segments in the background music where rap can be performed and automatically finding strong beats in the pure music segments.

[0111] It is understandable that in actual applications, the above step S23 may be executed first and then the above step S22. The embodiment of the present disclosure does not specifically limit the execution order of step S22 and step S23.

[0112] In step S24, a correspondence between the word speech fragments in the word speech fragment sequence and the strong beat intervals is established.

[0113] Specifically, each word speech segment may be sequentially allocated to a corresponding strong beat interval according to the arrangement order of the word speech segments in the word speech segment sequence.

[0114] In order to improve the accuracy of the strong beat points in the embodiment of the present disclosure, and thus improve the effect of rap music, it is required that the word voice segment cannot cross the strong beat points. Based on this, in an exemplary embodiment, the above step S24 may include when establishing the correspondence between the word voice segment and the strong beat interval in the word voice segment sequence. Figure 5 The following steps:

[0115] In step S241, the duration of each word speech segment in the word speech segment sequence is determined.

[0116] In step S242, the duration of the strong beat section is determined.

[0117] A strong beat interval is composed of two adjacent strong beat points, so the interval length of each strong beat interval can be obtained by the difference between the timestamps corresponding to the two strong beat points constituting the strong beat interval.

[0118] In step S243, the word speech segments in the word speech segment sequence are sequentially matched to the strong beat intervals according to the duration and the interval duration.

[0119] Each strong beat interval corresponds to at least one word speech segment, and the sum of the durations of the at least one word speech segment is less than the interval duration of the strong beat interval.

[0120] Specifically, starting from the first strong beat interval along the time axis, the word voice segments in the word voice segment can be sequentially matched to each strong beat interval, while ensuring that the sum of the duration of at least one word voice segment corresponding to each strong beat interval is less than the interval duration of the strong beat interval, so as to avoid the situation where the word voice segment crosses the strong beat point. It can be understood that at least one word voice segment corresponding to each strong beat interval is actually a word voice segment subsequence, and for each strong beat interval, the first strong beat point of the strong beat interval can be used as the starting time point of its corresponding word voice segment subsequence.

[0121] like Figure 6 FIG. 1 is an example of establishing a correspondence between a word speech segment and a strong beat interval in a word speech segment sequence. The word speech segment sequence is A1, A2, A3. Figure 6 In the figure, the length of the box represents the duration of the word speech segment, and the distance between adjacent strong beat points represents the duration of the corresponding strong beat interval. Starting from B1, the word speech segment sequence A1, A2, A3 is mapped to the strong beat intervals [B1, B2] and [B2, B3] in turn. Since T A1 +T A2 <T [B1,B2] <T A1 +T A2 +T A3 , where T A Indicates the duration of the speech segment of the word, T [,] Indicates the duration of the strong beat interval. Therefore, the word voice fragments A1 and A2 can be mapped to the strong beat interval [B1, B2]. If the word voice fragment A3 is also mapped to the strong beat interval [B1, B2], the word voice fragment A3 will cross the strong beat point B2. In order to improve the accuracy of the generated rap music, the word voice fragment A3 can be mapped to the strong beat interval [B2, B3], so that the correspondence between the word voice fragment sequence A1, A2, A3 and the strong beat interval can be obtained, that is, A1+A2->[B1, B2], A3->[B2, B3].

[0122] When establishing the above-mentioned corresponding relationship, the embodiment of the present disclosure avoids the word voice fragments crossing the strong beat points by making the sum of the duration of at least one word voice fragment corresponding to each strong beat interval smaller than the interval duration of the strong beat interval, thereby improving the accuracy of the strong beat points and further improving the accuracy of the generated rap music.

[0123] In step S25, the word speech segment sequence and the background music are fused according to the corresponding relationship to obtain the target rap music.

[0124] Specifically, the word speech fragment sequence and the background music are mixed according to the correspondence between each word speech fragment and the strong beat interval in the word speech fragment sequence, so as to obtain the target rap music, wherein the mixing processing is to integrate the word speech fragment and the background music into a stereo track or a mono track.

[0125] It can be seen from the above technical solutions of the embodiments of the present disclosure that the embodiments of the present disclosure realize automatic recognition of passages in background music that can be rapped, and automatically find strong beat points to associate word voice fragments, so that the word voice fragments can accurately hit the strong beat points in the pure music passages, thereby improving the accuracy and efficiency of rap music generation.

[0126] In order to improve the generation flexibility of rap music, in an exemplary embodiment, Figure 7 A flowchart of another music generation method is provided, which may include:

[0127] In step S71, the text to be synthesized is obtained.

[0128] The text to be synthesized may be a lyric text selected by the user from a lyric text list, or may be a text of other content input by the user.

[0129] In step S72, the text to be synthesized is input into a speech synthesis model to obtain output synthesized audio.

[0130] Among them, the speech synthesis model is a pre-trained machine learning model, which can convert the input text into corresponding synthesized audio.

[0131] Exemplarily, the speech synthesis model can be an end-to-end speech synthesis model, which includes a Mel spectrum prediction model part and a speech waveform conversion model part, wherein the Mel spectrum prediction model part can use an encoder-decoder model structure with an attention mechanism to predict the Mel spectrum according to the text to be synthesized, and the speech waveform conversion model part can use a convolutional unit (Convolutional Bank with Highway networks and Grated recurrent unit, CBHG) module with a multi-layer highway network and a bidirectional gated recurrent unit to convert the Mel spectrum into a spectrum amplitude, and then use the Griffin-Lim algorithm to perform phase prediction based on the obtained spectrum amplitude to reconstruct the speech waveform to obtain synthesized audio.

[0132] In step S73, the text to be synthesized is segmented to obtain a word sequence corresponding to the text to be synthesized.

[0133] Specifically, a word segmentation tool can be used to perform word segmentation on the text to be synthesized. The word segmentation tool can be, for example, the JieBa word segmentation tool. After the word segmentation processing, a word sequence corresponding to the text to be synthesized can be obtained. For example, after the word segmentation processing of the text "I want to eat hot pot", the word sequence "I", "want", and "eat hot pot" can be obtained.

[0134] In step S74, the start time and the end time of each word in the word sequence in the synthesized audio are determined.

[0135] Specifically, the synthesized audio can be input into a speech recognition model, and the start and end times of each word in the word sequence in the synthesized audio can be determined by the speech recognition model. The speech recognition model can be a hidden Markov model, a deep neural network model, etc.

[0136] In step S75, the synthesized audio is segmented according to the start time and end time of each word in the synthesized audio to obtain word voice segments corresponding to each word.

[0137] For example, in the above word sequence, the starting time and ending time of "I" in the synthesized audio are 0 second and 1 second respectively, the starting time and ending time of "want" in the synthesized audio are 1 second and 3 seconds respectively, and the starting time and ending time of "eat hot pot" in the synthesized audio are 3 seconds and 6 seconds respectively. When the synthesized audio is segmented according to the starting time and ending time of each word in the synthesized audio, the voice segment corresponding to the word "I" can be segmented into 0-1 second voice segments corresponding to the word "want", 1-3 second voice segments corresponding to the word "eat hot pot", and 3-6 second voice segments corresponding to the word "eat hot pot", so that a sequence of word voice segments sorted according to the word sequence can be obtained.

[0138] In step S76, the strong beat points in the acquired background music located in the pure music section are determined, and the adjacent strong beat nodes constitute a strong beat interval.

[0139] In step S77, a correspondence between the word speech fragments in the word speech fragment sequence and the strong beat intervals is established.

[0140] In step S78, the word speech segment sequence and the background music are fused according to the corresponding relationship to obtain the target rap music.

[0141] The detailed implementation of the above steps S76 to S78 can be found in the above Figure 2 The relevant contents of the method embodiment shown will not be repeated here.

[0142] The embodiments of the present disclosure can automatically generate a set of rap music according to the text provided by the user, which not only improves the accuracy of the rap music generation, but also improves the efficiency and flexibility of the rap music generation.

[0143] In order to further improve the generation flexibility of rap music, in another exemplary embodiment, Figure 8 A flowchart of another music generation method is provided, which may include:

[0144] In step S81, the input user voice is acquired.

[0145] Specifically, the user's voice when reading the text may be collected by a sound collection device to obtain the user's voice.

[0146] In step S82, speech recognition is performed on the user speech to obtain a recognition text corresponding to the user speech.

[0147] In step S83, the recognized text is segmented to obtain a word sequence corresponding to the recognized text.

[0148] Specifically, a word segmentation processing tool such as JieBa may be used to perform word segmentation processing on the recognition text to obtain a word sequence corresponding to the recognition text.

[0149] In step S84, the start time and the end time of each word in the word sequence in the user's voice are determined.

[0150] Specifically, the user voice can be input into a voice recognition model, and the start and end time of each word in the word sequence in the user voice can be determined by the voice recognition model. The voice recognition model can be a hidden Markov model, a deep neural network model, etc.

[0151] In step S85, the user voice is segmented according to the start time and end time of each word in the user voice to obtain a word voice segment corresponding to each word.

[0152] For details, please refer to the relevant description in the aforementioned step S75, which will not be repeated here.

[0153] In step S86, the strong beat points in the acquired background music located in the pure music section are determined, and the adjacent strong beat nodes constitute a strong beat interval.

[0154] In step S87, a correspondence between the word speech fragments in the word speech fragment sequence and the strong beat intervals is established.

[0155] In step S88, the word speech segment sequence and the background music are fused according to the corresponding relationship to obtain the target rap music.

[0156] Specifically, the detailed implementation of the above steps S86 to S88 can be found in the above Figure 2 The relevant contents of the method embodiment shown will not be repeated here.

[0157] The embodiments of the present disclosure can also accurately generate rap music when the user provides a read-aloud version of the text, so that the user can automatically generate his or her own rap music using his or her own voice, which greatly improves the flexibility of generating rap music.

[0158] Fig. 9 is a block diagram of a music generating device according to an exemplary embodiment. Fig. 9 The music generating device 900 includes a first acquiring unit 910, a segmenting unit 920, a strong beat point determining unit 930, a corresponding relationship establishing unit 940 and a fusion unit 950.

[0159] The first acquisition unit 910 is configured to acquire target voice and background music;

[0160] The segmentation unit 920 is configured to segment the speech segment corresponding to each word from the target speech to obtain a sequence of word speech segments;

[0161] The strong beat point determination unit 930 is configured to determine the strong beat points in the pure music section of the background music, and the adjacent strong beat nodes constitute a strong beat interval;

[0162] The corresponding relationship establishing unit 940 is configured to establish a corresponding relationship between the word speech segment in the word speech segment sequence and the strong beat interval;

[0163] The fusion unit 950 is configured to perform fusion of the word speech segment sequence and the background music according to the corresponding relationship to obtain the target rap music.

[0164] In an exemplary implementation, the correspondence establishing unit 940 includes:

[0165] A first duration determination unit is configured to determine the duration of each word speech segment in the word speech segment sequence;

[0166] A second duration determining unit is configured to determine the duration of the strong beat interval;

[0167] A corresponding unit is configured to sequentially correspond the word speech segments in the word speech segment sequence to the strong beat intervals according to the duration and the interval duration;

[0168] Each strong beat interval corresponds to at least one word speech segment, and the sum of the duration of the at least one word speech segment is smaller than the interval duration of the strong beat interval.

[0169] In an exemplary implementation, the first acquisition unit 910 includes:

[0170] A text acquisition unit, configured to acquire the text to be synthesized;

[0171] The audio synthesis unit is configured to input the text to be synthesized into a speech synthesis model to obtain an output synthesized audio; and use the synthesized audio as the target speech.

[0172] In an exemplary implementation, the first acquisition unit 910 includes:

[0173] The user voice acquisition unit is configured to acquire the input user voice and use the user voice as the target voice.

[0174] In an exemplary embodiment, the segmentation unit 920 includes:

[0175] A first word segmentation unit is configured to perform word segmentation processing on the text to be synthesized to obtain a word sequence corresponding to the text to be synthesized;

[0176] A first time determination unit is configured to determine the start time and the end time of each word in the word sequence in the synthesized audio;

[0177] The first segmentation unit is configured to segment the synthesized audio according to the start time and end time of each word in the synthesized audio to obtain a speech segment corresponding to each word.

[0178] In an exemplary embodiment, the segmentation unit 920 includes:

[0179] A recognition unit, configured to perform speech recognition on the user's speech to obtain a recognition text corresponding to the user's speech;

[0180] A second word segmentation unit is configured to perform word segmentation processing on the recognized text to obtain a word sequence corresponding to the recognized text;

[0181] A second time determination unit is configured to determine the start time and the end time of each word in the word sequence in the user's voice;

[0182] The second segmentation subunit is configured to segment the user voice according to the start time and the end time of each word in the user voice to obtain a voice segment corresponding to each word.

[0183] In an exemplary embodiment, the strong beat point determination unit 930 includes:

[0184] a pure music segment determination unit, configured to perform audio event detection on the background music to determine the pure music segment in the background music;

[0185] A beat detection unit, configured to perform beat detection on the background music and determine a strong beat point of the background music;

[0186] The strong beat point marking unit is configured to mark the strong beat points located in the pure music section.

[0187] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0188] The rap music generation device of the disclosed embodiment obtains target speech and background music, separates the speech segment corresponding to each word from the target speech, obtains a word speech segment sequence, and determines the strong beat points in the pure sound segment of the background music, and the adjacent strong beat points constitute a strong beat area, and then establishes a correspondence between the word speech segments and the strong beat intervals in the word speech segment sequence, and fuses the word speech segments and the background music according to the correspondence to obtain the target rap music, thereby realizing automatic recognition of the paragraphs in the background music that can be rapped, and automatically finding the strong beat points to associate the word speech segments, so that the word speech segments can be accurately stuck on the strong beat points, thereby improving the accuracy and efficiency of rap music generation.

[0189] In an exemplary embodiment, an electronic device is also provided, including a processor; a memory for storing processor executable instructions; wherein, when the processor is configured to execute the instructions stored in the memory, the music generation method provided in any of the above embodiments is implemented.

[0190] The electronic device may be a terminal, a server or a similar computing device. For example, the electronic device is a terminal. Fig.10 is a block diagram of an electronic device for generating music according to an exemplary embodiment. Specifically:

[0191] The terminal may include components such as an RF (Radio Frequency) circuit 1010, a memory 1020 including one or more computer-readable storage media, an input unit 1030, a display unit 1040, a sensor 1050, an audio circuit 1060, a WiFi (wireless fidelity) module 1070, a processor 1080 including one or more processing cores, and a power supply 1090. Those skilled in the art will appreciate that Fig.10 The terminal structure shown in the figure does not constitute a limitation on the terminal, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0192] The RF circuit 1010 can be used for receiving and sending signals during information transmission or calls. In particular, after receiving the downlink information of the base station, it is handed over to one or more processors 1080 for processing; in addition, the data related to the uplink is sent to the base station. Generally, the RF circuit 1010 includes but is not limited to an antenna, at least one amplifier, a tuner, one or more oscillators, a user identity module (SIM) card, a transceiver, a coupler, an LNA (Low Noise Amplifier), a duplexer, etc. In addition, the RF circuit 1010 can also communicate with the network and other terminals through wireless communication. The wireless communication can use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), LTE (Long Term Evolution), email, SMS (Short Messaging Service), etc.

[0193] The memory 1020 can be used to store software programs and modules, and the processor 1080 executes various functional applications and data processing by running the software programs and modules stored in the memory 1020. The memory 1020 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, application programs required for functions, etc.; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 1020 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 1020 may also include a memory controller to provide the processor 1080 and the input unit 1030 with access to the memory 1020.

[0194] The input unit 1030 can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control. Specifically, the input unit 1030 may include a touch-sensitive surface 1031 and other input devices 1032. The touch-sensitive surface 1031, also known as a touch display screen or a touch pad, can collect user touch operations on or near it (such as operations performed by the user using any suitable object or accessory such as a finger, stylus, etc. on or near the touch-sensitive surface 1031), and drive the corresponding connection device according to a pre-set program. Optionally, the touch-sensitive surface 1031 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch direction, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into touch point coordinates, and then sends it to the processor 1080, and can receive and execute commands sent by the processor 1080. In addition, the touch-sensitive surface 1031 may be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface 1031, the input unit 1030 may further include other input devices 1032. Specifically, the other input devices 1032 may include, but are not limited to, one or more of a physical keyboard, a function key (such as a volume control key, a switch key, etc.), a trackball, a mouse, a joystick, and the like.

[0195] The display unit 1040 can be used to display information input by the user or information provided to the user and various graphical user interfaces of the terminal, which can be composed of graphics, text, icons, videos and any combination thereof. The display unit 1040 may include a display panel 1041, and optionally, the display panel 1041 may be configured in the form of LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), etc. Further, the touch-sensitive surface 1031 may cover the display panel 1041, and when the touch-sensitive surface 1031 detects a touch operation on or near it, it is transmitted to the processor 1080 to determine the type of the touch event, and then the processor 1080 provides corresponding visual output on the display panel 1041 according to the type of the touch event. Among them, the touch-sensitive surface 1031 and the display panel 1041 can be two independent components to realize the input and output functions, but in some embodiments, the touch-sensitive surface 1031 and the display panel 1041 can also be integrated to realize the input and output functions.

[0196] The terminal may also include at least one sensor 1050, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor may adjust the brightness of the display panel 1041 according to the brightness of the ambient light, and the proximity sensor may turn off the display panel 1041 and / or the backlight when the terminal is moved to the ear. As a type of motion sensor, the gravity acceleration sensor can detect the magnitude of acceleration in each direction (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that identify the terminal posture (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc. that can be configured in the terminal, they will not be repeated here.

[0197] The audio circuit 1060, the speaker 1061, and the microphone 1062 can provide an audio interface between the user and the terminal. The audio circuit 1060 can transmit the received audio data to the speaker 1061 after converting the received audio data into an electrical signal, which is converted into a sound signal for output by the speaker 1061; on the other hand, the microphone 1062 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1060 and converted into audio data, and then the audio data is output to the processor 1080 for processing, and then sent to another terminal through the RF circuit 1010, or the audio data is output to the memory 1020 for further processing. The audio circuit 1060 may also include an earplug jack to provide communication between an external headset and the terminal.

[0198] WiFi is a short-range wireless transmission technology. The terminal can help users send and receive emails, browse web pages, and access streaming media through the WiFi module 1070, which provides users with wireless broadband Internet access. Fig.10 A WiFi module 1070 is shown, but it is understandable that it is not an essential component of the terminal and can be omitted as needed without changing the essence of the invention.

[0199] The processor 1080 is the control center of the terminal, and uses various interfaces and lines to connect various parts of the entire terminal. By running or executing software programs and / or modules stored in the memory 1020, and calling data stored in the memory 1020, the processor 1080 performs various functions of the terminal and processes data, thereby monitoring the terminal as a whole. Optionally, the processor 1080 may include one or more processing cores; preferably, the processor 1080 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 1080.

[0200] The terminal also includes a power supply 1090 (such as a battery) for supplying power to each component. Preferably, the power supply can be logically connected to the processor 1080 through a power management system, so that the power management system can manage charging, discharging, and power consumption management. The power supply 1090 can also include any components such as one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, and power status indicators.

[0201] Although not shown, the terminal may also include a camera, a Bluetooth module, etc., which will not be described in detail herein. Specifically in this embodiment, the terminal further includes a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors. The one or more programs include instructions for executing the music generation method provided in the above method embodiment.

[0202] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 1020 including instructions, and the instructions can be executed by a processor 1080 of the device 1000 to complete the above method. Alternatively, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0203] In an exemplary embodiment, a computer program product is also provided, including a computer program / instruction, which, when executed by a processor, implements the music generation method provided in any of the above embodiments.

[0204] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0205] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A music generation method, It is characterized in that include: Get the target voice and background music; Segmenting the speech segment corresponding to each word from the target speech to obtain a sequence of word speech segments; Determine a strong beat point in the pure music section of the background music, where adjacent strong beat points constitute a strong beat interval; the strong beat point indicates the starting time of a strong beat in the background music; Establishing a correspondence between the word speech fragments in the word speech fragment sequence and the strong beat intervals; According to the corresponding relationship, the word voice segment sequence and the background music are merged to obtain the target rap music; Wherein, the establishing of the correspondence between the word speech segments in the word speech segment sequence and the strong beat intervals includes: Determine the duration of each word speech segment in the word speech segment sequence; determine the interval duration of the strong beat interval; according to the duration and the interval duration, correspond the word speech segments in the word speech segment sequence to the strong beat intervals in sequence; wherein each strong beat interval corresponds to at least one word speech segment, and the sum of the durations of the at least one word speech segment is less than the interval duration of the strong beat interval.

2. The music generation method according to claim 1, It is characterized in that The acquiring of the target voice comprises: Get the text to be synthesized; Inputting the text to be synthesized into a speech synthesis model to obtain an output synthesized audio; The synthesized audio is used as the target speech.

3. The music generating method according to claim 1, It is characterized in that The acquiring of the target voice comprises: Get the user's voice input; The user voice is used as the target voice.

4. The music generating method according to claim 2, It is characterized in that The step of extracting a speech segment corresponding to each word from the target speech comprises: Performing word segmentation processing on the text to be synthesized to obtain a word sequence corresponding to the text to be synthesized; Determine the start time and end time of each word in the word sequence in the synthesized audio; The synthesized audio is segmented according to the start time and the end time of each word in the synthesized audio to obtain a speech segment corresponding to each word.

5. The music generating method according to claim 3, It is characterized in that The step of extracting a speech segment corresponding to each word from the target speech comprises: Performing speech recognition on the user's speech to obtain a recognition text corresponding to the user's speech; Performing word segmentation processing on the recognized text to obtain a word sequence corresponding to the recognized text; Determine the start time and end time of each word in the word sequence in the user's speech; The user voice is segmented according to the starting time and the ending time of each word in the user voice to obtain a voice segment corresponding to each word.

6. The music generating method according to claim 1, It is characterized in that Determining the strong beat point in the pure music section of the background music includes: Performing audio event detection on the background music to determine pure music segments in the background music; Performing beat detection on the background music to determine the strong beat points of the background music; The markers are located at the strong beat points of the pure music passage.

7. A music generating device, It is characterized in that include: A first acquisition unit is configured to acquire target voice and background music; A segmentation unit is configured to segment the speech segment corresponding to each word from the target speech to obtain a sequence of word speech segments; A strong beat point determination unit is configured to determine the strong beat points in the pure music section of the background music, and the adjacent strong beat points constitute a strong beat interval; The strong beat point indicates the starting moment of the strong beat in the background music; A correspondence establishing unit, configured to establish a correspondence between the word speech fragments in the word speech fragment sequence and the strong beat intervals; A fusion unit is configured to fuse the word speech segment sequence with the background music according to the corresponding relationship to obtain target rap music; Wherein, the corresponding relationship establishing unit includes: A first duration determination unit is configured to determine the duration of each word speech segment in the word speech segment sequence; A second duration determining unit is configured to determine the duration of the strong beat interval; A corresponding unit is configured to sequentially correspond the word speech segments in the word speech segment sequence to the strong beat intervals according to the duration and the interval duration; Each strong beat interval corresponds to at least one word speech segment, and the sum of the duration of the at least one word speech segment is smaller than the interval duration of the strong beat interval.

8. The music generating device according to claim 7, It is characterized in that The first acquiring unit includes: A text acquisition unit, configured to acquire the text to be synthesized; The audio synthesis unit is configured to input the text to be synthesized into a speech synthesis model to obtain an output synthesized audio; and use the synthesized audio as the target speech.

9. The music generating device according to claim 7, It is characterized in that The first acquiring unit includes: The user voice acquisition unit is configured to acquire the input user voice and use the user voice as the target voice.

10. The music generating device according to claim 8, It is characterized in that The segmentation unit comprises: A first word segmentation unit is configured to perform word segmentation processing on the text to be synthesized to obtain a word sequence corresponding to the text to be synthesized; A first time determination unit is configured to determine the start time and the end time of each word in the word sequence in the synthesized audio; The first segmentation unit is configured to segment the synthesized audio according to the start time and end time of each word in the synthesized audio to obtain a speech segment corresponding to each word.

11. The music generating device according to claim 9, It is characterized in that The segmentation unit comprises: A recognition unit, configured to perform speech recognition on the user's speech to obtain a recognition text corresponding to the user's speech; A second word segmentation unit is configured to perform word segmentation processing on the recognized text to obtain a word sequence corresponding to the recognized text; A second time determination unit is configured to determine the start time and the end time of each word in the word sequence in the user's voice; The second segmentation subunit is configured to segment the user voice according to the start time and the end time of each word in the user voice to obtain a voice segment corresponding to each word.

12. The music generating device according to claim 7, It is characterized in that The strong beat point determination unit comprises: a pure music segment determination unit, configured to perform audio event detection on the background music to determine the pure music segment in the background music; A beat detection unit, configured to perform beat detection on the background music and determine a strong beat point of the background music; The strong beat point marking unit is configured to mark the strong beat points located in the pure music section.

13. An electronic device, It is characterized in that include: processor; a memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the music generation method as described in any one of claims 1 to 6.

14. A computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to execute the music generation method as described in any one of claims 1 to 6.

15. A computer program product comprising a computer program / instructions, It is characterized in that When the computer program / instructions are executed by a processor, the music generation method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Device and method for conversing voice to be rap music

    CN101399036A

  • Music automatic generation system based on machine learning technology

    CN105976802A