Method and device for generating media data, equipment and medium
By providing a method and device for generating media data for music melody, the complex operation of existing music production tools is solved, and the need for simple music creation of ordinary users is realized. The generated media data has designated timbre and melody.
Patent Information
- Application Number
- CN202311618525.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
The existing music production tools are complex in operation and require users to have professional music knowledge, which is difficult to meet the simple music creation needs of ordinary users.
A method and device are provided for generating media data including a musical melody. The method includes responsive to receiving a request to create music, obtaining the first media data, and generating the second media data based on it. The device includes a data acquisition module, a template acquisition module and a generation module for processing media data and music templates uploaded by the user, and generating media data with specified tone and melody.
By simplifying the music creation process, ordinary users can generate music works with desired tones and melodies in a simpler and more effective way, meeting the simple music creation needs of ordinary users.
Smart Images

Figure CN120071872A_ABST
Abstract
Description
Technical Field
[0001] Exemplary implementations of the present disclosure generally relate to data processing, and in particular, to methods, apparatuses, devices, and computer-readable storage media for generating media data including music. Background Art
[0002] In the field of music production, digital synthesis technology has been proposed to create music works. For example, musicians can use tools such as samplers to collect sounds and add these sounds to music works through digital synthesis technology to make the music works have richer auditory effects. However, the operation of existing music production tools is complex and requires users to have rich professional music knowledge, which is not user-friendly for ordinary users. Summary of the Invention
[0003] In a first aspect of the present disclosure, a method for generating media data is provided. In this method, in response to receiving a creation request for creating music, first media data is acquired. A music template is acquired, and the music template includes melody data for specifying a music melody. Based on the first media data, second media data including the music melody is generated.
[0004] In a second aspect of the present disclosure, an apparatus for generating media data is provided. The apparatus includes: a data acquisition module configured to acquire first media data in response to receiving a creation request for creating music; a template acquisition module configured to acquire a music template, the music template including melody data for specifying a music melody; and a generation module configured to generate second media data including the music melody based on the first media data.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory, at least one memory being coupled to at least one processing unit and storing instructions for execution by at least one processing unit, the instructions causing the electronic device to execute the method according to the first aspect of the present disclosure when executed by at least one processing unit.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, the computer program causing a processor to implement the method according to the first aspect of the present disclosure when executed by the processor.
[0007] It should be understood that the content described in this content part is not intended to limit the key features or important features of the implementations of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Brief Description of the Drawings
[0008] In the following, with reference to the accompanying drawings and in conjunction with the following detailed description, the above and other features, advantages, and aspects of various implementations of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0009] Figure 1 A block diagram of an application environment according to an exemplary implementation of the present disclosure is shown;
[0010] Figure 2 A block diagram of a media data generation system according to some implementations of the present disclosure is shown;
[0011] Figure 3 A block diagram of melody data according to some implementations of the present disclosure is shown;
[0012] Figure 4 A block diagram of a process for dividing first media data into multiple audio segments according to some implementations of the present disclosure is shown;
[0013] Figure 5 A block diagram of a system for establishing a mapping relationship between notes and audio segments according to some implementations of the present disclosure is shown;
[0014] Figure 6 A block diagram of a process for processing multiple audio segments according to some implementations of the present disclosure is shown;
[0015] Figure 7 A block diagram of a process for synthesizing melody audio and accompaniment audio according to some implementations of the present disclosure is shown;
[0016] Figure 8 A block diagram of a process for generating video data according to some implementations of the present disclosure is shown;
[0017] Figure 9A and 9B Block diagrams of pages for generating media data according to some implementations of the present disclosure are shown, respectively;
[0018] Figure 10 A flowchart of a method for generating media data according to some implementations of the present disclosure is shown;
[0019] Figure 11 A block diagram of an apparatus for generating media data according to some implementations of the present disclosure is shown; and
[0020] Figure 12 A block diagram of a device capable of implementing multiple implementations of the present disclosure is shown. Detailed implementation manners
[0021] The implementation of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain implementations of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. On the contrary, these implementations are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and implementations of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0022] In the description of the implementations of the present disclosure, the term "including" and its like should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". There may also be other explicit and implicit definitions below. As used herein, the term "model" may represent the association relationship between various data. For example, the above-mentioned association relationship can be obtained based on various technical solutions known currently and / or to be developed in the future.
[0023] It can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0024] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner according to the relevant laws and regulations.
[0025] For example, when receiving the user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require the acquisition and use of the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0026] As an optional but non-limiting implementation, the manner of sending a prompt message to the user in response to receiving the user's active request can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0027] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementations of the present disclosure. Other ways that meet the relevant laws and regulations can also be applied to the implementations of the present disclosure.
[0028] As used herein, the term "responsive to" refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of the subsequent actions performed in response to the event or condition may not necessarily be strongly correlated with the time when the event occurs or the condition is established. For example, in some cases, the subsequent actions may be performed immediately when the event occurs or the condition is established; while in other cases, the subsequent actions may be performed after a period of time after the event occurs or the condition is established.
[0029] Example environment
[0030] In the field of music production, digital synthesis techniques have been proposed to create music works. For example, a musician can collect sounds and use professional music production tools to create music works. See Figure 1 Describe the application environment according to an exemplary implementation of the present disclosure, which Figure 1 shows a block diagram 100 of an application environment according to an exemplary implementation of the present disclosure. As Figure 1 shown, a user 110 (e.g., a professional musician) can create a music work 130 with the assistance of a music production tool 120.
[0031] It should be understood that the music production tools 120 herein can be professional tools, the operations of which are complex and require the user 110 to have solid professional music knowledge and rich software usage skills. For example, the user 110 can collect sounds, use the music production tool 120 to process the sounds, and then add corresponding sound effects to the music work 130, and so on.
[0032] However, the existing music production tools 120 are not user-friendly for ordinary users (e.g., ordinary users who do not have professional music knowledge and software usage skills). For example, a user may hope to generate a piece of music with the voice of their pet, or sing a hit song with their own voice, and so on. The existing music production tools cannot meet the simple music creation needs of ordinary users. At this time, it is desirable to provide auxiliary tools to ordinary users and / or professional users in a simpler and more effective way to facilitate the generation of the desired music works.
[0033] Summary of generating media data
[0034] To at least partially address the deficiencies in the prior art, according to an exemplary implementation of the present disclosure, a method for generating media data is proposed. See Figure 2 Describe the overview according to an exemplary implementation of the present disclosure, which Figure 2FIG. 200 is a block diagram showing a media data generation apparatus according to some implementations of the present disclosure. As Figure 2 shown, a music production tool may be provided. In response to receiving a creation request for creating music, the music production tool may present a page 210. The page 210 may include: a control 220 for uploading media data (which may be referred to as first media data), and one or more controls 230, 232, 234, etc. for specifying a music template.
[0035] Specifically, a user may upload the first media data via the control 220. The first media data may include a timbre that the user desires to specify. It should be understood that timbre refers to the essential characteristics of a sound and is the most fundamental feature that differentiates one sound from another. Timbre depends on the waveform of the sound, and different waveforms result in different timbres. For example, different people's voices have different timbres, and the voices of humans and animals have different timbres, etc. If the user wishes the generated music work to include the timbre of a pet, a segment of audio (and / or video) including the barking of the pet may be uploaded. Another example is that if the user wishes the generated music work to include their own timbre, a segment of audio (and / or video) including the user's speech or singing, etc. may be uploaded.
[0036] Furthermore, the user may specify the melody of the music work that they desire to generate via a music template. For example, a music template may be obtained via the control 230, etc. The music template includes melody data 240 for specifying a music melody. It should be understood that in the context of the present disclosure, the melody data 240 may include the main melody of a song, a segment of the main melody, or any melody that the user desires to specify, etc. The user may interact with the various controls in the page 210 to specify the desired media data and music template. At this time, the music production tool may receive the media data and music template specified by the user and then generate new media data (which may be referred to as second media data). At this time, the new media data may have the timbre specified in the media data and the melody specified by the music template.
[0037] According to an example implementation of the present disclosure, the second media data and the first media data may have the same timbre, and the timbre is specified by the first media data. In this way, it is convenient for users to generate a musical composition with a desired timbre and a desired melody in a simpler and more effective manner. Specifically, assume that the user uploads media data of a pet's barking, and the selected template includes the melody data of Song A. Then the generated media data at this time may include the melody of Song A sung by the pet's voice. Another example, assume that the user uploads media data including their own speaking voice, and the selected template includes the melody data of Song B. Then the generated media data at this time may include the melody of Song B sung by the user's own voice. Using the example implementation of the present disclosure, a music creation tool can be provided for ordinary users who do not have professional music knowledge, thereby meeting the simple music creation needs of ordinary users.
[0038] Detailed process of generating media data
[0039] has been referred to Figure 2 A summary of an example implementation according to the present disclosure has been described. In the following, more details regarding the generation of media data will be provided. It should be understood that the above only shows examples of music templates that specify the media data to be generated in an example manner. Alternatively and / or additionally, music templates can be used to specify other aspects of the media data to be generated. For example, the music template may further include accompaniment data 242, and the generated second media data will further include the accompaniment data. For example, different data can be processed through mixing and other means, so that the musical composition has a richer effect.
[0040] According to an example implementation of the present disclosure, the music template further includes style data 244 for specifying a music style, and the generated second media data will further have a music style. For example, a pop style, a rock style, a classical style, etc. can be specified. In this way, a musical composition with a richer effect can be generated.
[0041] According to an example implementation of the present disclosure, the melody data 240 may include a set of notes, and each note in the set of notes corresponds to its respective time length. At this time, the melody data 240 can specify the main melody to be used. Each melody can include a plurality of notes arranged in chronological order, and there is only a single note at the same time point. Refer to Figure 3 Describe more details about the melody data 240, which Figure 3 shows a block diagram 300 of melody data according to some implementations of the present disclosure. As Figure 3As shown, the melody data 240 can be represented in various ways. For example, the melody data 240 can be represented using audio data 310. At this time, the audio data 310 can include, for example, MIDI format data with speed information (e.g., beats per minute, abbreviated as BPM), or data in MusicXML format, and so on.
[0042] Alternatively and / or additionally, the melody data 240 can be represented using note data 320 with speed information. For example, the staff notation, numbered musical notation, or any other recognizable way can be used to represent the note data 320. According to an example implementation of the present disclosure, a page for specifying the melody data 240 can be provided to the user. For example, the audio data and / or note data imported by the user can be received to determine the main melody of the media data to be generated.
[0043] According to an example implementation of the present disclosure, in the case where the melody data 240 has been determined, the corresponding audio segments for each note in the melody data 240 can be found from the first media data, and then the second media data can be generated. Specifically, the second media data can be generated in the following way: based on the pitch information in the first media data, the first media data is divided into multiple audio segments; and the second media data is generated using the multiple audio segments. It should be understood that at this time, each audio segment has a specified timbre, so the second media data generated using each audio segment will have a specified timbre.
[0044] In the following, refer to Figure 4 Describe more details about dividing the audio segments. The Figure 4 shows a block diagram 400 of the process of dividing the first media data into multiple audio segments according to some implementations of the present disclosure. According to an example implementation of the present disclosure, the pitch at different time points in the first media data 410 can be detected to divide the first media data into multiple audio segments. It should be understood that pitch refers to the high or low of a sound, which depends on the vibration frequency of the sounding body. The faster the vibration frequency, the higher the pitch, and vice versa, the slower the vibration frequency, the lower the pitch.
[0045] According to an example implementation of the present disclosure, before the partitioning operation, preprocessing can also be performed on the received first media data. For example, noise reduction processing can be performed to eliminate ambient noise, audio track separation can be performed to extract melody data, reverberation cancellation can be performed to obtain clean melody data, and so on. According to an example implementation of the present disclosure, the pitch at each time point can be detected based on digital signal processing algorithms, and then the partitioning process can be performed. The portion between the initial time point and time point 430 can be used as audio segment 420, the portion between time point 430 and 432 can be used as audio segment 422, the portion between time point 432 and 434 can be used as audio segment 424, and so on. In this way, multiple audio segments 420, 422, 424, …, and 426 can be obtained in a simple and effective manner.
[0046] According to an example implementation of the present disclosure, a large number of audio segments can be obtained based on pitch detection, and a part of high-quality audio segments can be selected from the large number of audio segments. For example, the audio segments can be selected based on the following conditions: the time length of the target audio segment meets a predetermined length condition; the energy of the target audio segment meets a predetermined energy condition; and the pitch difference of the target audio segment meets a predetermined pitch condition, and so on.
[0047] Specifically, audio segments with too short a time length can be discarded, and only those with a time length meeting the predetermined length condition (for example, not less than 0.3 seconds or other values) are retained. Audio segments with too small an energy (for example, volume) can be discarded, and only those with an energy meeting the predetermined energy condition (for example, the root mean square of the audio segment exceeds a predetermined threshold) are retained. Audio segments with too large a pitch difference (that is, the pitch range spanned by the audio segment) can be discarded, and only those meeting the predetermined pitch condition are retained. According to an example implementation of the present disclosure, the multiple audio segments can be sorted in ascending order according to the pitch difference of the multiple audio segments, and then the first K (for example, 10 or other values) audio segments can be selected. At this time, the pitch differences of the selected audio segments are small, so they can be mapped to the notes in the melody data 240 in a more accurate manner.
[0048] According to an example implementation of the present disclosure, the second media data can be generated based on: selecting a set of audio segments corresponding to a set of notes from multiple audio segments, where the target note in the set of notes corresponds to the target audio segment in the set of audio segments; and using the set of audio segments to create the second media data. In this way, each note in the melody data can be mapped to a set of audio segments among the multiple audio segments, and thus each note in the melody data can be converted into an audio segment with a specified timbre in a simple and effective manner.
[0049] See Figure 5 for more details, the Figure 5 block diagram 500 for establishing a mapping relationship between musical notes and audio segments according to some implementations of the present disclosure is shown. As Figure 5 shown, it is assumed that the melody data 240 includes a plurality of musical notes, and each musical note may correspond to its respective time length. For example, the musical note 510 (do) may correspond to the time period t0 - t1, the musical note 512 (re) may correspond to the time period t1 - t3, the musical note 514 (mi) may correspond to the time period t3 - t4, the musical note 516 (fa) may correspond to the time period t4 - t5, and so on.
[0050] At this time, for a target musical note among the plurality of musical notes, the target musical note may be mapped to a certain audio segment among the plurality of audio segments. For example, the musical note 510 may be mapped to the audio segment 420, the musical note 512 may be mapped to the audio segment 422, the musical note 514 may be mapped to the audio segment 426, the musical note 516 may be mapped to the audio segment 420, and so on.
[0051] According to an example implementation of the present disclosure, the above mapping relationship may be established based on various methods. For example, the corresponding target audio segment may be selected for the target musical note based on a random selection method. At this time, the audio segments corresponding to each musical note 510 are randomly selected. For another example, the corresponding target audio segment may be selected for the target musical note based on a polling selection method. Assuming that the partitioning operation generates 10 audio segments, the first audio segment may be selected for the first musical note in the melody data in sequence, the second audio segment may be selected for the second musical note, …, the first audio segment may be selected for the 11th musical note, and so on.
[0052] Alternatively and / or additionally, the time length corresponding to the target musical note and the time length of the target audio segment may be compared, and then the audio segment with the closest time length may be selected for the target musical note. Using the example implementation of the present disclosure, an audio segment with a matching length is selected for each musical note, thereby reducing the amplitude of the time scaling operation performed on the audio segment in subsequent operations. Alternatively and / or additionally, the pitch corresponding to the target musical note and the pitch of the target audio segment may be compared, and then the audio segment with the closest pitch may be selected for the target musical note. Using the example implementation of the present disclosure, an audio segment with a matching pitch is selected for each musical note, thereby reducing the amplitude of the time pitch adjustment performed on the audio segment in subsequent operations.
[0053] Continue to refer to Figure 5, corresponding audio segments have been assigned to each note. At this time, the lengths of the audio segments can be adjusted according to the time lengths of each note. For example, a part of the length can be intercepted to obtain a shorter audio segment, or for another example, a longer audio segment can be obtained by copying. Alternatively and / or additionally, the pitch of the audio segment can be adjusted according to the pitch of each note. In this way, the pitch and time length of a group of audio segments can be adjusted respectively to match a group of notes, and then the second media data can be generated by combining the adjusted group of audio segments. In this way, the pitch and length of each audio segment can be made to match the pitch and length of each note, so as to obtain a musical work with a specified timbre and a specified melody.
[0054] See Figure 6 for more details. The Figure 6 shows a block diagram 600 of a process for processing multiple audio segments according to some implementations of the present disclosure. As Figure 6 shown, only the audio segment 420 is used as an example of multiple audio segments to describe the process of adjusting the pitch. At block 610, the pitch 620 of the audio segment 420 can be determined. For example, the pitch 620 can be determined based on the mean 612 of the pitches at each time point in the audio segment 420. Alternatively and / or additionally, the pitch 620 can be determined based on the median 614 of the pitches at each time point in the audio segment 420.
[0055] It should be understood that the pitch 620 is determined by the frequency of the sound wave. Thus, at block 630, the frequency of the sound can be adjusted by resampling 632 and / or time scaling 634. Specifically, resampling 632 can include upsampling or downsampling, and in this way, the frequency of the sound can be adjusted. Further, stretching or compressing the time length of the audio segment can also change the frequency of the sound, so that a sound with a desired pitch can be obtained in an accurate manner. For example, the pitch 650 of the note 510 can be obtained, and then the same pitch 650 can be obtained by resampling 632 and / or time scaling 634. At this time, the pitch of the obtained audio segment 640 is the same as the pitch of the note 510.
[0056] It should be understood that although Figure 6 only the process of adjusting the pitch of the audio segment 420 is used as an example to describe the process of performing pitch adjustment. Alternatively and / or additionally, the pitch of other audio segments can be adjusted in a similar manner. Returning Figure 5 , the pitch of the audio segment 422 can be adjusted to match the note 512, the pitch of the audio segment 426 can be adjusted to match the note 514, the pitch of the audio segment 424 can be adjusted to match the note 516, and so on. At this time, the adjusted audio segments can be connected to obtain the second media data.
[0057] Alternatively and / or additionally, smoothing processing may be performed between respective audio segments to obtain a more smooth music work. By using an exemplary implementation of the present disclosure, by replacing each note in the melody data with an audio segment having a specified timbre and corresponding pitch, a music work having a specified timbre and melody can be generated in an accurate manner.
[0058] According to an exemplary implementation of the present disclosure, various sound effect processes may be performed on the generated second media data to improve the auditory perception of the music work. For example, special effect processing may be performed to add special sound effects to the second media data; or gain processing may be performed to adjust the volume, and so on. Specifically, various audio processing technologies known currently and / or to be developed in the future may be used to make the music work present a better auditory effect. According to an exemplary implementation of the present disclosure, the music template may further include accompaniment data. At this time, the accompaniment data may be added to the music work.
[0059] See Figure 7 for more details describing an exemplary implementation of the present disclosure, the Figure 7 shows a block diagram 700 of a process for synthesizing a melody audio and an accompaniment audio according to some implementations of the present disclosure. As Figure 7 shown, the melody audio 710 may be generated in the manner described above. At this time, the melody audio 710 has a specified timbre and melody. Further, the specified accompaniment audio 720 in the music template may be obtained, and the final output audio 740 may be generated based on both the melody audio 710 and the accompaniment audio 720. It should be understood that the accompaniment audio 720 and the melody data should have the same tempo to generate a more harmonious output audio.
[0060] As Figure 7 shown, special effect processing 712 and gain processing 714 may be performed on the melody audio 710. Similarly, special effect processing 722 and gain processing 724 may be performed on the accompaniment audio 720. Further, mixing 730 processing may be performed to synthesize the melody audio and the accompaniment audio together. Then, based on various ways known currently and / or to be developed in the future, special effect processing 732, gain processing 734, limiter 736 processing, and so on may be performed to obtain the final output audio 740. At this time, the output audio 740 may have a specified timbre and melody, and have excellent accompaniment data.
[0061] It should be understood that the above has described the process of generating a music work only by taking audio data as a specific example of the first media data and the second media data. Alternatively and / or additionally, the first media data and the second media data may further include video data. At this time, the audio part in the media data can be processed in the manner described above. Further, the video part in the media data can be processed in a similar manner.
[0062] See Figure 8 For more details, the Figure 8 block diagram 800 showing a process for generating video data according to some implementations of the present disclosure is shown. As Figure 8 shown, video data 840 can be received and video data 842 can be generated. The audio part 810 in the video data 840 can be processed in the manner described above. Specifically, a set of audio segments can be obtained, such as audio segments 420, 422, 426, 424, and so on. Further, the timestamps of the audio part 810 and the video part 820 can be aligned, and then a set of video segments corresponding to the set of audio segments can be obtained, for example, video segments 830, 832, 836, 834, and so on. Subsequently, the video part 820 in the second media data can be generated using the set of video segments. At this time, by combining both the audio part 810 and the video part 820, media data 842 including more rich information can be generated.
[0063] Using the example implementation of the present disclosure, assume that the media data 840 is a video including the barking of a puppy, and the specified melody data by the user is song A. The generated media data 842 includes song A sung in the voice of the puppy, and the lip movement of the puppy in the video frame will match the lip movement during singing. In this way, personalized music creation can be achieved in a simpler and more effective manner, thereby providing more creation tools for ordinary users who do not have professional music knowledge.
[0064] See Figure 9A and Figure 9B For more details regarding generating a page, Figure 9A block diagrams 900A showing pages for generating media data according to some implementations of the present disclosure are shown respectively. As Figure 9A shown, a page 910 can be provided. The user can click on the control 912 to specify the desired melody data (for example, specify song A). Subsequently, the user can use the control 914 or 916 to select the desired timbre. In the case where the user selects the control 914, multiple music styles can be provided, such as pop 920, rock 922, classical 924, and so on.
[0065] At this time, the user can select the desired style, and media data of Song A with the specified style sung by a human voice will be generated. Suppose the user selects the Rock 922 style, then the generated media data can have a rock style. For example, the accompaniment music can have a distinct sense of rhythm and use instruments such as guitars, basses, and drums for accompaniment. Suppose the user selects the Rock 9924 style, then the generated media data can have a classical style. For example, a piano can be used for accompaniment.
[0066] Figure 9B Block diagrams 900B of pages for generating media data according to some implementations of the present disclosure are respectively shown. As Figure 9B shown, suppose the user selects an instrument 916, then controls for selecting multiple instruments can be provided. The user can, for example, select a piano 940, a violin 942, or others 944, and so on. At this time, the timbre of the corresponding instrument will be automatically specified, and media data with a specified melody will be generated using the audio including the playing of the specified instrument. Suppose the user specifies Song A and the piano 940, the audio played by the piano (for example, the first media data) can be automatically obtained. Subsequently, the audio of Song A including the playing of the piano timbre (for example, the second media data) can be generated in the manner described above. In this way, multiple selection methods can be provided to ordinary users, and then media data including richer and more flexible content can be generated.
[0067] Using the example implementations of the present disclosure, multiple music creation scenarios can be supported. For example, the user can adapt an existing song using a specified timbre, or the user can create a song from scratch. For example, the user can use the voice of their own pet at home to adapt an existing song. Again, for example, the user can be supported in creating a new song, and the user only needs to upload an audio and / or video with voice. For example, the user can edit the note data in the melody data to generate a new melody. Subsequently, by selecting different styles and / or different instruments, a brand-new song can be created.
[0068] Example process
[0069] Figure 10 A flowchart of a method 1000 for generating media data according to some implementations of the present disclosure is shown. At block 1010, in response to receiving a creation request for creating music, first media data is obtained. At block 1020, a music template is obtained, and the music template includes melody data for specifying a music melody. At block 1030, based on the first media data, second media data including the music melody is generated.
[0070] According to an example implementation of the present disclosure, the second media data and the first media data have the same timbre, and the timbre is specified by the first media data.
[0071] According to an example implementation of the present disclosure, the music template further includes accompaniment data, and the second media data further includes accompaniment data.
[0072] According to an example implementation of the present disclosure, the music template further includes style data for specifying a music style, and the second media data further has a music style.
[0073] According to an example implementation of the present disclosure, the melody data includes a set of notes, each note in the set of notes corresponding to a respective time length, and the second media data is generated by: dividing the first media data into a plurality of audio segments based on pitch information in the first media data; and using the plurality of audio segments to generate the second media data.
[0074] According to an example implementation of the present disclosure, a target audio segment among the plurality of audio segments satisfies the following conditions: the time length of the target audio segment satisfies a predetermined length condition; the energy of the target audio segment satisfies a predetermined energy condition; and the pitch difference of the target audio segment satisfies a predetermined pitch condition.
[0075] According to an example implementation of the present disclosure, the second media data is generated based on: selecting a set of audio segments respectively corresponding to a set of notes from the plurality of audio segments, a target note in the set of notes corresponding to a target audio segment in the set of audio segments; and using the set of audio segments to create the second media data.
[0076] According to an example implementation of the present disclosure, the target audio segment is selected based on at least any one of the following: a random selection method; a polling selection method; comparing the time length corresponding to the target note and the time length of the target audio segment; and comparing the pitch corresponding to the target note and the pitch of the target audio segment.
[0077] According to an example implementation of the present disclosure, the second media data is created based on: adjusting the pitch and time length of a set of audio segments respectively to match a set of notes; and combining the adjusted set of audio segments to generate the second media data.
[0078] According to an example implementation of the present disclosure, the pitch of a target audio segment in the set of audio segments is adjusted based on at least any one of the following: performing resampling on the target audio segment, and scaling the time length of the target audio segment.
[0079] According to an example implementation of the present disclosure, the first media data and the second media data include video data, and the video portion in the second media data is generated based on: obtaining a set of video segments respectively corresponding to a set of audio segments; and using the set of video segments to generate the video portion in the second media data.
[0080] According to an example implementation of the present disclosure, the melody data is represented by at least any one of the following: audio data and note data.
[0081] According to an example implementation of the present disclosure, the creation request further specifies the musical instrument for creating music, and the first media data and the second media data are played using the musical instrument.
[0082] Example device and equipment
[0083] Figure 11 A block diagram of a device 1100 for generating media data according to some implementations of the present disclosure is shown. The device 1100 includes: a data acquisition module 1110 configured to acquire first media data in response to receiving a creation request for creating music; a template acquisition module 1120 configured to acquire a music template, the music template including melody data for specifying a music melody; and a generation module 1130 configured to generate second media data including the music melody based on the first media data.
[0084] According to an example implementation of the present disclosure, the second media data and the first media data have the same timbre, and the timbre is specified by the first media data.
[0085] According to an example implementation of the present disclosure, the music template further includes accompaniment data, and the second media data further includes accompaniment data.
[0086] According to an example implementation of the present disclosure, the music template further includes style data for specifying a music style, and the second media data further has a music style.
[0087] According to an example implementation of the present disclosure, the melody data includes a set of notes, each note in the set of notes corresponding to its respective time length, and the generation module is further configured to: divide the first media data into multiple audio segments based on the pitch information in the first media data; and use the multiple audio segments to generate the second media data.
[0088] According to an exemplary implementation of the present disclosure, a target audio segment among a plurality of audio segments satisfies the following conditions: the time length of the target audio segment satisfies a predetermined length condition; the energy of the target audio segment satisfies a predetermined energy condition; and the pitch difference of the target audio segment satisfies a predetermined pitch condition.
[0089] According to an exemplary implementation of the present disclosure, the generation module is further configured to: select a set of audio segments corresponding to a set of musical notes respectively from the plurality of audio segments, where a target musical note in the set of musical notes corresponds to the target audio segment in the set of audio segments; and use the set of audio segments to create second media data.
[0090] According to an exemplary implementation of the present disclosure, the target audio segment is selected based on at least any one of the following: a random selection method; a polling selection method; comparing the time length corresponding to the target musical note and the time length of the target audio segment; and comparing the pitch corresponding to the target musical note and the pitch of the target audio segment.
[0091] According to an exemplary implementation of the present disclosure, the generation module is further configured to: adjust the pitch and time length of the set of audio segments respectively to match the set of musical notes; and combine the adjusted set of audio segments to generate second media data.
[0092] According to an exemplary implementation of the present disclosure, the pitch of the target audio segment in the set of audio segments is adjusted based on at least any one of the following: performing resampling on the target audio segment and scaling the time length of the target audio segment.
[0093] According to an exemplary implementation of the present disclosure, the first media data and the second media data include video data, and the generation module is further configured to: obtain a set of video segments corresponding to the set of audio segments respectively; and use the set of video segments to generate the video part in the second media data.
[0094] According to an exemplary implementation of the present disclosure, the melody data is represented by at least any one of the following: audio data and musical note data.
[0095] According to an exemplary implementation of the present disclosure, the creation request further specifies the musical instrument for creating music, and the first media data and the second media data are played by the musical instrument.
[0096] Figure 12 The block diagram of the device 1200 capable of implementing multiple implementations of the present disclosure is shown. It should be understood that Figure 12 The shown computing device 1200 is merely exemplary and should not constitute any limitation to the functions and scope of the implementations described herein. Figure 12The computing device 1200 shown can be used to implement the methods described above.
[0097] As Figure 12 shown, the computing device 1200 is in the form of a general-purpose computing device. The components of the computing device 1200 may include, but are not limited to, one or more processors or processing units 1210, a memory 1220, a storage device 1230, one or more communication units 1240, one or more input devices 1250, and one or more output devices 1260. The processing unit 1210 may be a physical or virtual processor and be capable of performing various processes according to the programs stored in the memory 1220. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing ability of the computing device 1200.
[0098] The computing device 1200 generally includes multiple computer storage media. Such media can be any available media accessible to the computing device 1200, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 1220 may be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 1230 may be removable or non-removable media and may include machine-readable media, such as flash drives, magnetic disks, or any other media that can be used to store information and / or data (such as training data for training) and can be accessed within the computing device 1200.
[0099] The computing device 1200 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 12 it, a disk drive for reading from and writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from and writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 1220 may include a computer program product 1225 having one or more program modules configured to perform the various methods or actions of the various implementations of the present disclosure.
[0100] The communication unit 1240 enables communication with other computing devices via a communication medium. Additionally, the functions of the components of the computing device 1200 can be implemented with a single computing cluster or multiple computer machines that are capable of communicating via a communication link. Thus, the computing device 1200 can operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.
[0101] The input device 1250 can be one or more input devices such as a mouse, keyboard, trackball, etc. The output device 1260 can be one or more output devices such as a display, speaker, printer, etc. The computing device 1200 can also communicate with one or more external devices (not shown) as needed via the communication unit 1240, the external devices such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the computing device 1200, or communicate with any device that enables the computing device 1200 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0102] According to an exemplary implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, there is also provided a computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, there is provided a computer program product having a computer program stored thereon, and the program implements the method described above when executed by a processor.
[0103] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0104] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processing unit of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions that implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0105] The computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0106] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive boxes may, in fact, be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each box of the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.
[0107] The foregoing has described various implementations of the present disclosure. The description is illustrative, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or improvements made to the technology in the marketplace, or to enable other ordinary skill in the art to understand the implementations disclosed herein.
Claims
1. A method for generating media data, comprising: acquiring first media data in response to receiving a creation request for creating music; acquiring a music template, the music template including melody data for specifying a music melody; and generating second media data including the music melody based on the first media data.
2. The method according to claim 1, wherein the second media data and the first media data have the same timbre, and the timbre is specified by the first media data.
3. The method according to claim 1, wherein the music template further includes accompaniment data, and the second media data further includes the accompaniment data.
4. The method according to claim 1, wherein the music template further includes style data for specifying a music style, and the second media data further has the music style.
5. The method according to claim 1, wherein the melody data includes a set of notes, each note in the set of notes corresponding to a respective time length, and the second media data is generated by: dividing the first media data into a plurality of audio segments based on pitch information in the first media data; and using the plurality of audio segments to generate the second media data.
6. The method according to claim 5, wherein a target audio segment among the plurality of audio segments satisfies the following conditions: the time length of the target audio segment satisfies a predetermined length condition; the energy of the target audio segment satisfies a predetermined energy condition; and the pitch difference of the target audio segment satisfies a predetermined pitch condition.
7. The method according to claim 5, wherein the second media data is generated based on: selecting a set of audio segments respectively corresponding to the set of notes from the plurality of audio segments, a target note in the set of notes corresponding to a target audio segment in the set of audio segments; and using the set of audio segments to create the second media data.
8. The method according to claim 7, wherein the target audio segment is selected based on at least any one of the following: random selection method; polling selection method; comparing the time length corresponding to the target note and the time length of the target audio segment; and comparing the pitch corresponding to the target note and the pitch of the target audio segment.
9. The method according to claim 7, wherein the second media data is created based on: respectively adjusting the pitch and time length of the set of audio segments to match the set of notes; and combining the adjusted set of audio segments to generate the second media data.
10. The method according to claim 9, wherein the pitch of a target audio segment in the set of audio segments is adjusted based on at least any one of the following: performing resampling on the target audio segment and scaling the time length of the target audio segment.
11. The method according to claim 7, wherein the first media data and the second media data include video data, and the video portion in the second media data is generated based on: obtaining a set of video segments respectively corresponding to the set of audio segments; and using the set of video segments to generate the video portion in the second media data.
12. The method according to claim 1, wherein the melody data is represented by at least any one of: audio data and note data.
13. The method according to claim 1, wherein the creation request further specifies an instrument for creating the music, and the first media data and the second media data are played by the instrument.
14. An apparatus for generating media data, comprising: a data acquisition module configured to obtain first media data in response to receiving a creation request for creating music; a template acquisition module configured to obtain a music template, the music template including melody data for specifying a music melody; and a generation module configured to generate second media data including the music melody based on the first media data.
15. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform the method according to any one of claims 1 to 13.
16. A computer-readable storage medium having stored thereon a computer program, which when executed by a processor causes the processor to implement the method according to any one of claims 1 to 13.