Music generation method and device, equipment, storage medium and product
By receiving music demand information and using a trained large model to generate music that conforms to the set format, this technology solves the problems of existing technologies that rely on databases and have insufficient learning capabilities, and achieves efficient and flexible music generation and a user-friendly interactive experience.
Patent Information
- Application Number
- CN202411364620.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-03-27
AI Technical Summary
Existing music generation solutions rely on built-in databases, have poor learning capabilities, and cannot output standard music formats, thus failing to meet the diverse needs of users.
By receiving music demand information, identifying and generating instructions, and using a trained large model combined with music standard information, music that conforms to the set format is generated, including encoding the output information to generate playable music files.
It enables users to obtain music files that conform to the set format through simple interaction, improving the user experience, and supports voice interaction and feedback correction, thereby improving the flexibility and quality of music generation.
Smart Images

Figure CN121747494A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, storage medium, and product for generating music. Background Technology
[0002] Creating music using models has become a hot research topic and is changing the way music is composed. Compared to traditional methods, model-based music creation is more efficient and faster, eliminating the need for years of learning instruments and music theory, as well as lengthy composition and revision processes. A piece can be generated quickly with just simple descriptions and adjustments, which is undoubtedly a huge boon for those who want to quickly experiment with different musical styles or urgently need musical compositions.
[0003] In existing technologies, music generation methods primarily utilize recurrent neural networks. These methods rely on music stored in a built-in database, generating music based on user-input text or through randomization. However, existing methods heavily depend on the database, exhibit poor learning capabilities, and cannot output standard music formats. Summary of the Invention
[0004] This invention provides a method, apparatus, device, storage medium, and product for generating music, enabling users to obtain music that meets their needs through simple interaction with the model.
[0005] According to one aspect of the present invention, a method for generating music is provided, comprising:
[0006] Receive music request information and identify music generation instructions based on the music request information;
[0007] Obtain preset music standard information, and input the music standard information and the music generation instruction as input information into a trained large model to obtain output information; wherein, the music standard information includes a set music format and task prompt information, and the task prompt information is used to instruct the large model to generate music that conforms to the set music format;
[0008] The output information is encoded to obtain a playable music file.
[0009] Furthermore, the music demand information includes voice description information, and music generation instructions are identified based on the music demand information, including:
[0010] The speech description information is recognized using the set speech recognition technology to obtain natural language text;
[0011] The natural language text is converted into the music generation instructions.
[0012] Further, the output information is encoded, including:
[0013] The output information is converted into at least one audio segment, and the time order of each audio segment and the time length occupied by each audio segment are determined.
[0014] The at least one audio segment is converted into a music data file according to the time sequence and the time length.
[0015] Furthermore, the training method for the large model includes:
[0016] Obtain a music dataset for training; wherein the music dataset includes lyrics corpus data and song melody data;
[0017] A large model is trained based on the lyrics corpus data to obtain the preliminary training results of the large model;
[0018] The music dataset is converted into a standard training set that conforms to the set music format, and the standard training set is used as the output dataset of the large model. The input-output combination of the large model is constructed based on the output dataset.
[0019] Based on the initial training results of the large model, the large model is trained according to the combination of the model input and output until the training results of the large model meet the set requirements.
[0020] Furthermore, constructing a large model input-output combination based on the output dataset includes:
[0021] For each output data in the output dataset, a matching music generation instruction is determined, and the matching music generation instruction and the music standard information are used as the corresponding input data.
[0022] Each output data point is combined with its corresponding input data to form a single input-output combination for the larger model.
[0023] Furthermore, after combining each output data point with its corresponding input data as a single input-output combination for the large model, the model also includes:
[0024] Obtain at least two of the input-output combinations of the large model;
[0025] The input-output combinations of each large model are concatenated sequentially, and the part of the concatenated information string excluding the output data at the end is used as new input data. The output data at the end of the concatenated information string is used as the new input data to match the new output data, thus obtaining a new large model input-output combination.
[0026] Furthermore, after obtaining the playable music file, it also includes:
[0027] Play the music file, obtain feedback information about the music file, and regenerate a new music file based on the feedback information.
[0028] Furthermore, based on the feedback information, a new music file is regenerated, including:
[0029] Based on the feedback information, a music correction instruction is identified;
[0030] The input information, the output information, and the music correction instruction are used as new input information and input into the trained large model to obtain new output information.
[0031] The new output information is encoded to obtain the new music file.
[0032] According to another aspect of the present invention, a music generation apparatus is provided, comprising:
[0033] The music generation instruction recognition module is used to receive music demand information and identify music generation instructions based on the music demand information.
[0034] The large model output module is used to acquire preset music standard information, input the music standard information and the music generation instruction as input information into the trained large model, and obtain output information; wherein, the music standard information includes a set music format and task prompt information, and the task prompt information is used to instruct the large model to generate music that conforms to the set music format;
[0035] The encoding module is used to encode the output information to obtain a playable music file.
[0036] Optionally, the music demand information includes voice description information, and the music generation instruction recognition module is further used for:
[0037] The speech description information is recognized using the set speech recognition technology to obtain natural language text;
[0038] The natural language text is converted into the music generation instructions.
[0039] Optionally, the encoding module is also used for:
[0040] The output information is converted into at least one audio segment, and the time order of each audio segment and the time length occupied by each audio segment are determined.
[0041] The at least one audio segment is converted into a music data file according to the time sequence and the time length.
[0042] Optionally, the device also includes a large model training module for:
[0043] Obtain a music dataset for training; wherein the music dataset includes lyrics corpus data and song melody data;
[0044] A large model is trained based on the lyrics corpus data to obtain the preliminary training results of the large model;
[0045] The music dataset is converted into a standard training set that conforms to the set music format, and the standard training set is used as the output dataset of the large model. The input-output combination of the large model is constructed based on the output dataset.
[0046] Based on the initial training results of the large model, the large model is trained according to the combination of the model input and output until the training results of the large model meet the set requirements.
[0047] Optionally, the large model training module is also used for:
[0048] For each output data in the output dataset, a matching music generation instruction is determined, and the matching music generation instruction and the music standard information are used as the corresponding input data.
[0049] Each output data point is combined with its corresponding input data to form a single input-output combination for the larger model.
[0050] Optionally, the large model training module is also used for:
[0051] Obtain at least two of the input-output combinations of the large model;
[0052] The input-output combinations of each large model are concatenated sequentially, and the part of the concatenated information string excluding the output data at the end is used as new input data. The output data at the end of the concatenated information string is used as the new input data to match the new output data, thus obtaining a new large model input-output combination.
[0053] Optionally, the device further includes a music file regeneration module for playing the music file, obtaining feedback information on the music file, and regenerating a new music file based on the feedback information.
[0054] Optionally, the music file regeneration module is also used for:
[0055] Based on the feedback information, a music correction instruction is identified;
[0056] The input information, the output information, and the music correction instruction are used as new input information and input into the trained large model to obtain new output information.
[0057] The new output information is encoded to obtain the new music file.
[0058] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0059] At least one processor; and
[0060] A memory communicatively connected to the at least one processor; wherein,
[0061] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the music generation method according to any embodiment of the present invention.
[0062] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the music generation method according to any embodiment of the present invention.
[0063] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program / instructions, which, when executed by a processor, implement the steps of the music generation method described in any embodiment of the present invention.
[0064] The music generation method disclosed in this invention first receives music demand information and identifies music generation instructions based on the music demand information; then, it acquires preset music standard information and inputs the music standard information and music generation instructions as input information into a trained large-scale model to obtain output information. The music standard information includes a set music format and task prompt information, the task prompt information being used to instruct the large-scale model to generate music conforming to the set music format; finally, the output information is encoded to obtain a playable music file. This music generation method utilizes a trained large-scale model to generate music based on received music demand information and sets the music format for model output during model input. This allows users to obtain standard-format music files output by the model through simple interaction, improving the user experience.
[0065] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0066] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0067] Figure 1 This is a flowchart of a music generation method according to Embodiment 1 of the present invention;
[0068] Figure 2 This is a flowchart of a music generation method according to Embodiment 2 of the present invention;
[0069] Figure 3 This is a schematic diagram of the structure of a music generation device according to Embodiment 3 of the present invention;
[0070] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the music generation method of Embodiment 4 of the present invention. Detailed Implementation
[0071] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0072] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0073] Example 1
[0074] Figure 1This is a flowchart of a music generation method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where music is created using large models. The method can be executed by a music generation device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0075] S110: Receive music request information and identify music generation instructions based on the music request information.
[0076] Among them, music demand information refers to the user's requirements for music description, such as style, emotion, and genre; music generation instructions are computer instructions identified based on the music demand information.
[0077] In this embodiment, the format of the music demand information may include, but is not limited to, text format, voice format, etc. The method for identifying music generation instructions based on the music demand information may be: performing semantic segmentation and keyword extraction on the text-formatted music demand information, and generating corresponding music generation instructions based on the extracted keywords; or, directly using the text-formatted music demand information as the music generation instructions. Another method for identifying music generation instructions based on the music demand information may be: performing speech recognition using a speech recognition algorithm on the voice-formatted music demand information, converting the voice information into text information, then performing semantic segmentation and keyword extraction on the text information, and generating corresponding music generation instructions based on the extracted keywords.
[0078] For example, a user can input music request information in voice format through a voice input device: "I want a children's song, and I hope it is cheerful and bright." After recognizing the music request information, the user can obtain the music generation instruction: create a cheerful and bright children's song.
[0079] S120. Obtain preset music standard information, input the music standard information and music generation instructions as input information into the trained large model, and obtain output information.
[0080] The music standard information includes setting the music format and task prompts. The task prompts are used to instruct the large model to generate music that conforms to the set music format.
[0081] In this embodiment, the preset music standard information defines the output format of the large model, ensuring that the output information of the large model conforms to the set music format, which can be recognized by the music decoder after processing. By inputting the preset music standard information and music generation instructions together as input information into the trained large model, output information conforming to the set music format can be obtained.
[0082] Optionally, the preset music standard information can be a prompt message that includes setting the music format. The specific content of the prompt message can be customized and adjusted according to the performance of the large model. The music format can be a musical score format.
[0083] For example, the preset music standard information could be a text prompt based on a large model: "You are a music creation master, skilled in generating various types of music, including but not limited to pop, jazz, classical, electronic, and more. You need to create corresponding sheet music according to the user's music generation instructions. For example, the music generation instruction—'Create a cheerful and bright children's song,' requires outputting sheet music in the following format:"
[0084] First line<F2|4> spring <33> Where <31> <5.>ya<5.0>\n
[0085] The second line is spring. <33> Where <31> inside? <30> \n
[0086] The third line of spring <55> There <31> Verdant <5.5.> Mountain <6.7.> Village <13> inside.
[0087] in<F2|4> The key in solfège, spring. <33> This indicates the musical notation scale corresponding to this word.
[0088] The user's music generation command is...
[0089] The aforementioned preset music standard information includes the music format settings:
[0090] "First line"<F2|4> spring <33> Where <31> <5.>ya<5.0>\n
[0091] The second line is spring. <33> Where <31> inside? <30> \n
[0092] The third line of spring <55> There <31> Verdant <5.5.> Mountain <6.7.> Village <13> inside.
[0093] in<F2|4> The key in solfège, spring. <33> This indicates the musical notation scale corresponding to this word.
[0094] Furthermore, it also includes the task prompt message "You need to create the corresponding score according to the user's music generation instructions," instructing the large model to generate music that conforms to the set music format according to the user's music generation instructions.
[0095] By combining the aforementioned preset music standard information with the music generation instructions as input information for the large model, the output information of the large model that conforms to the preset music standard information can be obtained.
[0096] S130. Encode the output information to obtain a playable music file.
[0097] Among them, music files are playable audio files, such as MID (Musical Instrument Digital Interface) files.
[0098] In this embodiment, after obtaining the output information of the large model that conforms to the set music format, the music file that can be played directly is obtained through encoding and then played.
[0099] Specifically, when encoding the output information, it can be processed into multiple audio segments based on the output information of the large model, and the time order and number of beats occupied by each audio segment can be set. For example, if the output information is "spring"... <33> Where <31> ", can then be processed into two audio segments: "Spring" <33> "1, 2" and "where" <31> , 2, 2". Here, the first audio segment has a time sequence of 1 and a duration of 2 beats; the second audio segment has a time sequence of 2 and a duration of 2 beats.
[0100] Furthermore, after encoding all the output information of the large model, the music file of the entire song is obtained. The entire song is then sorted in chronological order and input into the music decoder in sequence to play the song.
[0101] In this embodiment, the training method for the large model can be:
[0102] Obtain a music dataset for training; the music dataset includes lyrics data and song melody data; train a large model based on the lyrics data to obtain the initial training results of the large model; convert the music dataset into a standard training set that conforms to the set music format, and use the standard training set as the output dataset of the large model; construct the input-output combination of the large model based on the output dataset; based on the initial training results of the large model, train the large model according to the input-output combination until the training results of the large model meet the set requirements.
[0103] The music dataset is a dataset obtained from existing music collections, and the lyrics corpus data is the lyrics text extracted from the music dataset.
[0104] Specifically, when training the large model, initial training is performed using lyrics data from the music dataset. During this initial training, the lyrics data can be segmented into words, generating a vocabulary list which serves as the training set. In this initial training, the large model predicts the next word for each input word, adjusting the model parameters based on the prediction results until the predictions meet the requirements. Then, the music dataset can be converted into a standard training set conforming to the specified music format provided in this embodiment. Based on the initial training results, this standard training set is used as the output dataset for the large model. Input-output combinations are constructed based on the output dataset, and the input information from these combinations is input into the large model. The model parameters are adjusted based on the deviation between the output information from the input-output combinations and the actual model output until the training results meet the specified requirements.
[0105] Preferably, the SFT (Supervised Fine-Tuning) method can be used to further train the large model using labeled data, based on the initial training results of the large model.
[0106] Optionally, a method for constructing a large model input-output combination based on the output dataset can be as follows: for each output data in the output dataset, determine the matching music generation instruction, and use the matching music generation instruction and music standard information as the corresponding input data; combine each output data with the corresponding input data as a large model input-output combination.
[0107] Specifically, after using the standard training set as the output dataset for the large model, corresponding music generation instructions can be matched for each output data point in the output dataset. For example, for the output data:
[0108] "First line"<F2|4> spring <33> Where <31> <5.>ya<5.0>\n
[0109] The second line is spring. <33> Where <31> inside? <30> \n
[0110] The third line of spring <55> There <31> Verdant <5.5.> Mountain <6.7.> Village <13> inside",
[0111] It can be matched with the music generation command "Generate a cheerful and bright song". If the music standard information is:
[0112] "You are a master music composer, skilled in generating various genres of music, including but not limited to pop, jazz, classical, electronic, and more. You need to create corresponding sheet music according to the user's music generation instructions. For example, the music generation instruction—'Create a cheerful and bright children's song'—needs to output sheet music in the following format:"
[0113] First line<F2|4> spring <33> Where <31> <5.>ya<5.0>\n
[0114] The second line is spring. <33> Where <31> inside? <30> \n
[0115] The third line of spring <55> There <31> Verdant <5.5.> Mountain <6.7.> Village <13> inside.
[0116] in<F2|4> The key in solfège, spring. <33> This indicates the musical notation scale corresponding to this word.
[0117] The user's music generation command is…
[0118] The input data corresponding to the output data is:
[0119] "You are a master music composer, skilled in generating various genres of music, including but not limited to pop, jazz, classical, electronic, and more. You need to create corresponding sheet music according to the user's music generation instructions. For example, the music generation instruction—'Create a cheerful and bright children's song'—needs to output sheet music in the following format:"
[0120] First line<F2|4> spring <33> Where <31> <5.>ya<5.0>\n
[0121] The second line is spring. <33> Where <31> inside? <30> \n
[0122] The third line of spring <55> There <31> Verdant <5.5.> Mountain <6.7.> Village <13> inside.
[0123] in<F2|4> The key in solfège, spring. <33> This indicates the musical notation scale corresponding to this word.
[0124] The user's music generation command is to generate a cheerful and bright song.
[0125] By matching each output data in the output dataset in the above manner, each output data is then combined with the corresponding input data as a large model input-output combination.
[0126] Furthermore, after combining each output data with its corresponding input data as a large model input-output combination, it is also possible to: obtain at least two large model input-output combinations; concatenate the large model input-output combinations in sequence, and use the part of the concatenated information string excluding the output data at the end as new input data, and use the output data at the end of the concatenated information string as new input data to match new output data, thereby obtaining a new large model input-output combination.
[0127] Specifically, based on the large model input-output combinations obtained in the above steps, new large model input-output combinations containing multi-round input-output data can be constructed by concatenating at least two large model input-output combinations. For example, if there are n large model input-output combinations... <x1y1>, <X2 Y2>, <X3 Y3>, …, <Xn Yn>, wherein X represents input data, Y represents output data, two large model input and output combinations <X1 Y1> and <X5 Y5> are randomly obtained, the information string <X1 Y1 X5 Y5> is obtained after sequentially splicing, the output data at the end of the information string is Y5, and the part except the output data at the end is X1 Y1 X5, then X1 Y1 X5 is taken as new input data, and Y5 is taken as new output data, to obtain a new large model input and output combination.
[0128] The music generation method provided in the embodiment of the application first receives music demand information, identifies music generation instructions according to the music demand information, then obtains preset music standard information, inputs the music standard information and the music generation instructions into a trained large model as input information, and obtains output information, wherein the music standard information includes set music format and task prompt information, and the task prompt information is used to instruct the large model to generate music conforming to the set music format, and finally, the output information is encoded to obtain a playable music file. The music generation method provided in the embodiment of the application generates music according to the received music demand information by using the trained large model, and sets the music format of the model output when the model is input, so that the user only needs to interact with the model simply, and can obtain the music file of the standard format output by the model, thereby improving the user experience.
[0129] Embodiment two
[0130] Figure 2 A flowchart of a music generation method provided in the embodiment two of the application, and the embodiment is a refinement of the above-mentioned embodiment, as shown in Figure 2 The method comprises the following steps.
[0131] In S210, the voice description information is identified according to a set voice recognition technology to obtain natural language text.
[0132] The music demand information includes voice description information.
[0133] In the embodiment, the voice description information is a description of the music demand in a voice format, such as emotion, style and atmosphere, and the user can input the voice description information through a voice input device such as a microphone. After receiving the voice description information, the voice recognition algorithm can be used for voice recognition to convert the voice information into natural language text information. The natural language text refers to a text written in human natural language (including spoken language and written language).
[0134] Optionally, the voice recognition can be performed by using an ASR (Automatic Speech Recognition) method. In the process of language recognition, the ASR first converts the sound signal into a vector through coding, and then converts the vector into text through decoding. In the coding process, the ASR cuts the sound signal into small segments according to a very short time interval, each segment being a frame. For each frame, the features in the signal are extracted by a certain rule to become a multi-dimensional vector, and each dimension in the vector is a feature of the frame. In the decoding process, the vector obtained by coding is first processed by an acoustic model to combine adjacent frames to obtain phonemes, and then to further combine to obtain single words or Chinese characters. Then, the language model adjusts the logical words obtained by the acoustic model to make the recognition result smooth.
[0135] S220, converting the natural language text into music generation instructions.
[0136] The music generation instructions are computer-recognizable and executable instructions for instructing the large model to generate music that meets the music demand information input by the user.
[0137] In this embodiment, after obtaining the natural language text recognized from the voice description information, the natural language text can be converted into music generation instructions through processing of the natural language text.
[0138] Optionally, the processing method of the natural language text can be semantic division and keyword extraction of the text information, and generating corresponding music generation instructions according to the extracted keywords.
[0139] S230, obtaining preset music standard information, inputting the music standard information and the music generation instructions into the trained large model as input information to obtain output information.
[0140] The music standard information includes setting music format and task prompt information, and the task prompt information is used to instruct the large model to generate music that meets the set music format.
[0141] In this embodiment, the preset music standard information specifies the output format of the large model, so that the output information of the large model meets the set music format, and the format can be recognized by a music decoder. By inputting the preset music standard information and the music generation instructions into the trained large model as input information of the large model, output information that meets the set music format can be obtained.
[0142] Optionally, the preset music standard information can be prompt information containing the set music format, and the specific content of the prompt information can be customized and adjusted according to the performance of the large model, and the set music format can be a music tablature format.
[0143] The preset music standard information is combined with the music generation instruction as input information of the large model, so that output information of the large model in a set music format in the preset music standard information can be obtained.
[0144] In S240, the output information is encoded to obtain a playable music file.
[0145] The music file is a playable audio file, such as a MIDI (Musical Instrument Digital Interface) file.
[0146] In this embodiment, the method of encoding the output information can be: converting the output information into at least one audio segment, and determining the time sequence of each audio segment and the time length occupied by each audio segment; and converting the at least one audio segment into a music data file according to the time sequence and the time length.
[0147] Specifically, when encoding the output information, the output information can be processed into one or more audio segments according to the output information of the large model, and the time sequence of each audio segment and the time length occupied by each audio segment are determined. For example, if the output information is "spring <33> where <31>", it can be processed into two audio segments: "spring <33>, 1, 2" and "where <31>, 2, 2". The time sequence of the first audio segment is 1, and the time length occupied is 2 beats; the time sequence of the second audio segment is 2, and the time length occupied is 2 beats.
[0148] Further, after encoding all the output information of the large model, the music file of the whole song is obtained, and the whole song is sorted according to the time sequence and input into the music decoder in turn, so that the song can be played.
[0149] After obtaining the output information of the large model in the set music format, the music file that can be directly played is obtained through encoding and played.
[0150] Further, after obtaining the playable music file, the music file can also be played, feedback information of the music file is obtained, and a new music file is generated according to the feedback information.
[0151] The feedback information is evaluation feedback of the user input for the played music file, such as "not enough happy" and "rhythm too slow".
[0152] In this embodiment, after the music file is played, if the user is not satisfied, the feedback information can be continuously input, so that the large model continues to create on the basis of the original music file until no feedback information of the user is obtained or confirmation information of the user for the music file is obtained.
[0153] Optionally, the method of regenerating a new music file according to the feedback information can be: identifying a music correction instruction according to the feedback information; inputting the input information, the output information and the music correction instruction into the trained large model as new input information to obtain new output information; and encoding the new output information to obtain a new music file.
[0154] Specifically, after obtaining the feedback information of the user on the generated music file, a music correction instruction can be identified according to the feedback information. Similarly, the feedback information can include but is not limited to text format, voice format, etc. Through natural language processing, automatic speech recognition and other technologies, the feedback information can be converted into a computer-recognizable music correction instruction. When the large model modifies the generated music file, the historical dialogue will be retained, that is, the input information and output information of the last round and the music correction instruction will be input into the large model as the new input information of this round in the next round of output. After obtaining the new output information, the new output information is encoded to obtain a new music file. The above operation is repeated until no feedback information of the user is obtained or confirmation information of the user on the music file is obtained.
[0155] The music generation method provided by the embodiment of the application generates music according to received music demand information by using a trained large model, sets the music format of the model output when the model is input, so that the user only needs to interact with the model to obtain the music file in the standard format output by the model, and the user experience is improved. Moreover, the method can recognize the user's voice and provide a voice interaction mode for the user. The method can also modify the generated music file on the basis of the last round of input and output according to the user's feedback after playing the music file, until a satisfactory result is obtained, and the user experience can be further improved.
[0156] Embodiment three
[0157] Figure 3 A structural schematic diagram of a music generation device provided by the embodiment three of the application is shown in Figure 3 The device includes a music generation instruction identification module 310, a large model output module 320 and an encoding module 330.
[0158] The music generation instruction identification module 310 is configured to receive music demand information and identify a music generation instruction according to the music demand information.
[0159] The large model output module 320 is configured to obtain preset music standard information, input the music standard information and the music generation instruction into a trained large model as input information, and obtain output information.
[0160] The music standard information includes a set music format and task prompt information, and the task prompt information is used to instruct the large model to generate music in the set music format.
[0161] The encoding module 330 is configured to encode the output information to obtain a playable music file.
[0162] Optionally, the music demand information includes voice description information, and the music generation instruction recognition module 310 is further configured to:
[0163] According to the set voice recognition technology, the voice description information is recognized to obtain natural language text, and the natural language text is converted into music generation instructions.
[0164] Optionally, the encoding module 330 is further configured to:
[0165] The output information is converted into at least one audio segment, and the time sequence of each audio segment and the time length occupied by each audio segment are determined; and the at least one audio segment is converted into a music data file according to the time sequence and the time length.
[0166] Optionally, the device further includes a large model training module 340 configured to:
[0167] obtain a music data set for training; the music data set includes song lyrics data and song melody data; the large model is trained according to the song lyrics data to obtain a preliminary training result of the large model; the music data set is converted into a standard training set conforming to a set music format, and the standard training set is used as an output data set of the large model, and a large model input-output combination is constructed according to the output data set; and the large model is trained according to the model input-output combination on the basis of the preliminary training result of the large model until the training result of the large model meets the set requirements.
[0168] Optionally, the large model training module 340 is further configured to:
[0169] For each output data in the output data set, a matching music generation instruction is determined, and the matching music generation instruction and the music standard information are used as corresponding input data; and each output data and the corresponding input data are used as a large model input-output combination.
[0170] Optionally, the large model training module 340 is further configured to:
[0171] obtain at least two large model input-output combinations; sequentially splice the large model input-output combinations, and use the part of the spliced information string except the last output data as new input data, and use the last output data in the spliced information string as new output data matched with the new input data to obtain a new large model input-output combination.
[0172] Optionally, the apparatus further comprises a music file regeneration module 350 configured to play the music file, acquire feedback information of the music file, and regenerate a new music file according to the feedback information.
[0173] Optionally, the music file regeneration module 350 is further configured to:
[0174] identify a music correction instruction according to the feedback information, input the input information, the output information and the music correction instruction into the trained large model as new input information, obtain new output information, encode the new output information, and obtain a new music file.
[0175] The music generation apparatus provided in the embodiments of the present application can execute the music generation method provided in any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.
[0176] Embodiment four
[0177] Figure 4 A structural schematic diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.
[0178] As shown in Figure 4 The electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which are communicatively connected to the at least one processor 11, wherein the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0179] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0180] The processor 11 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the generation method of music.
[0181] In some embodiments, the generation method of music can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the generation of music described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the generation method of music by any other appropriate means, such as by means of firmware.
[0182] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0183] Computer programs for implementing the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program, when executed, can cause instructions defined in the flow charts and / or block diagrams to be implemented. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, and partially on a remote machine or entirely on a remote machine or server.
[0184] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0185] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0186] The systems and techniques described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described herein, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0187] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.
[0188] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, and the present disclosure is not limited in this regard.
[0189] The specific embodiments described above are not intended to limit the scope of the present disclosure. Those skilled in the art will understand that various modifications, combinations, sub-combinations, and alternatives can be made to the specific embodiments without departing from the spirit and principles of the present disclosure. Any further modifications, equivalents, and / or alternatives come within the scope of the present disclosure as recited by the claims.
Claims
1. A method for generating music, characterized in that, include: Receive music request information and identify music generation instructions based on the music request information; Obtain preset music standard information, and input the music standard information and the music generation instruction as input information into a trained large model to obtain output information; wherein, the music standard information includes a set music format and task prompt information, and the task prompt information is used to instruct the large model to generate music that conforms to the set music format; The output information is encoded to obtain a playable music file.
2. The method according to claim 1, characterized in that, The music demand information includes voice description information, and music generation instructions are identified based on the music demand information, including: The speech description information is recognized using the set speech recognition technology to obtain natural language text; The natural language text is converted into the music generation instructions.
3. The method according to claim 1, characterized in that, Encoding the output information includes: The output information is converted into at least one audio segment, and the time order of each audio segment and the time length occupied by each audio segment are determined. The at least one audio segment is converted into a music data file according to the time sequence and the time length.
4. The method according to claim 1, characterized in that, The training methods for the large model include: Obtain a music dataset for training; wherein the music dataset includes lyrics corpus data and song melody data; A large model is trained based on the lyrics corpus data to obtain the preliminary training results of the large model; The music dataset is converted into a standard training set that conforms to the set music format, and the standard training set is used as the output dataset of the large model. The input-output combination of the large model is constructed based on the output dataset. Based on the initial training results of the large model, the large model is trained according to the input-output combination until the training results of the large model meet the set requirements.
5. The method according to claim 4, characterized in that, Constructing a large model input-output combination based on the output dataset includes: For each output data in the output dataset, a matching music generation instruction is determined, and the matching music generation instruction and the music standard information are used as the corresponding input data. Each output data point is combined with its corresponding input data to form a single input-output combination for the larger model.
6. The method according to claim 5, characterized in that, After combining each output data point with its corresponding input data as the input-output of the large model, it also includes: Obtain at least two of the input-output combinations of the large model; The input-output combinations of each large model are concatenated sequentially, and the part of the concatenated information string excluding the output data at the end is used as new input data. The output data at the end of the concatenated information string is used as the new input data to match the new output data, thus obtaining a new large model input-output combination.
7. The method according to claim 1, characterized in that, After obtaining the playable music file, it also includes: Play the music file, obtain feedback information about the music file, and regenerate a new music file based on the feedback information.
8. The method according to claim 7, characterized in that, Based on the feedback information, a new music file is regenerated, including: Based on the feedback information, a music correction instruction is identified; The input information, the output information, and the music correction instruction are used as new input information and input into the trained large model to obtain new output information. The new output information is encoded to obtain the new music file.
9. A music generation device, characterized in that, include: The music generation instruction recognition module is used to receive music demand information and identify music generation instructions based on the music demand information. The large model output module is used to acquire preset music standard information, input the music standard information and the music generation instruction as input information into the trained large model, and obtain output information; wherein, the music standard information includes a set music format and task prompt information, and the task prompt information is used to instruct the large model to generate music that conforms to the set music format; The encoding module is used to encode the output information to obtain a playable music file.
10. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the music generation method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the method for generating music according to any one of claims 1-8.
12. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the music generation method as described in any one of claims 1-8.