Generation method for playing music, computer equipment and storage medium
By dynamically managing the performance model pool and feedback score-driven elimination and evolution mechanism, the problem that the stylized performance model in the existing technology cannot adapt to the music trend is solved, and the long-term adaptive maintenance and efficient innovation of the performance model are achieved.
Patent Information
- Application Number
- CN202510473716.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing static stylized performance model cannot capture and adapt to the latest music trends, resulting in the generated music gradually lagging behind the times, limiting its long-term application value and user experience.
By building a dynamically managed performance model pool, combining feedback score-driven differentiated elimination and evolution mechanisms, dynamically update and optimize the style model to ensure that the models in the model pool always match current music trends and user needs.
Long-term adaptive maintenance of the performance model is realized, the music quality is stable and the style is diversified, and while efficient exploration and innovation are explored and innovative, the model is avoided redundant or outdated, and the user experience and application value are enhanced.
Smart Images

Figure CN120015000A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of generation of performance music, and in particular to a method for generating performance music, a computer device and a storage medium. Background Art
[0002] With the development of digital music technology and artificial intelligence, stylized performance models have gradually become a hot topic in the field of music technology. These models analyze and understand the music information in paper or electronic format, imitate the performance style of a specific musician or the characteristics of a specific music genre, and provide users with a unique music experience.
[0003] For example, an invention patent application with patent application publication number CN113936625A discloses a system and method for simulating the automatic performance of a musician's style, wherein the system includes a storage module, a processing module and an interaction module, wherein the processing module is used to generate performance style models corresponding to different musicians through machine learning training of the music data stored in the storage module; query the corresponding performance style model according to the designated musician information transmitted by the interaction module, input the designated music, and output the simulated performance of the music, thereby realizing automatic performance.
[0004] The invention patent application with the patent application publication number CN111554255A discloses a MIDI performance style automatic conversion system based on a recurrent neural network, including a MIDI analysis module, a data preprocessing module, an autoencoder module, a style network module and a music generation module. Among them, the MIDI analysis module reads the user input and merges the multi-track MIDI into a single-track MIDI as a score; the data preprocessing module is used to extract note features from the score; the autoencoder module encodes and decodes the note features; the style network module learns the performance style of the score and predicts the dynamic vector; the music generation module is used to configure the dynamic vector predicted by the style network module for the score and convert it into an expressive music.
[0005] However, as music trends continue to evolve, static stylized performance models are too rigid to capture and adapt to the latest music trends, causing the generated music to gradually lag behind the times, limiting its long-term application value and reducing user experience. Summary of the invention
[0006] The main purpose of this application is to provide a method for generating music for performance, a computer device and a storage medium. In order to solve the above-mentioned technical problems, this application specifically adopts the following technical solutions: A first aspect of the present application is to provide a method for generating performance music, the method comprising: S101, obtaining a music score to be performed and a performance model pool, wherein the performance model pool includes a benchmark model and a preset number of style models, and each of the style models is associated with a style label, and the style models include a core model and an exploration model; S102, inputting the music score to be performed into the benchmark model to generate basic music data; S103, inputting the basic music data into a plurality of the style models respectively to generate a plurality of target music data; S104, generating a feedback score for each of the style models according to a number of the target music data; calculating a comprehensive feedback score for each of the core models based on a first iteration cycle; calculating a comprehensive feedback score for each of the exploration models based on a second iteration cycle; wherein the first iteration cycle is greater than the second iteration cycle; S105. When the comprehensive feedback score of any of the style models is continuously lower than a preset score threshold, the corresponding style model is removed from the performance model pool, and a replacement style model is imported from a model cache library to update the performance model pool, wherein the model cache library stores a plurality of pre-trained style models.
[0007] In some embodiments, the feedback score includes a first score and / or a second score, and the method includes: scoring the target music data of each of the style models based on preset scoring rules to obtain a first score for each of the style models; and / or obtaining user behavioral feedback data and generating a second score for each of the style models based on the behavioral feedback data.
[0008] In some embodiments, the method also includes: determining a number of high-scoring style models based on comprehensive feedback scores, and historical target music data historically output by the high-scoring style models; training a first newly added exploration model based on at least two of the high-scoring style models with different style labels and the corresponding historical target music data, and storing the first newly added exploration model in the model cache; and / or, training a second newly added exploration model based on at least two of the high-scoring style models with the same style label and the corresponding historical target music data, and storing the second newly added exploration model in the model cache.
[0009] In some embodiments, before importing the replacement style model from the model cache library, it also includes: according to the style label of the removed style model, counting the current total number of models with corresponding style labels in the performance model pool; if the current total number of models is greater than the preset number of models, using the first newly added exploration model as the replacement style model; if the current total number of models is less than the preset number of models, using the second newly added exploration model as the replacement style model.
[0010] In some embodiments, the style model also includes a resident model; the method also includes: based on the third iteration cycle, counting the comprehensive feedback score of each of the resident models; when the style model removed from the performance model pool is a resident model, storing the removed resident model in a historical version library; importing a newly added resident model from the model cache library, wherein the newly added resident model is obtained by training the resident model in the historical version library based on the historical target music data.
[0011] In some embodiments, the method further includes: when the removed style model is a core model, based on the style label of the removed core model and the comprehensive feedback scores of each exploration model, selecting a target exploration model from the exploration models, and updating the target exploration model to a newly added core model.
[0012] In some embodiments, the method also includes: comparing a number of the target music data, and when the similarity between any two of the target music data is higher than a first similarity threshold, respectively comparing the first historical music data and the second historical music data of the two corresponding style models; if the similarity between the first historical music data and the second historical music data is higher than a second similarity threshold, removing one of the style models from the performance model resource library.
[0013] In some embodiments, the method further includes: detecting a growth trend of comprehensive feedback scores of a plurality of core models; and retraining the baseline model when a growth trend of a fourth preset number of comprehensive feedback scores is abnormal.
[0014] A second aspect of the present application is to provide a computer device, the device comprising: Memory for storing computer programs; A processor is used to execute the computer program and implement the steps of the method for generating music for performance as provided in any embodiment of the present application when executing the computer program.
[0015] The third aspect of the present application is to provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor performs the steps of the method for generating music as provided in any embodiment of the present application.
[0016] Beneficial effects: This application proposes a method for generating performance music, a computer device and a storage medium. By constructing a dynamically managed performance model pool and combining a differentiated elimination and evolution mechanism driven by feedback scoring, it can stabilize the music quality, maintain style diversity, and efficiently explore and innovate, while avoiding model redundancy or obsolescence in the performance model pool, thereby achieving long-term adaptive maintenance of the performance model.
[0017] The performance model pool consists of a baseline model and a style model (including a core model, an exploration model, and a resident model), which separates the basic generation of music performance from the style transfer task. The baseline model is responsible for generating technical basic data to ensure the underlying technical correctness of music generation; the style model focuses on style diversification and achieves accurate simulation of different styles, avoiding the waste of computing power caused by repeated calculation of basic data. It only needs to update the style model to introduce new music styles or optimize the existing style expression without retraining the baseline model, which improves flexibility and effectively saves computing power resources.
[0018] Furthermore, different types of style models in the performance model pool each have specific responsibilities: the core model is used to maintain high-quality, stable output of music of various styles; the exploration model is used to try new trends or integrate the features of existing high-scoring models to promote style innovation and performance improvement; the resident model is used for long-term optimization of specific classic styles to prevent the loss of classic styles due to algorithm iteration. Through division of labor and cooperation, the performance model pool can not only cover a wide range of performance needs, but also continue to stimulate innovation possibilities while ensuring basic experience.
[0019] In addition, the elimination and evolution mechanism of the model is flexibly set based on the characteristics and functions of different types of style models (core models, exploration models, and resident models). For example, different style models have different iteration cycles and new models come from different sources. These differentiated settings can form a benign synergistic effect, so that several models within the performance model pool at different stages have existence value, maintain a balance between stability and innovation in the performance model pool, avoid homogenization or obsolescence of the content of the performance model pool, and adaptively maintain the vitality of the performance model pool in the long term.
[0020] Specifically, the core model uses a low-frequency iteration cycle for evaluation and update to ensure that the music data it generates maintains high quality and high satisfaction, while eliminating outdated core models to inherit the innovative results of the exploration model, ensuring that each exploration model promoted to a core model has a stable high-quality score for at least a longer probation period. In contrast, the exploration model uses a high-frequency iteration cycle to quickly trial and error and capture new trends, keep the model pool active, and by integrating high-scoring model features, it reduces computing power consumption while ensuring quality. The resident model focuses on the long-term maintenance of specific classic styles, while retaining the classic style, it can be appropriately adjusted according to user preferences. Furthermore, when the comprehensive feedback scores of a large number of core models are abnormal, it will also trigger the update of the baseline model. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, each element or part is not necessarily drawn according to the actual scale. Obviously, the drawings described below are some embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without paying creative labor.
[0022] Figure 1 is a schematic flow chart of a method for generating music for performance provided in an embodiment of the present application; Figure 2 is a schematic flow chart of another method for generating music for performance provided in an embodiment of the present application; Figure 3 is a schematic diagram of another dynamic update of a performance model pool provided by an embodiment of the present application; Figure 4 It is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0024] Herein, suffixes such as "module", "component" or "unit" used to represent elements are only used to facilitate the description of the present application, and have no specific meanings by themselves. Therefore, "module", "component" or "unit" can be used mixedly.
[0025] In this document, the terms "upper", "lower", "inner", "outer", "front", "back", "one end", "the other end" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the present application. In addition, the terms "first" and "second" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.
[0026] In this document, unless otherwise clearly specified and limited, the terms "installed", "provided with", "connected", etc. should be understood in a broad sense. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0027] Herein "and / or" includes any and all combinations of one or more of the associated listed items.
[0028] Herein, "plurality" means two or more than two, ie, it includes two, three, four, five, etc.
[0029] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0030] Traditional music generation technology often has static defects and lacks the ability to learn new trends in real time, resulting in the generation of content limited to the music style during training and difficulty in capturing emerging elements. In addition, model updates rely on manual intervention or full retraining, with long iteration cycles and high costs, and are unable to respond to users' demand for freshness in a timely manner. As music trends continue to evolve, the generated music gradually becomes out of touch with user expectations. In long-term applications, it is easy to lose competitiveness due to content homogeneity and lag, ultimately affecting user experience and actual application value.
[0031] Based on this, the present application proposes a method for generating performance music, a computer device and a storage medium. By constructing a dynamically managed performance model pool and combining it with a differentiated elimination and evolution mechanism driven by feedback scoring, it can stabilize the quality of music, maintain style diversity, and efficiently explore and innovate, while avoiding model redundancy or obsolescence in the performance model pool, thereby achieving long-term adaptive maintenance of the performance model.
[0032] Some embodiments of the present application are described in detail below in conjunction with the accompanying drawings. In the absence of conflict, the following embodiments and features of the embodiments can be combined with each other. Figure 1 , Figure 1 is a schematic flow chart of a method for generating a musical performance provided in an embodiment of the present application, such as Figure 1As shown, an embodiment of the present application provides a method for generating performance music, and the method includes S101 to S105.
[0033] S101, obtaining a music score to be performed and a performance model pool, wherein the performance model pool includes a benchmark model and a preset number of style models, and each style model is associated with a style label, and the style models include a core model and an exploration model.
[0034] The music score to be performed refers to the standardized digital music score file input by the user, which contains basic information such as the note sequence, rhythm, chord progression, etc. of the music, and is the basic input for the subsequent generation of music.
[0035] Among them, the performance model pool includes two types of music generation models, namely baseline models and style models, and the number of style models is a preset number to ensure that the performance model pool is eliminated and optimized within a reasonable number range. The specific number can be flexibly set according to actual needs and is not limited here.
[0036] Among them, the baseline model serves as the basic generation unit, which is used to parse the music score and generate basic music data without stylization to ensure accuracy and consistency at the technical level; the style model further stylizes the basic music data, and its quantity and type can be pre-set according to needs, and each style model has an associated style label to identify the music style it is good at or suitable for processing, such as "classical style", "jazz style", "impressionist style", "mixed style", etc.
[0037] Furthermore, the style model is subdivided into the core model and the exploration model. The core model is a high-quality style model that has been verified for a long time and is responsible for stably outputting high-quality music data in various styles. The exploration model is an experimental style model for optimizing model performance or innovating performance styles. It is responsible for style innovation and trend capture, and improving the fun of output music.
[0038] In some embodiments, the style model includes a first preset number of core models and a second preset number of exploration models. The first preset number and the second preset number can be flexibly set according to actual needs and are not limited here. For example, the second preset number can be less than the first preset number, and the ratio of the preset numbers of the two can be flexibly set to meet the differentiated needs of different application scenarios for style stability and innovation.
[0039] Therefore, the style model, through the distinction between style labels and functional positioning, together constitutes a dynamic hierarchical structure of the performance model pool, which not only ensures the stability of music data under different styles, but also provides flexible space for the exploration of new styles.
[0040] In some embodiments, in response to a user's selection operation and / or import operation, a number of core models selected by the user are determined, wherein the selection operation refers to an operation instruction for the user to select the required baseline model and style model in a preset model selection interface, and the import operation refers to an operation instruction for the user to import the required style model in a preset model import interface; the number of the selected core models is counted to obtain a specific value of a first preset number, the specific value of the second preset number is determined based on a preset number ratio between the first preset number and the second preset number and the specific value of the first preset number, the second preset number of style models is randomly selected from the model selection interface as the second preset number, and a performance model pool is constructed based on the first preset number of core models and the second preset number of the second preset number.
[0041] It should be understood that the performance model pool can be customized by the user, that is, it can be selected from the preset model library through the model selection interface, or the model to be used can be imported from the outside, and subsequent updates to the performance model pool are also based on user feedback, but the benchmark model only supports selection from the preset model library to ensure the correctness of the underlying technology of music generation, and restricts the core architecture and elimination optimization mechanism of the performance model pool, taking into account the controllability and scalability of music generation.
[0042] S102: Input the music score to be performed into the benchmark model to generate basic music data.
[0043] The benchmark model is a pre-trained deep neural network that converts the input music score to be played into basic music data standardized at the technical level. Exemplarily, the training data of the benchmark model is a large amount of standardized music scores and corresponding non-stylized data, which may include the annotation data of note sequence, note start and end time, beat information, chords, pedals, etc. corresponding to the music score.
[0044] Furthermore, the benchmark model can adopt an encoder-decoder architecture based on Transformer or LSTM, in which the encoder parses the input music score to be performed into a structured feature vector containing pitch, duration, and intensity; the decoder reconstructs basic music data without style tendencies based on the attention mechanism, including precise note start and end times, standard intensity values, standard beat information, etc.
[0045] It should be understood that the basic music data output by the benchmark model needs to be accurate in technical specifications and avoid basic errors, while maintaining style neutrality and avoiding bias towards any specific performance style, ensuring that all style models are stylized based on a unified technical foundation, providing a reliable foundation for diverse style expressions.
[0046] S103, inputting the basic music data into a plurality of the style models respectively to generate a plurality of target music data.
[0047] The style model is a deep neural network that has been trained in style, and is used to convert the basic music data generated by the baseline model into target music data with a specific style. For example, the training data of each style model includes a large number of annotated data sets of the corresponding style, such as a classical style data set that includes the typical rhythm changes, improvisational ornaments and other features of the style.
[0048] Furthermore, the style model can use a deep neural network with a style embedding layer, the input layer receives the basic music data, the style embedding layer loads the predefined style label vector, and the output layer generates the target music data containing the target style expression. It should be understood that in this process, each style model performs differential processing on the same basic music data, the core models of different styles generate stable and high-quality target music data; the exploration model generates innovative target music data. Moreover, all generated results maintain the main structure framework of the original score (such as more than 90% of the note sequence is not adapted), and superimpose stylized performance parameters (such as the ornamentation pattern of a specific genre), stylize the performance dimension, and finally output the target music data set corresponding to the style label.
[0049] S104. Generate a feedback score for each of the style models based on a number of the target music data; calculate a comprehensive feedback score for each of the core models based on a first iteration cycle; calculate a comprehensive feedback score for each of the exploration models based on a second iteration cycle; wherein the first iteration cycle is greater than the second iteration cycle.
[0050] Among them, the feedback score is an indicator for quantitatively evaluating the performance of the style model. The weighted average or sliding window mean of all feedback scores of the style model in a specific iteration cycle is taken to obtain a comprehensive feedback score.
[0051] Specifically, the core model is evaluated and updated in the first iteration cycle, which is set as a low-frequency evaluation, such as one month or one quarter, and the feedback scores of all target music data of the core model in the cycle are counted to generate a comprehensive feedback score. It should be understood that the core model is used to maintain high-quality and stable output of music of various styles. Its model performance has been verified for a long time and has a high degree of reliability. Therefore, it is only necessary to regularly verify whether the core model maintains high-quality and high-scoring performance, so as to eliminate outdated core models in a timely manner.
[0052] Specifically, the exploration model is evaluated and updated in the second iteration cycle. The second iteration cycle is set as a high-frequency evaluation, such as one week or three days. The feedback scores of all target music data in the exploration model within the cycle are statistically analyzed to generate a comprehensive feedback score. It should be understood that the exploration model is used to try new trends or integrate the features of existing high-scoring models to promote style innovation and performance improvement. Its core purpose is to try innovation. The actual performance of the model fluctuates greatly. It may perform exceptionally well, or it may perform mediocre or extremely poorly. High-frequency evaluation is required to accelerate elimination or optimization and quickly verify the positivity of innovation.
[0053] In some embodiments, the feedback score includes a first score and / or a second score, and the method includes: scoring the target music data of each of the style models based on preset scoring rules to obtain a first score for each of the style models; and / or obtaining user behavioral feedback data and generating a second score for each of the style models based on the behavioral feedback data.
[0054] Specifically, the preset scoring rules refer to the scoring rules set in advance according to technical compliance dimensions such as the rationality of the range, the logic of the chord progression, and the consistency of the dynamic performance. When the target music data is generated, the target music data will be comprehensively generated according to the above technical qualification dimensions to evaluate the technical compliance of the generated music. Exemplarily, different preset scoring rules can be set according to different styles, and different preset scoring rules can also be set according to different model types (such as core models, exploration models). Correspondingly, the corresponding first scoring model can be trained based on the music data and the scoring label data set under different preset scoring rules, and the target music data can be analyzed and scored based on different first scoring models.
[0055] Specifically, the behavioral feedback data refers to the behavioral records generated by the user in the process of interacting with the music content corresponding to the target music data, including but not limited to: the user's direct rating, the music play completion rate, the frequency of repeated play, the like rate, the collection rate, the sharing rate, the download rate, and the play skip rate, reflecting the user's preference, interest and satisfaction with the target music. When the target music data is generated, a second rating is generated based on the above-mentioned behavioral feedback data to evaluate the market acceptance of the generated music. For example, the second rating values corresponding to different behavioral feedback data can be pre-set, and a rating mapping table can be formed to quickly convert the behavioral feedback data into the second rating.
[0056] In some embodiments, the feedback score may also include: model performance score and expert sampling score. The method includes: scoring the target music data of each style model based on dimensions such as generation delay, resource occupancy rate, and audio quality to obtain the model performance score of each style model; expert sampling score, randomly extracting the target music data of the style model, and scoring based on dimensions such as style feature restoration, melody harmony, etc., to obtain the expert sampling score of the corresponding style model.
[0057] S105. When the comprehensive feedback score of any of the style models is continuously lower than a preset score threshold, the corresponding style model is removed from the performance model pool, and a replacement style model is imported from a model cache library to update the performance model pool, wherein the model cache library stores a plurality of pre-trained style models.
[0058] Specifically, based on the iteration cycle of different style models, the comprehensive feedback score of the style model will be generated multiple times, and each generated comprehensive feedback score will be compared with the preset score threshold. When the comprehensive feedback score of a style model continues to be lower than the preset score threshold within the preset cycle, such as the core model for three consecutive first iteration cycles and the exploration model for two consecutive second iteration cycles, it is judged that its technical compliance is low or the user satisfaction is low, and it is removed from the performance model pool. Subsequently, the replacement style model is imported from several pre-trained style models to update the performance model pool and fill the vacancies left by the model elimination.
[0059] It should be understood that by continuous screening and updating, eliminating inefficient models and introducing potential high-quality candidates, while maintaining the total amount of the performance model pool, the quality of output music is continuously improved, so that the multiple style models in the performance model pool always match the current music trends and user needs.
[0060] In some embodiments, the feedback score of the style model is stored in a time series database, and the score aggregation calculation is triggered at the end of the first iteration cycle and / or the second iteration cycle to obtain the comprehensive feedback score of each core model and / or the comprehensive feedback score of each exploration model, and the ranking list of several style models is updated according to the comprehensive feedback score. It should be noted that if the first iteration cycle is a multiple of the second iteration cycle, the first iteration cycle and the second iteration cycle end at the same time, and the comprehensive feedback score of each core model and the comprehensive feedback score of each exploration model in the cycle will be obtained at the same time.
[0061] In some embodiments, the initial value of the preset scoring threshold can be flexibly set according to actual needs to ensure that the comprehensive feedback score of the style model that has existed in the performance model pool for a long time remains at a high level, thereby ensuring that the music quality and user satisfaction output by the model maintain a high average level.
[0062] In some embodiments, based on a preset expected number of eliminations and a ranking list of several style models, an expected number of elimination models are obtained, and the average value of the comprehensive feedback scores of the expected elimination models is compared with a preset score threshold. If the score difference between the average value and the preset score threshold is greater than the preset score difference, the preset score threshold is updated according to the highest score of the comprehensive feedback score of the expected elimination model. Furthermore, the preset score threshold is restored to the initial value in the next cycle.
[0063] For example, when the preset score threshold is 90 points, the preset score difference is 10 points, and the expected number of eliminations is 2, the two expected elimination models with the lowest comprehensive feedback scores are screened out from the ranking list, and the comprehensive feedback scores are 60 points and 70 points respectively. At this time, the score difference between the average value and the preset score threshold is greater than the preset score difference, and 70 points is updated as the preset score threshold. Therefore, the adaptability of the preset score threshold is evaluated in each iteration cycle, and the preset score threshold is dynamically optimized and adjusted according to the actual scenario to adapt to the score fluctuations in different cycles, and to avoid eliminating a large number of style models at one time and failing to replenish the vacancies in the performance model pool in a timely manner, thereby improving the accuracy and practicality of the model elimination and optimization mechanism.
[0064] In some embodiments, the model cache library is a dynamically managed model warehouse for storing style models that are not directly deployed to the performance model pool but have potential value, such as style models retrained based on user surveys, market trends or social hot spots, continuously optimized versions of existing high-scoring style models, new style models generated by fusing two or more style models, etc., for rapid replenishment after the style model is removed from the performance model pool. Exemplarily, the model cache library is also provided with a quantity limit for the first newly added exploration model, the second newly added exploration model, and the newly added resident model. When the actual quantity is lower than the quantity limit, the adaptive generation of the corresponding model will be triggered or the administrator will be prompted to import the corresponding model from the outside world. The specific quantity limit can be flexibly determined according to the frequency and quantity of style models eliminated by the performance model pool in actual applications, so as to provide sufficient model supply for the performance model pool and avoid redundancy.
[0065] In some embodiments, several high-scoring style models with the highest comprehensive feedback scores and historical target music data historically output by the high-scoring style models are obtained; based on the several high-scoring style models and the historical target music data, a new style model is generated and stored in a model cache.
[0066] Exemplarily, several high-scoring style models are determined based on comprehensive feedback scores, as well as historical target music data historically output by the high-scoring style models; a first newly added exploration model is trained based on at least two of the high-scoring style models with different style labels and the corresponding historical target music data, and the first newly added exploration model is stored in the model cache; and / or, a second newly added exploration model is trained based on at least two of the high-scoring style models with the same style label and the corresponding historical target music data, and the second newly added exploration model is stored in the model cache.
[0067] Specifically, the style models with the highest recent comprehensive feedback scores (e.g., top 10%) are selected as high-scoring style models. At the same time, the historical target music data generated by these high-scoring style models are collected (e.g., target music data generated in the past three months). It should be understood that the high-scoring style models selected may be any type of resident models, core models, or exploration models. Therefore, the basic model architecture or training data used in subsequent model training has a high degree of randomness and flexibility to ensure the continuous advancement of the innovative evolution of the model.
[0068] Specifically, an innovative mixed style model can be generated based on the existing style model, and at least two high-scoring style models of different styles and the corresponding historical target music data can be selected from the high-scoring style models to train a style model with a unique style, i.e., the first newly added exploration model. At this time, the style label of "mixed style" can be uniformly associated with such models; the style labels of the source style model can also be comprehensively considered, and multiple style labels can be associated at the same time, such as "popular style" and "Beethoven style"; intuitive and descriptive names can also be reused or regenerated, such as "popular style-Beethoven style".
[0069] Specifically, an optimized iterative style model is generated for a certain style, at least two high-scoring style models of the same style are selected from the high-scoring style models, and the corresponding historical target music data are trained to obtain a style model with optimized performance, and the first style label is inherited to associate with the new style model, thereby maintaining the core features of the original style and promoting the improvement of model performance.
[0070] In some embodiments, machine learning can also be used to automatically analyze the audio features of music samples generated by the new style model, and automatically generate or recommend the most appropriate style labels based on these audio features to ensure the consistency and accuracy of the labels. For example, a sub-model specifically for style classification is trained based on standard music samples under different style labels, and the most suitable style label is predicted by evaluating the new music samples.
[0071] It should be noted that the style label of a style model is not completely equivalent to the name of the style model. The style label is for standardized management and calling of style models, so it has overlap. The name needs to be generated differently so that users can quickly select the style model of interest and the corresponding music.
[0072] It should be noted that for the generation of new style models, various model training techniques can be flexibly selected. For example, for the generation of iterative style models, a combination of knowledge distillation and data enhancement can be used to use the output results of two high-scoring style models as "teacher signals" to guide the new model to learn more robust style features. For another example, for the generation of hybrid style models, parameter reorganization can be used to cross-reorganize some parameter modules (such as rhythm processing units and timbre rendering layers) of the network layers of two high-scoring style models to form a new architecture that is both stable and innovative. The specific model training technology used can be flexibly set according to the specific model architecture of the selected high-scoring style model and the current experimental direction, and is not specifically limited here.
[0073] In some embodiments, the method further includes: when the removed style model is a core model, based on the style label of the removed core model and the comprehensive feedback scores of each exploration model, selecting a target exploration model from the exploration models, and updating the target exploration model to a newly added core model.
[0074] Specifically, when a core model is removed, a targeted promotion mechanism is initiated. First, based on the style label of the removed core model, candidate objects with the same style label in the exploration model are screened, and multiple comprehensive feedback scores of the candidate models in the recent period are extracted. The target exploration model with the highest average score is selected and promoted to the newly added core model. At this time, the target exploration model is given the attributes of the core model, and its iteration cycle is automatically switched to a long-term evaluation mode. For example, after the core model of "mixed style" is eliminated, a new core model will be selected from the exploration model with the style label of "mixed style".
[0075] It should be understood that the newly promoted core model has been doubly verified by users and technical scores during the exploration stage, and its style performance has proven the reliability of its model performance through short-term high-frequency iterations. The verified high-quality exploration model will inherit the style positioning of the original core model, fill the gaps in the core model, and maintain the stability of the performance model pool output. This avoids quality fluctuations caused by the introduction of unverified new models, and enables the performance model pool to continuously absorb innovative results while maintaining the reliability of the core style, forming a virtuous evolutionary cycle.
[0076] In some embodiments, before importing the replacement style model from the model cache library, it also includes: according to the style label of the removed style model, counting the current total number of models with corresponding style labels in the performance model pool; if the current total number of models is greater than the preset number of models, using the first newly added exploration model as the replacement style model; if the current total number of models is less than the preset number of models, using the second newly added exploration model as the replacement style model.
[0077] Specifically, the preset number of models is the model base that needs to be maintained for each style label, such as "retain up to 3 exploration models for each style" or "retain at least 1 core / exploration model", which is used to balance the style diversity and coverage of the performance model pool. The specific value can be flexibly set according to the actual scenario and is not limited here.
[0078] When an exploration model is removed or promoted, there will be a vacancy in the exploration model. The corresponding statistics are the existing number of style labels of the removed or promoted exploration models in the performance model pool, that is, the current total number of models. If the current total number of models exceeds the preset number of models, and the number of models of the corresponding style in the performance model pool is sufficient, the first cross-style exploration model is imported from the model cache library to avoid excessive concentration of the same style; if the current total number of models is lower than the preset number of models, and the number of models of the corresponding style in the performance model pool is relatively vacant, the second enhanced exploration model of the same style is imported from the model cache library to avoid the loss of specific style output capabilities due to model elimination.
[0079] It should be understood that by presetting the number of models, on the one hand, the basic coverage of each style is ensured, avoiding redundancy or vacancies in various style models in the performance model pool; on the other hand, space for exploration is reserved for emerging styles, breaking homogeneity through cross-style models, and maintaining the diversity and innovation of styles in the performance model pool.
[0080] In some embodiments, the style model in the performance model pool is preset with a model type. When the music style output by the style model is a style that is widely accepted and popular among the public, the model type of the style model is a professional model. When the music style output by the style model is a new exploratory music style, the model type of the style model is an innovative model. For example, the first newly added exploratory model is an innovative model, and the second newly added exploratory model is a growth model.
[0081] In some embodiments, importing a replacement style model from a model cache library further includes: determining the type of model to be imported from the removed or promoted model based on the model type of the removed style model, that is, the type of the promoted or removed exploration model (that is, an innovative model or a growing model). Exemplarily, when an exploration model is removed or promoted to a core model, when the promoted or removed exploration model is an innovative model, the first newly added exploration model is used as the replacement style model; when the promoted or removed exploration model is a growing model, the second newly added exploration model is used as the replacement style model.
[0082] In some embodiments, the style model also includes a resident model; the method also includes: based on the third iteration cycle, counting the comprehensive feedback score of each of the resident models; when the style model removed from the performance model pool is a resident model, storing the removed resident model in a historical version library; importing a newly added resident model from the model cache library, wherein the newly added resident model is obtained by training the resident model in the historical version library based on the historical target music data.
[0083] Among them, the resident model is a special type of style model in the performance model pool, which is used to retain classic and representative performances and avoid style loss due to algorithm iteration, such as "Beethoven style" and "Baroque style". Exemplarily, the style model includes a third preset number of resident models, which is a fixed number set for the resident models, and is used to maintain specific classic or core styles for a long time, ensuring that classic resident models are retained in the performance model pool.
[0084] Specifically, the third iteration cycle can be an extremely long evaluation window. When a resident model is removed due to insufficient scoring, its complete parameters and historical generation data are stored in the historical version library for future version rollback or feature reuse. At the same time, a new resident model in the model cache library is introduced.
[0085] In some embodiments, historical target music data of a high-scoring style model is selected from the current performance model pool, and the removed resident model, or the most initial resident model in the historical version library, or the resident model with the highest average comprehensive feedback score in the historical version library is trained to ensure that the newly added resident model inherits the essence of the classic style and integrates the latest optimization results and user preferences.
[0086] It should be understood that by taking the resident model as the basis for training, it is ensured that the essential characteristics of the classic style are preserved, and the style degradation caused by algorithm iteration is prevented. Then, the modern high-scoring data is mixed to inject innovative elements appropriately, ensuring that the newly added resident model inherits the essence of the classic style and integrates the latest optimization results and user preferences. In addition, when training the newly added resident model, the original resident model or high-scoring resident model with more accurate classic style features can be called based on the historical version library. The longer third iteration cycle can also avoid frequent iterations of the resident model, thereby avoiding the loss of style caused by rapid updates and iterations of the resident model, and realizing incremental innovation of the resident model.
[0087] It should be noted that the resident model is the reproduction and inheritance of classical music styles, such as "Beethoven-style piano sonata", "Baroque polyphonic style", and "Chopin nocturne style". In specific application scenarios, the music generated by it is mostly used to provide users with reference, learning and basic functional support for traditional music styles. Even if the user feedback score is low (for example, the like rate, collection rate and other behavioral feedback data are low), the existence of the resident model ensures the professionalism and comprehensiveness of the performance model pool. The core model can point to high-quality, high-usage, and multi-dimensional style models in specific application scenarios. Taking the core model with the style label "classical style" as an example, it can be a classical style model that caters to the public's aesthetics, or a classical style model that integrates the performance habits of many famous artists, or a classical style model for film soundtracks, or a classical style model that integrates the pedal control dimension in addition to the basic pitch, time value, and intensity control dimensions. The exploration model can be an iterative optimization version of the classical style model for film soundtracks, or a very interesting and innovative cyberpunk jazz style model, or a randomly collaged jazz and classical style model.
[0088] In some embodiments, it includes: importing a first newly added exploration model and / or a second newly added exploration model from a model cache library, using the imported first newly added exploration model and / or the second newly added exploration model as a new exploration model, thereby updating the performance model pool; importing a newly added resident model from the model cache library, using the imported newly added resident model as a new resident model, thereby updating the performance model pool.
[0089] In some embodiments, the method also includes: comparing a number of the target music data, and when the similarity between any two of the target music data is higher than a first similarity threshold, respectively comparing the first historical music data and the second historical music data of the two corresponding style models; if the similarity between the first historical music data and the second historical music data is higher than a second similarity threshold, removing one of the style models from the performance model resource library.
[0090] Among them, the first similarity threshold is a set music data similarity judgment standard, which is used to determine whether the real-time outputs of two target music data are highly identical; the first historical music data and the second historical music data respectively refer to the music data sets generated in the past by the corresponding two style models; the second similarity threshold is also a set music data similarity judgment standard, which is used to determine whether the historical data between the models show a high degree of similarity for a long time. It can be greater than, less than or equal to the first similarity threshold. The specific values of the first similarity threshold and the second similarity threshold can be set according to the sensitivity of the actual similarity judgment, and are not limited here.
[0091] Specifically, if it is detected that the similarity of the target music data generated by any two style models exceeds the first similarity threshold, the historical music data of the two style models are further retrieved, and the historical music data of the two style models are the first historical music data and the second historical music data, respectively, to calculate the historical similarity between the two. If the historical similarity also exceeds the second similarity threshold, it indicates that the long-term style performances of the two are highly similar, and one of the style models will be removed randomly or according to the comprehensive feedback score to avoid redundancy in the performance model pool.
[0092] It should be understood that through the dual similarity verification of real-time and history, long-term homogeneous models can be accurately identified to ensure the differentiation of musical styles output by the performance model pool and avoid misjudgment caused by coincidence of single generation. After similar models are eliminated, resource space is released to be allocated to new exploration models, accelerating the attempt of emerging styles, promoting innovation, and improving the vitality of the performance model pool.
[0093] In some embodiments, the method further includes: detecting a growth trend of comprehensive feedback scores of a plurality of core models; and retraining the baseline model when a growth trend of a fourth preset number of comprehensive feedback scores is abnormal.
[0094] Specifically, the growth trend refers to the pattern of change of the comprehensive feedback score of the core model over time (such as continuous increase, stability or decrease), while the abnormal growth trend refers to the significant decrease or fluctuation of the comprehensive feedback score of the core model within a preset period, such as the growth trend exceeds the preset interval. When such anomalies are detected in a fourth preset number of core models (such as 8 or more) at the same time, it is presumed that it may be due to fundamental defects in the basic music data output by the benchmark model. The abnormal reasons for the growth trend of the comprehensive feedback score are obtained through channels such as user surveys, market trends or social hot spots, and the training data of the benchmark model is updated based on the abnormal reasons, and / or the model architecture of the benchmark model is adjusted, and the benchmark model is retrained.
[0095] Among them, the abnormal reason may be that the benchmark model's ability to parse music scores has degraded or the generated parameters are inaccurate, or the benchmark model cannot output the benchmark music data of the dimensions required by the current emerging music trends. For example, the basic music data output by the benchmark model of the current experiment does not include pedal-related data, and the music score input by the user contains a large number of instructions for the use of the sustain pedal or other pedals, in which case the benchmark music data will be inaccurate.
[0096] Exemplarily, the baseline model can be retrained using the original training data after updating the architecture, or using the updated training data (for example, updating the data content or data accuracy contained in the non-stylized data), or using the updated training data to train the baseline model without updating the architecture.
[0097] It should be understood that the basic music data output by the benchmark model is the basic input for all style models. Its defects may lead to global errors. Timely repairs can avoid the collective failure of style models due to problems at the base layer. A stable benchmark model provides a reliable technical foundation for the style model, allowing it to focus on style innovation rather than error correction. Furthermore, since the music data output by the benchmark model is difficult to directly evaluate from dimensions such as technical compliance and user interaction feedback, the update of the benchmark model is triggered through the linkage of the health of the core model to ensure the overall adaptability and long-term stability of the performance model pool.
[0098] In some embodiments, the growth trend of the comprehensive feedback scores of several core models is detected; when the growth trend of the comprehensive feedback scores of the fifth preset number of the same style label is abnormal, it is inferred that it may be due to the decline in demand for the corresponding style or the existence of fundamental defects in the target music data output by the style model, and the style abnormality reasons for the growth trend of the comprehensive feedback scores are obtained based on channels such as user surveys, market trends or social hot spots, and the training data of the corresponding style model is updated based on the style abnormality reasons, and / or the model architecture of the corresponding style model is adjusted, and the corresponding style model is retrained. Furthermore, if the style abnormality reason is the decline in demand for the corresponding style, the number of preset models of the corresponding style can be lowered to reduce the generation of style models of the corresponding style under the subsequent elimination mechanism, release resources, and achieve dynamic resource optimization.
[0099] See also Figure 2 , Figure 2 FIG. 1 is a schematic flow chart of another method for generating music for performance provided in an embodiment of the present application. Figure 2As shown, the music score to be performed is input into the benchmark model to convert the music score information into basic music data standardized at the technical level, and then the basic music data generated by the benchmark model is input into multiple style models. The style model includes three types: core model, exploration model and resident model, which convert the basic music data into target music data with a specific style from different style dimensions. It should be understood that the basic generation of music performance is separated from the style transfer task. The benchmark model is responsible for generating technical basic data to ensure the correctness of the underlying technology of music generation; the style model focuses on style diversification, and realizes accurate simulation of different styles through diversified models, avoiding the waste of computing power caused by repeated calculation of basic data, and only needs to update the style model to introduce new music styles or optimize existing style expressions without retraining the benchmark model, which improves flexibility and effectively saves computing power resources.
[0100] See also Figure 3 , Figure 3 Schematic diagram of a dynamic update of a performance model pool provided by an embodiment of the present application. Figure 3 As shown, the performance model pool consists of a baseline model and a preset number of style models (including a first preset number of core models, a second preset number of exploration models, and a third preset number of resident models). Different types of style models in the performance model pool each have specific responsibilities: the core model is used to maintain high-quality, stable output of music of various styles; the exploration model is used to try new trends or integrate the features of existing high-scoring models to promote style innovation and performance improvement; the resident model is used for long-term optimization of specific classic styles to prevent the loss of classic styles due to algorithm iteration. Through division of labor and cooperation, the performance model pool can not only cover a wide range of performance needs, but also continue to stimulate innovation possibilities while ensuring the basic experience.
[0101] like Figure 3 As shown in the figure, the dotted line indicates the source and use of the training data in the dynamic update of the performance model pool. It can also be understood as the data flow of the data (including basic music data and target music data) output by each model in the performance model pool during use. The basic music data generated by the baseline model are input into the core model, the exploration model and the resident model respectively. These style models output the target music data respectively. These target music data are used to output the music required by the user on the one hand, and can also be used as Figure 3 The high-scoring style model or resident model shown as training data is used for training to obtain a newly added exploration model (including a first newly added exploration model and a second newly added exploration model) and a newly added resident model, thereby reducing computing power consumption while ensuring quality.
[0102] For the newly added exploratory model, the existing style model is trained with several existing target music data as training data, and the features of the high-scoring model are integrated to obtain an optimized iterative style model or an innovative mixed style model and store it in the model cache. For the resident model, the existing resident model (the resident model of the historical version can be called) is trained with several existing target music data as training data, and the features of the high-scoring model are integrated to obtain a new resident model and store it in the model cache, ensuring that the new resident model inherits the essence of the classic style and integrates the latest optimization results and user preferences.
[0103] like Figure 3 As shown, the solid line represents the dynamic update process of the model in the performance model pool. For the core model, after the outdated and inefficient core models are eliminated, new core models with excellent performance are selected from the exploration model to inherit the innovative results of the exploration model and ensure that the music data generated by each core model maintains high quality and high satisfaction. For the exploration model, when the exploration model is removed or promoted, the new exploration model is imported from the model cache library. For the resident model, when the resident model is removed, the new resident model is imported from the model cache library. Furthermore, when the comprehensive feedback scores of a large number of core models are abnormal, it will also trigger the update of the baseline model. Therefore, through adaptive elimination and evolution driven by feedback scores, the high vitality of the performance model pool is maintained.
[0104] Furthermore, the entire performance model pool is designed with differences in iteration cycles. Since the core model is relatively reliable, it can be evaluated and updated using a low-frequency iteration cycle to ensure that each exploration model promoted to a core model has a stable high-quality score for at least a longer probation period; on the contrary, the exploration model uses a high-frequency iteration cycle to quickly trial and error and capture new trends to keep the performance model pool active. The corresponding resident model also has a specific iteration cycle to avoid the loss of style due to rapid updates and iterations of the resident model, and to achieve gradual innovation of the resident model.
[0105] It should be understood that the adaptive elimination and evolution mechanism is flexibly set based on the characteristics and functions of different types of style models (core model, exploration model, resident model). For example, different style models have different iteration cycles. While ensuring the quality and stability of music, they continue to promote model innovation and integrate the latest optimization results and user preferences. The sources of new models are different, which not only ensures the stability of classic styles and the coverage of various styles, but also improves the exploration efficiency of emerging styles. These differentiated settings can form a benign synergy, maintain the vitality of the performance model pool in a long-term and adaptive manner, and avoid the performance model pool from losing competitiveness due to content homogeneity and lag, and finally achieve a balance between the stability and innovation of the performance model pool.
[0106] In addition, the model iteration is completed based on the high-scoring music data and high-scoring style models in actual use, so as to learn emerging elements in real time and avoid relying on fixed training data and rigid model architecture. In addition, the basic generation of music performance and the style transfer task are decoupled into two models. Style adjustment only requires retraining the style model, avoiding model iteration from relying on manual intervention or full retraining, greatly reducing computing resource consumption, training cycle, and training cost, thereby responding to users' demand for freshness in a timely manner.
[0107] In some embodiments, the first newly added exploration model, the second newly added exploration model, and the newly added resident model are all pre-evaluated for technical qualifications. If the first score is higher than the preset score, they are stored in the model cache library after ensuring that the quality meets the standard, ensuring that the model cache library continues to introduce style models with style innovation or performance optimization, and providing diversified choices for subsequent deployment.
[0108] See also Figure 4 , Figure 4 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device may be a terminal device or a server.
[0109] Exemplarily, the above method and apparatus may be implemented in the form of a computer program. The computer program may be implemented in Figure 4 Runs on the computer device shown.
[0110] like Figure 4 As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.
[0111] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any method for generating music for performance.
[0112] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.
[0113] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any method for generating music for performance.
[0114] The network interface is used for network communication, such as sending assigned tasks.
[0115] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0116] In one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps: S101, obtaining a music score to be performed and a performance model pool, wherein the performance model pool includes a benchmark model and a preset number of style models, and each of the style models is associated with a style label, and the style models include a core model and an exploration model; S102, inputting the music score to be performed into the benchmark model to generate basic music data; S103, inputting the basic music data into a plurality of the style models respectively to generate a plurality of target music data; S104, generating a feedback score for each of the style models according to a number of the target music data; calculating a comprehensive feedback score for each of the core models based on a first iteration cycle; calculating a comprehensive feedback score for each of the exploration models based on a second iteration cycle; wherein the first iteration cycle is greater than the second iteration cycle; S105. When the comprehensive feedback score of any of the style models is continuously lower than a preset score threshold, the corresponding style model is removed from the performance model pool, and a replacement style model is imported from a model cache library to update the performance model pool, wherein the model cache library stores a plurality of pre-trained style models.
[0117] Exemplarily, the processor is used to run a computer program stored in the memory, and is also used to implement the steps of the method for generating performance music provided in any embodiment of the present application, which will not be repeated here.
[0118] A computer-readable storage medium is also provided in an embodiment of the present application, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and the processor executes the program instructions to implement the steps of any one of the methods for generating music for performance provided in the embodiments of the present application.
[0119] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc., equipped on the computer device.
[0120] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A method for generating music for performance, characterized in that: The method comprises: S101, obtaining a music score to be performed and a performance model pool, wherein the performance model pool includes a benchmark model and a preset number of style models, and each of the style models is associated with a style label, and the style models include a core model and an exploration model; S102, inputting the music score to be performed into the benchmark model to generate basic music data; S103, inputting the basic music data into a plurality of the style models respectively to generate a plurality of target music data; S104, generating a feedback score for each of the style models according to a number of the target music data; calculating a comprehensive feedback score for each of the core models based on a first iteration cycle; calculating a comprehensive feedback score for each of the exploration models based on a second iteration cycle; wherein the first iteration cycle is greater than the second iteration cycle; S105. When the comprehensive feedback score of any of the style models is continuously lower than a preset score threshold, the corresponding style model is removed from the performance model pool, and a replacement style model is imported from a model cache library to update the performance model pool, wherein the model cache library stores a plurality of pre-trained style models.
2. The method according to claim 1, characterized in that The feedback score includes a first score and / or a second score, and the method includes: Scoring the target music data of each style model based on a preset scoring rule to obtain a first score for each style model; and / or, Behavior feedback data of the user is obtained, and a second score of each of the style models is generated based on the behavior feedback data.
3. The method according to claim 1, characterized in that: The method further comprises: Determine a plurality of high-scoring style models based on the comprehensive feedback scores, and historical target music data historically outputted by the high-scoring style models; Based on at least two of the high-scoring style models with different style labels and corresponding historical target music data, a first newly added exploration model is trained and the first newly added exploration model is stored in the model cache library; and / or, Based on at least two of the high-scoring style models with the same style tag and the corresponding historical target music data, a second newly added exploration model is trained and stored in the model cache.
4. The method according to claim 3, characterized in that Before importing the replacement style model from the model cache library, the method further includes: According to the style label of the removed style model, counting the total number of current models corresponding to the style label in the performance model pool; If the total number of current models is greater than the preset number of models, using the first newly added exploration model as the replacement style model; If the total number of current models is less than the preset number of models, the second newly added exploration model is used as the replacement style model.
5. The method according to claim 3, characterized in that: The style model also includes a resident model; the method also includes: Calculating the comprehensive feedback score of each of the resident models based on the third iteration cycle; When the style model removed from the performance model pool is a resident model, storing the removed resident model into a historical version library; A newly added resident model is imported from the model cache library, wherein the newly added resident model is obtained by training the resident model in the historical version library based on the historical target music data.
6. The method according to claim 1, characterized in that The method further comprises: When the removed style model is a core model, based on the style label of the removed core model and the comprehensive feedback scores of each exploration model, a target exploration model is selected from the exploration models, and the target exploration model is updated to the newly added core model.
7. The method according to claim 1, characterized in that The method further comprises: Comparing a plurality of the target music data, when the similarity between any two of the target music data is higher than a first similarity threshold, the first historical music data and the second historical music data of the two corresponding style models are respectively compared; If the similarity between the first historical music data and the second historical music data is higher than a second similarity threshold, one of the style models is removed from the performance model resource library.
8. The method according to claim 1, characterized in that The method further comprises: Detect the growth trend of the comprehensive feedback scores of several core models; When a growth trend of a fourth preset number of comprehensive feedback scores is abnormal, the benchmark model is retrained.
9. A computer device, characterized in that: The device comprises: Memory for storing computer programs; A processor, configured to execute the computer program and implement the method for generating performance music as claimed in any one of claims 1 to 8 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to implement the method for generating performance music according to any one of claims 1 to 8.
Citation Information
Patent Citations
MIDI playing style automatic conversion system based on recurrent neural network
CN111554255A
Music family style simulation automatic playing system and method
CN113936625A
Customized music generation method and device based on common semantic space
CN106898341A
Music style merging method based on coupled generative adversarial networks
CN110085203A
Method and device for synthesizing music
CN111724764A
Cited By
Playing data display method and display system
CN120260527A
Method and system for inserting pedal data in playing file
CN121640970A
Method for generating performance music, method for training and using performance model, computer device, and storage medium
WO2026179727A1