A method for generating music performance, a computer device, and a storage medium
By building a dynamically managed performance model pool, combined with feedback score-driven differentiated elimination and evolution mechanism, the problem that static stylized performance models cannot adapt to music trends is solved, long-term adaptive maintenance and style diversification of the performance models are achieved, and user experience and resource utilization efficiency are improved.
Patent Information
- Application Number
- CN202510473716.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing static stylized performance model cannot capture the latest music trends, resulting in the generated music gradually lagging behind the times, affecting the user experience and application value.
Build a dynamically managed performance model pool, including benchmark models and multiple style models, through feedback score-driven differentiated elimination and evolutionary mechanisms, ensure long-term adaptability and diversity of models, and divide labor and cooperation to maintain music quality and style innovation.
Long-term adaptive maintenance of the performance model pool is realized, avoiding model redundancy or obsoleteness, maintaining music quality stability and style diversity, and improving flexibility and computing resource utilization efficiency.
Smart Images

Figure CN120015000B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of generating performance music, and particularly to a method for generating performance music, a computer device, and a storage medium. Background Art
[0002] With the development of digital music technology and artificial intelligence, stylized performance models have gradually become a hot topic in the field of music technology. These models analyze and understand musical score information in paper or electronic format, imitate the performance styles of specific musicians or the characteristics of specific music genres, and provide users with unique music experiences.
[0003] For example, the invention patent application with the publication number CN113936625A discloses an automatic performance system and method for simulating a musician's style. The system includes a storage module, a processing module, and an interaction module. The processing module is used to generate performance style models corresponding to different musicians through machine learning training on the music data stored in the storage module; query the corresponding performance style model according to the specified musician information transmitted by the interaction module, input the specified music, and output the music of the simulated performance, thereby realizing automatic performance.
[0004] The invention patent application with the publication number CN111554255A discloses a MIDI performance style automatic conversion system based on a recurrent neural network, including a MIDI analysis module, a data preprocessing module, an autoencoder module, a style network module, and a music generation module. Among them, the MIDI analysis module reads the user input, merges the multi-track MIDI into a single-track MIDI as the musical score; the data preprocessing module is used to extract note features from the musical score; the autoencoder module encodes and decodes the note features; the style network module learns the performance style of the musical score and predicts the intensity vector; the music generation module is used to configure the intensity vector predicted by the style network module for the musical score and convert it into an expressive music.
[0005] However, with the continuous evolution of music trends, static stylized performance models are too rigid to capture and adapt to the latest music trends, resulting in the generated music gradually falling behind the times, limiting its long-term application value, and reducing the user experience. Summary of the Invention
[0006] The main purpose of this application is to provide a method for generating performance music, a computer device, and a storage medium. To solve the above-mentioned technical problems, this application specifically adopts the following technical solutions:
[0007] In the first aspect of this application, a method for generating performance music is provided. The method includes:
[0008] S101. Obtain the music score to be played and a performance model pool. The performance model pool includes a benchmark model and a preset number of style models, and each style model is associated with a style label. The style model includes a core model and an exploration model;
[0009] S102. Input the music score to be played into the benchmark model to generate basic music data;
[0010] S103. Input the basic music data into several style models respectively to generate several target music data;
[0011] S104. Generate a feedback score for each style model according to the several target music data; statistically calculate the comprehensive feedback score of each core model based on the first iteration period; statistically calculate the comprehensive feedback score of each exploration model based on the second iteration period; wherein, the first iteration period is greater than the second iteration period;
[0012] S105. When the comprehensive feedback score of any style model continuously falls below a preset score threshold, remove the corresponding style model from the performance model pool, and import a replacement style model from the model cache library to update the performance model pool. A number of pre-trained style models are stored in the model cache library.
[0013] In some embodiments, the feedback score includes a first score and / or a second score. The method includes: scoring the target music data of each style model based on a preset scoring rule to obtain the first score of each style model; and / or, obtaining the behavioral feedback data of the user, and generating the second score of each style model based on the behavioral feedback data.
[0014] In some embodiments, the method further includes: determining several high-score style models based on the comprehensive feedback score, and the historical target music data output by the high-score style models in history; training a first newly added exploration model based on at least two high-score style models with different style labels and the corresponding historical target music data, and storing the first newly added exploration model in the model cache library; and / or, training a second newly added exploration model based on at least two high-score style models with the same style label and the corresponding historical target music data, and storing the second newly added exploration model in the model cache library.
[0015] In some embodiments, before importing the replacement style model from the model cache library, the method further includes: according to the style label of the removed style model, counting the total number of current models in the performance model pool corresponding to the style label; if the total number of current models is greater than the preset number of models, using the first newly added exploration model as the replacement style model; if the total number of current models is less than the preset number of models, using the second newly added exploration model as the replacement style model.
[0016] In some embodiments, the style model further includes a resident model; the method further includes: based on the third iteration cycle, counting the comprehensive feedback score of each resident model; when the removed style model from the performance model pool is a resident model, storing the removed resident model in the historical version library; importing a newly added resident model from the model cache library, where the newly added resident model is obtained by training the resident model in the historical version library based on the historical target music data.
[0017] In some embodiments, the method further includes: when the removed style model is a core model, selecting a target exploration model from the exploration models based on the style label of the removed core model and the comprehensive feedback scores of each exploration model, and updating the target exploration model to a newly added core model.
[0018] In some embodiments, the method further includes: comparing several pieces of the target music data, and when the similarity between any two pieces of the target music data is higher than the first similarity threshold, the first historical music data and the second historical music data corresponding to the two respective style models; if the similarity between the first historical music data and the second historical music data is higher than the second similarity threshold, removing one of the style models from the performance model resource library.
[0019] In some embodiments, the method further includes: detecting the growth trend of the comprehensive feedback scores of several core models; when the growth trends of the comprehensive feedback scores of the fourth preset number are abnormal, retraining the benchmark model.
[0020] The second aspect of the present application is to provide a computer device, which includes:
[0021] A memory for storing a computer program;
[0022] A processor for executing the computer program and implementing the steps of the method for generating performance music provided in any embodiment of the present application when executing the computer program.
[0023] In a third aspect of the present application, there is also correspondingly provided a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to perform the steps of the music performance generation method provided in any embodiment of the present application.
[0024] Beneficial effects:
[0025] The present application proposes a music performance generation method, a computer device, and a storage medium. By constructing a dynamic management performance model pool and combining a differential elimination and evolution mechanism driven by feedback scoring, while stabilizing the music quality, maintaining style diversification, and efficiently exploring innovation, it avoids model redundancy or obsolescence in the performance model pool and realizes the long-term adaptive maintenance of the performance model.
[0026] The performance model pool consists of a benchmark model and style models (including core models, exploration models, and resident models), separating the basic generation and style transfer tasks of music performance. The benchmark model is responsible for generating technical basic data to ensure the correctness of the underlying technology of music generation; the style models focus on style diversification, achieving accurate simulation of different styles, avoiding waste of computing power caused by repeated calculation of basic data, and only need to update the style models to introduce new music styles or optimize existing style expressions without retraining the benchmark model, improving flexibility and effectively saving computing resources.
[0027] Furthermore, different types of style models in the performance model pool each undertake specific responsibilities: the core model is used to maintain the stable output of high-quality music in various styles; the exploration model is used to try new trends or integrate the features of existing high-score models to promote style innovation and performance improvement; the resident model is used for the long-term optimization of specific classic styles to prevent the loss of classic styles caused by algorithm iteration. Through division of labor and cooperation, the performance model pool can not only cover a wide range of performance requirements but also continuously stimulate innovation possibilities while ensuring the basic experience.
[0028] In addition, based on the characteristics and functions of different types of style models (core models, exploration models, resident models), the elimination and evolution mechanisms of the models are flexibly set. For example, the iteration cycles of different style models are different, and the sources of new models are different. These differential settings can form a beneficial synergy, making several models inside the performance model pool in different stages all have value for existence, maintaining the balance between stability and innovation in the performance model pool, avoiding the homogenization or obsolescence of the content of the performance model pool, and adaptively maintaining the vitality of the performance model pool in the long term.
[0029] Specifically, the core model is evaluated and updated using a low-frequency iteration cycle to ensure that the music data it generates maintains high quality and high satisfaction. At the same time, outdated core models are eliminated to inherit the innovative achievements of the exploration models, ensuring that each exploration model promoted to the core model has at least a stable high-quality score during a relatively long probationary period. On the contrary, the exploration model adopts a high-frequency iteration cycle to quickly trial and error and capture new trends, maintaining the vitality of the model pool. By integrating the features of high-scoring models, it ensures quality while reducing computing power consumption. The resident model focuses on the long-term maintenance of specific classic styles and can be appropriately adjusted according to user preferences. Further, when the comprehensive feedback scores of a large number of core models are abnormal, it will also trigger an update of the benchmark model. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale. Obviously, the following-described drawings are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0031] Figure 1 is a schematic flowchart of a method for generating performance music provided by an embodiment of the present application;
[0032] Figure 2 is a schematic flowchart of another method for generating performance music provided by an embodiment of the present application;
[0033] Figure 3 is a schematic diagram of dynamic update of another performance model pool provided by an embodiment of the present application;
[0034] Figure 4 is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0036] In this document, suffixes such as "module", "component", or "unit" used to represent elements are only for the convenience of explaining this application, and have no specific meaning in themselves. Therefore, "module", "component", or "unit" can be used interchangeably.
[0037] In this document, the orientation or positional relationship indicated by terms such as "upper", "lower", "inner", "outer", "front", "rear", "one end", "the other end", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation on this application. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0038] In this document, unless otherwise clearly specified and defined, terms such as "installed", "provided with", "connected", etc. should be understood in a broad sense. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium, and can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0039] In this document, "and / or" includes any and all combinations of one or more of the listed related items.
[0040] In this document, "a plurality of" means two or more, that is, it includes two, three, four, five, etc.
[0041] It should be noted that in this document, the term "comprising", "including", or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article, or device comprising that element.
[0042] Traditional music generation techniques often have the defect of staticization and lack the ability to learn new trends in real time, resulting in the generated content being limited to the music styles during training and making it difficult to capture emerging elements. In addition, model updates rely on manual intervention or full-scale retraining, with a long iteration cycle and high cost, and cannot respond in a timely manner to users' demand for freshness. As music trends continue to evolve, the generated music gradually becomes out of touch with users' expectations, and is prone to losing competitiveness due to content homogenization and lag in long-term applications, ultimately affecting the user experience and practical application value.
[0043] Based on this, the present application proposes a method for generating performance music, a computer device, and a storage medium. By constructing a dynamically managed performance model pool and combining a differential elimination and evolution mechanism driven by feedback scoring, while stabilizing the music quality, maintaining style diversity, and efficiently exploring innovation, it avoids model redundancy or obsolescence in the performance model pool and realizes the long-term adaptive maintenance of the performance models.
[0044] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other. Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for generating performance music provided by an embodiment of the present application. As Figure 1 shown, an embodiment of the present application provides a method for generating performance music, and the method includes S101 to S105.
[0045] S101. Obtain a score to be performed and a performance model pool. The performance model pool includes a benchmark model and a preset number of style models, and each style model is associated with a style label. The style models include a core model and an exploration model.
[0046] Among them, the score to be performed refers to a standardized digital score file input by the user, which contains basic information such as the note sequence, rhythm, and chord progression of the music piece, and is the basic input for generating music subsequently.
[0047] Among them, the performance model pool includes two types of music generation models, namely a benchmark model and style models. The number of style models is a preset number to ensure that the performance model pool is eliminated and optimized within a reasonable number range. The specific number can be flexibly set according to actual needs and is not limited here.
[0048] Among them, the benchmark model serves as a basic generation unit for parsing the score and generating non-stylized basic music data to ensure technical accuracy and consistency; the style models further perform stylization processing on the basic music data. The number and type of the style models can be preset according to requirements, and each style model is associated with a style label for identifying the music style that it is good at processing or applicable to, such as "classical style", "jazz style", "impressionist style", "mixed style", etc.
[0049] Further, the style models are subdivided into a core model and an exploration model. The core model is a high-quality style model that has been verified for a long time and is responsible for stably outputting high-quality music data under various styles. The exploration model is an experimental style model for optimizing model performance or innovating performance styles, and is responsible for style innovation and trend capture to enhance the interest of the output music.
[0050] In some embodiments, the style model includes a first preset number of core models and a second preset number of exploration models. The first preset number and the second preset number can be flexibly set according to actual needs and are not limited herein. Exemplarily, the second preset number can be less than the first preset number, and the preset number ratio between the two can be flexibly set to meet the different requirements for style stability and innovation in different application scenarios.
[0051] Thus, through the distinction between style tags and functional positioning, the style models jointly constitute a dynamic hierarchical structure of the performance model pool, which not only ensures the stability of music data under different styles but also provides a flexible space for the exploration of new styles.
[0052] In some embodiments, in response to a user's selection operation and / or import operation, a plurality of core models selected by the user are determined. The selection operation refers to an operation instruction for the user to select the required reference model and style model on a preset model selection interface, and the import operation refers to an operation instruction for the user to import the required style model on a preset model import interface. The number of the selected plurality of core models is counted to obtain the specific value of the first preset number. Based on the preset number ratio between the first preset number and the second preset number and the specific value of the first preset number, the specific value of the second preset number is determined. A second preset number of style models are randomly selected from the model selection interface as the second preset number, and a performance model pool is constructed based on the first preset number of core models and the second preset number of style models.
[0053] It should be understood that the performance model pool can be customized by the user. That is, it can be selected from a preset model library through the model selection interface or imported from the outside as needed. And the subsequent update of the performance model pool is also based on user feedback. However, the reference model only supports selection from the preset model library to ensure the correctness of the underlying technology of music generation, and the core architecture and elimination and optimization mechanism of the performance model pool are restricted, taking into account both the controllability and expandability of music generation.
[0054] S102: Input the to-be-performed music score into the reference model to generate basic music data.
[0055] Among them, the reference model is a pre-trained deep neural network for converting the input to-be-performed music score into technically standardized basic music data. Exemplarily, the training data of the reference model is a large number of standardized music scores and corresponding non-style data. The non-style data can include annotation data such as the note sequence, note start and end times, beat information, chords, pedals, etc. corresponding to the music score.
[0056] Furthermore, the benchmark model can adopt an encoder-decoder architecture based on Transformer or LSTM. The encoder parses the input musical score to be played into a structured feature vector containing pitch, duration, and dynamics. The decoder reconstructs the style-neutral basic music data based on the attention mechanism, including precise note start and end times, standard dynamics values, standard beat information, etc.
[0057] It should be understood that the basic music data output by the benchmark model needs to be accurate and error-free in terms of technical specifications, avoiding basic errors, while maintaining style neutrality, avoiding bias towards any specific playing style, ensuring that all style models are stylized based on a unified technical foundation, and providing a reliable basis for diverse style expressions.
[0058] S103. Input the basic music data into a number of the style models respectively to generate a number of target music data.
[0059] Among them, the style model is a deep neural network trained for stylization, which is used to transform the basic music data generated by the benchmark model into target music data with a specific style. Exemplarily, the training data of each style model contains a large number of labeled data sets corresponding to the style. For example, the data set using the classical style contains features such as typical rhythm changes and improvised ornaments of this style.
[0060] Furthermore, the style model can adopt a deep neural network with a style embedding layer. The input layer receives the basic music data, the style embedding layer loads the predefined style label vector, and the output layer generates the target music data containing the target style expressiveness. It should be understood that in this process, each style model processes the same basic music data differently. The core models of different styles generate stable and high-quality target music data; the exploration model generates innovative target music data. And all the generated results maintain the main structure framework of the original musical score (such as more than 90% of the note sequences are not adapted), and superimpose stylized performance parameters (such as ornament patterns of a specific genre), and stylistically reconstruct the performance dimension, and finally output a set of target music data corresponding to the style label.
[0061] S104. Generate a feedback score for each of the style models according to the number of the target music data; statistically calculate the comprehensive feedback score of each of the core models based on the first iteration cycle; statistically calculate the comprehensive feedback score of each of the exploration models based on the second iteration cycle; where the first iteration cycle is greater than the second iteration cycle.
[0062] Among them, the feedback score is an index for quantitatively evaluating the performance of the style model. The weighted average or the moving window average of all the feedback scores of the style model within a specific iteration cycle is taken to obtain the comprehensive feedback score.
[0063] Specifically, the core model is evaluated and updated in the first iteration cycle, which is set to low-frequency evaluation, such as once a month or once a quarter. The feedback scores of all target music data within the cycle are statistically analyzed for the core model to generate a comprehensive feedback score. It should be understood that the core model is used to maintain the stable output of high-quality music in various styles. Its model performance has been verified over a long time and has a high degree of reliability. Therefore, it is only necessary to regularly verify whether the core model maintains high-quality and high-score performance to promptly eliminate outdated core models.
[0064] Specifically, the exploration model is evaluated and updated in the second iteration cycle, which is set to high-frequency evaluation, such as once a week or once every three days. The feedback scores of all target music data within the cycle are statistically analyzed for the exploration model to generate a comprehensive feedback score. It should be understood that the exploration model is used to attempt new trends or integrate the characteristics of existing high-score models to promote style innovation and performance improvement. Its core purpose is to attempt innovation, and the actual model performance fluctuates greatly. It may either perform exceptionally well or mediocrely or extremely poorly. High-frequency evaluation is required to accelerate elimination or optimization and quickly verify the positive nature of innovation.
[0065] In some embodiments, the feedback score includes a first score and / or a second score, and the method includes: scoring the target music data of each style model based on a preset scoring rule to obtain the first score of each style model; and / or, obtaining the behavioral feedback data of users and generating the second score of each style model based on the behavioral feedback data.
[0066] Specifically, the preset scoring rule refers to a scoring rule set in advance according to technical compliance dimensions such as the rationality of the pitch range, the logic of chord progressions, and the consistency of dynamic performance. When the target music data is generated, the first score will be comprehensively generated for the target music data according to the above technical compliance dimensions to evaluate the technical compliance of the generated music. Exemplarily, different preset scoring rules can be set according to different styles, and different preset scoring rules can also be set according to different model types (such as the core model, exploration model). Correspondingly, the corresponding first scoring models can be trained based on the music data and the scoring marker datasets under different preset scoring rules, and the target music data can be analyzed and scored based on different first scoring models.
[0067] Specifically, the behavior feedback data refers to the behavior records generated during the user's interaction with the music content corresponding to the target music data, including but not limited to: the user's direct rating, the music playback completion rate, the repeat playback frequency, the like rate, the collection rate, the sharing rate, the download rate, and the playback skip rate, which reflect the user's preferences, interests, and satisfaction with the target music. After the target music data is generated, a second score is comprehensively generated based on the above-mentioned behavior feedback data to evaluate the market acceptance of the generated music. Exemplarily, the second score values corresponding to different behavior feedback data can be preset in advance and a score mapping table can be formed to quickly convert the behavior feedback data into the second score.
[0068] In some embodiments, the feedback score may further include: a model performance score and an expert sampling score. The method includes: scoring the target music data of each style model based on dimensions such as generation latency, resource occupancy rate, and audio quality to obtain the model performance score of each style model; for the expert sampling score, randomly selecting the target music data of the style model and scoring it based on dimensions such as style feature restoration degree and melody harmony to obtain the expert sampling score of the corresponding style model.
[0069] S105. When the comprehensive feedback score of any one of the style models continuously falls below the preset score threshold, remove the corresponding style model from the performance model pool and import a replacement style model from the model cache library to update the performance model pool. A number of pre-trained style models are stored in the model cache library.
[0070] Specifically, based on the iteration cycles of different style models, the comprehensive feedback scores of the style models will be generated multiple times. Each generated comprehensive feedback score will be compared with the preset score threshold. When the comprehensive feedback score of a certain style model continuously falls below the preset score threshold within a preset cycle, for example, the core model for 3 consecutive first iteration cycles and the exploration model for 2 consecutive second iteration cycles, it is determined that its technical compliance or user satisfaction is relatively low, and it will be removed from the performance model pool. Subsequently, a replacement style model is imported from a number of pre-trained style models to update the performance model pool and fill the vacancy after the model is eliminated.
[0071] It should be understood that through continuous screening and updating, inefficient models are eliminated and potential high-quality candidates are introduced. While maintaining the total amount of the performance model pool, the output music quality is continuously improved, so that multiple style models in the performance model pool always match the current music trends and user needs.
[0072] In some embodiments, the feedback scores of the style models are stored in a time-series database. At the end of the first iteration cycle and / or the second iteration cycle, the score aggregation calculation is triggered to obtain the comprehensive feedback scores of each core model and / or the comprehensive feedback scores of each exploration model, and the ranking list of several style models is updated according to the comprehensive feedback scores. It should be noted that if the first iteration cycle is a multiple of the second iteration cycle, there will be a situation where the first iteration cycle and the second iteration cycle end at the same time, and the comprehensive feedback scores of each core model and each exploration model within this cycle will be obtained simultaneously.
[0073] In some embodiments, the initial value of the preset score threshold can be flexibly set according to actual needs to ensure that the comprehensive feedback scores of the style models that have long existed in the performance model pool remain at a relatively high level, thereby ensuring that the music quality and user satisfaction output by the models remain at a relatively high average level.
[0074] In some embodiments, according to the preset expected elimination quantity and the ranking list of several style models, the expected elimination quantity of expected elimination models is obtained. The average value of the comprehensive feedback scores of the expected elimination models is compared with the preset score threshold. If the score difference between the average value and the preset score threshold is greater than the preset score difference, the preset score threshold is updated according to the highest score value of the comprehensive feedback scores of the expected elimination models. Further, in the next cycle, the preset score threshold is restored to the initial value.
[0075] For example, when the preset score threshold is 90 points, the preset score difference is 10 points, and the expected elimination quantity is 2, the two expected elimination models with the lowest comprehensive feedback scores are selected from the ranking list, and the comprehensive feedback scores are 60 points and 70 points respectively. At this time, the score difference between the average value and the preset score threshold is greater than the preset score difference, and 70 points is updated as the preset score threshold. Thus, in each iteration cycle, the adaptability evaluation of the preset score threshold is carried out, and the preset score threshold is dynamically optimized and adjusted according to the actual scenario to adapt to the score fluctuations in different cycles, and to avoid eliminating a large number of style models at one time and being unable to timely fill the vacancies in the performance model pool, thereby improving the accuracy and practicality of the model elimination and optimization mechanism.
[0076] In some embodiments, the model cache library is a dynamically managed model repository for storing style models that are not directly deployed to the performance model pool but have potential value, such as style models retrained based on user research, market trends, or social hotspots, continuously optimized versions of existing high-score style models, new style models generated by fusing two or more style models, etc., for quickly replenishing the style models removed from the performance model pool. Exemplarily, the model cache library is also set with quantity limits for the first newly added exploration model, the second newly added exploration model, and the newly added resident model. When the actual quantity is lower than the quantity limit, it will trigger the adaptive generation of the corresponding model or prompt the administrator to import the corresponding model from the outside. The specific quantity limit can be flexibly determined according to the frequency and quantity of the style models eliminated from the performance model pool in actual applications, so as to provide sufficient model supply for the performance model pool and avoid redundancy.
[0077] In some embodiments, obtain several high-score style models with the highest comprehensive feedback scores, and the historical target music data output by the high-score style models historically; based on the several high-score style models and the historical target music data, generate a new style model and store it in the model cache library.
[0078] Exemplarily, determine several high-score style models based on the comprehensive feedback score, and the historical target music data output by the high-score style models historically; based on at least two of the high-score style models with different style tags and the corresponding historical target music data, train to obtain the first newly added exploration model, and store the first newly added exploration model in the model cache library; and / or, based on at least two of the high-score style models with the same style tag and the corresponding historical target music data, train to obtain the second newly added exploration model, and store the second newly added exploration model in the model cache library.
[0079] Specifically, screen out the style models with the highest comprehensive feedback scores recently (for example, the top 10%) as high-score style models. At the same time, collect the historical target music data generated by these high-score style models (such as the target music data generated in the recent three months). It should be understood that the high-score style models obtained by screening may be any type of resident model, core model, or exploration model. Therefore, the basic model architecture or training data used in subsequent model training has high randomness and flexibility, ensuring the continuous advancement of model innovation and evolution.
[0080] Specifically, an innovative hybrid style model can be generated based on existing style models. At least two high-score style models with different styles and corresponding historical target music data are selected from the high-score style models, and a style model with a unique style, that is, the first newly added exploration model, is trained. At this time, a style label of "hybrid style" can be uniformly associated with such models; or the style labels of the source style models can be comprehensively considered, and multiple style labels, such as "pop style" and "Beethoven style", can be associated at the same time; or an intuitive and highly descriptive name can be reused or regenerated, such as "pop style - Beethoven style".
[0081] Specifically, for a certain style, an optimized and iterative style model is generated. At least two high-score style models with the same style and corresponding historical target music data are selected from the high-score style models, and a style model with optimized performance is trained, and the first style label is inherited to associate with the new style model, promoting the progress of the model performance while maintaining the core features of the original style.
[0082] In some embodiments, machine learning can also be used to automatically analyze the audio features of the music samples generated by the new style model, and based on these audio features, the most suitable style label can be automatically generated or recommended to ensure the consistency and accuracy of the label. Exemplarily, a sub-model dedicated to style classification is trained based on standard music samples under different style labels, and the most suitable style label is predicted by evaluating the new music samples.
[0083] It should be noted that the style label of the style model is not exactly the same as the name of the style model. The style label is for the standardized management and invocation of the style model, so there is overlap, while the name needs to be generated differently to enable users to quickly select the style model and corresponding music they are interested in.
[0084] It should be noted that for the generation of new style models, various model training techniques can be flexibly selected. For example, for the generation of iterative style models, the combination of knowledge distillation and data augmentation can be adopted, and the output results of two high-score style models are used as the "teacher signal" to guide the new model to learn more robust style features. Another example is that for the generation of hybrid style models, the parameter recombination method can be adopted, and some parameter modules of the network layers of two high-score style models (such as the rhythm processing unit and the timbre rendering layer) are cross-recombined to form a new architecture with both stability and innovation. The specific model training techniques used can be flexibly set according to the specific model architecture of the selected high-score style models and the current experimental direction, and are not specifically limited here.
[0085] In some embodiments, the method further includes: when the removed style model is a core model, based on the style label of the removed core model and the comprehensive feedback scores of each exploration model, selecting a target exploration model from the exploration models, and updating the target exploration model to a newly added core model.
[0086] Specifically, when the core model is removed, a targeted promotion mechanism is initiated. First, according to the style label of the removed core model, candidate objects with the same style label in the exploration models are screened, multiple comprehensive feedback scores of the candidate models in the recent period are extracted, the target exploration model with the highest average score is selected and promoted to a newly added core model. At this time, the target exploration model is given the attributes of a core model, and its iteration period is automatically switched to a long-period evaluation mode. For example, after the "mixed style" core model is eliminated, a newly added core model will be selected from the exploration models with the style label of "mixed style".
[0087] It should be understood that the newly promoted core model has been double-verified by users and technical scores during the exploration stage. Its style performance has proven the reliability of its model performance through short-term high-frequency iteration. Inheriting the style positioning of the original core model by the verified high-quality exploration model to fill the vacancy of the core model and maintain the stability of the output of the performance model pool. Thus, introducing untested new models that may cause quality fluctuations is avoided, and while maintaining the reliability of the core style, the performance model pool continuously absorbs innovative achievements, forming a virtuous evolutionary cycle.
[0088] In some embodiments, before importing the replacement style model from the model cache library, it further includes: according to the style label of the removed style model, counting the current total number of models with the corresponding style label in the performance model pool; if the current total number of models is greater than the preset number of models, using the first newly added exploration model as the replacement style model; if the current total number of models is less than the preset number of models, using the second newly added exploration model as the replacement style model.
[0089] Specifically, the preset number of models is the model base number that each style label needs to maintain, such as "at most 3 exploration models are reserved for each style" or "at least 1 core / exploration model is reserved", which is used to balance the style diversity and coverage rate of the performance model pool. The specific value can be flexibly set according to the actual scenario and will not be limited here.
[0090] When the exploration model is removed or promoted, a vacancy will occur in the exploration model, and the statistics of the existing quantity in the performance model pool of the style label to which the exploration model being removed or promoted belongs will be removed or promoted, that is, the current total number of models. If the current total number of models exceeds the preset number of models, at this time, the number of models of the corresponding style in the performance model pool is sufficient, then the first cross-style exploration model is imported from the model cache library to avoid over-concentration of the same style; if the current total number of models is lower than the preset number of models, at this time, the number of models of the corresponding style in the performance model pool is relatively vacant, then the second exploration model of the same style enhancement type is imported from the model cache library to avoid the loss of the output ability of a specific style due to model elimination.
[0091] It should be understood that by setting the preset number of models, on the one hand, the basic coverage rate of each style is ensured, and the redundancy or vacancy of various style models in the performance model pool is avoided. On the other hand, space for exploration is reserved for emerging styles, and homogenization is broken through cross-style models to maintain the diversity and innovation of styles in the performance model pool.
[0092] In some embodiments, the style models in the performance model pool are preset with model types. When the music style output by the style model is a widely accepted and popular style among the public, the model type of the style model is a professional model. When the music style output by the style model is a newly explored music style, the model type of the style model is an innovative model. For example, the first newly added exploration model is an innovative model, and the second newly added exploration model is a growth model.
[0093] In some embodiments, importing a replacement style model from the model cache library further includes: determining the model type to be imported from the removed or promoted model according to the model type of the removed style model, that is, the type of the exploration model being promoted or removed (i.e., an innovative model or a growth model). Exemplarily, when the exploration model is removed or promoted to a core model, when the exploration model being promoted or removed is an innovative model, the first newly added exploration model is used as the replacement style model; when the exploration model being promoted or removed is a growth model, the second newly added exploration model is used as the replacement style model.
[0094] In some embodiments, the style model further includes a resident model; the method further includes: statistically calculating the comprehensive feedback score of each resident model based on the third iteration cycle; when the style model removed from the performance model pool is a resident model, storing the removed resident model in the historical version library; importing a newly added resident model from the model cache library, and the newly added resident model is obtained by training the resident model in the historical version library based on the historical target music data.
[0095] Among them, the resident model is a special type of style model in the performance model pool, which is used to retain classic and representative performance expressions and avoid the loss of styles caused by algorithm iteration, such as "Beethoven style" and "Baroque style". Exemplarily, the style model includes a third preset number of resident models, and the third preset number is a fixed number set for the resident models, which is used to maintain specific classic or core styles in the long term to ensure that there are classic resident models in the performance model pool.
[0096] Specifically, the third iteration period can be an ultra-long evaluation window. When a resident model is removed due to insufficient scoring, its complete parameters and historical generation data are stored in the historical version library for version rollback or feature reuse when needed in the future. At the same time, a newly added resident model in the model cache library is introduced.
[0097] In some embodiments, historical target music data of high-score style models are selected from the current performance model pool to train the removed resident model, or the initial resident model in the historical version library, or the resident model with the highest average comprehensive feedback score in the historical version library, so as to ensure that the newly added resident model not only inherits the essence of the classic style but also integrates the latest optimization results and user preferences.
[0098] It should be understood that by using the resident model as the basis for training, the essential features of the classic style are ensured to be retained, and the style degradation caused by algorithm iteration is prevented. Then, modern high-score data is mixed and appropriate innovative elements are injected to ensure that the newly added resident model not only inherits the essence of the classic style but also integrates the latest optimization results and user preferences. Moreover, when training the newly added resident model, the original resident model or high-score resident model with more accurate classic style features can be called based on the historical version library, and the relatively long third iteration period can also avoid the frequent iteration of the resident model. Thus, the loss of styles caused by the rapid update and iteration of the resident model is avoided, and the progressive innovation of the resident model is achieved.
[0099] It should be noted that the resident model is a reproduction and inheritance of classical music styles, such as "Beethoven-style piano sonatas", "Baroque polyphonic style", "Chopin nocturne style". In specific application scenarios, the music generated by it is mostly used to provide users with reference, learning, and basic functional support for traditional music styles. Even if the user feedback scores are relatively low (such as low behavior feedback data like the like rate and collection rate), the existence of the resident model ensures the professionalism and comprehensiveness of the performance model pool. The core model can point to high-quality, high-usage, and multi-dimensional style models in specific application scenarios. Taking the core model with the style label "classical style" as an example, it can be a classical style model that caters to the public aesthetic, a classical style model that integrates the playing habits of multiple famous musicians, a classical style model for film scoring, or a classical style model that integrates a pedal control dimension in addition to the basic control dimensions such as pitch, time value, and dynamics. The exploration model can be an iterative optimized version of the classical style model for film scoring, or a highly interesting and innovative cyberpunk jazz style model, a randomly collaged jazz and classical style model.
[0100] In some embodiments, it includes: importing the first newly added exploration model and / or the second newly added exploration model from the model cache library, and using the imported first newly added exploration model and / or the second newly added exploration model as the new exploration model, thereby updating the performance model pool; importing the newly added resident model from the model cache library, and using the imported newly added resident model as the new resident model, thereby updating the performance model pool.
[0101] In some embodiments, the method further includes: comparing a plurality of the target music data, and when the similarity between any two of the target music data is higher than the first similarity threshold, the first historical music data and the second historical music data of the two corresponding style models; if the similarity between the first historical music data and the second historical music data is higher than the second similarity threshold, removing one of the style models from the performance model resource library.
[0102] Among them, the first similarity threshold is a set criterion for judging the similarity of music data, used to determine whether the real-time outputs of two target music data are highly the same; the first historical music data and the second historical music data respectively refer to the music data sets generated by the two corresponding style models in the past; the second similarity threshold is also a set criterion for judging the similarity of music data, used to determine whether the historical data between models shows a high degree of similarity for a long time, and it can be greater than, less than, or equal to the first similarity threshold. The specific values of the first similarity threshold and the second similarity threshold can be set according to the sensitivity of the actual similarity judgment, and are not limited here.
[0103] Specifically, when the similarity of the target music data generated by any two style models is detected to exceed the first similarity threshold, the historical music data of the two style models are further retrieved. The historical music data of the two style models are the first historical music data and the second historical music data respectively, and the historical similarity between them is calculated. If the historical similarity also exceeds the second similarity threshold, it indicates that their long-term style performances are highly convergent, and one of the style models will be randomly or removed according to the comprehensive feedback score to avoid redundancy in the performance model pool.
[0104] It should be understood that through the dual similarity verification of the immediate and historical aspects, long-term homogeneous models can be accurately identified to ensure the differentiation of the music styles output by the performance model pool, and to avoid misjudgment caused by single-generation coincidence. After eliminating similar models, the resource space is released and allocated to new exploration models, accelerating the attempt of emerging styles, promoting innovation, and improving the vitality of the performance model pool.
[0105] In some embodiments, the method further includes: detecting the growth trend of the comprehensive feedback scores of several core models; when the growth trends of the comprehensive feedback scores of the fourth preset number are abnormal, retraining the benchmark model.
[0106] Specifically, the growth trend refers to the change law of the comprehensive feedback score of the core model over time (such as continuous increase, stability or decrease), and the abnormal growth trend means that the comprehensive feedback score of the core model shows a significant decrease or fluctuation within a preset period, for example, the growth trend exceeds the preset interval range. When it is detected that such abnormalities occur simultaneously in the fourth preset number of core models (such as 8 or more), it is presumed that there may be a basic defect in the basic music data output by the benchmark model. Obtain the abnormal reasons for the growth trend of the comprehensive feedback score through channels such as user research, market trends or social hotspots, update the training data of the benchmark model based on the abnormal reasons, and / or adjust the model architecture of the benchmark model, and retrain the benchmark model.
[0107] Among them, the abnormal reasons may be that the benchmark model's ability to parse musical scores degenerates or the generation parameters are inaccurate, or the benchmark model cannot output the benchmark music data required for the current emerging music trends. For example, the basic music data output by the current experimental benchmark model does not contain data related to pedals, while the input performance score contains a large number of instructions for using sustain pedals or other pedals, and at this time the benchmark music data will be inaccurate.
[0108] Exemplarily, the retraining of the benchmark model can be performed using the original training data after updating the architecture, or using the updated training data for training (such as updating the data content or data accuracy included in the non-stylized data), or training the benchmark model with the updated training data without updating the architecture.
[0109] It should be understood that the basic music data output by the benchmark model is the basic input for all style models. Its defects may lead to global errors. Timely repair can avoid the collective failure of style models caused by problems at the basic layer. A stable benchmark model provides a reliable technical foundation for style models, enabling them to focus on style innovation rather than error correction. Further, since it is difficult to directly evaluate the music data output by the benchmark model from dimensions such as technical compliance and user interaction feedback, the update of the benchmark model is triggered through the linkage of the health of the core model to ensure the overall adaptability and long-term stability of the performance model pool.
[0110] In some embodiments, the growth trend of the comprehensive feedback scores of several core models is detected; when the growth trend of the comprehensive feedback scores of the fifth preset number of the same style labels is abnormal, it is presumed that it may be due to a decrease in the demand for the corresponding style or a basic defect in the target music data output by the style model. The reasons for the abnormal style of the growth trend of the comprehensive feedback scores are obtained through channels such as user research, market trends, or social hotspots. Based on the reasons for the abnormal style, the training data of the corresponding style model is updated, and / or the model architecture of the corresponding style model is adjusted, and the corresponding style model is retrained. Further, if the reason for the abnormal style is a decrease in the demand for the corresponding style, the preset model number of the corresponding style can be lowered to reduce the generation of style models of the corresponding style under the subsequent elimination mechanism, release resources, and achieve dynamic resource optimization.
[0111] Please refer to Figure 2 , Figure 2 is a schematic flowchart of another method for generating performance music provided by an embodiment of the present application. As Figure 2 shown, the music score to be performed is input into the benchmark model to convert the music score information into standardized basic music data at the technical level. Then, the basic music data generated by the benchmark model is input into multiple style models. The style models include three types: core models, exploration models, and resident models, which respectively convert the basic music data into target music data with specific styles from different style dimensions. It should be understood that the basic generation of music performance and the style transfer task are separated. The benchmark model is responsible for generating technical basic data to ensure the correctness of the underlying technology of music generation; the style models focus on style diversification and accurately simulate different styles through diverse models, avoiding the waste of computing power caused by repeated calculation of basic data, and only need to update the style models to introduce new music styles or optimize existing style expressions without retraining the benchmark model, improving flexibility and effectively saving computing power resources.
[0112] Please refer to Figure 3 , Figure 3 is a schematic diagram of the dynamic update of a performance model pool provided by an embodiment of the present application. As Figure 3As shown, the performance model pool consists of a benchmark model and a preset number of style models (including a first preset number of core models, a second preset number of exploration models, and a third preset number of resident models). Different types of style models in the performance model pool each undertake specific responsibilities: the core models are used to maintain a stable output of high-quality music in various styles; the exploration models are used to try new trends or integrate the features of existing high-score models to promote style innovation and performance improvement; the resident models are used for the long-term optimization of specific classic styles to prevent the loss of classic styles due to algorithm iteration. Through division of labor and cooperation, the performance model pool can not only cover a wide range of performance requirements but also continuously stimulate innovation possibilities while ensuring the basic experience.
[0113] As Figure 3 shown, the dotted lines represent the sources and uses of the training data during the dynamic update of the performance model pool, and can also be understood as the data flow of the data output by each model during the use of the performance model pool (including basic music data and target music data). The basic music data generated by the benchmark model are respectively input into the core models, exploration models, and resident models. These style models respectively output target music data. On the one hand, these target music data are used to output the music required by users, and on the other hand, they can also be used as training data for training high-score style models or resident models as Figure 3 shown, to obtain new exploration models (including a first new exploration model and a second new exploration model) and new resident models, while reducing the computing power consumption and ensuring the quality.
[0114] For the new exploration models, a number of existing target music data are used as training data to train the existing style models, integrate the features of high-score models, and obtain optimized and iterated style models or innovative hybrid style models, which are stored in the model cache library. For the resident models, a number of existing target music data are used as training data to train the existing resident models (the historical versions of the resident models can be called), integrate the features of high-score models, and obtain new resident models, which are stored in the model cache library, ensuring that the new resident models not only inherit the essence of classic styles but also integrate the latest optimization results and user preferences.
[0115] As Figure 3As shown, the solid line represents the dynamic update process of the models in the performance model pool. For the core models, after eliminating the outdated and inefficient core models, excellent new core models are selected from the exploration models to inherit the innovative achievements of the exploration models and ensure that the music data generated by each core model maintains high quality and high satisfaction. For the exploration models, when an exploration model is removed or promoted, a new exploration model is imported from the model cache library. For the resident models, when a resident model is removed, a new resident model is imported from the model cache library. Further, when the comprehensive feedback scores of a large number of core models are abnormal, the benchmark model is also triggered for update. Thus, through the adaptive elimination and evolution driven by the feedback scores, the high vitality of the performance model pool is maintained.
[0116] Furthermore, the entire performance model pool is designed with different iteration cycles. Since the core models are relatively reliable, a low-frequency iteration cycle can be used for evaluation and update to ensure that each exploration model promoted to a core model has stable high-quality scores in at least one long test period. On the contrary, the exploration models adopt a high-frequency iteration cycle to quickly try out errors and capture new trends to maintain the vitality of the performance model pool. The corresponding resident models also have specific iteration cycles to avoid the loss of style caused by the rapid update and iteration of the resident models and achieve the progressive innovation of the resident models.
[0117] It should be understood that the adaptive elimination and evolution mechanism is flexibly set based on the characteristics and functions of different types of style models (core models, exploration models, resident models). For example, the iteration cycles of different style models are different. While ensuring the music quality and stability, it continuously promotes model innovation, integrates the latest optimization results and user preferences. The sources of new models are different, which not only ensures the stability of classic styles and the coverage of various styles but also improves the exploration efficiency of emerging styles. These differentiated settings can form a benign synergy, adaptively maintaining the vitality of the performance model pool in the long term, avoiding the performance model pool losing competitiveness due to content homogenization and lag, and ultimately achieving the balance between the stability and innovation of the performance model pool.
[0118] Moreover, the iteration of the model is completed based on the high-score music data and high-score style models in the actual use process to learn emerging elements in real time and avoid relying on fixed training data and rigid model architectures. And the basic generation and style transfer tasks of music performance are decoupled and completed in two models. Style adjustment only requires retraining the style model, avoiding model iteration relying on manual intervention or full-scale retraining, greatly reducing the computational resource consumption, training cycle, and training cost, and thus timely responding to the user's demand for freshness.
[0119] In some embodiments, the first newly added exploration model, the second newly added exploration model, and the newly added resident model have all been pre-evaluated for technical eligibility. For example, if the first score is higher than the preset score, after ensuring that the quality meets the standard, they are stored in the model cache library to ensure that the model cache library continuously introduces style models with style innovation or performance optimization, providing diverse choices for subsequent deployment.
[0120] Please refer to Figure 4 , Figure 4 which is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. The computer device can be a terminal device or a server.
[0121] Exemplarily, the above method and device can be implemented in the form of a computer program, and the computer program can run on a computer device as shown in Figure 4 .
[0122] As shown in Figure 4 , the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a non-volatile storage medium and an internal memory.
[0123] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any method for generating music performance.
[0124] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0125] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any method for generating music performance.
[0126] The network interface is used for network communication, such as sending assigned tasks, etc.
[0127] It should be understood that the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor, or the processor can also be any conventional processor, etc.
[0128] Among them, in one embodiment, the processor is used to run a computer program stored in a memory to implement the following steps:
[0129] S101. Obtain a music score to be played and a performance model pool. The performance model pool includes a benchmark model and a preset number of style models, and each style model is associated with a style label. The style models include a core model and an exploration model;
[0130] S102. Input the music score to be played into the benchmark model to generate basic music data;
[0131] S103. Input the basic music data into several of the style models respectively to generate several target music data;
[0132] S104. Generate a feedback score for each style model according to the several target music data; statistically calculate the comprehensive feedback score of each core model based on a first iteration period; statistically calculate the comprehensive feedback score of each exploration model based on a second iteration period; wherein, the first iteration period is greater than the second iteration period;
[0133] S105. When the comprehensive feedback score of any style model continuously falls below a preset score threshold, remove the corresponding style model from the performance model pool, and import a replacement style model from a model cache library to update the performance model pool. The model cache library stores several pre-trained style models.
[0134] Exemplarily, the processor is used to run a computer program stored in a memory, and is also used to implement the steps of the method for generating performance music provided in any embodiment of the present application, which will not be elaborated here.
[0135] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. The processor executes the program instructions to implement the steps of the method for generating performance music provided in any one of the embodiments of the present application.
[0136] Among them, the computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.
[0137] As described above, it is only the specific implementation manner of the present application. However, the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for generating music performance, characterized in that, The method includes: S101. Obtain the music score to be played and a performance model pool. The performance model pool includes a benchmark model and a preset number of style models. Each style model is associated with a style label. The style models include a core model and an exploration model; S102. Input the music score to be played into the benchmark model to generate basic music data; S103. Input the basic music data into several of the style models respectively to generate several target music data; S104. Generate a feedback score for each style model based on the several target music data; statistically calculate the comprehensive feedback score of each core model based on the first iteration period; statistically calculate the comprehensive feedback score of each exploration model based on the second iteration period; wherein, the first iteration period is greater than the second iteration period; S105. When the comprehensive feedback score of any style model continuously drops below a preset score threshold, remove the corresponding style model from the performance model pool, and import a replacement style model from the model cache library to update the performance model pool. The model cache library stores several pre-trained style models.
2. The method according to claim 1, characterized in that, The feedback score includes a first score and / or a second score. The method includes: Scoring the target music data of each style model based on a preset scoring rule to obtain the first score of each style model; and / or, Obtain the user's behavioral feedback data, and generate the second score of each style model based on the behavioral feedback data.
3. The method according to claim 1, wherein The method further includes: Determine several high-score style models based on the comprehensive feedback score, and the historical target music data output by the high-score style models in history; Train a first newly added exploration model based on at least two of the high-score style models with different style labels and the corresponding historical target music data, and store the first newly added exploration model in the model cache library; and / or, Train a second newly added exploration model based on at least two of the high-score style models with the same style label and the corresponding historical target music data, and store the second newly added exploration model in the model cache library.
4. The method according to claim 3, wherein Before importing the replacement style model from the model cache library, it further includes: According to the style label of the removed style model, statistically calculate the current total number of models with the corresponding style label in the performance model pool; If the current total number of models is greater than the preset number of models, use the first newly added exploration model as the replacement style model; If the current total number of models is less than the preset number of models, use the second newly added exploration model as the replacement style model.
5. The method according to claim 3, characterized in that, The style model further includes a resident model. The method further includes: Statistically calculate the comprehensive feedback score of each resident model based on the third iteration period; When the style model removed from the performance model pool is a resident model, store the removed resident model in the historical version library; Import a newly added resident model from the model cache library. The newly added resident model is obtained by training the resident model in the historical version library based on the historical target music data.
6. The method according to claim 1, wherein The method further includes: When the removed style model is the core model, a target exploration model is selected from the exploration models based on the style tags of the removed core model and the comprehensive feedback scores of each exploration model, and the target exploration model is updated to a newly added core model.
7. The method according to claim 1, characterized in that, The method further includes: Comparing a plurality of the target music data, and when the similarity between any two of the target music data is higher than a first similarity threshold, the first historical music data and the second historical music data of the two corresponding style models respectively; If the similarity between the first historical music data and the second historical music data is higher than a second similarity threshold, one of the style models is removed from the performance model repository.
8. The method according to claim 1, wherein The method further includes: Detecting the growth trend of the comprehensive feedback scores of a plurality of core models; When the growth trends of the comprehensive feedback scores of a fourth preset number are abnormal, retraining the benchmark model.
9. A computer device, characterized in that, The device includes: A memory for storing a computer program; A processor for executing the computer program and implementing the method for generating performance music according to any one of claims 1 to 8 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the method for generating performance music according to any one of claims 1 to 8.
Citation Information
Patent Citations
MIDI playing style automatic conversion system based on recurrent neural network
CN111554255A
Music family style simulation automatic playing system and method
CN113936625A
Customized music generation method and device based on common semantic space
CN106898341A
Music style merging method based on coupled generative adversarial networks
CN110085203A