A training method, system and medium for a multimodal, multitasking large model
By integrating multimodal music data to train users' music preference models and conducting comprehensive evaluation, the problem of high efficiency and low training cost of multimodal multitasking models in the existing technology is solved, and efficient training and optimization of personalized music services is achieved.
Patent Information
- Application Number
- CN202510318958.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-03-18
AI Technical Summary
Most existing music recommendation systems and analysis methods rely on single mode or simple mode combinations, and cannot fully mine multimodal data, resulting in high training costs and low efficiency, and the inability of each model to share knowledge and features, making it difficult to train an efficient multimodal multitasking model under limited resources.
By integrating music data from multiple modalities for model training, including user basic information, operation records, playback history and audio feature data, training the user's music preference model, and obtaining model performance evaluation index through model universality and computing efficiency evaluation, providing targeted optimization strategies.
It improves the overall performance and practicality of the model, provides a personalized and high-quality music service experience, and reduces training costs and time.
Smart Images

Figure CN119848289B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of model training technology, and more specifically, to a training method, system, and medium for a multi-modal, multi-task large model. Background Art
[0002] With the booming music industry and widespread adoption of digital music platforms, users are increasingly demanding personalized music experiences. While music itself possesses a rich multimodal resource encompassing data such as audio waveforms, music spectra, lyrics, and diverse user behaviors, existing music recommendation systems and music analysis methods mostly rely on a single modality or a simple combination of modalities, failing to fully tap into and utilize the comprehensive information contained in multimodal data. Furthermore, in real-world applications, music-related tasks are often diverse and complex, such as music classification (by genre, decade, emotion, etc.), music recommendation, and music style transfer. Currently, training multiple independent models for these diverse tasks is often required. This not only results in high training costs and low efficiency, but also prevents the sharing of knowledge and features between models, hindering the full synergy of multimodal data across different tasks. Furthermore, with the increasing complexity of models and the continued expansion of data sizes, the computational and time costs of model training have also skyrocketed. In real-world applications, efficiently training high-performing multimodal, multi-task models within limited computational resources and time constraints has become a critical challenge. Summary of the Invention
[0003] The purpose of this application is to provide a training method, system and medium for a multimodal, multi-task large model. By integrating music data from multiple modalities for model training, and conducting a comprehensive analysis of the performance indicators of the trained model in multiple dimensions such as accuracy, computational efficiency, and generalization ability to new data, key feedback information is provided for model training optimization, thereby achieving targeted model training optimization, improving the overall performance and practicality of the model, and helping to provide users with a more personalized and high-quality music service experience.
[0004] This application also provides a method for training a multi-modal, multi-task large model, comprising the following steps:
[0005] Collect multimodal data related to music users, perform model training based on the multimodal data, and obtain a user music preference model;
[0006] Testing the user music preference model to obtain model test data, including model universality test data and model calculation efficiency data;
[0007] Obtaining a model universality evaluation index according to the model universality test data processing;
[0008] Obtaining a computational efficiency evaluation index according to computational efficiency data of the model;
[0009] The model performance evaluation index is obtained according to the model universality evaluation index and the computational efficiency evaluation index, and the corresponding model correction strategy data is obtained.
[0010] Optionally, in the multimodal multitask large model training method described in the present application, collecting multimodal data related to music users, performing model training based on the multimodal data, and obtaining a user music preference model include:
[0011] The multimodal data includes user basic information data, user operation record data, music playing history data, music basic attribute data and music audio feature data;
[0012] Model training is performed based on the user basic information data, user operation record data, music playback history data, music basic attribute data and music audio feature data to obtain a user music preference model.
[0013] Optionally, in the multimodal multitask large model training method described in the present application, the user music preference model is tested to obtain model test data, including model universality test data and model computational efficiency data, including:
[0014] The model universality test data includes model accuracy test data, cross-dataset learning accuracy data and task performance consistency data;
[0015] The model accuracy test data includes music preference prediction accuracy, music classification accuracy and music sentiment analysis accuracy;
[0016] The model calculation efficiency data includes memory occupancy, CPU usage and model inference time.
[0017] Optionally, in the multimodal multitask large model training method described in the present application, obtaining a model universality evaluation index based on the model universality test data processing includes:
[0018] Obtaining a model accuracy evaluation index based on the music preference prediction accuracy, music classification accuracy, and music sentiment analysis accuracy;
[0019] The model universality evaluation index is obtained based on the model accuracy evaluation index, cross-dataset learning accuracy data and task performance consistency data.
[0020] Optionally, in the multimodal multitask large model training method described in the present application, obtaining a computational efficiency evaluation index based on the model computational efficiency data processing includes:
[0021] The computational efficiency evaluation index is obtained based on the memory occupancy rate, CPU usage rate and model inference time.
[0022] Optionally, in the multimodal multi-task large model training method described in the present application, obtaining a model performance evaluation index based on the model universality evaluation index and the computational efficiency evaluation index, and obtaining corresponding model correction strategy data, includes:
[0023] Obtaining a model performance evaluation index according to the model universality evaluation index and the computational efficiency evaluation index;
[0024] Comparing the model performance evaluation index with a preset model performance evaluation index threshold, and determining the model performance evaluation level according to the range level to which the threshold comparison result belongs;
[0025] The model performance evaluation level is input into the preset model performance evaluation platform database for matching and identification to obtain corresponding model correction strategy data.
[0026] Optionally, the multimodal multitask large model training method described in this application further includes:
[0027] Recommending music to the user based on the user music preference model, and obtaining the click rate, skip rate, number of plays, and play time of the recommended music;
[0028] Obtaining a recommendation effect evaluation index based on the click rate, skip rate, number of plays, and play time;
[0029] The recommendation effect evaluation index is compared with a preset recommendation effect evaluation index threshold. If the threshold comparison result does not meet the preset threshold comparison result requirement, feedback of poor model recommendation effect is given.
[0030] In a second aspect, the present application provides a training system for a multimodal, multitask large model, the system comprising: a memory and a processor, wherein the memory stores a program for a training method for a multimodal, multitask large model, and when the program for the training method for the multimodal, multitask large model is executed by the processor, the following steps are implemented:
[0031] Collect multimodal data related to music users, perform model training based on the multimodal data, and obtain a user music preference model;
[0032] Testing the user music preference model to obtain model test data, including model universality test data and model calculation efficiency data;
[0033] Obtaining a model universality evaluation index according to the model universality test data processing;
[0034] Obtaining a computational efficiency evaluation index according to computational efficiency data of the model;
[0035] The model performance evaluation index is obtained according to the model universality evaluation index and the computational efficiency evaluation index, and the corresponding model correction strategy data is obtained.
[0036] Optionally, in the multimodal multitask large model training system described in the present application, the collecting of multimodal data related to music users, performing model training based on the multimodal data, and obtaining a user music preference model include:
[0037] The multimodal data includes user basic information data, user operation record data, music playing history data, music basic attribute data and music audio feature data;
[0038] Model training is performed based on the user basic information data, user operation record data, music playback history data, music basic attribute data and music audio feature data to obtain a user music preference model.
[0039] In a third aspect, the present application also provides a computer-readable storage medium, which stores a training method program for a multimodal multi-task large model. When the training method program for a multimodal multi-task large model is executed by a processor, the steps of the training method for a multimodal multi-task large model as described in any one of the above items are implemented.
[0040] From the above, it can be seen that the present application provides a multimodal, multi-task large model training method, system and medium, which integrates music data of multiple modalities for model training, and conducts a comprehensive analysis of the performance indicators of the trained model in multiple dimensions such as accuracy, computational efficiency and generalization ability of new data, providing key feedback information for model training optimization, thereby achieving targeted model training optimization, improving the overall performance and practicality of the model, and helping to provide users with a more personalized and high-quality music service experience.
[0041] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or understood by practicing the embodiments of the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures particularly pointed out in the written description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0043] Figure 1 A flowchart of a method for training a multi-modal, multi-task large model provided in an embodiment of the present application;
[0044] Figure 2 A flowchart of obtaining a user music preference model for a multimodal, multitask large model training method provided in an embodiment of the present application;
[0045] Figure 3 A flowchart of obtaining a model universality evaluation index for a training method of a multimodal, multitask large model provided in an embodiment of the present application;
[0046] Figure 4 A flowchart of obtaining corresponding model correction strategy data for the training method of the multimodal multi-task large model provided in an embodiment of the present application. DETAILED DESCRIPTION
[0047] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work fall within the scope of protection of the present application.
[0048] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0049] Please refer to Figure 1 , Figure 1 This is a flowchart of a method for training a multimodal, multitask large model in some embodiments of the present application. The multimodal, multitask large model training method is used in a terminal device, such as a computer, a mobile phone terminal, etc. The multimodal, multitask large model training method includes the following steps:
[0050] S11. Collect multimodal data related to music users, perform model training based on the multimodal data, and obtain a user music preference model;
[0051] S12. Testing the user music preference model to obtain model test data, including model universality test data and model calculation efficiency data;
[0052] S13, obtaining a model universality evaluation index according to the model universality test data;
[0053] S14, obtaining a computational efficiency evaluation index according to the computational efficiency data of the model;
[0054] S15. Obtain a model performance evaluation index according to the model universality evaluation index and the computational efficiency evaluation index, and obtain corresponding model correction strategy data.
[0055] It should be noted that this application integrates multiple modalities of music data such as user basic information data, user operation record data, music playback history data, music basic attribute data and music audio feature data to train the model, and evaluates the versatility and computational efficiency of the trained model to obtain a comprehensive evaluation result of the model performance, and finally obtains the corresponding model correction strategy data to achieve the purpose of improving model training accuracy and computational efficiency.
[0056] Please refer to Figure 2 , Figure 2 This is a flow chart of a method for training a multimodal, multitask large model in some embodiments of the present application to obtain a user music preference model. According to an embodiment of the present invention, collecting multimodal data related to music users, performing model training based on the multimodal data, and obtaining a user music preference model include:
[0057] S21, the multimodal data includes user basic information data, user operation record data, music playing history data, music basic attribute data and music audio feature data;
[0058] S22. Perform model training based on the user basic information data, user operation record data, music playback history data, music basic attribute data, and music audio feature data to obtain a user music preference model.
[0059] It should be noted that basic user information data includes age, gender, and region; user operation record data includes favorites, likes, shares, and comments; music playback history data includes song title data, play counts, and play duration; basic music attribute data includes song type data, artist information data, release date, lyrics content data, and music style label data; and music audio feature data includes rhythm feature data, melody feature data, and timbre feature data. During model training, emotional tendency data and topic keyword data are extracted from song title data and lyrics content data as text feature data for model training to obtain a user music preference model.
[0060] According to an embodiment of the present invention, the user music preference model is tested to obtain model test data, including model universality test data and model calculation efficiency data, including:
[0061] The model universality test data includes model accuracy test data, cross-dataset learning accuracy data and task performance consistency data;
[0062] The model accuracy test data includes music preference prediction accuracy, music classification accuracy and music sentiment analysis accuracy;
[0063] The model calculation efficiency data includes memory occupancy, CPU usage and model inference time.
[0064] It should be noted that the cross-dataset learning accuracy represents the accuracy of the model for different data sets, which can be expressed by the ratio of the number of correctly predicted sample sets to the total number of test sample sets. The model task performance consistency data is a quantitative data used to measure the degree of performance balance of the model in the three tasks of music preference prediction, music classification and music sentiment analysis, which can be expressed by calculating the variance of the music preference prediction accuracy, music classification accuracy and music sentiment analysis accuracy.
[0065] Please refer to Figure 3 , Figure 3 The flowchart of the method for training a multimodal, multitask large model in some embodiments of the present application is to obtain a model universality evaluation index. According to an embodiment of the present invention, the method of obtaining the model universality evaluation index based on the model universality test data includes:
[0066] S31, obtaining a model accuracy evaluation index according to the music preference prediction accuracy, music classification accuracy, and music sentiment analysis accuracy;
[0067] S32. Obtain a model universality evaluation index based on the model accuracy evaluation index, cross-dataset learning accuracy data, and task performance consistency data.
[0068] It should be noted that the model accuracy evaluation index is calculated based on the accuracy of the model in performing the three tasks of music preference prediction, music classification, and music sentiment analysis;
[0069] The calculation formula of the model accuracy evaluation index is:
[0070] ;
[0071] in, is the model accuracy evaluation index, 、 and They are respectively the music preference prediction accuracy, music classification accuracy and music sentiment analysis accuracy, 、 and is the preset characteristic coefficient (which can be obtained through the preset model performance evaluation platform);
[0072] The calculation formula of the model universality evaluation index is:
[0073] ;
[0074] in, is the model universality evaluation index, is the model accuracy evaluation index, and They are cross-dataset learning accuracy data and task performance consistency data, and is the preset characteristic coefficient (which can be obtained through the preset model performance evaluation platform).
[0075] According to an embodiment of the present invention, the step of obtaining a computing efficiency evaluation index based on the model computing efficiency data processing includes:
[0076] The computational efficiency evaluation index is obtained based on the memory occupancy rate, CPU usage rate and model inference time.
[0077] It should be noted that the calculation formula of the computational efficiency evaluation index is:
[0078] ;
[0079] in, is the computational efficiency evaluation index, 、 and They are memory usage, CPU usage and model inference time, 、 and is the preset characteristic coefficient (which can be obtained through the preset model performance evaluation platform).
[0080] Please refer to Figure 4 , Figure 4 This is a flowchart of obtaining corresponding model correction strategy data for a multimodal, multitask large model training method in some embodiments of the present application. According to an embodiment of the present invention, the processing of obtaining a model performance evaluation index based on the model universality evaluation index and the computational efficiency evaluation index, and obtaining corresponding model correction strategy data, includes:
[0081] S41, obtaining a model performance evaluation index according to the model universality evaluation index and the computational efficiency evaluation index;
[0082] S42, comparing the model performance evaluation index with a preset model performance evaluation index threshold, and determining the model performance evaluation level according to the range level to which the threshold comparison result belongs;
[0083] S43: Input the model performance evaluation level into a preset model performance evaluation platform database for matching and identification to obtain corresponding model correction strategy data.
[0084] It should be noted that the calculation formula of the model performance evaluation index is:
[0085] ;
[0086] in, is the model performance evaluation index, is the model universality evaluation index, is the computational efficiency evaluation index, is the preset characteristic coefficient (which can be obtained through the preset model performance evaluation platform).
[0087] According to an embodiment of the present invention, the further embodiment includes:
[0088] Recommending music to the user based on the user music preference model, and obtaining the click rate, skip rate, number of plays, and play time of the recommended music;
[0089] Obtaining a recommendation effect evaluation index based on the click rate, skip rate, number of plays, and play time;
[0090] The recommendation effect evaluation index is compared with a preset recommendation effect evaluation index threshold. If the threshold comparison result does not meet the preset threshold comparison result requirement, feedback of poor model recommendation effect is given.
[0091] It should be noted that music recommendations are made to users based on their music preference model, and the recommendation effects are evaluated;
[0092] The calculation formula of the recommendation effect evaluation index is:
[0093] ;
[0094] in, is the recommendation effect evaluation index, 、 、 and They are click rate, skip rate, number of plays and play time. 、 、 and is the preset characteristic coefficient (which can be obtained through the preset model performance evaluation platform).
[0095] The present invention also discloses a multimodal multitask large model training system, comprising a memory and a processor, wherein the memory stores a multimodal multitask large model training method program, and when the multimodal multitask large model training method program is executed by the processor, the following steps are implemented:
[0096] Collect multimodal data related to music users, perform model training based on the multimodal data, and obtain a user music preference model;
[0097] Testing the user music preference model to obtain model test data, including model universality test data and model calculation efficiency data;
[0098] Obtaining a model universality evaluation index according to the model universality test data processing;
[0099] Obtaining a computational efficiency evaluation index based on computational efficiency data of the model;
[0100] The model performance evaluation index is obtained according to the model universality evaluation index and the computational efficiency evaluation index, and the corresponding model correction strategy data is obtained.
[0101] It should be noted that this application integrates multiple modalities of music data such as user basic information data, user operation record data, music playback history data, music basic attribute data and music audio feature data to train the model, and evaluates the versatility and computational efficiency of the trained model to obtain a comprehensive evaluation result of the model performance, and finally obtains the corresponding model correction strategy data to achieve the purpose of improving model training accuracy and computational efficiency.
[0102] According to an embodiment of the present invention, collecting multimodal data related to music users, performing model training based on the multimodal data, and obtaining a user music preference model includes:
[0103] The multimodal data includes user basic information data, user operation record data, music playing history data, music basic attribute data and music audio feature data;
[0104] Model training is performed based on the user basic information data, user operation record data, music playback history data, music basic attribute data and music audio feature data to obtain a user music preference model.
[0105] It should be noted that basic user information data includes age, gender, and region; user operation record data includes favorites, likes, shares, and comments; music playback history data includes song title data, play counts, and play duration; basic music attribute data includes song type data, artist information data, release date, lyrics content data, and music style label data; and music audio feature data includes rhythm feature data, melody feature data, and timbre feature data. During model training, emotional tendency data and topic keyword data are extracted from song title data and lyrics content data as text feature data for model training to obtain a user music preference model.
[0106] According to an embodiment of the present invention, the user music preference model is tested to obtain model test data, including model universality test data and model calculation efficiency data, including:
[0107] The model universality test data includes model accuracy test data, cross-dataset learning accuracy data and task performance consistency data;
[0108] The model accuracy test data includes music preference prediction accuracy, music classification accuracy and music sentiment analysis accuracy;
[0109] The model calculation efficiency data includes memory occupancy, CPU usage and model inference time.
[0110] It should be noted that the cross-dataset learning accuracy represents the accuracy of the model for different data sets, which can be expressed by the ratio of the number of correctly predicted sample sets to the total number of test sample sets. The model task performance consistency data is a quantitative data used to measure the degree of performance balance of the model in the three tasks of music preference prediction, music classification and music sentiment analysis, which can be expressed by calculating the variance of the music preference prediction accuracy, music classification accuracy and music sentiment analysis accuracy.
[0111] According to an embodiment of the present invention, the step of obtaining a model universality evaluation index based on the model universality test data includes:
[0112] Obtaining a model accuracy evaluation index based on the music preference prediction accuracy, music classification accuracy, and music sentiment analysis accuracy;
[0113] The model universality evaluation index is obtained based on the model accuracy evaluation index, cross-dataset learning accuracy data and task performance consistency data.
[0114] It should be noted that the model accuracy evaluation index is calculated based on the accuracy of the model in performing the three tasks of music preference prediction, music classification, and music sentiment analysis;
[0115] The calculation formula of the model accuracy evaluation index is:
[0116] ;
[0117] in, is the model accuracy evaluation index, 、 and They are respectively the music preference prediction accuracy, music classification accuracy and music sentiment analysis accuracy, 、 and is the preset characteristic coefficient (which can be obtained through the preset model performance evaluation platform);
[0118] The calculation formula of the model universality evaluation index is:
[0119] ;
[0120] in, is the model universality evaluation index, is the model accuracy evaluation index, and They are cross-dataset learning accuracy data and task performance consistency data, and is the preset characteristic coefficient (which can be obtained through the preset model performance evaluation platform).
[0121] According to an embodiment of the present invention, the step of obtaining a computing efficiency evaluation index based on the model computing efficiency data processing includes:
[0122] The computational efficiency evaluation index is obtained based on the memory occupancy rate, CPU usage rate and model inference time.
[0123] It should be noted that the calculation formula of the computational efficiency evaluation index is:
[0124] ;
[0125] in, is the computational efficiency evaluation index, 、 and They are memory usage, CPU usage and model inference time, 、 and is the preset characteristic coefficient (which can be obtained through the preset model performance evaluation platform).
[0126] According to an embodiment of the present invention, the processing of obtaining a model performance evaluation index based on the model universality evaluation index and the computational efficiency evaluation index, and obtaining corresponding model correction strategy data, includes:
[0127] Obtaining a model performance evaluation index according to the model universality evaluation index and the computational efficiency evaluation index;
[0128] Comparing the model performance evaluation index with a preset model performance evaluation index threshold, and determining the model performance evaluation level according to the range level to which the threshold comparison result belongs;
[0129] The model performance evaluation level is input into the preset model performance evaluation platform database for matching and identification to obtain corresponding model correction strategy data.
[0130] It should be noted that the calculation formula of the model performance evaluation index is:
[0131] ;
[0132] in, is the model performance evaluation index, is the model universality evaluation index, is the computational efficiency evaluation index, is the preset characteristic coefficient (which can be obtained through the preset model performance evaluation platform).
[0133] According to an embodiment of the present invention, the further embodiment includes:
[0134] Recommending music to the user based on the user music preference model, and obtaining the click rate, skip rate, number of plays, and play time of the recommended music;
[0135] Obtaining a recommendation effect evaluation index based on the click rate, skip rate, number of plays, and play time;
[0136] The recommendation effect evaluation index is compared with a preset recommendation effect evaluation index threshold. If the threshold comparison result does not meet the preset threshold comparison result requirement, feedback of poor model recommendation effect is given.
[0137] It should be noted that music recommendations are made to users based on their music preference model, and the recommendation effects are evaluated;
[0138] The calculation formula of the recommendation effect evaluation index is:
[0139] ;
[0140] in, is the recommendation effect evaluation index, 、 、 and They are click rate, skip rate, number of plays and play time. 、 、 and is the preset characteristic coefficient (which can be obtained through the preset model performance evaluation platform).
[0141] The third aspect of the present invention provides a readable storage medium, which stores a training method program for a multimodal multi-task large model. When the training method program for a multimodal multi-task large model is executed by a processor, the steps of the training method for a multimodal multi-task large model as described in any one of the above items are implemented.
[0142] The present invention discloses a multimodal, multi-task large model training method, system, and medium. The method integrates music data from multiple modalities for model training, and conducts a comprehensive analysis of the performance indicators of the trained model in multiple dimensions, such as accuracy, computational efficiency, and generalization ability for new data. The method provides key feedback information for model training optimization, thereby achieving targeted model training optimization, improving the overall performance and practicality of the model, and helping to provide users with a more personalized, high-quality music service experience.
[0143] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0144] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0145] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0146] Those skilled in the art will appreciate that all or part of the steps of the above-mentioned method embodiments may be implemented by hardware related to program instructions, and the aforementioned program may be stored in a readable storage medium. When the program is executed, the program executes the steps of the above-mentioned method embodiments. The aforementioned storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0147] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as standalone products, they can also be stored on a readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This software product, stored on a storage medium, includes instructions for enabling a computer device (such as a personal computer, server, or network device) to execute all or part of the methods described in the various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as removable storage devices, ROM, RAM, magnetic disks, or optical disks.
Claims
1. A training method for a multi-modal, multi-task large model, characterized in that: The following steps are involved: Collect multimodal data related to music users, perform model training based on the multimodal data, and obtain a user music preference model; The multimodal data includes user basic information data, user operation record data, music playing history data, music basic attribute data and music audio feature data; User basic information data includes age data, gender data and regional data; User operation record data includes collection data, like data, sharing data, and comment data; Music playback history data includes song name data, number of plays, and playback duration; Music basic attribute data includes song type data, singer information data, release time, lyrics content data and music style label data; The music audio feature data includes rhythm feature data, melody feature data and timbre feature data; Performing model training based on the user basic information data, user operation record data, music playback history data, music basic attribute data, and music audio feature data to obtain a user music preference model; In the model training process, emotional tendency data and theme keyword data are extracted from song name data and lyrics content data as text feature data to train the model and obtain the user music preference model; Testing the user music preference model to obtain model test data, including model universality test data and model calculation efficiency data; The model universality test data includes model accuracy test data, cross-dataset learning accuracy data and task performance consistency data; The cross-dataset learning accuracy indicates the accuracy of the model for different data sets, expressed as the ratio of the number of correctly predicted sample sets to the total number of test sample sets; The model task performance consistency data is used to quantitatively measure the degree of performance balance of the model in the three tasks of music preference prediction, music classification, and music sentiment analysis. It is represented by calculating the variance of the music preference prediction accuracy, music classification accuracy, and music sentiment analysis accuracy. The model accuracy test data includes music preference prediction accuracy, music classification accuracy and music sentiment analysis accuracy; The model calculation efficiency data includes memory usage, CPU usage and model inference time; Obtaining a model universality evaluation index based on the model universality test data, including obtaining a model accuracy evaluation index based on the music preference prediction accuracy, the music classification accuracy, and the music sentiment analysis accuracy; Obtaining a model universality evaluation index based on the model accuracy evaluation index, cross-dataset learning accuracy data, and task performance consistency data; Obtaining a computational efficiency evaluation index according to computational efficiency data of the model; The model performance evaluation index is obtained according to the model universality evaluation index and the computational efficiency evaluation index, and the corresponding model correction strategy data is obtained.
2. The method for training a multimodal multitask large model according to claim 1, characterized in that: The step of obtaining a computational efficiency evaluation index based on the computational efficiency data of the model includes: The computational efficiency evaluation index is obtained based on the memory occupancy rate, CPU usage rate and model inference time.
3. The multi-modal multi-task large model training method according to claim 2, characterized in that: The step of obtaining a model performance evaluation index based on the model universality evaluation index and the computational efficiency evaluation index, and obtaining corresponding model correction strategy data, includes: Obtaining a model performance evaluation index according to the model universality evaluation index and the computational efficiency evaluation index; Comparing the model performance evaluation index with a preset model performance evaluation index threshold, and determining the model performance evaluation level according to the range level to which the threshold comparison result belongs; The model performance evaluation level is input into the preset model performance evaluation platform database for matching and identification to obtain corresponding model correction strategy data.
4. The method for training a multi-modal multi-task large model according to claim 3, characterized in that: Also includes: Recommending music to the user based on the user music preference model, and obtaining the click rate, skip rate, number of plays, and play time of the recommended music; Obtaining a recommendation effect evaluation index based on the click rate, skip rate, number of plays, and play time; The recommendation effect evaluation index is compared with a preset recommendation effect evaluation index threshold. If the threshold comparison result does not meet the preset threshold comparison result requirement, feedback of poor model recommendation effect is given.
5. A multi-modal, multi-task large model training system, characterized by: The system includes a memory and a processor, wherein the memory stores a program of a training method for a multimodal multitask large model, and when the program of the training method for the multimodal multitask large model is executed by the processor, the following steps are implemented: Collect multimodal data related to music users, perform model training based on the multimodal data, and obtain a user music preference model; The multimodal data includes user basic information data, user operation record data, music playing history data, music basic attribute data and music audio feature data; User basic information data includes age data, gender data and regional data; User operation record data includes collection data, like data, sharing data, and comment data; Music playback history data includes song name data, number of plays, and playback duration; Music basic attribute data includes song type data, singer information data, release time, lyrics content data and music style label data; The music audio feature data includes rhythm feature data, melody feature data and timbre feature data; Performing model training based on the user basic information data, user operation record data, music playback history data, music basic attribute data, and music audio feature data to obtain a user music preference model; In the model training process, emotional tendency data and theme keyword data are extracted from song name data and lyrics content data as text feature data to train the model and obtain the user music preference model; Testing the user music preference model to obtain model test data, including model universality test data and model calculation efficiency data; The model universality test data includes model accuracy test data, cross-dataset learning accuracy data and task performance consistency data; The cross-dataset learning accuracy indicates the accuracy of the model for different data sets, expressed as the ratio of the number of correctly predicted sample sets to the total number of test sample sets; The model task performance consistency data is used to quantitatively measure the degree of performance balance of the model in the three tasks of music preference prediction, music classification, and music sentiment analysis. It is represented by calculating the variance of the music preference prediction accuracy, music classification accuracy, and music sentiment analysis accuracy. The model accuracy test data includes music preference prediction accuracy, music classification accuracy and music sentiment analysis accuracy; The model calculation efficiency data includes memory usage, CPU usage and model inference time; Obtaining a model universality evaluation index based on the model universality test data, including obtaining a model accuracy evaluation index based on the music preference prediction accuracy, the music classification accuracy, and the music sentiment analysis accuracy; Obtaining a model universality evaluation index based on the model accuracy evaluation index, cross-dataset learning accuracy data, and task performance consistency data; Obtaining a computational efficiency evaluation index based on computational efficiency data of the model; The model performance evaluation index is obtained according to the model universality evaluation index and the computational efficiency evaluation index, and the corresponding model correction strategy data is obtained.
6. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a training program for a multimodal, multitask large model. When the training program for the multimodal, multitask large model is executed by a processor, the steps of the training method for a multimodal, multitask large model as described in any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Intelligent analysis method and system based on big data
CN118132856A
Large language model evaluation system and method
CN119558412A