Training method of song information auditing model and song information auditing method
By screening and training a benchmark large language model, a second review result that meets the preset conditions is generated, which solves the reliability problem of the song information review model and achieves a more accurate and interpretable review effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
- Filing Date
- 2025-12-01
- Publication Date
- 2026-04-21
AI Technical Summary
Existing BERT-based song information review models suffer from difficulties in interpreting the decision-making process, leading to low reliability in the review process.
Multiple first review results are generated by pre-training a benchmark large language model. Second review results that meet the preset conditions are selected. The benchmark large language model is then trained based on the similarity between the second review results and the sample review results to generate a target song information review model.
This improves the reliability of song information review, avoids the shortcomings of the "black box" model's decision-making process that is difficult to explain, and achieves a more accurate review effect.
Smart Images

Figure CN121902897A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of text review technology, and in particular to a training method for a song information review model, a song information review method, an apparatus, a computer device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] With the booming development of the internet content ecosystem, text-based content has become a core carrier of user production and consumption, covering a wide range of topics, such as song lyrics and user comments on music platforms. During the dissemination of text-based content, inappropriate information may be included; therefore, timely and accurate review of text-based content is crucial.
[0003] In related technologies, discriminative classification methods based on BERT (Bidirectional Encoder Representations from Transformers) are commonly used to determine whether song information contains illegal information. However, BERT is a "black box" model, and its decision-making process is difficult to explain. Song information review needs to follow certain review standards and criteria, which leads to low reliability of song information review. Summary of the Invention
[0004] Therefore, it is necessary to address the aforementioned technical problem of low reliability in song information verification by providing a training method, song information verification method, apparatus, computer equipment, computer-readable storage medium, and computer program product for a song information verification model that can improve the reliability of song information verification.
[0005] Firstly, this application provides a training method for a song information verification model, including:
[0006] Acquire sample data; the sample data includes sample song information and the sample review results of the sample song information; the sample review results include review results under multiple preset dimensions;
[0007] The sample song information is reviewed using a pre-trained benchmark large language model, generating multiple first review results for the sample song information; each first review result includes multiple review results under the preset dimensions.
[0008] Based on the sample review results, the first review result that satisfies the preset difference condition with the sample review results among the first review results is determined as the second review result;
[0009] Based on the similarity between the second review result and the sample review result under each preset dimension, the benchmark large language model is trained to obtain the target song information review model.
[0010] Secondly, this application also provides a method for verifying song information, including:
[0011] Obtain song information for songs awaiting review;
[0012] The song information is input into the target song information review model to obtain the review result of the song information; the target song information review model is trained by the training method of the song information review model in any embodiment;
[0013] Display the audit results.
[0014] Thirdly, this application also provides a training device for a song information verification model, comprising:
[0015] The sample data acquisition module is used to acquire sample data; the sample data includes sample song information and the sample review results of the sample song information; the sample review results include review results under multiple preset dimensions;
[0016] The review result generation module is used to review the sample song information using a pre-trained benchmark large language model and generate multiple first review results for the sample song information; each first review result includes multiple review results under the preset dimensions;
[0017] The audit result filtering module is used to determine, based on the sample audit results, a first audit result whose difference from the sample audit results meets a preset difference condition and is used as a second audit result;
[0018] The review model training module is used to train the benchmark large language model based on the similarity between the second review result and the sample review result under each preset dimension to obtain the target song information review model.
[0019] Fourthly, this application also provides a song information verification device, including:
[0020] The song information acquisition module is used to acquire song information for songs pending review.
[0021] The song information review module is used to input the song information into the target song information review model to obtain the review result of the song information; the target song information review model is trained by the training method of the song information review model in any embodiment;
[0022] The audit results display module is used to display the audit results.
[0023] Fifthly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0024] Acquire sample data; the sample data includes sample song information and the sample review results of the sample song information; the sample review results include review results under multiple preset dimensions;
[0025] The sample song information is reviewed using a pre-trained benchmark large language model, generating multiple first review results for the sample song information; each first review result includes multiple review results under the preset dimensions.
[0026] Based on the sample review results, the first review result that satisfies the preset difference condition with the sample review results among the first review results is determined as the second review result;
[0027] Based on the similarity between the second review result and the sample review result under each preset dimension, the benchmark large language model is trained to obtain the target song information review model.
[0028] Sixthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0029] Acquire sample data; the sample data includes sample song information and the sample review results of the sample song information; the sample review results include review results under multiple preset dimensions;
[0030] The sample song information is reviewed using a pre-trained benchmark large language model, generating multiple first review results for the sample song information; each first review result includes multiple review results under the preset dimensions.
[0031] Based on the sample review results, the first review result that satisfies the preset difference condition with the sample review results among the first review results is determined as the second review result;
[0032] Based on the similarity between the second review result and the sample review result under each preset dimension, the benchmark large language model is trained to obtain the target song information review model.
[0033] In a seventh aspect, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0034] Acquire sample data; the sample data includes sample song information and the sample review results of the sample song information; the sample review results include review results under multiple preset dimensions;
[0035] The sample song information is reviewed using a pre-trained benchmark large language model, generating multiple first review results for the sample song information; each first review result includes multiple review results under the preset dimensions.
[0036] Based on the sample review results, the first review result that satisfies the preset difference condition with the sample review results among the first review results is determined as the second review result;
[0037] Based on the similarity between the second review result and the sample review result under each preset dimension, the benchmark large language model is trained to obtain the target song information review model.
[0038] The training method, apparatus, computer equipment, computer-readable storage medium, and computer program product of the aforementioned song information review model, through a pre-trained benchmark large language model, can generate multiple first review results for sample song information. Using these sample review results, second review results that meet preset difference conditions can be selected from the multiple first review results. Based on the similarity between the second review results and the sample review results in each preset dimension, the benchmark large language model can be trained. On the one hand, using a large language model as the song information review model avoids the difficulty in interpreting the decision-making process of "black box" models because the large language model can standardize the task execution process and output results based on prompt word engineering. On the other hand, training the benchmark large language model based on the similarity between the second review results that meet preset difference conditions and the sample review results enables it to perform song information review tasks more accurately. In summary, the training method for the song information review model based on the above process improves the reliability of song information review. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 This is a flowchart illustrating the training method of a song information verification model in one embodiment;
[0041] Figure 2 This is a flowchart illustrating a song information verification method in one embodiment;
[0042] Figure 3 This is a flowchart illustrating a text moderation method based on large-model reinforcement learning in one embodiment.
[0043] Figure 4 This is a schematic diagram illustrating the steps of data acquisition and fine-grained labeled dataset construction in one embodiment;
[0044] Figure 5 This is a structural block diagram of a training device for a song information verification model in one embodiment;
[0045] Figure 6 This is a structural block diagram of a song information verification device in one embodiment;
[0046] Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0048] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0049] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0050] In one exemplary embodiment, such as Figure 1 As shown, a training method for a song information review model is provided. This embodiment illustrates the application of this method to a server. It is understood that this method can also be applied to a terminal, or to a system including both a server and a terminal, and is implemented through interaction between the server and the terminal. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. In this embodiment, the method includes the following steps:
[0051] Step S102: Obtain sample data; the sample data includes sample song information and the sample review results of the sample song information.
[0052] The sample data consists of multiple records, with each record corresponding to a single sample song.
[0053] If each sample data corresponds to a sample song, then the sample song information in the sample data is the song information of the corresponding sample song. The song information includes the song lyrics, and further includes at least one of the following: song title, artist name, album name, album description, and song comments.
[0054] The sample review results include review results under multiple preset dimensions; the multiple preset dimensions include at least a baseline dimension and an extended dimension. The baseline dimension includes at least one of the anomaly category and the anomaly level, and the extended dimension includes at least one of the content summary, anomaly song information fragments, and anomaly reasons.
[0055] Among them, the song information is in text form; the anomaly category describes the category of the abnormal content in the song information; the anomaly level describes the level of the abnormal content in the song information; the content summary is a summary of the song information; the abnormal song information fragment is a fragment of the original text in the song information containing abnormal content; and the anomaly reason describes the reason why the song information is determined to contain abnormal content.
[0056] Specifically, the server obtains sample song information from multiple sample songs and determines the sample review result for each sample song information; then, for each sample song information, the server uses the sample review result of that sample song information as the tag information for that sample song information to construct a corresponding sample data.
[0057] Step S104: The sample song information is reviewed using the pre-trained benchmark large language model, generating multiple first review results for the sample song information.
[0058] The baseline large language model can be an open-source general-purpose large language model, or it can be a large language model obtained by performing domain-specific fine-tuning on an open-source general-purpose large language model.
[0059] Each first review result includes review results under multiple preset dimensions.
[0060] Specifically, the server reviews the information of each sample song using a benchmark large language model, and obtains multiple first review results for each sample song information.
[0061] In practical applications, users can change the prompt words for the benchmark large language model to make it output multiple first review results for a single sample song; users can also directly command the benchmark large language model to output multiple first review results for a single sample song by outputting a request.
[0062] In practical applications, the server selects at least one first sample data from each sample data set and inputs it into the benchmark large language model to obtain multiple corresponding first review results.
[0063] Step S106: Based on the sample audit results, determine the first audit result among the first audit results whose difference from the sample audit results meets the preset difference conditions as the second audit result.
[0064] The second review result is the first review result whose difference from the sample review result is within a preset range. Further, the second review result is the first review result where the review result under multiple specific dimensions is neither entirely the same as nor entirely different from the sample review result. In specific applications, the baseline dimension includes at least one of anomaly category and anomaly level, and the extended dimensions include at least one of content summary, anomaly song information fragment, and anomaly cause. The baseline dimension can more intuitively express whether the song information is abnormal compared to the extended dimensions, and the baseline dimension is easier to compare whether they are completely the same or completely different. Therefore, preferably, the specific dimension is the baseline dimension.
[0065] Specifically, the server selects the first audit results from each first audit result that are the same as the sample audit results in each benchmark dimension, as positive audit results; and selects the first audit results from each first audit result that are different from the sample audit results in each benchmark dimension, as negative audit results; then, the first audit results other than the positive and negative audit results are determined as second audit results.
[0066] Step S108: Based on the similarity between the second review result and the sample review result in each preset dimension, the benchmark large language model is trained to obtain the target song information review model.
[0067] Specifically, the server determines the similarity between each second review result and the sample review result under each preset dimension, and inputs the similarity into the preset reward function to calculate the reward function value of the benchmark large language model under the reward function. Then, the server iteratively strengthens the benchmark large language model based on the reward function value of the benchmark large language model until the preset strengthening training termination condition is met, and uses the benchmark large language model at the preset strengthening training termination condition as the target song information review model.
[0068] In practical applications, the preset termination condition for reinforcement training is that the similarity between the review results of each second review result and the sample review result in each preset dimension meets the preset similarity condition, or the number of iterations of reinforcement training reaches the preset number of iterations.
[0069] In the training method of the aforementioned song information review model, the server can generate multiple first review results for sample song information using a pre-trained benchmark large language model. Based on these sample review results, second review results whose differences from the sample review results are within a preset range can be selected from the multiple first review results. The benchmark large language model can be trained based on the similarity between the second review results and the sample review results across preset dimensions. On one hand, using a large language model as the song information review model avoids the difficulty in interpreting the decision-making process of "black box" models because the large language model can standardize the task execution process and output results based on prompt word engineering. On the other hand, training the benchmark large language model based on the similarity between the second review results and the sample review results, which satisfy preset difference conditions, enables it to perform the song information review task more accurately. In summary, the training method of the song information review model based on the above process improves the reliability of song information review.
[0070] In an exemplary embodiment, the preset dimensions include a baseline dimension and an extended dimension. The baseline dimension includes at least one of anomaly category and anomaly level, and the extended dimension includes at least one of content summary, anomaly song information fragment, and anomaly cause.
[0071] In this embodiment, step S102, obtaining sample data, specifically includes the following steps: obtaining sample song information and the manual review results of the sample song information; using the initial review prompt information and the sample song information as input to the review result optimization model, and using the manual review results of the sample song information as the supervisory information for the output of the review result optimization model, to optimize the initial review prompt information multiple times until a preset optimization stop condition is reached to obtain optimized review prompt information; using the model review result of the sample song information under the optimized review prompt information as the sample review result of the sample song information; and using the sample review result as the tag information of the sample song information to construct the sample data.
[0072] The results of manual review include those based on the baseline dimension.
[0073] The sample audit results include audit results under the benchmark dimension and audit results under the extended dimension.
[0074] The review result optimization model is an open-source general-purpose large language model; the review prompt information is the prompt words corresponding to the song information review task.
[0075] Specifically, the server obtains sample song information for multiple sample songs, as well as the results of manual review of each sample song.
[0076] Then, the server uses the information of each sample song as input to the review result optimization model, and combines it with the review prompts of the review result optimization model to output the review results of each sample song. The manual review results of each sample song are used as the supervision information for the output of the review result optimization model. Combined with multiple optimizations of the initial review prompts of the review result optimization model, the review result optimization model outputs the sample review results of each sample song. In specific applications, the optimization process of the initial review prompts ends when a preset optimization stopping condition is reached. The review prompts obtained at the end of the optimization process are the optimized review prompts. The review results output by the review result optimization model under the optimized review prompts are the sample review results.
[0077] Next, after obtaining the sample review results for each sample song information, the server uses the sample review results of that sample song information as the tag information for that sample song information to construct a corresponding sample data.
[0078] In this embodiment, the manual review results include the review results under the baseline dimension, and the sample review results include the review results under the baseline dimension and the review results under the extended dimension. That is to say, the sample review results are extended compared to the manual review results, containing richer information, and are therefore optimized. Constructing sample data based on the optimized sample review results can make the content of the sample data richer, thereby further improving the reliability of the large language model in the song information review task when training the baseline large language model based on the sample data.
[0079] In an exemplary embodiment, the above steps, using initial review prompts and sample song information as input to the review result optimization model, and using the manual review results of the sample song information as supervisory information for the output of the review result optimization model, optimize the initial review prompts multiple times until a preset optimization stop condition is reached to obtain optimized review prompts. Specifically, this includes the following steps: inputting the initial review prompts and sample song information into the review result optimization model to obtain the model review results of the sample song information output by the review result optimization model; updating the initial review prompts based on the difference between the model review results and the manual review results in the baseline dimension to obtain updated review prompts; using the updated review prompts as new initial review prompts, returning to the step of inputting the initial review prompts and sample song information into the review result optimization model, until the consistency between the output model review results and the manual review results in the baseline dimension meets a preset consistency condition, then determining that the preset optimization stop condition has been reached, and using the updated review prompts at the point where the preset optimization stop condition is reached as the optimized review prompts.
[0080] The model review results include review results under the baseline dimension and review results under the extended dimension.
[0081] The preset consistency condition means that the model review results and manual review results of each sample song information are completely consistent under the benchmark dimension.
[0082] Specifically, the server or user constructs initial review prompts for the song information review task. For each sample song, the server combines the initial review prompts with the sample song information and inputs this combination into the review result optimization model. The model then generates a review result for the sample song. Next, the server filters out model review results that are not entirely identical to the sample review results of the corresponding sample song information in the baseline dimension. Then, the server or user updates the initial review prompts based on the differences between the filtered model review results and the corresponding sample review results, obtaining updated review prompts. Finally, the server uses the updated review prompts as the new initial review prompts and repeats this process for each sample song. The process involves inputting song information into the review result optimization model, generating model review results for the sample song information, and so on, until the model review results and human review results for each sample song information are completely consistent in the baseline dimension. At this point, the consistency between the output model review results and the human review results in the baseline dimension is confirmed to meet the preset consistency condition, and the preset optimization stop condition is determined to have been reached. Finally, the server uses the updated review prompt information when the preset optimization stop condition is reached as the optimized review prompt information, and uses the model review results for each sample song information output by the review result optimization model under the optimized review prompt information—that is, the model review results that are completely consistent with the human review results in the baseline dimension—as the sample review results for each sample song.
[0083] In practical applications, if the preset consistency condition is still not met after the preset number of optimizations for the initial review prompt information, the server will issue a prompt to the user to confirm whether there is a deviation in the manual review result, whether the manual review result needs to be adjusted, and whether the optimization of the initial review prompt information needs to be manually ended.
[0084] In practical applications, the initial structure of the initial review prompt message is as follows:
[0085] 1. Define identity roles: Tell the large language model the role you want it to play.
[0086] 2. Define the task clearly: tell the large language model what it needs to do.
[0087] 3. Mindset Guidance: Break down large problems into several smaller steps, guiding the large language model to reason and output according to the presented path.
[0088] 4. Provide examples: Provide some examples to help the large language model understand the user's requirements.
[0089] 5. Clarify the points to note: Point out the important points that the large language model needs to consider when answering the questions.
[0090] For example, construct the following prompt words:
[0091] {
[0092] [Role]
[0093] As a seasoned content security review expert...
[0094] [Task]
[0095] Your task is to carefully read the input JSON data, analyze the content to be reviewed based on the review results and the background information I provided, and then output the results.
[0096] [Mind Chain Guidance]
[0097] First, briefly summarize the input content to understand its meaning. Then, carefully read the following descriptions of the exception categories and levels, and refer to the "Exception Category" and "Exception Level" fields to output the original exception text and cause analysis. If..., then..., otherwise...
[0098] [example]
[0099] Explanation of anomaly categories: ...
[0100] Violation level description: ...
[0101] [Important Notes]
[0102] - The output format must be JSON, containing...
[0103] - Don't include this in your output...
[0104] -……
[0105] }
[0106] In this embodiment, the anomaly categories and anomaly levels generated by the review result optimization model are required to be strictly aligned with the anomaly categories and anomaly levels manually annotated. Based on this, a content summary, anomaly text, and judgment reason are generated. This is to ensure the accuracy and logical consistency of the sample data while annotating the sample song information with higher quality and more refined details. This makes the content of the sample data richer and further improves the reliability of the benchmark large language model when training the benchmark large language model based on the sample data in the song information review task.
[0107] In one exemplary embodiment, the server uses multiple different general-purpose large language models as review results to optimize the model and generate model review results simultaneously, and adopts the consistent model review results of multiple models as sample review results through voting.
[0108] In an exemplary embodiment, before reviewing the sample song information using the pre-trained benchmark large language model in step S104, the following steps are further included to obtain the benchmark large language model: rewriting the constructed review prompt information to obtain multiple rewritten review prompt information; using each rewritten review prompt information and the sample song information as input information, and using the sample review results as supervision information, fine-tuning the general large language model to obtain the benchmark large language model.
[0109] Specifically, the server or user constructs review prompts corresponding to the song information review task, and rewrites these prompts using an open-source general-purpose language model to obtain multiple extended review prompts. The server identifies both the extended and constructed prompts as the rewritten prompts. Next, the server constructs multiple fine-tuning training datasets based on the sample data and the rewritten prompts. Each fine-tuning training dataset includes one sample and one rewritten prompt. Then, the server uses the rewritten prompts and sample song information from the fine-tuning training datasets as input to the open-source general-purpose language model, and the review results from the sample data as output supervision information, to fine-tune and train the general-purpose language model, enabling it to learn domain knowledge in the song information review field and obtain a baseline language model.
[0110] In practical applications, the general large language model used for fine-tuning training and the general large language model used to rewrite the constructed review prompt information can be the same model. However, the general large language model used for fine-tuning training and the general large language model used to generate sample review results are usually not the same model.
[0111] In practical applications, the server selects at least one second sample from each sample data to construct fine-tuning training data; there is no duplicate sample data between the first sample data and the second sample data; furthermore, the ratio of positive to negative sample data in the second sample data is the optimal ratio after verification, such as 1:1.
[0112] In this embodiment, by fine-tuning the general language model based on sample data, the general language model can be aligned from "general language understanding" to the specific task of "song information review" and learn the domain knowledge of the song information review task. In addition, by rewriting the review prompts to construct fine-tuned training data, the robustness and generalization ability of the general language model under different review prompt expressions can be enhanced through fine-tuning training. In summary, the reliability of the general language model under the song information review task is further improved.
[0113] In an exemplary embodiment, step S108, which trains a benchmark large language model based on the similarity between the second review result and the sample review result in each preset dimension to obtain a target song information review model, specifically includes the following steps: for each preset dimension, determining the similarity of the second review result in the preset dimension based on the review results of the second review result and the sample review result in each preset dimension; calculating the reward function value of the benchmark large language model under a preset reward function based on the similarity of the second review result in each preset dimension; and performing reinforcement training on the benchmark large language model based on the reward function value to obtain the target song information review model.
[0114] Specifically, for each sample song information and each preset dimension, the server determines the similarity of the second review result under the preset dimension based on the similarity between the second review result of the sample song information and the review result of the sample review result under the preset dimension. Then, the server merges the similarity of the second review results of each sample song information under each preset dimension according to a preset reward function to obtain the reward function value of the benchmark large language model under the current reward function. Finally, the server performs iterative reinforcement training on the benchmark large language model based on the reward function value to obtain the target song information review model.
[0115] In practical applications, the server uses the scores under each preset dimension as the similarity of the audit results under each preset dimension.
[0116] For the baseline dimension "abnormal category", the server compares the abnormal category in the second review result with the abnormal category in the sample review result. If the two are consistent, the server determines that the score under the baseline dimension "abnormal category" is 1 point; otherwise, it is 0 points.
[0117] For the baseline dimension "abnormality level", the server maps each level to a different value. The higher the level, the higher the corresponding value. The server calculates the absolute value of the difference between the value corresponding to the abnormality level in the second review result and the value corresponding to the abnormality level in the sample review result, and determines the score of the baseline dimension "abnormality level" based on the absolute value of the difference. The larger the absolute value of the difference, the lower the score, and the smaller the absolute value of the difference, the higher the score.
[0118] For the extended dimension "Abnormal Song Information Fragments," both the abnormal song information fragments in the second review result and the abnormal song information fragments in the sample review result consist of several keywords. The server uses the F1 score to evaluate the keyword overlap between the abnormal song information fragments in the second review result and the abnormal song information fragments in the sample review result, and uses the calculated F1 score as the score under the extended dimension "Abnormal Song Information Fragments." The formula for calculating the F1 score is shown in Formula 1:
[0119] (Formula 1)
[0120] Wherein, Precision is the precision of the second review result, defined as the proportion of correct keywords in the abnormal song information fragments in the second review result to all keywords in the abnormal song information fragments in the second review result; Recall is the recall rate, defined as the proportion of correct keywords in the abnormal song information fragments in the second review result to all keywords in the abnormal song information fragments in the sample review result; the correct keyword is any one of several keywords in the abnormal song information fragments in the sample review result.
[0121] For the extended dimension "Causes of Anomalies," the server selects ROUGE-L (Recall-Oriented Understudy for Gisting Evaluation-L, one of the metrics in the ROUGE metric family) to determine the score under "Causes of Anomalies." A higher ROUGE-L score results in a higher score, and a lower ROUGE-L score results in a lower score. Furthermore, the server can also select or combine ROUGE-N (one of the metrics in the ROUGE metric family), ROUGE-S (one of the metrics in the ROUGE metric family), ROUGE-W (one of the metrics in the ROUGE metric family), an embedding model to calculate semantic similarity, and a reranker model to calculate semantic similarity to determine the score under the extended dimension "Causes of Anomalies."
[0122] In this embodiment, the server uses a reward function to increase the probability of generating high-quality output while suppressing the probability of generating low-quality output, thereby further improving the reliability of the large language model obtained through reinforcement training in the song information review task.
[0123] In one exemplary embodiment, the second audit result is an audit result in text form.
[0124] Before calculating the reward function value of the benchmark large language model under the preset reward function based on the similarity of the second review results under each preset dimension, the above steps include: determining the text information of the second review results.
[0125] The text information includes at least one of the following: text format consistency, text length differences, and repeated word statistics.
[0126] Among them, text format consistency refers to the consistency between the text format of the second review result and the preset text format.
[0127] Among them, the text length difference is the difference between the text length of the second review result and the preset text length; further, the preset text length is the preset text length range.
[0128] The results of the duplicate word statistics were obtained by counting the words that appeared repeatedly in the second review results.
[0129] The above steps, based on the similarity of the second review results under each preset dimension, calculate the reward function value of the benchmark large language model under the preset reward function. Specifically, this includes the following steps: based on the similarity of the second review results under each preset dimension and each text information, calculate the current reward function value of the benchmark large language model.
[0130] In practical applications, there are usually requirements for the format of the review results of song information. To avoid issues such as reward leakage during model training, significant increases in text length, and meaningless repetition, the reward function involved in this application also includes terms related to text format, text length, and word repetition within the text. Specifically, for each sample song information, the server determines the text format and text length of the second review result for that sample song information, and statistically analyzes the repeated words in the second review result to obtain the repeated word statistics result. Then, the server combines the consistency between the text format of the second review result of each sample song information and the preset text format, the difference between the text length of the second review result of each sample song information and the preset text length range, the repeated word statistics result of the second review result of each sample song information, and the similarity of the review results of the second review result of each sample song information under each preset dimension to calculate the reward function value of the benchmark large language model under the current reward function.
[0131] In practical applications, the server uses text format consistency score, text length difference score, and repeated word score to participate in the calculation of the reward function value.
[0132] Regarding text format, if the text format of the second review result is consistent with the preset text format, the server determines the text format consistency score as 1 point; otherwise, it is 0 points.
[0133] Regarding text length, if the text length of the abnormal reason in the second review result is outside the preset text length, the server determines the text length difference score to be -1 point; otherwise, it is 0 points. That is, if the text length does not exceed the preset text length, 1 point is deducted; if it does not exceed the preset text length, no points are deducted.
[0134] Regarding word repetition, the server uses the n-grams algorithm to generate all n-gram phrases for the reasons for anomalies in the second review result. Then, it iterates through all n-gram phrases, counts the number of repeated n-grams, and determines the score for the repeated word based on the number of repeated n-grams. If there are no repeated n-grams, the score for the repeated word is 0 points; if there are repeated n-grams, the score for the repeated word is negative. The more repeated n-grams there are, the larger the absolute value of the score for the repeated word. That is, the more repeated n-grams there are, the more points are deducted; if there are no repetitions, no points are deducted.
[0135] In this embodiment, by introducing terms corresponding to text format, text length, and word repetition into the reward function, the output format of the target song information review model can be made to strictly meet the requirements, and the reward leakage of the benchmark large language model during the reinforcement training process can be avoided.
[0136] In an exemplary embodiment, the reward function includes reward items and penalty items; text format consistency and similarity of review results under each preset dimension correspond to reward items, while text length differences and repeated word statistics correspond to penalty items.
[0137] In this embodiment, the design of rewards and penalties increases the probability of generating high-quality outputs while suppressing the probability of generating low-quality outputs, thereby further improving the reliability of the large language model obtained through reinforcement training in the song information review task.
[0138] In an exemplary embodiment, step S106 above, based on the sample audit results, determines the first audit results among the first audit results whose differences from the sample audit results meet preset difference conditions as the second audit results, specifically including the following steps: from each first audit result, select the first audit results whose audit results under each benchmark dimension are the same as the sample audit results to obtain positive audit results; from each first audit result, select the first audit results whose audit results under each benchmark dimension are different from the sample audit results to obtain negative audit results; and determine the first audit results other than the positive and negative audit results among the first audit results as the second audit results.
[0139] Specifically, the server selects the first audit results from each first audit result that are the same as the sample audit results in each benchmark dimension, as positive audit results; and selects the first audit results from each first audit result that are different from the sample audit results in each benchmark dimension, as negative audit results; then, the first audit results other than the positive and negative audit results are determined as second audit results.
[0140] In this embodiment, by filtering out the second review results that are neither entirely correct nor entirely wrong, the relative advantage of the second review results that are neither entirely correct nor entirely wrong can be calculated, which facilitates the subsequent reinforcement training of the benchmark large language model.
[0141] In one exemplary embodiment, such as Figure 2As shown, a song information verification method is provided. This embodiment illustrates the application of this method to a server, but it is understood that the method can also be applied to a terminal, or to a system including both a server and a terminal, and is implemented through the interaction between the server and the terminal. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. In this embodiment, the method includes the following steps:
[0142] Step S202: Obtain the song information of the song to be reviewed.
[0143] Step S204: Input the song information into the target song information review model to obtain the review result of the song information.
[0144] Step S206: Display the audit results.
[0145] The specific limitations in this embodiment can be found in the limitations on the training method of the song information review model mentioned above, and will not be repeated here.
[0146] Specifically, the server obtains the song information of the song to be reviewed, then inputs the song information into a pre-trained target song information review model, reviews the song information based on the target song information review model, and generates and outputs the review results of the song information.
[0147] In this embodiment, the target song information review model can quickly and reliably review the song information.
[0148] To more clearly illustrate the training method and song information review method of the song information review model provided in this application's embodiments, the following specific embodiment will be used to describe the training method and song information review method of the song information review model. However, it should be understood that the embodiments of this application are not limited thereto. Figure 3 As shown, in one exemplary embodiment, this application also provides a text moderation method based on large model reinforcement learning, specifically including the following steps:
[0149] 1. Data collection and construction of fine-grained labeled datasets.
[0150] The raw data, which includes complete song information (singer name, song name, lyrics) and human review results, is obtained online. After deduplication and cleaning, structured labels (including content summary, violation category, violation level, original text of the violation, and reason for judgment) are generated using a general large language model based on the human review results (violation level and violation category). Finally, a high-quality, fine-grained training dataset is constructed.
[0151] Specifically, such as Figure 4 As shown, step 1 also includes the following steps:
[0152] ① Data filtering; Perform initial filtering, deduplication, and cleaning on the raw data.
[0153] ② Data annotation; Generate review tags using a general large language model.
[0154] ③ Tag comparison; compare the review tags generated by the general large language model with the human review tags.
[0155] ④ Optimize suggestion keywords; optimize suggestion keywords based on comparison results;
[0156] ⑤ Construct a training dataset; continuously improve the quality of the review tags based on the optimized prompt words, and finally construct a high-quality, fine-grained training dataset.
[0157] 2. Cold start: Full-scale supervised fine-tuning.
[0158] Based on the training dataset, a full-scale supervised fine-tuning was performed on another general-purpose large language model, enabling the large language model to quickly master the basic rules and patterns of the review task, thereby possessing the preliminary risk identification and violation classification capabilities in the security review vertical domain scenario.
[0159] 3. Reinforce learning to optimize large language models.
[0160] Building upon fully supervised fine-tuning, GRPO (Generalized Reinforcement Policy Optimization, a reinforcement learning algorithm) is introduced for reinforcement learning training to obtain a large-scale model for song information review within a specific domain. The focus of this stage is to further improve the discrimination accuracy and consistency of the large language model, align its output with human review standards, and optimize the accuracy and interpretability of the large language model's output by designing a reward function. Furthermore, the server can also employ other reinforcement learning algorithms for reinforcement training.
[0161] 4. Review song information based on a well-trained vertical domain model for song information review.
[0162] In this embodiment, compared to the general large model, the safety accuracy is comparable, but the absolute value of the violation recall rate is increased by 30%, the accuracy of violation categories is increased by 2%, and the accuracy of violation level determination is increased by 37%; 2) compared to the vertical category large model with full supervised fine-tuning, the absolute value of the violation recall rate is increased by 8%, the accuracy of violation categories is increased by 9%, and the accuracy of violation level determination is increased by 7%. The output includes content summary, violation category, violation level, original violation text, and reason for determination, facilitating rapid location and handling of violation fragments during manual review, thus improving human review efficiency. The higher precision and recall rate also enables partial automated review, reducing review costs and improving the timeliness of handling risky content.
[0163] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0164] Based on the same inventive concept, this application also provides a training device for a song information review model to implement the training method for the song information review model described above. The solution provided by this device is similar to the implementation described in the above method. Therefore, the specific limitations of one or more training device embodiments for song information review models provided below can be found in the limitations of the training method for song information review models described above, and will not be repeated here.
[0165] In one exemplary embodiment, such as Figure 5 As shown, a training device for a song information review model is provided, including: a sample data acquisition module 502, a review result generation module 504, a review result filtering module 506, and a review model training module 508, wherein:
[0166] The sample data acquisition module 502 is used to acquire sample data; the sample data includes sample song information and the sample review results of the sample song information; the sample review results include review results under multiple preset dimensions.
[0167] The review result generation module 504 is used to review the sample song information through a pre-trained benchmark large language model and generate multiple first review results for the sample song information; each first review result includes review results under multiple preset dimensions.
[0168] The audit result filtering module 506 is used to determine, based on the sample audit results, the first audit result whose difference from the sample audit result meets the preset difference condition as the second audit result.
[0169] The review model training module 508 is used to train the benchmark large language model based on the similarity between the review results of the second review result and the sample review result in each preset dimension, so as to obtain the target song information review model.
[0170] In an exemplary embodiment, the preset dimensions include a baseline dimension and an extended dimension. The baseline dimension includes at least one of anomaly category and anomaly level, and the extended dimension includes at least one of content summary, anomaly song information fragment, and anomaly cause.
[0171] The sample data acquisition module 502 is also used to acquire sample song information and the manual review results of the sample song information; the manual review results include the review results under the baseline dimension; the initial review prompt information and sample song information are used as input to the review result optimization model, and the manual review results of the sample song information are used as the supervision information for the output of the review result optimization model, so as to optimize the initial review prompt information multiple times until the preset optimization stopping condition is reached to obtain the optimized review prompt information; the model review result of the sample song information under the optimized review prompt information is used as the sample review result of the sample song information; the sample review result includes the review result under the baseline dimension and the review result under the extended dimension.
[0172] In an exemplary embodiment, the sample data acquisition module 502 is further configured to input the initial review prompt information and sample song information into the review result optimization model to obtain the model review result of the sample song information output by the review result optimization model; the model review result includes the review result under the baseline dimension and the review result under the extended dimension; based on the difference between the model review result and the manual review result under the baseline dimension, the initial review prompt information is updated to obtain the updated review prompt information; the updated review prompt information is used as the new initial review prompt information, and the step of inputting the initial review prompt information and sample song information into the review result optimization model is returned until the consistency between the output model review result and the manual review result under the baseline dimension meets the preset consistency condition, then the preset optimization stop condition is determined to be reached, and the updated review prompt information when the preset optimization stop condition is reached is used as the optimized review prompt information.
[0173] In an exemplary embodiment, the training device for the song information review model further includes a model fine-tuning training module, which is used to rewrite the constructed review prompt information to obtain multiple rewritten review prompt information; and to fine-tune the general large language model with each rewritten review prompt information and sample song information as input information and the sample review results as supervision information to obtain a benchmark large language model.
[0174] In an exemplary embodiment, the review model training module 508 is further configured to, for each preset dimension, determine the similarity of the review results of the second review result in the preset dimension based on the review results of the second review result and the sample review result in the preset dimension; calculate the reward function value of the benchmark large language model under the preset reward function based on the similarity of the review results of the second review result in each preset dimension; and perform reinforcement training on the benchmark large language model based on the reward function value to obtain the target song information review model.
[0175] In one exemplary embodiment, the second audit result is an audit result in text form.
[0176] The review model training module 508 is also used to determine the text information of the second review result; the text information includes at least one of text format consistency, text length difference, and repeated word statistics; text format consistency is the consistency between the text format of the second review result and the preset text format; text length difference is the difference between the text length of the second review result and the preset text length; repeated word statistics are obtained by counting the repeated words in the second review result; based on the similarity of the review results of the second review result under each preset dimension, the reward function value of the benchmark large language model under the preset reward function is calculated, including: based on the similarity of the review results of the second review result under each preset dimension and each text information, the current reward function value of the benchmark large language model is calculated.
[0177] In an exemplary embodiment, the reward function includes reward items and penalty items; text format consistency and similarity of review results under each preset dimension correspond to reward items, while text length differences and repeated word statistics correspond to penalty items.
[0178] In one exemplary embodiment, the preset dimension includes a baseline dimension, which includes at least one of anomaly category and anomaly level.
[0179] The audit result filtering module 506 is also used to filter out the first audit results from each first audit result that are the same as the sample audit results under each benchmark dimension, and obtain positive audit results; to filter out the first audit results from each first audit result that are different from the sample audit results under each benchmark dimension, and obtain negative audit results; and to determine the first audit results other than positive and negative audit results from each first audit result as second audit results.
[0180] Based on the same inventive concept, this application also provides a song information review device for implementing the song information review method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more song information review device embodiments provided below can be found in the limitations of the song information review method described above, and will not be repeated here.
[0181] In one exemplary embodiment, such as Figure 6 As shown, a song information verification device is provided, including: a song information acquisition module 602, a song information verification module 604, and a verification result display module 606, wherein:
[0182] The song information acquisition module 602 is used to acquire song information of songs to be reviewed.
[0183] The song information review module 604 is used to input song information into the target song information review model and obtain the review result of the song information; the target song information review model is trained by the training method of the song information review model in any embodiment.
[0184] The audit results display module 606 is used to display the audit results.
[0185] The training device and various modules of the aforementioned song information review model can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0186] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores song information and song review results, among other data. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a training method for a song information review model and a song information review method.
[0187] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0188] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0189] In one exemplary embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above-described method embodiments.
[0190] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0191] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0192] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0193] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A training method for a song information verification model, characterized in that, The method includes: Acquire sample data; the sample data includes sample song information and the sample review results of the sample song information; the sample review results include review results under multiple preset dimensions; The sample song information is reviewed using a pre-trained benchmark large language model to generate multiple first review results for the sample song information; each first review result includes multiple review results under the preset dimensions. Based on the sample review results, the first review result that satisfies the preset difference condition with the sample review results among the first review results is determined as the second review result; Based on the similarity between the second review result and the sample review result under each preset dimension, the benchmark large language model is trained to obtain the target song information review model.
2. The method according to claim 1, characterized in that, The preset dimensions include a baseline dimension and extended dimensions. The baseline dimension includes at least one of anomaly category and anomaly level. The extended dimensions include at least one of content summary, anomaly song information fragment, and anomaly cause. The acquisition of sample data includes: Obtain the sample song information and the manual review results of the sample song information; the manual review results include the review results under the benchmark dimension. The initial review prompt information and the sample song information are used as inputs to the review result optimization model. The manual review results of the sample song information are used as the supervision information of the output of the review result optimization model. The initial review prompt information is optimized multiple times until the preset optimization stop condition is reached to obtain the optimized review prompt information. The model review result of the sample song information under the optimized review prompt information is used as the sample review result of the sample song information; the sample review result includes the review result under the baseline dimension and the review result under the extended dimension. The sample data is constructed by using the sample review results as the tag information for the sample song information.
3. The method according to claim 2, characterized in that, The process of using the initial review prompt information and the sample song information as input to the review result optimization model, and using the manual review results of the sample song information as supervisory information as the output of the review result optimization model, to optimize the initial review prompt information multiple times until a preset optimization stopping condition is reached to obtain the optimized review prompt information, includes: The initial review prompt information and the sample song information are input into the review result optimization model to obtain the model review result of the sample song information output by the review result optimization model; the model review result includes the review result under the baseline dimension and the review result under the extended dimension; Based on the difference between the model review result and the manual review result under the benchmark dimension, the initial review prompt information is updated to obtain the updated review prompt information; The updated review prompt information is used as the new initial review prompt information. The process of inputting the initial review prompt information and the sample song information into the review result optimization model is repeated until the consistency between the output model review result and the manual review result in the baseline dimension meets the preset consistency condition. Then, the preset optimization stop condition is determined to have been reached, and the updated review prompt information at the time the preset optimization stop condition is reached is used as the optimized review prompt information.
4. The method according to claim 1, characterized in that, Before reviewing the sample song information using a pre-trained benchmark large language model, the process also includes: The constructed review prompt message is rewritten to obtain multiple rewritten review prompt messages; Using the rewritten review prompts and the sample song information as input, and the sample review results as supervision information, the general large language model is fine-tuned and trained to obtain the benchmark large language model.
5. The method according to any one of claims 1 to 4, characterized in that, The step of training the benchmark large language model based on the similarity between the second review result and the sample review result under each preset dimension to obtain the target song information review model includes: For each preset dimension, the similarity of the second review result under the preset dimension is determined based on the review results of the second review result and the sample review result under the preset dimension; Based on the similarity of the second review result under each preset dimension, calculate the reward function value of the benchmark large language model under the preset reward function; The benchmark large language model is reinforced and trained based on the reward function value to obtain the target song information review model.
6. The method according to claim 5, characterized in that, The second review result is in text format; Before calculating the reward function value of the benchmark large language model under the preset reward function based on the similarity of the second review result under each preset dimension, the method further includes: The text information of the second review result is determined; the text information includes at least one of text format consistency, text length difference, and repeated word statistics; the text format consistency is the consistency between the text format of the second review result and the preset text format; the text length difference is the difference between the text length of the second review result and the preset text length; the repeated word statistics are obtained by counting the repeated words in the second review result; The step of calculating the reward function value of the benchmark large language model under the preset reward function based on the similarity of the second review result under each preset dimension includes: Based on the similarity of the second review results under each preset dimension and each text information, the current reward function value of the benchmark large language model is calculated.
7. The method according to claim 6, characterized in that, The reward function includes reward items and penalty items; the text format consistency and the similarity of the review results under each preset dimension correspond to the reward items, and the text length difference and the repeated word statistics correspond to the penalty items.
8. The method according to any one of claims 1 to 4, characterized in that, The preset dimension includes a baseline dimension, which includes at least one of anomaly category and anomaly level; The step of determining, based on the sample review results, a first review result whose difference from the sample review results meets a preset difference condition as a second review result includes: From each of the first audit results, select the first audit results that are the same as the sample audit results under each of the benchmark dimensions to obtain the positive audit results; From each of the first audit results, the first audit results that are different from the sample audit results under each of the benchmark dimensions are selected to obtain negative audit results; The first audit result other than the positive audit result and the negative audit result among the first audit results is determined as the second audit result.
9. A method for verifying song information, characterized in that, The method includes: Obtain song information for songs awaiting review; The song information is input into the target song information review model to obtain the review result of the song information; the target song information review model is trained by the training method of the song information review model according to any one of claims 1 to 8; Display the audit results.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the training method for the song information review model according to any one of claims 1 to 8 or the song information review method according to claim 9.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the training method for the song information review model according to any one of claims 1 to 8 or the song information review method according to claim 9.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the training method for the song information review model according to any one of claims 1 to 8 or the song information review method according to claim 9.