A song processing method, a song recommendation method, a device, a medium and a product

By guiding a general language model to annotate song data and training the model in two stages, the problem of misidentification in the age group recognition of songs in traditional machine learning was solved, achieving higher recognition accuracy and stability.

CN122364500APending Publication Date: 2026-07-10TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
Filing Date
2026-04-10
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Traditional machine learning methods have limitations in identifying the appropriate age range for songs, making it difficult to accurately identify emerging expressions and semantically complex lyrics, leading to misidentification.

Method used

The system uses prompt templates to guide a general language model to automatically label song data. Through two-stage model training, it constructs target sample sets and boundary sample sets to optimize the identification of song age groups.

Benefits of technology

It improves the accuracy and generalization ability of song recognition for different age groups, reduces the reliance on manual annotation, and ensures the stability of the model's recognition on data of songs of all ages and songs that are difficult to distinguish.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122364500A_ABST
    Figure CN122364500A_ABST
Patent Text Reader

Abstract

This application discloses a song processing method, song recommendation method, device, medium, and product, relating to the field of song recognition. The method includes: using a prompt template to guide a general model to annotate the original song; determining training data based on the original song and the corresponding annotation results; the annotation results include the age range suitable for the song; performing a first-stage training of the age range recognition model based on the training data; updating the prompt template based on difficult-to-distinguish data determined from the training data; using the prompt template to guide the general model to annotate the difficult-to-distinguish data; constructing a boundary sample set based on the difficult-to-distinguish data and the corresponding annotation results; constructing a target sample set in conjunction with the training data; and performing a second-stage training of the model after the first-stage training to finally obtain a trained age range recognition model, which is then used to identify the age range suitable for the song to be recognized. This application can improve the accuracy of identifying the suitable age range for a song.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of song recognition, and in particular to a song processing method, song recommendation method, device, medium, and product. Background Technology

[0002] With the rapid development of internet technology, music streaming platforms have become the main channel for users to obtain audio content. In order to improve user experience, personalized recommendation algorithms are widely used. However, due to the huge differences in the physiological development, psychological cognition and social roles of users of different ages, resulting in a gap in music preferences, especially for underage users, whose auditory perception and emotional understanding abilities are significantly different from those of adults, using the same algorithm logic may lead to unsuitable content (such as premature exposure to adult emotional theme songs) or monotonous content (such as only pushing songs for toddlers, lacking positive and uplifting content that matches the growth needs of adolescents). As a result, there is a need to recommend songs according to the appropriate age group.

[0003] For identifying the appropriate age range for a song, conventional techniques often employ age classification methods based on traditional machine learning. These techniques typically target text content such as lyrics. First, manual design and construction of vector representations of keywords, topic distributions, and syntactic features are undertaken. Then, traditional classification models such as logistic regression, support vector machines, and random forests are used to predict the appropriate age range for the lyrics. However, due to its reliance on extensive manual annotation, it has significant limitations. Ultimately, traditional machine learning struggles to recognize lyrics involving emerging expressions or metaphors, and it fails to accurately identify semantically complex and vaguely defined lyrics, leading to misclassification of the song's appropriate age range. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a song processing method, a song recommendation method, a device, a medium, and a product, which can improve the accuracy of identifying the appropriate age group for a song. The specific solution is as follows: Firstly, this application provides a song processing method, including: The current prompt template guides the general language model to annotate the original song data, resulting in the first annotation result; the annotation result includes the age group suitable for the song data; Training data is determined based on the original song data and the corresponding first annotation results. The current age group recognition model is trained based on the training data to obtain the first trained model. The target difficult-to-distinguish data is determined from the training data; the target difficult-to-distinguish data is data that meets the preset difficult-to-distinguish judgment conditions; After updating the current prompt template based on the target indistinguishable data, the current prompt template is used to guide the general language model to annotate the target indistinguishable data to obtain a second annotation result, and a boundary sample set is constructed based on the target indistinguishable data and the corresponding second annotation result; A target sample set is constructed based on the boundary sample set and the training data, and the first trained model is trained using the target sample set to obtain a trained target age group recognition model, so as to use the target age group recognition model to identify the age group suitable for the song data to be identified.

[0005] Optionally, the current prompt template is a prompt template based on the guidance example, age definitions for each age group, appropriate lyrics themes, and lyrics themes to avoid.

[0006] Optionally, the step of using the current prompt template to guide the general language model to annotate the original song data and obtain the first annotation result includes: Using the current prompt template to guide the general language model, the target lyrics of the original song data are analyzed to determine the target avoidance lyrics themes involved in the avoidance lyrics themes of each age group; Extract target lyric fragments that correspond to the theme of the target avoidance lyrics from the target lyrics, and determine the target age group suitable for the original song data based on the target lyric fragments and the theme of the target avoidance lyrics; The first annotation result corresponding to the original song data is determined based on the target information; the target information includes any one or a combination of the target age group, the target lyric fragment, the target avoidance lyric theme, and the target reason; the target reason is the reason why the general language model determines that the original song data is suitable for the target age group.

[0007] Optionally, before determining the training data based on the original song data and the corresponding first annotation results, the method further includes: The format of the first annotation result corresponding to the original song data is validated, and the song data that fails the first annotation result validation is removed from the original song data to obtain the original song data after the first deletion. The correctness of the first annotation result corresponding to the original song data after the first clearing is checked to obtain the corresponding check result; If the inspection result does not meet the preset inspection requirements, the current prompt template is updated based on the inspection result, and the user is redirected back to the step of using the current prompt template to guide the general language model to annotate the original song data. If the inspection result meets the preset inspection requirements, the step of determining training data based on the original song data and the corresponding first annotation result is triggered.

[0008] Optionally, determining the training data based on the original song data and the corresponding first annotation results includes: Logical verification is performed on the first annotation result corresponding to the original song data, and the song data that fails the first annotation result verification is removed from the original song data to obtain the second removed original song data; Training data is determined based on the original song data after the second clearing and the corresponding first annotation results.

[0009] Optionally, determining the target-difficult-to-distinguish data from the training data includes: At least two models are used to identify the age group suitable for the training data and determine the corresponding recognition confidence; the at least two models include at least two of the general language model, the current age group recognition model, and a third-party model. From the training data, identify target indistinguishable data that meet the preset indistinguishability criteria; The preset difficulty-to-distinguish criteria include differences in the age groups identified by the at least two models in the training data, and / or, the difference in confidence level between the at least two models in the training data is greater than a preset threshold.

[0010] Optionally, updating the current prompt template based on the target indistinguishable data includes: Generate indistinguishable examples based on the aforementioned target indistinguishable data; The difficult-to-distinguish examples are added to the current prompt template, and the age recognition rules in the current prompt template are updated using the target difficult-to-distinguish data.

[0011] Optionally, training the current age group recognition model based on the training data to obtain a first trained model includes: The target training data corresponding to each age group are sampled evenly from the training data to construct a training sample set; The current age group recognition model is trained based on the training sample set to obtain the first trained model.

[0012] Optionally, constructing the target sample set based on the boundary sample set and the training data includes: The target training data corresponding to each specified age group in the training sample set are upsampled to obtain a sampled sample set; the lower limit of the specified age group is greater than the upper limit of the other age groups; the other age groups are the age groups other than the specified age group in each age group. The target sample set is constructed based on the boundary sample set and the sampled sample set.

[0013] Optionally, training the first trained model using the target sample set to obtain a trained target age group recognition model includes: The first trained model is trained using the target sample set to obtain the second trained model; If the second trained model does not meet the preset termination condition, the current prompt template is updated, and the second trained model is determined as the new current age group recognition model, and the process jumps back to the step of determining the training data based on the original song data and the corresponding first annotation result; If the second trained model meets the preset termination condition, then the trained target age group recognition model is determined based on the second trained model.

[0014] Secondly, this application provides a song recommendation method based on an age group recognition model, wherein the age group recognition model is a target age group recognition model trained based on the aforementioned method, and the method includes: Obtain the song data to be identified; The age group recognition model is used to perform age recognition on the song data to be recognized, so as to determine the target age group suitable for the song data to be recognized; The song data to be identified is recommended to the target user terminal; the actual age of the user corresponding to the target user terminal is within the target age range.

[0015] Thirdly, this application provides an electronic device, comprising: Memory, used to store computer programs; A processor for executing the computer program to implement the aforementioned method.

[0016] Fourthly, this application provides a computer-readable storage medium for storing a computer program that, when executed by a processor, implements the aforementioned method.

[0017] Fifthly, this application provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the aforementioned method.

[0018] In this application, a general language model is guided by a current prompt template to annotate the original song data, obtaining a first annotation result. The annotation result includes the age range suitable for the song data. Training data is determined based on the original song data and the corresponding first annotation result. The current age range recognition model is trained based on the training data to obtain a first trained model. Target difficult-to-distinguish data is determined from the training data. The target difficult-to-distinguish data is data that meets preset difficult-to-distinguish judgment conditions. After updating the current prompt template based on the target difficult-to-distinguish data, the general language model is guided by the current prompt template to annotate the target difficult-to-distinguish data, obtaining a second annotation result. A boundary sample set is constructed based on the target difficult-to-distinguish data and the corresponding second annotation result. A target sample set is constructed based on the boundary sample set and the training data. The first trained model is trained using the target sample set to obtain a trained target age range recognition model, which is then used to identify the age range suitable for the song data to be recognized.

[0019] Therefore, this application utilizes prompt templates to guide a general language model to automatically annotate the original song data. Based on the original song data and the corresponding annotation results, training data is determined. Then, difficult-to-distinguish data is identified from the training data, and the prompt template is updated based on this data to guide the general language model in recognizing the appropriate age group for the difficult-to-distinguish data, thus achieving accurate annotation of the difficult-to-distinguish data. In this way, this application uses continuously updated prompt templates to guide the general language model to automatically annotate song data, eliminating the need for manual annotation. Furthermore, through continuous optimization and learning by the model itself, the accuracy of the age group annotation for songs can be improved, and it possesses strong generalization ability. Subsequently, using the song data annotated by the general language model to train the age group recognition model, the age group recognition model can also achieve high recognition accuracy and strong generalization ability, thereby achieving accurate identification of the appropriate age group for songs.

[0020] Furthermore, this application utilizes training data and a target sample set to perform two-stage training on the age group recognition model. The first stage of training, based on the training data, allows the model to learn basic recognition capabilities across all age groups, enabling it to stably predict the appropriate age group for a song. This yields a basic recognition model that is consistent in label distribution and judgment criteria, and covers all age groups. The second stage of training, based on the target sample set, allows the model to focus more on difficult-to-distinguish song data, improving the accuracy of recognizing the appropriate age group for such data. This avoids the problem of the model's overall prediction distribution being severely biased towards difficult-to-distinguish song data if the model initially focuses on it. In this way, this application can improve the model's accuracy in recognizing difficult-to-distinguish song data while ensuring the model's stability across all age groups. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0022] Figure 1 A system framework diagram provided for an embodiment of this application; Figure 2 A flowchart of a song processing method provided in an embodiment of this application; Figure 3 A flowchart of a general language model annotation provided for embodiments of this application; Figure 4 A flowchart of an age group recognition model training process provided in this application embodiment; Figure 5 A flowchart of a song recommendation method based on an age group recognition model is provided for embodiments of this application; Figure 6 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] Traditional machine learning-based age classification techniques rely heavily on manual annotation, which has significant limitations. This results in poor recognition of lyrics involving emerging or metaphorical expressions, and difficulty in accurately identifying semantically complex and vaguely defined lyrics, leading to misclassification of the appropriate age range for a song. To address this, this application provides a song processing method that uses prompt templates to guide a general language model in automatically annotating song data. The annotated data is then used to train an age group recognition model in two stages to improve its accuracy and achieve accurate identification of the appropriate age range for a song.

[0025] The system framework used in the method disclosed in this application can be found in [reference needed]. Figure 1 As shown, it may specifically include: a backend server 01 and one or more user terminals 02 connected to the backend server 01. The user terminal 02 can be a mobile device, such as a smartphone, tablet, laptop, or smartwatch. Smartwatches include children's smartwatches and adult smartwatches. The user terminal 02 can also be a web page; no limitation is made here.

[0026] In this application, the backend server 01 is mainly used to train the age group recognition model of the song audience and to recommend songs based on the age group recognition model.

[0027] The backend server 01, when training the age group recognition model for the song's target audience, executes the following steps of the song processing method: It guides a general language model using the current prompt template to annotate the original song data, obtaining a first annotation result; the annotation result includes the age group suitable for the song data; it determines training data based on the original song data and the corresponding first annotation result, and trains the current age group recognition model based on the training data to obtain a first trained model; it identifies target difficult-to-distinguish data from the training data that meets preset difficult-to-distinguish criteria; after updating the current prompt template based on the target difficult-to-distinguish data, it guides the general language model using the current prompt template to annotate the target difficult-to-distinguish data, obtaining a second annotation result, and constructs a boundary sample set based on the target difficult-to-distinguish data and the corresponding second annotation result; it constructs a target sample set based on the boundary sample set and the training data, and trains the first trained model using the target sample set to obtain a trained target age group recognition model, so that the target age group recognition model can be used to identify the age group suitable for the song data to be recognized.

[0028] The backend server 01 is used to implement the song recommendation method based on the age group recognition model. The steps include obtaining the song data to be identified; using the age group recognition model to identify the age of the song data to be identified in order to determine the target age group suitable for the song data to be identified; and recommending the song data to be identified to the target user terminal 02. The actual age of the user corresponding to the target user terminal 02 is within the target age group.

[0029] Furthermore, the backend server 01 can also be used to deploy the trained target age group recognition model to the third-party platform according to the configuration of the third-party platform. After the third-party platform has deployed the target age group recognition model, it can also use the target age group recognition model to perform age recognition on the song data to be recognized, so as to determine the target age group suitable for the song data to be recognized, and then recommend the song data to be recognized to the corresponding target user terminal 02.

[0030] See Figure 2 As shown, an embodiment of the present invention discloses a song processing method, including: Step S11: Use the current prompt template to guide the general language model to annotate the original song data and obtain the first annotation result; wherein, the annotation result includes the age group suitable for the song data.

[0031] In this embodiment of the application, in order to achieve automatic annotation of song data without the use of manual annotation, a large-parameter general language model with strong general understanding and reasoning capabilities is first selected to automatically reason and annotate the song data, thereby eliminating the need to rely on manual annotation and reducing labor and maintenance costs.

[0032] In order to enable the general language model to have good annotation capabilities, this application constructs a prompt template for the general language model to guide it to automatically annotate song data.

[0033] The prompt template is based on the guidance example, age definitions for each age group, and the matching and avoiding of lyric themes.

[0034] Guided examples include any one or a combination of positive examples, negative examples, and indistinguishable examples. Guided examples can constrain the age group recognition criteria of the general language model.

[0035] The age definition of each age group specifically refers to the age range corresponding to each age group, such as 0-2, 3-7, 8-11, 12-15, 16-17, and 18+.

[0036] The age-appropriate lyric themes refer to the lyric themes suitable for each age group. Different age groups have their own corresponding age-appropriate lyric themes. These age-appropriate lyric themes include, but are not limited to, lullabies, natural sound effects, fun nursery rhymes, parent-child interaction, animation soundtracks, children's stories, ancient poems, classic melodies, children's songs, classical masterpieces, animation and film soundtracks, public welfare education, classic Chinese literature, early childhood education, positive and popular content, popular science, textbooks and supplementary materials, and mental health.

[0037] The lyrics themes to avoid for each age group refer to the lyrics themes that need to be avoided at each age level. Different age groups have their own specific lyrics themes to avoid. Among these, the lyrics themes to avoid are mainly those that are difficult to understand and are likely to have a negative impact, such as love, negative social phenomena, and negative emotions.

[0038] The theme of the lyrics refers to the core idea, emotion, or central content expressed in the lyrics of a song.

[0039] After obtaining the general language model and constructing the prompt template, the annotation process for the original song data can specifically include: using the current prompt template to guide the general language model to analyze the target lyrics of the original song data to determine the target avoidance lyric topics involved in the avoidance lyric topics of each age group; then extracting the target lyric fragments corresponding to the target avoidance lyric topics from the target lyrics, and determining the target age group suitable for the original song data based on the target lyric fragments and the target avoidance lyric topics; and then determining the first annotation result corresponding to the original song data based on the target information; wherein, the target information includes any one or a combination of several of the following: target age group, target lyric fragments, target avoidance lyric topics, and target reasons; the target reasons are the reasons why the general language model determines that the original song data is suitable for the target age group.

[0040] It should be noted that the annotation results corresponding to the song data should at least include the target age group that the song data is suitable for, the target avoidance lyric theme involved in the song data, and the target reason why the general language model determines that the song data is suitable for the target age group, so as to facilitate subsequent verification of whether the annotation results are correct. In addition, the annotation results corresponding to the song data may also include the target lyric fragments in the song data that correspond to the target avoidance lyric theme.

[0041] Furthermore, this application sets the output format of the annotation results in the prompt template to guide the general language model to output the annotation results corresponding to the song data according to the preset output format; wherein, the output format can adopt JSON (JavaScript Object Notation) format to ensure the convenience of subsequent training.

[0042] In this way, the prompt template adopts a distributed structured design, guiding the general language model to sequentially perform the determination of the lyrics theme, the extraction of the corresponding lyrics fragments, and the identification of the appropriate age group for the song. This improves the inference stability of the general language model and provides reasons for the general language model to determine that the song data is suitable for the target age group, thus having a certain degree of interpretability.

[0043] After labeling the original song data and obtaining the first labeling result, it is necessary to perform a rough cleaning and random check on the first labeling result corresponding to the original song data to obtain the corresponding check result. Then, it is determined whether the check result meets the preset check requirements to determine whether to trigger the step of determining training data based on the original song data and the corresponding first labeling result.

[0044] Specifically, such as Figure 3 As shown, after using the current prompt template to guide the general language model to annotate the original song data and obtain the first annotation result, it is necessary to perform format validation on the first annotation result corresponding to the original song data and remove the song data that fails the first annotation result validation from the original song data to obtain the first cleaned original song data; then, the correctness of the first annotation result corresponding to the first cleaned original song data is checked to obtain the corresponding check result; if the check result does not meet the preset check requirements, the current prompt template is updated based on the check result, and the process jumps back to the step of using the current prompt template to guide the general language model to annotate the original song data; if the check result meets the preset check requirements, the step of determining the training data based on the original song data and the corresponding first annotation result is triggered.

[0045] In one example, format validation is performed on the annotation results, including but not limited to validating the syntax and structure of the annotation results, verifying the completeness of each target information in the annotation results, and verifying the correctness of the information type of each target information in the annotation results. In this way, format validation can effectively remove noisy data, improve the quality of data used to train the age group recognition model, and prevent the training of the age group recognition model from being affected by low-quality data.

[0046] It should be noted that song data that successfully passes the first annotation verification should be retained in the original song data, while song data that fails the first annotation verification should be removed from the original song data, thus obtaining the original song data after the first removal.

[0047] According to one example, the correctness check is performed on the first annotation result corresponding to the original song data after the first clearing. This can be done either by checking all the song data after the first clearing or by extracting a portion of the song data from the original song data after the first clearing for the correctness check of the first annotation result.

[0048] It should be noted that when extracting some song data from the original song data after the first clearing, we can focus on the song data that is difficult to distinguish. Song data that is difficult to distinguish refers to song data that is difficult to classify correctly during the model training process. It usually has the characteristics of low classification confidence and large deviation between the classification results and the actual results. For example, song data that is difficult to distinguish can include fuzzy samples near the boundaries of each age group, samples with inconsistent judgment criteria or obviously insufficient reasons in the annotation results, etc.

[0049] Furthermore, the inspection results can include whether the first annotation results corresponding to the original song data are correct and the corresponding accuracy rate. Accordingly, the preset inspection requirements can include that all the first annotation results corresponding to the original song data are correct, or that the accuracy rate reaches a preset value.

[0050] According to one example, updating the current prompt template based on the inspection results may include: generating a guidance example based on the song data that failed the correctness check of the first annotation result, adding the guidance example to the current prompt template, and updating the age recognition rules in the current prompt template using the inspection results.

[0051] It should be noted that the current update of the prompt template includes, but is not limited to, adjusting the lyrics theme and adjusting the guidance examples (including positive examples, negative examples, and difficult-to-distinguish examples), thereby gradually tightening the content boundaries for each age group and improving the consistency and stability of the structured annotation results.

[0052] Step S12: Determine training data based on the original song data and the corresponding first annotation results, and train the current age group recognition model based on the training data to obtain the first trained model.

[0053] In the process of determining training data based on the original song data and the corresponding first annotation results, logical verification is performed on the first annotation results corresponding to the original song data, and song data that fails the first annotation result verification is removed from the original song data to obtain the second cleaned original song data; then, based on the second cleaned original song data and the corresponding first annotation results, the training data is determined. Finally, the current age group recognition model is trained in the first stage based on the training data to obtain the first trained model.

[0054] In one example, the first annotation result undergoes logical verification. This includes using a preset rule engine or the current age group recognition model to verify the judgment logic of each target information in the first annotation result. For example, the judgment logic between the target age group, target lyric fragments, target avoidance lyric theme, and target reason is verified. Song data that fails the first annotation result verification is song data where the judgment logic of each target information in the first annotation result is flawed. Examples include obvious self-contradiction among target information, obvious discrepancies between the target avoidance lyric theme and the target age group, obvious insufficiency of the target reason, and obvious discrepancy between the target reason and the target age group.

[0055] It should be noted that song data that successfully passes the first annotation verification should be retained in the original song data, while song data that fails the first annotation verification should be removed from the original song data, thus obtaining the second, cleared original song data.

[0056] According to one example, training a current age group recognition model based on training data to obtain a first post-trained model can specifically include: sampling target training data corresponding to each age group from the training data in a balanced manner to construct a training sample set, and training the current age group recognition model based on the training sample set to obtain the first post-trained model.

[0057] The training data is sampled evenly from the training data for each age group. This balance involves two aspects: first, sampling is made as evenly as possible across age groups to ensure that the amount of target training data for each age group is the same, so as to avoid the age group recognition model from overfitting to a certain age group; second, within each age group, training data involving different avoidance lyric themes are sampled evenly so that the age group recognition model has the basic ability to recognize various avoidance lyric themes and avoids bias towards a certain avoidance lyric theme.

[0058] Furthermore, by training the current age group recognition model using the training sample set, the age group recognition model can learn the structured output format of the labeled results, stably output structured labeled results, and learn the basic age recognition ability for all age groups, thus exhibiting superior recognition ability across all age groups. In addition, it can also learn the basic recognition and classification ability for various lyrics themes that avoid the target audience. Thus, through the first stage of training, a basic age group recognition model that is relatively consistent in label distribution and judgment criteria and covers all age groups can be obtained.

[0059] Step S13: Determine the target difficult-to-distinguish data from the training data; the target difficult-to-distinguish data is data that meets the preset difficult-to-distinguish judgment conditions.

[0060] For identifying target difficult-to-distinguish data from training data, the specific methods may include: using at least two models to identify the age group suitable for the training data and determining the corresponding recognition confidence; wherein, the at least two models include at least two of a general language model, a current age group recognition model, and a third-party model; identifying target difficult-to-distinguish data from the training data that meets preset difficult-to-distinguish criteria; wherein, the preset difficult-to-distinguish criteria include that the age groups identified by the at least two models differ, and / or, the difference in recognition confidence between the at least two models for the training data is greater than a preset threshold.

[0061] It should be noted that the age group recognition model is a small-parameter model, meaning that the model parameters of the age group recognition model are smaller than those of the general language model. Furthermore, third-party models can employ a large language model (LLM).

[0062] Furthermore, the general language model, age group recognition model, and third-party model in the embodiments of this application are not limited to a specific large model. They can be selected according to the actual needs of the user. For example, different series of large models can be used, such as Qwen2.5, GPT (Generative Pre-trained Transformer) model, etc. The model size can be 7B, 13B, 70B, etc.

[0063] Specifically, for the same training data, at least two models are invoked to perform age recognition on the training data to determine the suitable age range for the training data and the corresponding recognition confidence level. The recognition confidence level ranges from 0 to 1; the closer to 1, the higher the recognition confidence and the more accurate the age range recognition. After determining the suitable age range for the training data and the corresponding recognition confidence level, in one scenario, training data where the age ranges identified by at least two models differ is defined as target data that is difficult to distinguish; in another scenario, training data where the difference in recognition confidence levels between at least two models is greater than a preset threshold is defined as target data that is difficult to distinguish; and in yet another scenario, training data where the difference in age ranges identified by at least two models and the difference in recognition confidence levels is greater than a preset threshold is defined as target data that is difficult to distinguish.

[0064] When there are at least two models, the difference in the recognition confidence of the training data between the two models is the confidence difference between the recognition confidence of the training data between the two models. When the difference is greater than a preset threshold, the training data can be determined to be the target data that is difficult to distinguish.

[0065] When there are at least two models or three models, the confidence difference between the at least two models in identifying the training data includes three confidence differences. These three confidence differences correspond to three model combinations, which are the three possible combinations obtained by pairwise pairwise combinations of the three models. When a predetermined number of the three confidence differences exceeds a predetermined threshold, the training data is determined to be target data that is difficult to distinguish. The predetermined number can be 1, 2, or 3, with 1 being preferred, and is not specifically limited here.

[0066] In this way, by cross-labeling multiple models, it is possible to determine whether there are differences in the age groups recognized by multiple models for the same training data and whether the difference in recognition confidence exceeds a preset threshold. This helps to identify difficult-to-distinguish data from the training data, reduce the randomness of using a single model, improve the robustness of the overall labeled data, and increase the accuracy of identifying difficult-to-distinguish data.

[0067] Step S14: After updating the current prompt template based on the target indistinguishable data, the current prompt template is used to guide the general language model to annotate the target indistinguishable data to obtain a second annotation result, and a boundary sample set is constructed based on the target indistinguishable data and the corresponding second annotation result.

[0068] In this embodiment of the application, after identifying the target data that is difficult to distinguish from the training data, it is necessary to update the current prompt template based on the target data that is difficult to distinguish, and then use the current prompt template to guide the general language model to re-annotate the target data that is difficult to distinguish in order to obtain the second annotation result. After that, the boundary sample set can be constructed based on the target data that is difficult to distinguish and the corresponding second annotation result.

[0069] For updating the current prompt template based on the target indistinguishable data, the specific steps may include: generating indistinguishable examples based on the target indistinguishable data; adding the indistinguishable examples to the current prompt template; and updating the age recognition rules in the current prompt template using the target indistinguishable data.

[0070] Among these, target-difficult-to-distinguish data typically refers to easily confused and error-prone song data, such as songs that offer guidance on negative emotions versus songs that express negative emotions. For target-difficult-to-distinguish data, we can identify the appropriate age group by distinguishing its specific scenarios and details, determine the relevant and avoidable lyric themes, and identify the specific reasons for identifying it as belonging to that age group. This generates difficult-to-distinguish examples, which are then added to the current prompt template to guide the general language model to better identify the age of difficult-to-distinguish data.

[0071] In this way, the embodiments of this application continuously update the prompt template to form a continuously optimized rule system, thereby improving the recognition accuracy of the general language model for the appropriate age group of the song data without adjusting the model parameters of the general language model, and avoiding the problems of long cycle and large workload caused by adjusting model parameters.

[0072] After updating the current prompt template, the annotation process for target-indistinguishable data can specifically include: using the current prompt template to guide a general language model to analyze the target lyrics of the target-indistinguishable data to determine the target avoidance lyric themes involved in the avoidance lyric themes of each age group; extracting target lyric fragments corresponding to the target avoidance lyric themes from the target lyrics, and determining the target age group suitable for the target-indistinguishable data based on the target lyric fragments and the target avoidance lyric themes; determining the second annotation result corresponding to the target-indistinguishable data based on any one or more combinations of the target age group, target lyric fragments, target avoidance lyric themes, and target reasons; wherein, the target reasons are the reasons why the general language model determines that the target-indistinguishable data is suitable for the target age group. Finally, based on the target-indistinguishable data and the corresponding second annotation result, the boundary sample set can be constructed.

[0073] Step S15: Construct a target sample set based on the boundary sample set and the training data, and use the target sample set to train the first trained model to obtain a trained target age group recognition model, so as to use the target age group recognition model to identify the age group suitable for the song data to be identified.

[0074] In this embodiment, after obtaining a first trained model by completing the first stage of training of the current age group recognition model based on training data, the first trained model is trained in the second stage using a target sample set to obtain a second trained model. If the second trained model meets the preset termination condition, the trained target age group recognition model is determined based on the second trained model. If the second trained model does not meet the preset termination condition, the current prompt template is updated, and the second trained model is determined as the new current age group recognition model. The process then jumps back to the step of determining the training data based on the original song data and the corresponding first annotation results to continue iterative optimization training of the age group recognition model.

[0075] The preset termination conditions include, but are not limited to, the model accuracy of the second trained model reaching a preset accuracy, the model loss of the second trained model being less than a preset loss, and the cumulative number of training iterations of the second trained model reaching a preset number.

[0076] Furthermore, the construction of the target sample set based on the boundary sample set and training data can specifically include: upsampling the target training data corresponding to each specified age group in the training sample set to obtain the sampled sample set; wherein, the lower limit of the specified age group is greater than the upper limit of the other age groups, and the other age groups are the age groups other than the specified age group; then the target sample set is constructed based on the boundary sample set and the sampled sample set.

[0077] For example, when the age ranges include 0-2, 3-7, 8-11, 12-15, 16-17, and 18+, the specified age range can include 16-17 and 18+. The above is just one example, and other age ranges can also be selected as the specified age range.

[0078] It should be noted that, by upsampling the training data corresponding to a specified age group in this embodiment, the proportion of training data corresponding to the specified age group in the training sample set can be increased. This makes the decision boundary of the age group recognition model tend to shift towards a higher age group when there is uncertainty or insufficient evidence in age recognition, because the higher age group is more inclusive than the lower age group. This achieves a conservative age recognition strategy for song data that is difficult to distinguish, and maintains higher sensitivity and recall for lyrics themes that are avoided for higher age groups.

[0079] Of course, if you want the decision boundary of the age group recognition model to tend to shift towards lower age groups, you can also set the specified age group to a lower age group, that is, the upper age limit of the specified age group is lower than the lower age limit of other age groups. For example, when the age groups include 0-2, 3-7, 8-11, 12-15, 16-17, and 18+, the specified age group can include 0-2, 3-7, 8-11, and 12-15.

[0080] In other words, the specified age range can be selected based on the user's actual preference for the decision boundary of the age range recognition model, with higher age ranges being preferred.

[0081] For constructing a target sample set based on a boundary sample set and a sampled sample set, in one case, the target sample set can be directly constructed based on the boundary sample set and the sampled sample set. In another case, the boundary sample set can be upsampled to increase the proportion of the boundary sample set, and then the target sample set can be constructed based on the upsampled boundary sample set and the sampled sample set.

[0082] In this way, based on the completion of the first stage of model training, by adding boundary sample sets, the age group recognition model can further learn difficult-to-distinguish data during the second stage of model training without changing the loss function and model parameters of the age group recognition model. This allows the model to obtain more contrast learning signals in the boundary regions of each age group, thereby improving its ability to recognize age in difficult-to-distinguish data.

[0083] For the two-stage training of the age group recognition model, the specific details are as follows: Figure 4 As shown, in the process of determining training data based on the original song data and the corresponding first annotation results, this application firstly performs logical verification on the first annotation results corresponding to the original song data, and removes the song data that fails the first annotation result verification from the original song data to obtain the second cleaned original song data. Then, the training data is determined based on the second cleaned original song data and the corresponding first annotation results. Afterwards, the current age group recognition model is trained in the first stage using the training data to obtain the first trained model. Further, this application uses multi-model cross-labeling to determine target difficult-to-distinguish data from the training data and constructs a boundary sample set. Based on the boundary sample set and the training data, a target sample set is constructed to train the first trained model in the second stage using the target sample set to obtain the second trained model. If the second trained model meets the preset termination condition, the trained target age group recognition model is obtained based on the second trained model. If the second trained model does not meet the preset termination condition, the current prompt template is updated, and the process returns to the step of determining training data based on the original song data and the corresponding first annotation results to continue iterative optimization training of the age group recognition model.

[0084] See Figure 3 and Figure 4 As can be seen, the embodiments of this application construct a closed-loop iterative mechanism of "general language model annotation - two-stage training of age group recognition model - prompt template update - next round of annotation and training" to regard the general language model as the teacher model and the age group recognition model as the student model. This allows the student model to learn the same structured annotation results as the teacher model after multiple rounds of iteration. Furthermore, the structured annotation results enable the student model to perform multi-task learning during training, including age group recognition, lyric theme recognition, and reason generation, forming a richer internal representation than single age group labels. This improves the model's understanding of lyric themes and enhances the interpretability of the annotation results, providing support for format verification, correctness verification, and logical verification.

[0085] Taking the R1 model for the general language model and the Qwen2.5-7B model for the age group recognition model as examples, the age group recognition model trained using the method of this application achieved a recall rate of 96.8% across all age groups, enabling high coverage and high interception of songs unsuitable for each age group. Furthermore, compared to the general language model, the trained age group recognition model has a single inference cost of approximately 1 / 20 and an inference speed approximately 10 times faster, significantly reducing inference costs while maintaining the accuracy of age group recognition and greatly improving overall service throughput.

[0086] Therefore, this application utilizes prompt templates to guide a general language model to automatically annotate the original song data. Based on the original song data and the corresponding annotation results, training data is determined. Then, difficult-to-distinguish data is identified from the training data, and the prompt template is updated based on this data to guide the general language model in recognizing the appropriate age group for the difficult-to-distinguish data, thus achieving accurate annotation of the difficult-to-distinguish data. In this way, this application uses continuously updated prompt templates to guide the general language model to automatically annotate song data, eliminating the need for manual annotation. Furthermore, through continuous optimization and learning by the model itself, the accuracy of the age group annotation for songs can be improved, and it possesses strong generalization ability. Subsequently, using the song data annotated by the general language model to train the age group recognition model, the age group recognition model can also achieve high recognition accuracy and strong generalization ability, thereby achieving accurate identification of the appropriate age group for songs.

[0087] Furthermore, this application utilizes a training sample set and a target sample set to perform two-stage training on the age group recognition model. The first stage of training, based on the training sample set, allows the model to learn basic recognition capabilities across all age groups, enabling it to stably predict the appropriate age group for a song. This results in a basic recognition model that is consistent in label distribution and judgment criteria, and covers all age groups. The second stage of training, based on the target sample set, allows the model to focus more on difficult-to-distinguish song data, improving the accuracy of recognizing the appropriate age group for such data. This avoids the problem of the model's overall prediction distribution being severely biased towards difficult-to-distinguish song data from the outset. In this way, this application can improve the model's accuracy in recognizing difficult-to-distinguish song data while ensuring the model's stability across all age groups.

[0088] See Figure 5 As shown in the figure, this invention discloses a song recommendation method based on an age group recognition model, wherein the age group recognition model is a target age group recognition model trained based on the aforementioned embodiments, and the song recommendation method specifically includes: Step S21: Obtain the song data to be identified; Step S22: Use the age group recognition model to perform age recognition on the song data to be recognized, so as to determine the target age group suitable for the song data to be recognized; Step S23: Recommend the song data to be identified to the target user terminal; the actual age of the user corresponding to the target user terminal is within the target age range.

[0089] In this embodiment, after training the age group recognition model based on the previous embodiment, the data of the song to be identified can be obtained, and the age group recognition model can be used to identify the age of the song data to determine the target age group suitable for the song data. This allows for the construction of a song library suitable for different age groups, meaning that different age groups have their own song libraries. This enables the recommendation of songs from the corresponding song libraries to users of different age groups, thereby achieving differentiated song exposure and control for users of different age groups and reducing the risk of minors being exposed to inappropriate songs.

[0090] Furthermore, in terms of actual products, this application can also set up a minor mode and an adult mode, and apply the song recommendation method proposed in the embodiments of this application to the minor mode. The adult mode can continue to use the conventional song recommendation method, thereby meeting the diverse needs of different users. In the minor mode, a healthy, suitable, and safe music environment can be provided for minors, prioritizing the recommendation of songs suitable for their age group, improving the practicality of the minor mode, and allowing minor users to enjoy music with peace of mind in a controlled environment.

[0091] Furthermore, embodiments of this application also disclose an electronic device, Figure 6 This is a structural diagram of an electronic device 10 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0092] Figure 6 This is a schematic diagram of the structure of an electronic device 10 provided in an embodiment of this application. Specifically, the electronic device 10 may include: at least one processor 11, at least one memory 12, a power supply 13, a communication interface 14, an input / output interface 15, and a communication bus 16. The memory 12 stores a computer program, which is loaded and executed by the processor 11 to implement the relevant steps in the methods disclosed in any of the foregoing embodiments. Furthermore, the electronic device 10 in this embodiment may specifically be an electronic computer.

[0093] In this embodiment, the power supply 13 is used to provide operating voltage for each hardware device on the electronic device 10; the communication interface 14 can create a data transmission channel between the electronic device 10 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 15 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0094] In addition, the memory 12, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 121, computer program 122, etc., and the storage method can be temporary storage or permanent storage.

[0095] The operating system 121 is used to manage and control the various hardware devices on the electronic device 10 and the computer program 122, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including computer programs capable of performing the methods executed by the electronic device 10 as disclosed in any of the foregoing embodiments, the computer program 122 may further include computer programs capable of performing other specific tasks.

[0096] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned disclosed method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0097] Furthermore, this application also discloses a computer program product, including a computer program / instructions, wherein the computer program / instructions, when executed by a processor, implement the aforementioned disclosed method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0098] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0099] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0100] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0101] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0102] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A song processing method, characterized in that, include: The current prompt template guides the general language model to annotate the original song data, resulting in the first annotation result; the annotation result includes the age group suitable for the song data; Training data is determined based on the original song data and the corresponding first annotation results. The current age group recognition model is trained based on the training data to obtain the first trained model. The target difficult-to-distinguish data is determined from the training data; the target difficult-to-distinguish data is data that meets the preset difficult-to-distinguish judgment conditions; After updating the current prompt template based on the target indistinguishable data, the current prompt template is used to guide the general language model to annotate the target indistinguishable data to obtain a second annotation result, and a boundary sample set is constructed based on the target indistinguishable data and the corresponding second annotation result; A target sample set is constructed based on the boundary sample set and the training data, and the first trained model is trained using the target sample set to obtain a trained target age group recognition model, so as to use the target age group recognition model to identify the age group suitable for the song data to be identified.

2. The song processing method according to claim 1, characterized in that, The current prompt template is based on the guidance example, age definitions for each age group, and prompts for appropriate and inappropriate lyrics themes.

3. The song processing method according to claim 2, characterized in that, The step of using the current prompt template to guide the general language model to annotate the original song data and obtain the first annotation result includes: Using the current prompt template to guide the general language model, the target lyrics of the original song data are analyzed to determine the target avoidance lyrics themes involved in the avoidance lyrics themes of each age group; Extract target lyric fragments that correspond to the theme of the target avoidance lyrics from the target lyrics, and determine the target age group suitable for the original song data based on the target lyric fragments and the theme of the target avoidance lyrics; The first annotation result corresponding to the original song data is determined based on the target information; the target information includes any one or a combination of the target age group, the target lyric fragment, the target avoidance lyric theme, and the target reason; the target reason is the reason why the general language model determines that the original song data is suitable for the target age group.

4. The song processing method according to claim 3, characterized in that, Before determining the training data based on the original song data and the corresponding first annotation results, the method further includes: The format of the first annotation result corresponding to the original song data is validated, and the song data that fails the first annotation result validation is removed from the original song data to obtain the original song data after the first deletion. The correctness of the first annotation result corresponding to the original song data after the first clearing is checked to obtain the corresponding check result; If the inspection result does not meet the preset inspection requirements, the current prompt template is updated based on the inspection result, and the user is redirected back to the step of using the current prompt template to guide the general language model to annotate the original song data. If the inspection result meets the preset inspection requirements, the step of determining training data based on the original song data and the corresponding first annotation result is triggered.

5. The song processing method according to claim 4, characterized in that, The step of determining training data based on the original song data and the corresponding first annotation results includes: Logical verification is performed on the first annotation result corresponding to the original song data, and the song data that fails the first annotation result verification is removed from the original song data to obtain the second removed original song data; Training data is determined based on the original song data after the second clearing and the corresponding first annotation results.

6. The song processing method according to claim 1, characterized in that, The step of identifying target-difficult-to-distinguish data from the training data includes: At least two models are used to identify the age group suitable for the training data and determine the corresponding recognition confidence; the at least two models include at least two of the general language model, the current age group recognition model, and a third-party model. From the training data, identify target indistinguishable data that meet the preset indistinguishability criteria; The preset difficulty-to-distinguish criteria include differences in the age groups identified by the at least two models in the training data, and / or, the difference in confidence level between the at least two models in the training data is greater than a preset threshold.

7. The song processing method according to claim 1, characterized in that, The step of updating the current prompt template based on the target data that is difficult to distinguish includes: Generate indistinguishable examples based on the aforementioned target indistinguishable data; The difficult-to-distinguish examples are added to the current prompt template, and the age recognition rules in the current prompt template are updated using the target difficult-to-distinguish data.

8. The song processing method according to claim 1, characterized in that, The step of training the current age group recognition model based on the training data to obtain a first trained model includes: The target training data corresponding to each age group are sampled evenly from the training data to construct a training sample set; The current age group recognition model is trained based on the training sample set to obtain the first trained model.

9. The song processing method according to claim 8, characterized in that, The construction of the target sample set based on the boundary sample set and the training data includes: The target training data corresponding to each specified age group in the training sample set are upsampled to obtain a sampled sample set; the lower limit of the specified age group is greater than the upper limit of the other age groups; the other age groups are the age groups other than the specified age group in each age group. The target sample set is constructed based on the boundary sample set and the sampled sample set.

10. The song processing method according to any one of claims 1 to 9, characterized in that, The step of training the first trained model using the target sample set to obtain a trained target age group recognition model includes: The first trained model is trained using the target sample set to obtain the second trained model; If the second trained model does not meet the preset termination condition, the current prompt template is updated, and the second trained model is determined as the new current age group recognition model, and the process jumps back to the step of determining the training data based on the original song data and the corresponding first annotation result; If the second trained model meets the preset termination condition, then the trained target age group recognition model is determined based on the second trained model.

11. A song recommendation method based on an age group recognition model, characterized in that, The age group recognition model is a target age group recognition model trained based on the method described in any one of claims 1 to 10, the method comprising: Obtain the song data to be identified; The age group recognition model is used to perform age recognition on the song data to be recognized, so as to determine the target age group suitable for the song data to be recognized; The song data to be identified is recommended to the target user terminal; the actual age of the user corresponding to the target user terminal is within the target age range.

12. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the method as claimed in any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, implements the method as described in any one of claims 1 to 11.

14. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method as described in any one of claims 1 to 11.