Song feature processing method, server, program product and storage medium

By matching reference songs with similar audio characteristics in the music application and generating posterior features of new or unpopular songs, the problem of these songs ranking low in the rankings is solved, and their exposure and promotional effects are improved.

CN120067386APending Publication Date: 2025-05-30TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510102771.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Due to the lack of user operation records, new songs or unpopular songs rank low on the rankings of music applications, and lack of exposure, which is not conducive to song promotion.

Method used

By matching reference songs with similar audio features in the song library, the posterior features of the reference song are used to generate the posterior features of the target song, thereby making up for the problem that the target song lacks posterior features.

Benefits of technology

It has increased the exposure of new or unpopular songs in song recommendations and rankings, helping these songs gain more user attention and promotion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067386A_ABST
    Figure CN120067386A_ABST
Patent Text Reader

Abstract

The invention provides a song feature processing method and device and a storage medium, and relates to the technical field of audio, and the method comprises the steps: determining a target song, and determining a plurality of candidate songs whose posterior features meet a preset condition from a song library, the posterior features being used for representing the condition that the candidate songs are operated by a user; according to the similarity between the audio feature of each candidate song and the audio feature of the target song, determining at least one candidate song meeting a similarity condition from the plurality of candidate songs as a reference song; and according to the posterior feature of the at least one reference song, generating the posterior feature of the target song. By adopting the method and the device, the posterior features of the songs without the posterior features can be automatically generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio technology, and particularly relates to a method for processing song features, a method for training a click-through rate prediction model, a server, a computer program product, and a storage medium. Background Art

[0002] Various song rankings in music applications are the main ways for song promotion and exposure. Usually, the sorting of songs in the rankings is based on song features. Song features are obtained by integrating posterior features and prior features. Posterior features refer to the situations where songs are operated by users (such as clicks, favorites, plays, etc.), and prior features refer to the features of the songs themselves (such as song names, singer names, song styles, etc.). However, for new songs or unpopular songs, they may not have been operated by users at all. Then, their rankings in the rankings may be very low, which may lead to no exposure of such songs and is not conducive to song promotion. Summary of the Invention

[0003] Embodiments of the present application provide a method for processing song features, a server, a program product, and a storage medium, which can estimate and generate posterior features for songs without posterior features, thereby facilitating the exposure of these songs in subsequent services. The technical solutions are as follows:

[0004] In a first aspect, a method for processing song features is provided. The method includes:

[0005] Determine a target song and determine multiple candidate songs in a song library whose posterior features meet a preset condition, where the posterior features are used to characterize the situations where the candidate songs are operated by users;

[0006] According to the similarity between the audio features of each candidate song and the audio features of the target song, determine at least one candidate song that meets the similarity condition as a reference song among the multiple candidate songs; and generate the posterior features of the target song according to the posterior features of at least one reference song.

[0007] In a second aspect, a device for processing song features is provided. The device includes:

[0008] An acquisition module, configured to determine a target song and determine multiple candidate songs in a song library whose posterior features meet a preset condition, where the posterior features are used to characterize the situations where the candidate songs are operated by users; and determine at least one candidate song that meets the similarity condition as a reference song among the multiple candidate songs according to the similarity between the audio features of each candidate song and the audio features of the target song.

[0009] A generation module, configured to generate posterior features of the target song according to posterior features of at least one of the reference songs.

[0010] In a third aspect, a method for training a click-through rate prediction model is provided, the method comprising:

[0011] Obtaining a click label of a target sample song and audio features of the target sample song; the click label is used to indicate whether the target sample song has been operated on by a user;

[0012] Determining, in a song library, a reference sample song whose similarity between audio features and the audio features of the target sample song meets a similarity condition, and obtaining posterior features of the reference sample song, where the posterior features are used to indicate the situation of the reference sample song being operated on by a user;

[0013] Training a click-through rate prediction model according to the audio features of the target sample song, the audio features of the reference sample song, the posterior features of the reference sample song, the similarity between the audio features of the reference sample song and the audio features of the sample song, and the click label of the target sample song, to obtain a trained click-through rate prediction model.

[0014] In a fourth aspect, a server is provided, the terminal includes a processor and a memory, and at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the operations performed by the song feature processing method as described in the first aspect above and any possible implementation of the first aspect, or the click-through rate prediction model training method as described in the third aspect above and any possible implementation of the first aspect.

[0015] In a fifth aspect, a computer-readable storage medium is provided, and at least one instruction is stored in the storage medium, and the instruction is loaded and executed by a processor to implement the operations performed by the song feature processing method as described in the first aspect above and any possible implementation of the first aspect, or the click-through rate prediction model training method as described in the third aspect above and any possible implementation of the first aspect.

[0016] In a sixth aspect, a computer program product is provided, and at least one instruction is stored in the computer program product, and the instruction is loaded and executed by a processor to implement the operations performed by the song feature processing method as described in the first aspect above and any possible implementation of the first aspect, or the click-through rate prediction model training method as described in the third aspect above and any possible implementation of the first aspect.

[0017] The beneficial effects brought by the technical solution provided in this application are:

[0018] In the technical solution provided by this application, a reference song whose audio features match those of the target song and has posterior features is selected from the song library. Furthermore, based on the posterior features of the reference song, the posterior features of the target song are automatically generated, making up for the problem that the target song itself does not have posterior features, which is beneficial to increasing the exposure of the target song in subsequent business scenarios such as song recommendation and song ranking. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 is a flowchart of a method for processing song features provided by an embodiment of this application;

[0021] Figure 2 is a schematic diagram of posterior feature generation provided by an embodiment of this application;

[0022] Figure 3 is a schematic structural diagram of a device for processing song features provided by an embodiment of this application;

[0023] Figure 4 is a schematic structural diagram of a terminal provided by an embodiment of this application;

[0024] Figure 5 is a schematic structural diagram of a server provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.

[0026] The embodiments of this application provide a method for processing song features. This method can be implemented by a computing device. Among them, the computing device can be a terminal or a server. The terminal can be a user-side device such as a mobile phone, a tablet computer, a laptop computer, or a desktop computer. The server can be a single server or a server cluster.

[0027] In music applications, there are usually various music charts. The sorting of songs in these music charts is usually based on a combination of posterior features and prior features of the songs. Some popular songs or songs that have been on the market for a long time may be frequently operated by users, giving them posterior features. Such songs will rank very high in the chart. For some newly released songs or unpopular songs, it is very likely that they have not been operated by users, resulting in no posterior features. Such songs will rank very low in the chart or even not appear in the chart at all. In this way, it will lead to fewer users paying attention to these songs, forming a vicious cycle and discouraging the enthusiasm of creators. To address this issue, in related technologies, operators adjust the sorting of these newly released songs or unpopular songs based on experience to enable them to obtain a relatively high ranking. However, this is overly dependent on the preferences of the operators themselves, resulting in high labor costs and inaccurate adjustment results due to the involvement of human factors.

[0028] Based on the above problems, the embodiments of the present application provide a method for processing song features. In this method, a computing device obtains the audio features of a target song (such as an unpopular song or a newly released song), and matches reference songs with similar audio features to the target song in the song library. Furthermore, based on the posterior features of the reference songs, the posterior features of the target song are automatically generated, making up for the problem that the target song itself has no posterior features, which is beneficial to increasing the exposure of the target song in subsequent business scenarios such as song recommendation and song ranking.

[0029] The following describes the method for processing song features provided by the embodiments of the present application in conjunction with the accompanying drawings. Refer to Figure 1 , the processing of this method may include the following steps:

[0030] Step 101: Determine the target song and obtain the audio features of the target song.

[0031] In implementation, the target song can be any song that has not been operated by users or has been operated by users less frequently (such as the number of operations is below a certain threshold), such as a newly released song, an unpopular song, etc.

[0032] In the case where posterior features of a target song need to be generated, a computing device first extracts features of the target song to obtain audio features of the target song. Among them, the feature extraction method may include one or more of the MFCC (Mel-scale Frequency Cepstral Coefficients) method, ZCR (Zero-Crossing Rate) method, STFT (Short-Time Fourier Transform), neural network model, and other methods. In a specific implementation, multiple feature extraction methods can be used to extract the audio features of the same target song respectively. For example, the MFCC, STFT, and neural network model are respectively used as three feature extraction methods to extract the audio features of the target song, and the audio features 1, audio features 2, and audio features 3 of the target song are obtained.

[0033] Step 102: Determine multiple candidate songs whose posterior features meet the preset conditions from the song library, and obtain the audio features of each candidate song.

[0034] Among them, the posterior features are used to characterize the situation (such as the number of operations, the frequency of operations, etc.) of the candidate song being operated by the user (such as clicking, commenting, playing, favoriting, sharing, etc.). It can be understood that the posterior features meeting the preset conditions means that the situation of the candidate song being operated by the user meets the preset conditions, such as the number of times the candidate song is operated by the user reaches the number threshold, the frequency of the candidate song being operated by the user reaches the frequency threshold, and so on.

[0035] Feature extraction is performed on the candidate songs to obtain the audio features of each candidate song. Among them, the feature extraction method may include one or more of the MFCC method, ZCR method, STFT, neural network model, and other methods. It should be noted here that the method for extracting features from the candidate songs needs to be the same as the method for extracting features from the target song. For example, in the case where the MFCC, STFT, and neural network model are used as three feature extraction methods to extract the audio features of the target song, each candidate song also needs to be extracted using the MFCC, STFT, and neural network model as three feature extraction methods. If the candidate songs include Song 1 and Song 2, then the above three feature extraction methods are respectively used to extract the audio features 11, audio features 21, and audio features 31 of Song 1, and the audio features 12, audio features 22, and audio features 32 of Song 2 are obtained by extracting features from Song 2.

[0036] Step 103: Determine at least one candidate song that meets the similarity condition as a reference song from the multiple candidate songs according to the similarity between the audio features of each candidate song and the audio features of the target song.

[0037] In implementation, for each candidate song, the similarity between the audio features of the candidate song and the audio features of the target song is calculated. It should be noted that, in the case where multiple feature extraction methods are used for feature extraction, the process of calculating the similarity can be to calculate the similarity of the audio features obtained by the same feature extraction method, and then the multiple similarities are combined to obtain the similarity obtained in step 103, which is explained below with reference to an example:

[0038] For example, the three feature extraction methods of MFCC, STFT, and neural network model are used to extract features of the target song, and audio feature 1, audio feature 2, and audio feature 3 of the target song are obtained; the three feature extraction methods of MFCC, STFT, and neural network model are used to extract features of song 1, and audio feature 11, audio feature 21, and audio feature 31 of song 1 are obtained. The method for calculating the similarity between the audio features of the target song and the audio features of song 1 can be: calculating similarity 1 between audio feature 1 and audio feature 11, calculating similarity 2 between audio feature 2 and audio feature 21, calculating similarity 3 between audio feature 3 and audio feature 31, and performing weighted averaging of similarity 1, similarity 2, and similarity 3 to obtain the similarity between the audio features of the target song and the audio features of song 1. Among them, the weights used in weighted averaging can be configured by relevant personnel according to actual needs, and the embodiments of the present application do not limit this.

[0039] There may be many similarity conditions, and several of them are described below as examples: Similarity condition 1: Determine the N songs with the highest similarity as reference songs, where N is a positive integer, such as N=3. Similarity condition 2: Determine songs whose similarity reaches a threshold value as reference songs.

[0040] Step 104: Generate a posteriori features of the target song based on the posteriori features of at least one reference song.

[0041] In implementation, the music library may record a posteriori features of songs, and after determining the reference song in step 103, the a posteriori features of the reference song may be further obtained. The a posteriori features may be vectors for characterizing the situations in which the song is operated by the user.

[0042] There are many methods for generating the posterior features of the target song based on the posterior features of the reference song, and several of them are described below as examples:

[0043] Method 1: Calculate the mean of the posterior features of the reference song as the posterior features of the target song. For example, if the reference songs are song 1 and song 2, calculate the mean of the posterior feature 1 of song 1 and the posterior feature 2 of song 2 as the posterior features of the target song.

[0044] Method 2: Based on the similarity between the audio features of the reference song and the target song, perform weighted averaging on the posterior features of the reference song to obtain the posterior features of the target song.

[0045] For example, if the reference songs are Song 1 and Song 2, the similarity between the audio features of Song 1 and the target song is Similarity A, and the similarity between the audio features of Song 2 and the target song is Similarity B. Then, use Similarity A as the weight for the posterior feature 1 of Song 1, and use Similarity B as the weight for the posterior feature 2 of Song 2. Then, perform weighted averaging on the posterior feature 1 and the posterior feature 2 to obtain the posterior features of the target song.

[0046] Method 3: Compose the similarity between the audio features of the reference song and the target song into a similarity vector feature, and input the composed similarity vector feature and the posterior features of the reference song into a feature fusion network to obtain the posterior features of the target song.

[0047] For example Figure 2 As shown, if the reference songs are Song 1 and Song 2, the similarity between the audio features of Song 1 and the target song is Similarity A, and the similarity between the audio features of Song 2 and the target song is Similarity B. Then, compose Similarity A and Similarity B into a similarity vector, and input the similarity vector, the posterior feature of Song 1 (i.e., posterior feature 1), and the posterior feature of Song 2 (i.e., posterior feature 2) into the feature fusion network. The feature fusion network outputs the posterior features of the target song.

[0048] In a possible implementation, the posterior features of the song can be used to predict the click-through rate of the song. The corresponding processing method can be as follows: The above-mentioned feature fusion network is a network layer in the click-through rate prediction model. For example, it can be a DNN (Deep Neural Network). The click-through rate prediction model also includes a prediction network, and the prediction network can be a fully connected layer, etc. This application embodiment does not make any limitations in this regard. After the feature fusion network outputs the posterior features of the target song, input the posterior features of the target song into the prediction network to obtain the estimated click-through rate of the target song.

[0049] In another possible implementation, in order to improve the accuracy of click-through rate estimation, it is also possible to combine the prior features of the song for prediction. Correspondingly, the processing can be as follows: Concatenate the posterior features of the target song and the prior features of the target song to obtain a fused feature, and input the fused feature into the prediction network to obtain the estimated click-through rate of the target song. Among them, the prior features of the target song are used to represent the attributes of the target song itself, such as song name, singer, music style, genre, etc.

[0050] The training of the above click-through rate prediction model will be described below:

[0051] S1. Obtain the click label of the target sample song.

[0052] In implementation, the target sample song can be selected from the song library, and the click label of the target sample song can be obtained. If the target sample song has been operated by the user, its click label is 1; if the target sample song has not been operated by the user, its initial click label is 0. In the case where the initial click label is 0, the click label can be compensated and reconstructed. The compensation and reconstruction of the click label are described below:

[0053] Select at least one song in the song library whose audio features satisfy the similarity condition with the audio features of the target sample song, and determine the click label of the target sample song according to the click labels of the at least one selected song. There can be multiple determination methods, and several of them are listed below for description:

[0054] Method 1: Select the song with the highest similarity (i.e., the reference song with the highest similarity between the audio features and the audio features of the target sample song among the above reference songs), and determine the click label of this song as the click label of the target sample song.

[0055] Method 2: Select the M songs with the highest similarity (which can include some or all of the above reference songs), where M is an integer greater than 1. According to the similarities between these M songs and the target sample song respectively, the click labels of the M songs are weighted and averaged to obtain the click label of the target sample song.

[0056] Method 3: Select the songs whose similarities reach the threshold (which can include some or all of the above reference songs). According to the similarities between the songs whose similarities reach the threshold and the target sample song respectively, the click labels of these songs are weighted and averaged to obtain the click label of the target sample song.

[0057] In the case where the target sample song has not been operated on after it is launched, the initial click label of the target sample song is 0, indicating that the target sample song is a niche song or a newly launched song. Then, if the click-through rate prediction model is trained using this initial click label, the trained click-through rate prediction model will predict a very low click-through rate when predicting the click-through rate of unoperated songs such as niche songs or newly launched songs in the future. In this way, the ranking of such songs in song ranking scenarios such as song rough ranking, fine ranking, and re-ranking will be very low, seriously affecting the exposure of such songs. By using the above method to compensate and reconstruct the click label of the target sample song in this case, the click label after compensation and reconstruction is not 0. Based on the click-through rate prediction model trained thereby, even when predicting the click-through rate of niche songs or newly launched songs, the predicted click-through rate will not be too low, so that the ranking of such songs in song ranking scenarios such as song rough ranking, fine ranking, and re-ranking will not be very low, and the exposure of such songs is improved to a certain extent.

[0058] S2. Determine a reference sample song in the song library whose similarity between the audio features and the audio features of the target sample song meets the similarity condition, obtain the posterior features of the reference sample song, and train the click-through rate prediction model according to the audio features of the target sample song, the audio features of the reference sample song, the posterior features of the reference sample song, the similarity between the audio features of the reference sample song and the audio features of the target sample song, and the click label of the target sample song.

[0059] In implementation, the computing device inputs the audio features of the sample song, the audio features of the reference sample song, the posterior features of the reference sample song, and the similarity between the audio features of the reference sample song and the audio features of the sample song into the click-through rate prediction model. The click-through rate prediction model outputs an estimated click-through rate, calculates the loss value between the estimated click-through rate and the click label, and adjusts the parameters of the click-through rate prediction model according to the loss value. Among them, the loss value can be cross entropy. Correspondingly, the calculation method of the loss value can be as shown in the following formula (1):

[0060]

[0061] where loss is the loss value, y is the click label, represents the estimated click-through rate.

[0062] All the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present disclosure, which will not be elaborated one by one here.

[0063] In the technical solution provided by the embodiments of the present application, the audio features of a target song (such as an obscure song or a newly released song) are obtained, and reference songs similar to the audio features of the target song are matched in the song library. Furthermore, based on the posterior features of the reference songs, the posterior features of the target song are automatically generated, making up for the problem that the target song itself has no posterior features, which is beneficial to increasing the exposure of the target song in subsequent business scenarios such as song recommendation and song ranking.

[0064] The technical solution provided by the embodiments of the present application can be applied to song ranking scenarios such as song rough ranking, fine ranking, and re-ranking. In the ranking model, the above-mentioned solution for generating posterior features is introduced. When ranking songs, user features, user search requests, the list of songs to be ranked, the audio features of the songs in the list, and the posterior features are input into the ranking model. For obscure songs or new songs without posterior features, the ranking model can use the technical solution provided by the embodiments of the present application to generate their posterior features. The ranking model outputs the ranking scores of the songs in the list. In addition, the newly generated posterior features can also be output for subsequent use.

[0065] Based on the same technical concept, the embodiments of the present application also provide a processing device for song features. The device can be a terminal, such as Figure 3 As shown, the device includes an acquisition module 510 and a generation module 520, where:

[0066] The acquisition module 510 is configured to determine a target song and multiple candidate songs whose posterior features meet preset conditions from the song library, where the posterior features are used to characterize the situation of the candidate songs being operated by users; according to the similarity between the audio features of each candidate song and the audio features of the target song, at least one candidate song that meets the similarity condition is determined from the multiple candidate songs as a reference song.

[0067] The generation module 520 is configured to generate the posterior features of the target song according to the posterior features of at least one reference song.

[0068] In a possible implementation, the generation module 520 is configured to: generate the posterior features of the target song according to the similarity between the audio features of each reference song and the audio features of the target song and the posterior features of each reference song.

[0069] In a possible implementation, the generation module 520 is configured to: use the similarity between the audio features of each reference song and the audio features of the target song as a weight, and perform weighted averaging on the posterior features of at least one reference song to obtain the posterior features of the target song.

[0070] In a possible implementation, the generating module 520 is configured to: form a similarity vector feature from the similarities between the audio features of at least one of the reference songs and the audio features of the target song;

[0071] input the similarity vector feature and the posterior features of at least one of the reference songs into a feature fusion network of a pre-trained click-through rate prediction model to obtain the posterior features of the target song.

[0072] In a possible implementation, the apparatus includes a prediction module configured to: concatenate the posterior features of the target song and the prior features of the target song to obtain a fused feature of the target song, where the prior features of the target song are used to characterize the attributes of the target song itself; and input the fused feature of the target song into a prediction network of the click-through rate prediction model to obtain a predicted click-through rate of the target song.

[0073] In a possible implementation, the generating module 520 is configured to: calculate the mean of the posterior features of at least one of the reference songs as the posterior features of the target song.

[0074] In a possible implementation, the obtaining module 510 is configured to: determine the N candidate songs with the highest corresponding similarity in the multiple candidate songs as reference songs, where N is a positive integer; or,

[0075] determine the candidate songs with the corresponding similarity reaching a preset similarity threshold in the multiple candidate songs as reference songs.

[0076] In the technical solution provided in the embodiments of the present application, a computing device obtains the audio features of a target song (such as a niche song or a newly released song), and matches reference songs similar to the audio features of the target song in a song library. Furthermore, based on the posterior features of the reference songs, the posterior features of the target song are automatically generated, making up for the problem that the target song itself has no posterior features, which is beneficial to increasing the exposure of the target song in subsequent business scenarios such as song recommendation and song ranking.

[0077] It should be noted that: when the song feature processing apparatus provided in the above embodiments generates song features, only the above-mentioned functional modules are used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computing device is divided into different functional modules to complete all or part of the functions described above. In addition, the song feature processing apparatus provided in the above embodiments and the embodiments of the song feature processing method belong to the same concept, and the specific implementation process is detailed in the method embodiments and will not be repeated here.

[0078] Figure 4The structural block diagram of a terminal 600 provided by an exemplary embodiment of the present application is shown. The terminal 600 may be a portable mobile terminal, such as: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer or a desktop computer. The terminal 600 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.

[0079] Generally, the terminal 600 includes: a processor 601 and a memory 602.

[0080] The processor 601 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). The processor 601 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 601 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 601 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0081] The memory 602 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 602 may also include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices, flash storage devices.

[0082] In some embodiments, the terminal 600 may further optionally include: a peripheral device interface 603 and at least one peripheral device. The processor 601, the memory 602, and the peripheral device interface 603 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 603 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 604, a display screen 605, a camera assembly 606, an audio circuit 607, a positioning assembly 608, and a power supply 609.

[0083] In some embodiments, the terminal 600 further includes one or more sensors 610. The one or more sensors 610 include, but are not limited to: an acceleration sensor 611, a gyroscope sensor 612, a pressure sensor 613, a fingerprint sensor 614, an optical sensor 615, and a proximity sensor 616.

[0084] Those skilled in the art can understand that Figure 4 the structure shown in

[0085] Figure 5 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1000 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 1001 and one or more memories 1002. Among them, at least one instruction is stored in the memory 1002, and the at least one instruction is loaded and executed by the processor 1001 to implement the methods provided by the above-mentioned various method embodiments. Of course, this computing device may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. This computing device may further include other components for implementing the functions of the device, which will not be elaborated here.

[0086] In an exemplary embodiment, a computer-readable storage medium is also provided. For example, a memory including instructions, and the above instructions can be executed by a processor in the terminal to complete the method for processing song features in the above embodiments. The computer-readable storage medium may be non-transitory. For example, the computer-readable storage medium may be a ROM (Read-Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0087] In an exemplary embodiment, a computer program product is further provided. At least one instruction is stored in the computer program product, and the instruction is loaded and executed by a processor to implement the operations performed by the method for processing song features in the above embodiment.

[0088] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the song audio, click tags, posterior features, etc. involved in this application are all obtained under full authorization.

[0089] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiment can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a disk, an optical disc, or the like.

[0090] In this application, terms such as "first" and "second" are used to distinguish between identical or similar items with basically the same function and role. It should be understood that there is no logical or temporal dependence between "first" and "second", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms such as first and second to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of various examples, the first threshold can be referred to as the second threshold, and similarly, the second threshold can be referred to as the first threshold. The first threshold and the second threshold can both be collectively referred to as the threshold, and in some cases, they can be separate and different thresholds.

[0091] The meaning of the term "at least one" in this application refers to one or more, and the meaning of the term "multiple" in this application refers to two or more.

[0092] The above description is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A method for processing song features, characterized in that: The method comprises: Determine a target song, and determine from the music library a plurality of candidate songs whose a posteriori features meet preset conditions, wherein the a posteriori features are used to characterize the situation in which the candidate songs are operated by the user; According to the similarity between the audio features of each candidate song and the audio features of the target song, determining at least one candidate song that meets the similarity condition from among the multiple candidate songs as a reference song; and, The posterior features of the target song are generated according to the posterior features of at least one of the reference songs.

2. The method according to claim 1, characterized in that: The step of generating the posterior feature of the target song according to the posterior feature of at least one of the reference songs comprises: The posterior features of the target song are generated according to the similarity between the audio features of each of the reference songs and the audio features of the target song and the posterior features of each of the reference songs.

3. The method according to claim 2, characterized in that The generating the posterior features of the target song according to the similarity between the audio features of each of the reference songs and the audio features of the target song and the posterior features of each of the reference songs comprises: The similarity between the audio features of each reference song and the audio features of the target song is used as a weight, and the posterior features of at least one reference song are weighted averaged to obtain the posterior features of the target song.

4. The method according to claim 2, characterized in that: The generating the posterior features of the target song according to the similarity between the audio features of each of the reference songs and the audio features of the target song and the posterior features of each of the reference songs comprises: Combining similarity vector features with similarities between the audio features of at least one of the reference songs and the audio features of the target song; The similarity vector features and the posterior features of at least one of the reference songs are input into a feature fusion network of a pre-trained click rate prediction model to obtain the posterior features of the target song.

5. The method according to claim 4, characterized in that The method further comprises: splicing the posterior features of the target song and the prior features of the target song to obtain fusion features of the target song, wherein the prior features of the target song are used to characterize the properties of the target song itself; and, The fusion features of the target song are input into the prediction network of the click rate prediction model to obtain the estimated click rate of the target song.

6. The method according to claim 1, characterized in that The step of generating the posterior feature of the target song according to the posterior feature of at least one of the reference songs comprises: Calculate the mean of the posterior features of at least one of the reference songs as the posterior features of the target song.

7. The method according to any one of claims 1 to 6, characterized in that The step of determining at least one candidate song satisfying a similarity condition among the plurality of candidate songs as a reference song comprises: Determine N candidate songs with the highest corresponding similarity among the plurality of candidate songs as reference songs, where N is a positive integer; or, Determine, among the plurality of candidate songs, candidate songs whose corresponding similarities reach a preset similarity threshold as reference songs.

8. A click rate prediction model training method, characterized in that: The method comprises: Obtaining a click tag of a target sample song and an audio feature of the target sample song; the click tag is used to indicate whether the target sample song has been operated by a user; Determine in the music library a reference sample song whose audio features meet a similarity condition with the audio features of the target sample song, and obtain a posteriori features of the reference sample song, wherein the a posteriori features are used to characterize a situation in which the reference sample song is operated by a user; A click rate prediction model is trained according to the audio features of the target sample song, the audio features of the reference sample song, the posterior features of the reference sample song, the similarity between the audio features of the reference sample song and the audio features of the sample song, and the click label of the target sample song to obtain a trained click rate prediction model.

9. The method according to claim 8, characterized in that The step of obtaining the click tags of the target sample songs includes: If the click tag of the target sample song indicates that the target sample song has been operated by the user, determining that the click tag of the target sample song is 1; If the click tag of the target sample song indicates that the target sample song has not been operated by the user, the click tag of the target sample song is determined based on the click tag of the reference sample song.

10. A server, characterized in that: The computing device includes a processor and a memory, wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the operations performed by the song feature processing method described in any one of claims 1 to 7 or the operations performed by the click-through rate prediction model training method described in any one of claims 8-9.

11. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, which is loaded and executed by the processor to implement the operations performed by the song feature processing method described in any one of claims 1 to 7 or the operations performed by the click-through rate prediction model training method described in any one of claims 8-9.

12. A computer program product, characterized in that The computer program product stores at least one instruction, which is loaded and executed by a processor to implement the operations performed by the song feature processing method described in any one of claims 1 to 7 or the operations performed by the click-through rate prediction model training method described in any one of claims 8-9.