Recommendation method and device, storage medium and computing equipment

By combining multi-task model with multimedia objects and user characteristics to calculate the recommendation score of the playback clip, the problem of poor recommendation effect in the existing technology is solved, the accurate recommendation of the playback clip is achieved, and the conversion rate of users to enable preset permissions is improved.

CN120386920APending Publication Date: 2025-07-29HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411776023.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The recommendation scheme for playing clips of multimedia objects in the prior art lacks targetedness, resulting in poor recommendation results and inability to effectively increase the conversion rate of users to enable preset permissions.

Method used

Through the multi-task model, the clip characteristics and user characteristics of the multimedia object are comprehensively considered, and the probability of each playback clip attracting the user to enable preset permissions is calculated, and quantified into the recommended score. The higher the recommended score, the higher the probability of the user opening the preset permissions.

Benefits of technology

It realizes accurate recommendations for playing clips, improves the conversion rate of preset permissions, and attracts users to activate preset permissions for a better service experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386920A_ABST
    Figure CN120386920A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a recommendation method and device, a storage medium and computing equipment. Comprising the following steps: acquiring fragment characteristics related to each playing fragment of a multimedia object and user characteristics related to a to-be-recommended user; inputting the fragment features and the user features into a pre-trained multi-task model, and calculating a recommendation score of the playing fragment based on the fragment features and the user features; wherein the recommendation score represents a probability that the user opens a preset permission after the playing fragment is played; and sorting the recommendation scores of the multimedia fragments output by the multi-task model, and recommending the target playing fragment corresponding to the highest recommendation score to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology. More specifically, embodiments of the present disclosure relate to a recommendation method, apparatus, storage medium, and computing device. Background Art

[0002] This section aims to provide background or context for embodiments of the present disclosure. The descriptions herein are not admitted to be prior art merely because they are included in this section.

[0003] With the development of the Internet, people increasingly enjoy content services provided by the Internet through websites or applications, such as online listening, purchasing, streaming, and downloading multimedia objects.

[0004] In addition, in order to obtain a better service experience, users can also enable preset permissions. An account with preset permissions enabled can obtain additional service functions compared to an account without preset permissions enabled. For example, when playing some multimedia objects, an account without preset permissions enabled can only play content segments and cannot play the entire content completely; while an account with preset permissions enabled can play the entire content.

[0005] Since enabling preset permissions can bring benefits to the operators of multimedia objects, the operators expect more users to enable preset permissions. To this end, in the related art, recommendations are used to guide users to enable preset permissions. Specifically, the operators will recommend playback segments of some multimedia objects to users, hoping to attract users through the playback segments; if users are interested in the playback segments, then users will want to play the complete content of the multimedia objects, and thus will have the motivation to enable preset permissions.

[0006] Generally, the complete content of a multimedia object can generate multiple playback segments, and the attractiveness of each playback segment to different users varies. However, existing recommendation schemes generally adopt a "broadcast net" method and do not consider the matching degree between the playback segments and users, resulting in poor recommendation effects.

[0007] Therefore, how to accurately recommend playback segments suitable for users' preferences to users has become an urgent problem to be solved. Summary of the Invention

[0008] In a first aspect of the embodiments of the present disclosure, a recommendation method is provided. The method includes:

[0009] Obtaining segment features related to each playback segment of a multimedia object and user features related to a user to be recommended;

[0010] Input the segment feature and user feature into a pre-trained multi-task model, and calculate the recommendation score of the playing segment based on the segment feature and user feature; wherein, the recommendation score represents the probability that the user enables a preset permission after playing the playing segment.

[0011] Sort the recommendation scores of each multimedia segment output by the multi-task model, and recommend the target playing segment corresponding to the highest recommendation score to the user.

[0012] Optionally, the multi-task model includes a multi-layer perceptron for receiving input segment features and user features, and a main task module and a secondary task module respectively connected to the output of the multi-layer perceptron.

[0013] Among them, the secondary task module includes a first tower network corresponding to the secondary task, the main task module includes a second tower network corresponding to the main task, the main task module further includes a self-attention module for performing attention calculation based on the outputs of the first tower network and the second tower network, and an activation function for predicting the recommendation scores corresponding to each playing segment based on the output of the self-attention module.

[0014] Optionally, the step of inputting the segment feature and user feature into a pre-trained multi-task model and calculating the recommendation score of the playing segment based on the segment feature and user feature includes:

[0015] Concatenate the segment feature and user feature, and input the obtained concatenated feature into the multi-layer perceptron in the pre-trained multi-task model to calculate the shared feature vector in the concatenated feature.

[0016] Input the feature vector into the second tower network of the main task module and the first tower network of the secondary task module respectively for calculation.

[0017] Further input the outputs of the first tower network and the second tower network into the self-attention module for attention calculation.

[0018] Further input the output of the self-attention module into the activation function for regression calculation to obtain the recommendation score corresponding to the playing segment.

[0019] Optionally, both the first tower network and the second tower network include multi-layer perceptrons using fully connected layers.

[0020] Optionally, the multi-task model is iteratively trained in the following manner until the multi-task model converges:

[0021] Obtain a sample set for training; wherein, the sample set includes a number of training samples, and the training samples include sample features and true labels corresponding to pre-annotated recommendation scores.

[0022] Input the sample set into the multi-task model for supervised training, calculate the loss function of the training samples according to forward propagation, calculate the gradient of the loss function according to backward propagation, and update the model parameters of the multi-task model according to the gradient; wherein, the model parameters at least include the parameter weights in the multi-layer perceptron, the first tower network, the second tower network, the self-attention module, and the activation function.

[0023] Optionally, the loss function includes a cross-entropy loss function

[0024] Wherein, M represents the total number of the main task and the secondary tasks, N represents the total number of training samples, y i represents the true label of the i-th training sample, and p i represents the recommendation score of the i-th training sample predicted by the activation function in the main task module.

[0025] Optionally, the user features include user attribute features related to user personal information, and / or

[0026] user behavior features related to the user's operation behavior in the multimedia platform.

[0027] Optionally, the user behavior features include: whether the user has historically enabled a preset permission.

[0028] Optionally, the segment features include the playback position of the playback segment in the multimedia object, and / or

[0029] the proportion of the playback segment containing the key content of the multimedia object.

[0030] Optionally, the user to be recommended includes users who do not have a preset permission.

[0031] Optionally, the preset permission includes a membership permission, and the membership permission allows playing the complete multimedia object.

[0032] Optionally, the multimedia object includes a music song, and the playback segment includes a trial listening segment of the music song.

[0033] In the second aspect of the embodiments of the present disclosure, a recommendation device is provided, and the device includes:

[0034] An acquisition unit that acquires segment features related to each playback segment of a multimedia object, and user features related to a user to be recommended;

[0035] A computing unit that inputs the segment feature and the user feature into a pre-trained multi-task model, and calculates a recommendation score for the playback segment based on the segment feature and the user feature; wherein, the recommendation score represents the probability that the user enables a preset permission after playing the playback segment.

[0036] A recommendation unit that sorts the recommendation scores of each multimedia segment output by the multi-task model, and recommends the target playback segment corresponding to the highest recommendation score to the user.

[0037] Optionally, the multi-task model includes a multi-layer perceptron that receives input segment features and user features, and a main task module and a secondary task module respectively connected to the output of the multi-layer perceptron.

[0038] Wherein, the secondary task module includes a first tower network corresponding to the secondary task, the main task module includes a second tower network corresponding to the main task, the main task module further includes a self-attention module that performs attention calculation based on the outputs of the first tower network and the second tower network, and an activation function that predicts the recommendation scores corresponding to each playback segment based on the output of the self-attention module.

[0039] Optionally, the computing unit is further configured to splice the segment feature and the user feature, input the obtained spliced feature into the multi-layer perceptron in the pre-trained multi-task model, and calculate the feature vector shared in the spliced feature; input the feature vector into the second tower network of the main task module and the first tower network of the secondary task module respectively for calculation; further input the outputs of the first tower network and the second tower network into the self-attention module for attention calculation; further input the output of the self-attention module into the activation function for regression calculation to obtain the recommendation score corresponding to the playback segment.

[0040] Optionally, both the first tower network and the second tower network include a multi-layer perceptron using fully connected layers.

[0041] Optionally, the multi-task model is iteratively trained by a training unit until the multi-task model converges:

[0042] The training unit obtains a sample set for training; wherein, the sample set includes a number of training samples, and each training sample includes sample features and a true label corresponding to a pre-annotated recommendation score; the sample set is input into the multi-task model for supervised training, the loss function of the training samples is calculated according to forward propagation, the gradient of the loss function is calculated according to backward propagation, and the model parameters of the multi-task model are updated according to the gradient; wherein, the model parameters at least include the parameter weights in the multi-layer perceptron, the first tower network, the second tower network, the self-attention module, and the activation function.

[0043] Optionally, the loss function includes a cross-entropy loss function

[0044] where M represents the total number of the main task and the sub-tasks, N represents the total number of training samples, y i represents the true label of the i-th training sample, and p i represents the recommendation score of the i-th training sample predicted by the activation function in the main task module.

[0045] Optionally, the user features include user attribute features related to the user's personal information, and / or

[0046] user behavior features related to the user's operation behavior in the multimedia platform.

[0047] Optionally, the user behavior features include whether the user has historically enabled a preset permission.

[0048] Optionally, the segment features include the playing position of the playing segment in the multimedia object, and / or

[0049] the proportion of the playing segment that contains the key content of the multimedia object.

[0050] Optionally, the user to be recommended includes a user who does not have the preset permission.

[0051] Optionally, the preset permission includes a membership permission, and the membership permission allows playing the complete multimedia object.

[0052] Optionally, the multimedia object includes a music song, and the playing segment includes a trial listening segment of the music song.

[0053] In the third aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, including:

[0054] When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device can execute the recommendation method as described in any one of the preceding items.

[0055] In a fourth aspect of the embodiments of the present disclosure, a computing device is provided, including:

[0056] a processor;

[0057] a memory for storing executable instructions of the processor;

[0058] wherein the processor is configured to execute the executable instructions to implement the recommendation method as described in any of the preceding items.

[0059] According to the recommendation solution provided by the embodiments of the present disclosure, the multi-task model synthesizes the segment features of each playback segment and the user features of the user to be recommended, calculates the probability that each playback segment attracts the user to enable a preset permission, and quantifies it as the recommendation score of each playback segment; since the higher the recommendation score indicates the higher the probability that the user enables the preset permission, the target playback segment corresponding to the highest recommendation score can be recommended to the user, thereby attracting the user to activate the preset permission. In this way, accurate recommendation of playback segments can be achieved, and the conversion rate of the preset permission can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] By referring to the accompanying drawings and reading the following detailed description, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown by way of illustration and not limitation, wherein:

[0061] Figure 1 Schematically shows a schematic diagram of a recommendation system provided by the present disclosure;

[0062] Figure 2 Schematically shows a schematic diagram of a recommendation method provided by the present disclosure;

[0063] Figure 3 Schematically shows a schematic diagram of a multi-task model provided by the present disclosure;

[0064] Figure 4 Schematically shows a schematic diagram of a main task module and a secondary task module provided by the present disclosure;

[0065] Figure 5 Schematically shows a schematic diagram of a medium provided by the present disclosure;

[0066] Figure 6 Schematically shows a schematic diagram of a recommendation device provided by the present disclosure;

[0067] Figure 7 Schematically shows a schematic diagram of a computing device provided by the present disclosure.

[0068] In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed implementation manners

[0069] The principles and spirit of the present disclosure will be described below with reference to several exemplary implementation manners. It should be understood that these implementation manners are provided only to enable those skilled in the art to better understand and then implement the present disclosure, rather than limiting the scope of the present disclosure in any way. On the contrary, these implementation manners are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.

[0070] Those skilled in the art know that the implementation manners of the present disclosure can be implemented as a system, a device, an equipment, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0071] According to the implementation manners of the present disclosure, a recommendation method, a computer-readable storage medium, a device, and a computing device are provided.

[0072] In this article, it should be understood that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0073] The principles and spirit of the present disclosure will be elaborated below with reference to several representative implementation manners of the present disclosure.

[0074] The data involved in the present disclosure can be data authorized by users or fully authorized by all parties. The collection, dissemination, use, etc. of the data all comply with the requirements of relevant national laws and regulations. The implementation manners / embodiments of the present disclosure can be combined with each other. Summary of the Invention

[0076] The present disclosure aims to provide a recommendation scheme. By means of a multi-task model, the segment features of each playback segment and the user features of the user to be recommended are comprehensively considered, the probability that each playback segment attracts the user to enable a preset permission is calculated, and it is quantified as the recommendation score of each playback segment; since the higher the recommendation score indicates the higher the probability that the user enables the preset permission, the target playback segment corresponding to the highest recommendation score can be recommended to the user, so as to attract the user to activate the preset permission. In this way, accurate recommendation of playback segments can be realized, and the conversion rate of the preset permission can be improved.

[0077] After introducing the basic principles of the present disclosure, the various non-limiting implementation manners of the present disclosure will be specifically introduced below.

[0078] Overview of Application Scenarios

[0079] First, refer to Figure 1A recommended system architecture diagram is shown. In this system architecture diagram, various network nodes can achieve information communication through the network, and then complete interaction and data processing. The system architecture diagram may include an operation server 12 that conducts data communication with one or more clients 11 via network 13, and a database 14 that can be integrated into the operation server 12 or independent of the operation server 12.

[0080] The operation server 12 may store a pre-trained multi-task model. Through the multi-task model, the operation server 12 can have the ability to recommend specific content, specifically referring to automatically generating recommendation scores for each playback segment of a multimedia object with the help of the multi-task model.

[0081] The operation server 12 may include a service platform that provides services for media objects. Among them, the multimedia objects may include videos (such as short videos, long videos), audios (such as radio stations, music songs, user-recorded audios, audiobooks, etc.), text images (such as e-books, comic books, illustrated news, blogs, microblogs, etc.), and so on.

[0082] In some application scenarios, the operation server 12 may include a dedicated service platform that is used to specifically play a certain type of multimedia object. For example, a video platform can only be used to play videos and cannot play e-books; a music platform can only play audio and cannot play videos and e-books. In some other application scenarios, the operation server 12 may include a multi-functional service platform that generally supports at least two different types of multimedia objects. Regardless of the form of the service platform, it can have the function or service of recommending multimedia objects to users.

[0083] In the system architecture diagram, each network 13 may include wired or wireless telecommunication devices, and the network devices on which the clients 11 are based can exchange data through the wired or wireless telecommunication devices. For example, each network 13 may include a local area network (“LAN”), a wide area network (“WAN”), an intranet, the Internet, a mobile phone network, a virtual private network (VPN), a cellular or other mobile communication network, Bluetooth, NFC, or any combination thereof. In the discussion of exemplary embodiments, it should be understood that the terms “data” and “information” may be used interchangeably herein to refer to text, images, audio, video, or any other form of information that may exist in a computer-based environment.

[0084] Each network device on which a client 11 is based may include a device having a communication module capable of transmitting and receiving data via a network 13. For example, each network device on which a client 11 is based may include a server, a desktop computer, a laptop computer, a tablet computer, a smart phone, a handheld computer, a personal digital assistant ("PDA"), or any other wired or wireless processor-driven device.

[0085] In Figure 1 the illustrated exemplary embodiment, the network device on which the client 11 is based may be operated by a consumer user of the multimedia object.

[0086] Users (including individuals or organizations) may use an application such as a web browser application or a stand-alone application to view, download, upload, or otherwise access a service platform or a service website via the network 13, so as to obtain services of multimedia objects.

[0087] The application of the web browser application or the stand-alone application may interact with a web server connected to the network 13 to complete data interaction.

[0088] The data / relationships that need to be read or the processing that needs to be performed involved in the interaction process may need to be obtained from the connected database 14, and the data / relationships that need to be written or the processing results involved in the interaction process may need to be written into the connected database 14.

[0089] Figure 1 In, the computing device 15, which may be in an integrated relationship or a separate relationship with the operation server 12, especially in the latter case, may generally be connected through an internal network or a private network, or may also be connected through an encrypted public network. In particular, when in an integrated relationship, a connection in the form of a more efficient and faster transmission speed internal bus may be adopted. The computing device 15, whether in an integrated relationship or a separate relationship, may directly (not shown in the figure) or access the database 14 through the operation server 12.

[0090] By appropriately programming the computing device 15, the implementation of the method in this specification can be controlled by such instructions. In particular, when in an integrated relationship, the transactions processed by the computing device 15 may be regarded as the processing of the operation server 12 without special distinction.

[0091] Exemplary Method

[0092] The following combines Figure 1 the illustrated application scenarios and refers to Figure 2A method for recommendation according to an exemplary embodiment of the present disclosure will be described. It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0093] As Figure 2 shown, the recommendation method can be applied to the aforementioned operation server and may include the following steps:

[0094] Step 210: Obtain the segment features related to each playing segment of the multimedia object and the user features related to the user to be recommended.

[0095] The multimedia object in this specification can include, as described above, but is not limited to, videos (such as short videos, long videos), audios (such as radio stations, music songs, user-recorded audios, audiobooks, etc.), text images (such as e-books, comic books, illustrated news, blogs, microblogs, etc.). Below, music songs will be used as an example for introduction.

[0096] The standard duration of a music song usually ranges from 3 to 5 minutes. However, for audition purposes, its playing segment (also called audition segment) is usually limited within 1 minute. This means that for any song, multiple audition segments of different lengths can be generated based on different start and end points, such as 30 seconds, 35 seconds, or 40 seconds. The selection of these audition segments needs to ensure the coherence of music playback. Therefore, the exact content of the audition segment is determined by its start and end time points. The start point can be set at any moment in the beginning, prelude, chorus, or verse of the song, and the end point can also be set at any music node such as the verse, chorus, or others. Based on different start and end time points, audition segments reflecting different parts of the song can be constructed.

[0097] Since the structure of each song is different, the audition segment may include the transition from the prelude to the verse of the song, or from the verse to the chorus, or even only include parts of the verse or chorus. This diversity ensures that the audition segment can show the diverse aspects of the song.

[0098] Since the song parts covered by the audition clips vary, different experiences are generated after users listen. Some audition clips only contain the prelude of the song, which may confuse users and make it difficult for them to judge whether the whole song is pleasant or in line with their personal taste. Other clips include the chorus part of the whole song, which is the most fascinating part of the song. After listening to this part, users can usually quickly decide whether they like the song. There are also some audition clips that contain a part of the verse and a small part of the chorus. This incomplete display may leave users feeling unsatisfied. Especially when the audition clip happens to touch on the user's preferences, it is more likely to arouse their interest and prompt them to activate the preset permission to enjoy the full version of the song.

[0099] However, given that a song can generate multiple audition clips and the attractiveness of each audition clip varies for different users, how to accurately recommend audition clips suitable for the user's preferences to increase the likelihood of the user activating the preset permission becomes an urgent problem to be solved.

[0100] In view of this, the embodiments of this specification match the most suitable playback clip for the user by comprehensively considering the characteristics of each playback clip and the characteristics of the user to be recommended, so as to achieve accurate recommendation of the playback clip and thereby increase the likelihood of the user activating the preset permission.

[0101] The preset permission in this specification may include membership permissions, VIP permissions, etc. The user account that activates the preset permission can obtain the complete content of the multimedia object corresponding to the playback clip.

[0102] In addition, the preset permission can also obtain a better user experience. For example, when listening to music songs, better sound quality can be enabled, when playing videos, higher-definition video quality can be switched, and when posting comments on multimedia objects, the comment content can be highlighted (such as highlighted display, special fonts, special colors, etc.).

[0103] In step 210, obtain the clip features related to each playback clip of the multimedia object and the user features related to the user to be recommended.

[0104] Among them, each playback clip needs to be segmented and generated from the complete content of the multimedia object. This segmentation can be either manual segmentation or automatic segmentation based on preset rules. The preset rules can include relatively simple rules such as sequentially intercepting playback clips according to a fixed duration (such as 30 seconds), or relatively complex rules such as the starting point of the playback clip being in the prelude part and the ending point being in the verse part, etc. Since how to generate each playback clip from the multimedia object is not the focus of the embodiments of this specification, no limitation is imposed on how to segment here.

[0105] For each playback segment of the generated multimedia object, since the playback segment, as an audio file, has a large data volume, it is difficult to directly input the playback segment into the model for processing.

[0106] Therefore, the playback segment can be feature-extracted to obtain segment features related to each playback segment of the multimedia object.

[0107] Exemplarily, the segment feature may include the playback position of the playback segment in the multimedia object and / or the proportion of the key content of the multimedia object contained in the playback segment.

[0108] Still taking a music song as an example below, for each audition segment, the playback position of the audition segment in the music song can be considered first, such as whether it starts from the starting point of the song (i.e., 0 seconds) and whether it includes the start and end parts of the chorus;

[0109] Then, the proportion of the key content contained in the audition segment can be considered. In a music song, the key content may refer to the chorus part of the song, because the chorus part is usually the most fascinating part of the song, and the more chorus content there is, the more conducive it is to attracting users to want to listen to the complete song, thus being more conducive to stimulating users to activate the preset permission.

[0110] It can be seen that the playback position of the playback segment and / or the proportion of the key content can be used as a reference basis for predicting whether a user will activate the preset permission after consuming the playback segment.

[0111] In addition, since the same song can have multiple different audition segments, and the interception rules for each segment are not the same. For example, some may follow fixed start and end times (such as 0 - 30 seconds), and some may start from the beginning of the song until the end of the chorus. To effectively incorporate these different audition segments into the multi-task model analysis, this specification also proposes a systematic numbering method to uniquely identify all audition segments. For example, three different audition segments are numbered 0, 1, and 2 respectively.

[0112] Furthermore, in order to enable the multi-task model to distinguish and process these different audition segments, the one-hot encoding method can be used to convert the number into a format recognizable by the model. For example, the 1st segment is represented as [0, 1, 0] through one-hot encoding. Such an encoding method allows the model to perform end-to-end learning and analysis on the audition segments throughout the modeling process.

[0113] In addition to the segment features, step 210 may also obtain user features related to the user to be recommended.

[0114] First, considering that the purpose of the embodiments of this specification is to attract users to activate a preset permission, the users to be recommended can refer to users who do not have the preset permission. And users who do not have the preset permission can include both users who have never activated the preset permission and users who have activated the preset permission before but the preset permission has expired.

[0115] Secondly, the user characteristics may include, but are not limited to, user attribute characteristics related to user personal information, and / or user behavior characteristics related to the user's operations on the multimedia platform.

[0116] Among them, the user attribute characteristics may refer to the basic personal attributes of the user, such as age, gender, residential area, scale level of the city, device information used by the user, such as mobile phone model, operating system version, and network operator, etc.

[0117] The user behavior characteristics may refer to various behavior data of the user on the service platform, such as the user's platform level, registration time, activity frequency, and past consumption records, etc.

[0118] In particular, considering that the purpose of the embodiments of this specification is to attract users to activate a preset permission, the user behavior characteristics may further include: whether the user has historically activated the preset permission.

[0119] Analysis of historical data shows that there are obvious differences in the tendency to activate the preset permission between users who have not historically activated the preset permission and users who have historically activated the preset permission; for example, compared with users who have not historically activated the preset permission, the probability of those users who have historically activated the preset permission to activate the preset permission again is much higher.

[0120] Therefore, incorporating the user behavior characteristic of whether the user has historically activated the preset permission into the model calculation helps the model output more accurate calculation results.

[0121] In the embodiments of this specification, in addition to the above-mentioned segment characteristics and user characteristics, what can also be input into the multi-task model may include: multimedia object characteristics, cross characteristics, multi-modal characteristics, etc. These three characteristics will be introduced separately below.

[0122] [Multimedia Object Characteristics]

[0123] The multimedia object characteristics may refer to the characteristics related to the multimedia object. Different from the segment characteristics, the multimedia object can cover multiple aspects. For example, the basic attributes of the multimedia object, taking a music song as an example, may include music style, language type, etc.

[0124] In addition, multimedia object features can also include interaction data reflecting a user's recent interactions with the multimedia object, such as whether the user has recently played the multimedia object, whether they have performed effective playback actions, and whether they have played the multimedia object in its entirety. In short, multimedia object features can depict a user's familiarity with and preference for the multimedia object. Using multimedia object features, we can gain a deeper understanding of the user's interaction patterns with the multimedia object and assess their preferences and understanding of the multimedia object.

[0125] [Cross Features]

[0126] These cross-features can be used to measure the conversion rate of a segment at different user levels. For popular multimedia objects, the conversion rate of each segment can be evaluated and further segmented by user level, such as calculating the conversion rate for different mobile operating systems or specific user groups (e.g., new or returning users). The cross-features derived from these segmentations can help us better understand the impact of different factors on conversion rate.

[0127] In contrast, for less popular or less popular multimedia objects, due to limited historical data availability, a more general evaluation approach can be employed: calculating the overall playback conversion rate of a specific type of playback segment (e.g., segment number 1) for a specific segment (e.g., a user group). This approach effectively utilizes limited data to evaluate specific playback segments of less popular multimedia objects from a macro perspective and generate corresponding cross-features.

[0128] [Multimodal features]

[0129] The multi-model features can refer to key information extracted from the playback clip. This key information can be extracted using common industry techniques. For example, the open source tool MERT (Music Understanding Model with Large-Scale Self-supervised Training) can be used to generate an audio feature vector corresponding to the audition clip and key information extracted from the audio feature vector. Multi-model features can help the model better understand the playback clip and, therefore, better match it with the user.

[0130] Step 220: Input the segment features and user features into a pre-trained multi-task model, and calculate a recommendation score for the playback segment based on the segment features and user features; wherein the recommendation score represents the probability that the user will enable preset permissions after playing the playback segment.

[0131] The multi-task model (Multi-task Learning, MTL) in this specification is a machine learning model. The multi-task model improves the performance of the model by simultaneously learning multiple related tasks. In multi-task learning, the model is trained to solve multiple tasks, usually sharing some of the same network layers in order to enhance the effects of each task by sharing knowledge.

[0132] As Figure 3 shown, the multi-task model may include a multi-layer perceptron that receives model inputs (such as the input segment features and user features shown in the previous embodiments, or may also include multimedia object features, cross features, multi-modal features, etc.), and a primary task module and a secondary task module respectively connected to the output of the multi-layer perceptron.

[0133] The structure of the multi-task model will be introduced separately below:

[0134] [Multi-layer perceptron]

[0135] In this specification, the multi-layer perceptron (MLP) may refer to a typical feedforward neural network, usually composed of multiple layers of neurons. It is one of the basic architectures of deep learning and neural networks, and is widely used in solving problems involving classification, regression, etc. Given that the embodiments of this specification need to calculate the recommendation scores of each playback segment, and this calculation belongs to a regression problem, a multi-layer perceptron can be used.

[0136] The multi-layer perceptron generally may include an input layer, one or more hidden layers, and an output layer. Among them, the input layer is used to receive the original data input into the model; each hidden layer can be composed of several neurons, and each neuron is connected to the neurons of the previous layer, and information is transmitted through weights. The neurons in the hidden layer perform non-linear transformation through activation functions (such as ReLU, Sigmoid, Tanh, etc.) to obtain complex or implicit key features in the original data. The output layer can generate the final prediction result according to the person type, and for regression problems, the output layer usually can use an activation function to predict the result.

[0137] [Primary task module and secondary task module]

[0138] The primary task module is used to process the primary task, while the secondary task module is used to process secondary tasks (which can also be called auxiliary tasks), and the secondary task module ultimately serves the primary task. Different from the primary task, there can be multiple secondary tasks.

[0139] In this specification, in view of the goal of attracting users to enable a preset permission by recommending play segments, the goal can be decomposed into several tasks. For example, the four tasks are playing the segment in full, favoriting the multimedia object corresponding to the play segment, triggering the checkout counter, and purchasing the preset permission. Among them, purchasing the preset permission can be used as the main task, while playing the segment in full, favoriting the multimedia object corresponding to the play segment, and triggering the checkout counter can be used as secondary tasks.

[0140] The secondary task module calculates the secondary task calculation results of the three secondary tasks of playing the segment in full, favoriting the multimedia object corresponding to the play segment, and then triggering the checkout counter, and inputs the secondary task calculation results as auxiliary data into the main task module to help the main task module calculate the main task calculation result of purchasing the preset permission. Finally, the main task module quantifies and outputs the final recommendation score corresponding to the play segment based on the main task calculation result. This recommendation score can represent the probability that the user enables the preset permission after playing this play segment.

[0141] Furthermore, please refer to Figure 4 the schematic diagrams of the main task module and the secondary task module shown. As Figure 4 shown, both the main task and the secondary tasks can have a tower network (tower) dedicated to processing tasks. That is, the secondary task module includes a first tower network corresponding to the secondary task, the main task module includes a second tower network corresponding to the main task, the main task module also includes a self-attention module that performs attention calculation based on the outputs of the first tower network and the second tower network, and an activation function that predicts the recommendation scores corresponding to each play segment based on the output of the self-attention module.

[0142] It should be noted that since there can be multiple secondary tasks at the same time, in fact, each secondary task module can include a first tower network corresponding one-to-one to the secondary task. For example, when there are three secondary tasks, the secondary task module can include three parallel first tower networks, and each first tower network is used to process its respective secondary task. In addition, each first tower network and the second tower network can have the same network structure. For example, both can adopt a multi-layer perceptron with fully connected layers.

[0143] To prevent the model from overfitting, a dropout operation can be added to the fully connected layer, and to maintain parameter stability, a normalization operation such as Normalization can also be added after the fully connected layer.

[0144] In addition, the self-attention module in the main task module can learn the information of multiple tasks based on the attention mechanism, such as Figure 4As shown, the input of the self-attention module is the outputs of the first tower network and the second tower network. That is to say, the self-attention module can comprehensively calculate the attention based on the outputs of the first tower network and the second tower network, so as to learn the relevant information of the main task and the secondary task simultaneously.

[0145] Exemplarily, the self-attention module can perform attention calculation using the following formula:

[0146]

[0147] where {t1, t2,..., t M} represents the outputs of M tower networks corresponding to the main task and the secondary task, and M is the total number of the main task and the secondary task; w u represents the weights corresponding to the outputs of each tower network during attention calculation; v u represents the vector information obtained by dot-multiplying u after being transformed by h2 and h3; k represents the feature length after being transformed by h i ; h i represents a learnable projection mapping.

[0148] Based on the model structure of the multi-task model introduced above, for step 220 in this specification, it may further include:

[0149] Concatenate the segment feature and the user feature, and input the obtained concatenated feature into the multi-layer perceptron in the pre-trained multi-task model to calculate the shared feature vector in the concatenated feature;

[0150] Input the feature vector into the second tower network of the main task module and the first tower network of the secondary task module respectively for calculation;

[0151] Further input the outputs of the first tower network and the second tower network into the self-attention module for attention calculation;

[0152] Further input the output of the self-attention module into the activation function for regression calculation to obtain the recommendation score corresponding to the playing segment.

[0153] Please continue to refer to Figure 4 As shown, for the segment feature and the user feature obtained in step 210, before inputting into the model, first concatenate the segment feature and the user feature, and then input the concatenated feature obtained by concatenation into the model; the feature vectors obtained by multiple perceptron calculations of the concatenated feature can be input into the secondary task module and the main task module respectively. Among them, in the secondary task module, after the feature vector is calculated by the first tower network and the activation function 1 connected to the first tower network, the score y1 of the secondary task can be obtained.

[0154] In the main task module, after the feature vector is calculated by the second tower network, it is first input into the self-attention module for attention calculation, and then the attention calculation result is input into the activation function 2 for further calculation, and finally the score y2 of the main task is obtained.

[0155] Among them, the input of the attention module includes not only the output of the second tower network, but also the output of the first tower network. That is to say, the self-attention module comprehensively performs attention calculation by combining the output results of the first tower network and the second tower network. In this way, the sub-task of the first tower network assists the main task to obtain a more accurate score y2.

[0156] For a multi-task model, which is a machine learning model, in order to better obtain the model performance, that is, to calculate the recommended scores of each playback segment more accurately, usually before applying the model, it is necessary to use a large number of training samples to train the model to fully learn the relevant information required for calculating the recommended scores in the training samples.

[0157] Therefore, the present specification also provides the following embodiments related to model training.

[0158] In an exemplary embodiment, the multi-task model is iteratively trained in the following manner until the multi-task model converges:

[0159] Obtain a sample set for training; wherein, the sample set includes a certain number of training samples, and the training samples include sample features and true labels with pre-annotated recommended scores;

[0160] Input the sample set into the multi-task model for supervised training, calculate the loss function of the training samples according to forward propagation, calculate the gradient of the loss function according to backward propagation, and update the model parameters of the multi-task model according to the gradient; wherein, the model parameters at least include the parameter weights in the multi-layer perceptron, the first tower network, the second tower network, the self-attention module, and the activation function.

[0161] In the present specification, the sample features may be the same features as the input features that need to be input into the multi-task model in the foregoing step 210. For example, they may include segment features and user features, etc. These sample features can be segmented from real multimedia objects so that the trained model can adapt to and meet the recommendation requirements generated by real services.

[0162] Each training sample is a sample pair of a set of sample features and the corresponding true label. The true label is the pre-annotated recommended score for the sample features. The goal of model training is to learn the potential mapping relationship between each sample feature and the corresponding true label. And this mapping relationship is reflected in the model as the parameter values of the model parameters.

[0163] Since the default or initial model parameters of the multi-task model are usually not optimal, it is necessary to train the optimal model parameters with the help of training samples.

[0164] The model training adopts supervised training. The sample features of the training samples are used as the model input for calculation to obtain the predicted result (recommendation score) of the output. Usually, there is an error between the calculation result and the true result (pre-labeled recommendation score) corresponding to the sample features. The loss function is a function used to measure this error between the predicted result and the true result. Using the loss function, the model parameters can be adjusted according to the measured error.

[0165] Specifically, the model parameters in the loss function are calculated through forward propagation, and backpropagation is performed using optimization algorithms (including but not limited to gradient descent method, simulated annealing, Adam optimizer, etc.) to complete the iterative optimization of the model parameters. Then, based on the adjusted model parameters, the model training is performed again. After multiple iterations like this, the model parameters can be gradually optimized to minimize the error between the calculation result and the true result until the model converges or reaches the convergence condition.

[0166] In this specification, the loss function may include but is not limited to functions such as mean squared error loss, cross-entropy loss function, absolute error loss function, Kullback-Leibler divergence, etc.

[0167] Exemplarily, taking the cross-entropy loss function as an example, its function formula can be shown as follows:

[0168]

[0169] Among them, M represents the total number of the primary task and the secondary tasks, N represents the total number of training samples, y i represents the true label of the i-th training sample, and p i represents the recommendation score of the i-th training sample predicted by the activation function in the primary task module.

[0170] Step 230: Sort the recommendation scores of each multimedia segment output by the multi-task model, and recommend the target playback segment corresponding to the highest recommendation score to the user.

[0171] According to Figure 4As shown, although the multi-task model can obtain the score y2 of the primary task and the score y1 of the secondary task, when making content recommendations to users, only the score y2 of the primary task, that is, the recommendation score corresponding to the playback segment, needs to be used. Since the recommendation score represents the probability of the user enabling the preset permission after playing the playback segment, after obtaining the recommendation scores of each playback segment output by the multi-task model, by sorting these recommendation scores, since the highest recommendation score represents the highest probability of the user enabling the preset permission, the target playback segment corresponding to the highest recommendation score is recommended to the user, which can attract the user to enable the preset permission with the greatest possibility.

[0172] Exemplary Medium

[0173] After introducing the method of the exemplary embodiment of the present disclosure, next, reference is made to Figure 5 to describe the medium of the exemplary embodiment of the present disclosure.

[0174] In this exemplary embodiment, the above method can be implemented by a program product. For example, a portable compact disc read-only memory (CD-ROM) can be adopted and includes program code, and can be run on a device, such as a personal computer. However, the program product of the present disclosure is not limited to this. In this document, a readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0175] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0176] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.

[0177] The program code contained on a readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0178] The program code for performing the operations of the present disclosure can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the C language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0179] In summary, the present disclosure can provide a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can be enabled to execute the foregoing recommended method embodiments.

[0180] Exemplary Device

[0181] After introducing the medium of the exemplary embodiments of the present disclosure, next, reference is made to Figure 6 to describe the apparatus of the exemplary embodiments of the present disclosure.

[0182] Figure 6 A block diagram of a recommended apparatus according to an embodiment of the present disclosure is schematically shown, corresponding to the foregoing Figure 2 shown method embodiments. The recommended apparatus may include:

[0183] An acquisition unit 510 that acquires segment features related to each playback segment of a multimedia object and user features related to a user to be recommended;

[0184] A calculation unit 520 that inputs the segment features and user features into a pre-trained multi-task model and calculates a recommendation score for the playback segment based on the segment features and user features; wherein the recommendation score represents the probability that the user enables a preset permission after playing the playback segment;

[0185] A recommendation unit 530 that sorts the recommendation scores of each multimedia segment output by the multi-task model and recommends the target playback segment corresponding to the highest recommendation score to the user.

[0186] Optionally, the multi-task model includes a multi-layer perceptron that receives input segment features and user features, and a primary task module and a secondary task module respectively connected to the output of the multi-layer perceptron;

[0187] Among them, the secondary task module includes a first tower network corresponding to the secondary task, the primary task module includes a second tower network corresponding to the primary task, the primary task module further includes a self-attention module that calculates attention based on the outputs of the first tower network and the second tower network, and an activation function that predicts the recommended scores corresponding to each playing segment based on the output of the self-attention module.

[0188] Optionally, the calculation unit 520 is further configured to splice the segment features and user features, input the obtained spliced features into a multi-layer perceptron in a pre-trained multi-task model, and calculate the feature vectors shared in the spliced features; input the feature vectors into the second tower network of the primary task module and the first tower network of the secondary task module respectively for calculation; further input the outputs of the first tower network and the second tower network into the self-attention module for attention calculation; further input the output of the self-attention module into the activation function for regression calculation to obtain the recommended scores corresponding to the playing segments.

[0189] Optionally, both the first tower network and the second tower network include multi-layer perceptrons using fully connected layers.

[0190] Optionally, the multi-task model is iteratively trained by the training unit 500 until the multi-task model converges:

[0191] The training unit 500 obtains a sample set for training; wherein, the sample set includes a number of training samples, and the training samples include sample features and true labels with pre-annotated recommended scores; input the sample set into the multi-task model for supervised training, calculate the loss function of the training samples according to forward propagation, calculate the gradient of the loss function according to backward propagation, and update the model parameters of the multi-task model according to the gradient; wherein, the model parameters at least include the parameter weights in the multi-layer perceptron, the first tower network, the second tower network, the self-attention module, and the activation function.

[0192] Optionally, the loss function includes a cross-entropy loss function

[0193] Among them, M represents the total number of primary tasks and secondary tasks, N represents the total number of training samples, y i represents the true label of the i-th training sample, p iIndicates the recommended score of the i-th training sample predicted by the activation function in the main task module.

[0194] Optionally, the user features include user attribute features related to the user's personal information, and / or user behavior features related to the user's operation behavior in the multimedia platform.

[0195] Optionally, the user behavior features include whether the user has historically enabled a preset permission.

[0196] Optionally, the segment features include the playback position of the playback segment in the multimedia object, and / or the proportion of the key content of the multimedia object included in the playback segment.

[0197] Optionally, the user to be recommended includes users who do not have the preset permission.

[0198] Optionally, the preset permission includes a membership permission, and the membership permission allows playing the complete multimedia object.

[0199] Optionally, the multimedia object includes a music song, and the playback segment includes a trial listening segment of the music song.

[0200] Exemplary Computing Device

[0201] After introducing the methods, media, and devices of the exemplary embodiments of the present disclosure, next, reference is made to Figure 7 to describe the computing device of the exemplary embodiments of the present disclosure.

[0202] Figure 7 The computing device 1500 shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0203] As Figure 7 shown, the computing device 1500 is presented in the form of a general-purpose computing device. The components of the computing device 1500 may include, but are not limited to: at least one processing unit 1501, at least one storage unit 1502, and a bus 1503 connecting different system components (including the processing unit 1501 and the storage unit 1502).

[0204] The bus 1503 includes a data bus, a control bus, and an address bus.

[0205] The storage unit 1502 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 15021 and / or a cache memory 15022, and may further include a readable medium in the form of a non-volatile memory, such as a read-only memory (ROM) 15023.

[0206] The storage unit 1502 may also include a program / utilities 15025 having a set (at least one) of program modules 15024. Such program modules 15024 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0207] The computing device 1500 may also communicate with one or more external devices 1504 (such as a keyboard, a pointing device, etc.).

[0208] Such communication may be through an input / output (I / O) interface 1505. Also, the computing device 1500 may further communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 1506. As Figure 7 shown, the network adapter 1506 communicates with other modules of the computing device 1500 through a bus 1503. It should be understood that although not shown in the figures, other hardware and / or software modules may be used in conjunction with the computing device 1500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0209] Through the computing device 1500 as Figure 7 shown, the foregoing recommended method may be implemented. More specifically, the storage unit 1502 stores instructions executable by the processing unit 1501. When the processing unit 1501 executes the instructions, the foregoing recommended method is implemented.

[0210] It should be noted that although several units / modules or sub-units / modules of the recommended device are mentioned in the foregoing detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the units / modules described above may be embodied in one unit / module. Conversely, the features and functions of one unit / module described above may be further divided and embodied by multiple units / modules.

[0211] In addition, although the operations of the method of the present disclosure are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0212] Although the spirit and principles of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed, and the division of each aspect does not mean that the features in these aspects cannot be combined for benefit. This division is only for convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A recommendation method, comprising: Obtaining segment features related to each playing segment of a multimedia object, and user features related to a user to be recommended; Inputting the segment features and user features into a pre-trained multi-task model, and calculating a recommendation score for the playing segment based on the segment features and user features; wherein, the recommendation score represents the probability that the user enables a preset permission after playing the playing segment; Sorting the recommendation scores of each multimedia segment output by the multi-task model, and recommending the target playing segment corresponding to the highest recommendation score to the user.

2. The method according to claim 1, wherein the multi-task model comprises a multi-layer perceptron for receiving input segment features and user features, and a main task module and a sub-task module respectively connected to the output of the multi-layer perceptron; Among them, The sub-task module comprises a first tower network corresponding to the sub-task, the main task module comprises a second tower network corresponding to the main task, the main task module further comprises a self-attention module for performing attention calculation based on the output of the first tower network and the output of the second tower network, and an activation function for predicting the recommendation score corresponding to each playing segment based on the output of the self-attention module.

3. The method according to claim 2, wherein the step of inputting the segment features and user features into a pre-trained multi-task model and calculating a recommendation score for the playing segment based on the segment features and user features comprises: Concatenating the segment features and user features, and inputting the obtained concatenated features into the multi-layer perceptron in the pre-trained multi-task model to calculate a shared feature vector in the concatenated features; Inputting the feature vector into the second tower network of the main task module and the first tower network of the sub-task module respectively for calculation; Further inputting the outputs of the first tower network and the second tower network into the self-attention module for attention calculation; Further inputting the output of the self-attention module into the activation function for regression calculation to obtain a recommendation score corresponding to the playing segment.

4. The method according to claim 2, wherein both the first tower network and the second tower network comprise a multi-layer perceptron using fully connected layers.

5. The method according to claim 2, wherein the multi-task model is iteratively trained in the following manner until the multi-task model converges: Obtain a sample set for training; wherein, The sample set comprises a plurality of training samples, and each training sample comprises sample features and a true label with a pre-labeled recommendation score; Inputting the sample set into the multi-task model for supervised training, calculating a loss function of the training sample according to forward propagation, calculating a gradient of the loss function according to backward propagation, and updating model parameters of the multi-task model according to the gradient; wherein, the model parameters at least comprise parameter weights of the multi-layer perceptron, the first tower network, the second tower network, the self-attention module, and the activation function.

6. The method according to claim 5, wherein the loss function comprises a cross-entropy loss function Among them, Let \(M\) denote the total number of primary tasks and secondary tasks, and \(N\) denote the total number of training samples. Let \(y\) i denote the true label of the \(i\)-th training sample, and \(p\) i denote the recommended score of the \(i\)-th training sample predicted by the activation function in the primary task module.

7. The method according to claim 1, wherein the user features comprise user attribute features related to user personal information, and / or User behavior characteristics related to the user's operation behavior in the multimedia platform.

8. A recommendation device, the device comprising: An acquisition unit that acquires segment features related to each playing segment of a multimedia object and user features related to a user to be recommended; A calculation unit that inputs the segment features and user features into a pre-trained multi-task model and calculates a recommendation score for the playing segment based on the segment features and user features; wherein the recommendation score represents the probability that the user enables a preset permission after playing the playing segment; A recommendation unit that sorts the recommendation scores of each multimedia segment output by the multi-task model and recommends the target playing segment corresponding to the highest recommendation score to the user.

9. A computer-readable storage medium, comprising: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the recommendation method according to any one of claims 1-7.

10. A computing device, comprising: A processor; A memory for storing executable instructions of the processor; Wherein the processor is configured to execute the executable instructions to implement the recommendation method according to any one of claims 1-7.