Music Video Recommendation Method, Its Device, Equipment, Medium, and Product

By analyzing user historical behavior data and using the music library knowledge graph to recall and screen music videos, the problem of inaccurate recommendation of music videos in the existing technology is solved, and a deep understanding of user interests and the accuracy of recommendations is achieved.

CN114218426BActive Publication Date: 2025-05-30GUANGZHOU KUGOU COMP TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111547721.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-16
Publication Date
2025-05-30
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

The existing music video recommendation technology has systematic problems with information cocoon effect and content synonyms, making it difficult to effectively recommend music videos that users are interested in, and it is impossible to explore the user's interest boundaries.

Method used

By analyzing the user's historical behavior data, constructing user interest vectors, and using the music library's knowledge graph to recall a list of music videos with similar semantics. Combining long-term and short-term historical access data, a comprehensive encoded vector is constructed, a candidate music video list is filtered, and a deduplication is performed through the similarity calculation of deep semantic vectors to generate a recommended music video list.

Benefits of technology

It achieves accurate matching of user interests, breaks the constraints of information cocoon, can effectively explore the boundaries of users' interest, and improves the accuracy of recommendations and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114218426B_ABST
    Figure CN114218426B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of music information retrieval, and discloses a music video recommendation method, its device, equipment, medium, and product. The method includes: determining a list of historical accessed music videos of the user according to the user's historical behavior data; obtaining a user interest vector based on the list of historical accessed music videos, and recalling a list of relevant music videos that match the user interest vector; constructing a comprehensive coding vector according to the lists of historical accessed music videos corresponding to multiple historical spans of the user, and screening out a list of candidate music videos from the list of relevant music videos according to the deep semantic information of the comprehensive coding vector; calculating the similarity of the deep semantic vectors of pairwise candidate music videos in the list of candidate music videos, and performing duplicate removal processing on the candidate music videos that form similarities to obtain the list of recommended music videos. This application can break the constraints of the information cocoon and accurately obtain a list of recommended music videos that matches the user's behavior data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of music information retrieval, and in particular to a music video recommendation method, its corresponding device, computer device, computer-readable storage medium, and computer program product. Background Art

[0002] People's way of listening to music is no longer satisfied with static resources such as text and pictures. They are more willing to enjoy music through video, and satisfy the immersive audio-visual enjoyment through music videos. The characteristics of music videos are that they can provide richer information for the description of music content, including not only the conventional text information related to song descriptions, but also the image information related to song image content. These information can theoretically be used for data mining.

[0003] As a means to enrich and meet users' personalization of music video resources, existing music recommendation technologies can find the most likely and suitable music videos for users from a vast amount of content through recall and ranking means. However, in the case of insufficient information and different algorithm designs, the existing recommendation technologies are still not intelligent enough. For the recommendation of music videos, various existing technologies are stretched thin. While pursuing recommendation accuracy, there are still systematic problems such as the information cocoon effect and content synonyms, trapping users in the information cocoon, unable to explore the boundaries of users' interests, and unable to more scientifically associate with the changes in users' interests to provide appropriate music video recommendations.

[0004] It can be seen that solving this problem plays a crucial role in the healthy development of the product ecosystem. In view of this, the applicant of the present application has made corresponding explorations based on the need to improve music video recommendations. Summary of the Invention

[0005] The primary objective of the present application is to solve at least one of the above problems and provide a music video recommendation method, its corresponding device, computer device, computer-readable storage medium, and computer program product.

[0006] To meet the various objectives of the present application, the present application adopts the following technical solutions:

[0007] A music video recommendation method provided to meet one of the objectives of the present application includes the following steps:

[0008] Respond to the user's music video recommendation request, and determine the user's historical access music video list according to the user's historical behavior data;

[0009] Obtain the user interest vector according to the historical access music video list, and recall the relevant music video list that matches the user interest vector from the music library knowledge graph;

[0010] Construct a comprehensive coding vector based on the historical access music video lists corresponding to multiple historical spans of the user, and screen out a candidate music video list from the relevant music video list according to the deep semantic information of the comprehensive coding vector;

[0011] Calculate the similarity of the deep semantic vectors of pairwise candidate music videos in the candidate music video list, and perform deduplication processing on the candidate music videos that form similarities to obtain the recommended music video list for the recommendation request to answer the recommendation request.

[0012] In an extended embodiment, before the step of responding to the music video recommendation request of the user, the following steps are included:

[0013] Create a knowledge graph corresponding to the music library;

[0014] Extract the knowledge propagation paths between the historical music videos accessed by each user and the interest tags of each user from the historical behavior data of all users, where the interest tags are portrait tags used to label the music videos;

[0015] Represent the music video and the interest tag as entity nodes in the knowledge graph, and establish an association relationship between each entity node according to the knowledge propagation path;

[0016] Store the content information of the music video as an attribute item in its corresponding entity node, where the content information includes the access link of the corresponding music video and the deep semantic vector of the music video, and the deep semantic vector is a comprehensive feature representation of the cover information, audio information, and text information of the music video.

[0017] In a refined embodiment, obtaining a user interest vector according to the historical access music video list and recalling a relevant music video list that matches the user interest vector from the music library knowledge graph includes the following steps:

[0018] Use a multi-interest recall model pre-trained to a convergent state to perform representation learning based on the coding information corresponding to the historical access music video list to obtain corresponding deep semantic information, and perform multi-classification mapping according to the deep semantic information to obtain multiple interest tags;

[0019] According to the interest tags, recall all target music videos carrying any one of the multiple interest tags from the music library knowledge graph;

[0020] Construct all the target music videos into a relevant music video list.

[0021] In a refined embodiment, a comprehensive coding vector is constructed based on the historical access music video lists corresponding to multiple historical spans of a user. According to the deep semantic information of the comprehensive coding vector, a candidate music video list is filtered from the relevant music video list, including the following steps:

[0022] According to two historical spans representing long term and short term, the long-term historical access music video list and the short-term historical access music video list of the user are respectively obtained, where the long-term historical span covers the short-term historical span;

[0023] The long-term historical access music video list and the short-term historical access music video list are respectively vectorized and spliced into a high-dimensional vector to obtain a comprehensive coding vector;

[0024] A recommendation ranking model pre-trained to a converged state is used to perform representation learning on the comprehensive coding vector to obtain its deep semantic information. According to the deep semantic information, multi-classification mapping is performed to obtain a candidate music video list corresponding to multiple preset targets. Each candidate music video list is sorted according to the click access statistical quantity, effective play statistical quantity, and complete play statistical quantity of the candidate music videos.

[0025] In a refined embodiment, the similarity of the deep semantic vectors of pairwise candidate music videos in the candidate music video list is calculated, and the duplicate candidate music videos are removed to obtain the recommended music video list for answering the recommendation request, including the following steps:

[0026] According to the candidate music video list, the deep semantic vectors in the entity nodes corresponding to each candidate music video are obtained from the knowledge graph. The deep semantic vector is a comprehensive feature representation of the cover information, audio information, and text information of the corresponding music video;

[0027] For each candidate music video list, the similarity of the deep semantic vectors of pairwise candidate music videos is calculated to obtain the similarity values between pairwise candidate music videos;

[0028] For each candidate music video list, the similarity values are filtered using the maximum greedy algorithm to obtain the recommended music video list;

[0029] After associating the access link of the music video with the recommended music video list, the recommended music video list is pushed to the user.

[0030] In a specific embodiment, after associating the access link of the music video with the recommended music video list, the recommended music video list is pushed to the user, including the following steps:

[0031] Obtain the unique feature information corresponding to each music video in the recommended music video list from the knowledge graph;

[0032] Obtain the content information corresponding to each music video from the music library according to the unique feature information, where the content information includes the access link, cover information, audio information, and text information of the corresponding music video;

[0033] Format the content information and associate it with the corresponding music video in the recommended music video list;

[0034] Push the recommended music video list to the user.

[0035] A music video recommendation device provided for one of the purposes of this application includes: a request response module, an interest recall module, a data pre-screening module, and a re-ranking processing module. Among them, the request response module is used to respond to the user's music video recommendation request and determine the user's historical access music video list according to the user's historical behavior data; the interest recall module is used to obtain the user interest vector according to the historical access music video list and recall the relevant music video list that matches the user interest vector from the music library knowledge graph; the data pre-screening module is used to construct a comprehensive coding vector according to the historical access music video lists corresponding to multiple historical spans of the user, and screen out the candidate music video list from the relevant music video list according to the deep semantic information of the comprehensive coding vector; the re-ranking processing module is used to calculate the similarity of the deep semantic vectors of two candidate music videos in the candidate music video list, perform duplicate removal processing on the similar candidate music videos, and obtain the recommended music video list of the recommendation request to answer the recommendation request.

[0036] In an extended embodiment, the music video recommendation device of this application further includes: a graph creation module for creating a knowledge graph corresponding to the music library; a path sorting module for extracting the knowledge propagation paths between the historical music videos accessed by each user and the interest tags of each user from the historical behavior data of all users, where the interest tags are portrait tags used to label the music videos; a relationship representation module for representing the music video and the interest tag as entity nodes in the knowledge graph and establishing the association relationship between each entity node according to the knowledge propagation path; an information association module for storing the content information of the music video as an attribute item in its corresponding entity node, where the content information includes the access link of the corresponding music video and the deep semantic vector of the music video, and the deep semantic vector is a comprehensive feature representation of the cover information, audio information, and text information of the music video.

[0037] In a refined embodiment, the interest recall module includes: an interest classification sub-module, which is configured to perform representation learning based on the encoded information corresponding to the historical accessed music video list by using a multi-interest recall model pre-trained to a converged state to obtain corresponding deep semantic information, and perform multi-classification mapping based on the deep semantic information to obtain multiple interest tags; a data recall sub-module, which is configured to recall all target music videos carrying any one of the multiple interest tags from the music library knowledge graph according to the interest tags; and a list construction sub-module, which is configured to construct all the target music videos into a relevant music video list.

[0038] In a refined embodiment, the data pre-screening module includes: a duration division sub-module, which is configured to obtain a long-term historical accessed music video list and a short-term historical accessed music video list of the user respectively according to two historical spans representing long term and short term, wherein the long-term historical span covers the short-term historical span; a vector encoding sub-module, which is configured to vectorize the long-term historical accessed music video list and the short-term historical accessed music video list respectively and splice them into a high-dimensional vector to obtain a comprehensive encoded vector; and a recommendation ranking sub-module, which is configured to perform representation learning on the comprehensive encoded vector by using a recommendation ranking model pre-trained to a converged state to obtain its deep semantic information, and perform multi-classification mapping according to the deep semantic information to obtain a candidate music video list corresponding to multiple preset targets, and each candidate music video list is sorted according to the click access statistical quantity, effective play statistical quantity, and complete play statistical quantity of the candidate music videos.

[0039] In a refined embodiment, the rearrangement processing module includes: a candidate calling sub-module, which is configured to obtain the deep semantic vector in the entity node corresponding to each candidate music video from the knowledge graph according to the candidate music video list, and the deep semantic vector is a comprehensive feature representation of the cover information, audio information, and text information of the corresponding music video; a similarity calculation sub-module, which is configured to calculate the similarity of the deep semantic vectors of two candidate music videos in each candidate music video list to obtain the similarity value between two candidate music videos; a filtering processing sub-module, which is configured to filter the similarity value by adapting the maximum greedy algorithm for each candidate music video list to obtain a recommended music video list; and an associated recommendation sub-module, which is configured to push the recommended music video list to the user after associating the access link of the music video in the recommended music video list.

[0040] In a specific embodiment, the associated recommendation sub-module includes: a feature determination unit for obtaining the unique feature information corresponding to each music video in the recommended music video list from the knowledge graph; a content invocation unit for obtaining the content information corresponding to each music video from the music library according to the unique feature information, where the content information includes the access link, cover information, audio information, and text information of the corresponding music video; a format processing unit for formatting the content information and associating it with the corresponding music video in the recommended music video list; and a push execution unit for pushing the recommended music video list to the user.

[0041] A computer device provided to meet one of the purposes of the present application includes a central processing unit and a memory. The central processing unit is used to call and run a computer program stored in the memory to execute the steps of the music video recommendation method described in the present application.

[0042] A computer-readable storage medium provided to meet another purpose of the present application stores a computer program implemented according to the music video recommendation method in the form of computer-readable instructions. When the computer program is called and run by a computer, it executes the steps included in the method.

[0043] A computer program product provided to meet another purpose of the present application includes a computer program / instructions. When the computer program / instructions are executed by a processor, they implement the steps of the method described in any embodiment of the present application.

[0044] Compared with the prior art, the advantages of the present application are as follows:

[0045] First, based on the user behavior data, the present application determines the user's historical access audio and video list. First, it determines the user interest vector according to the historical access audio and video list to characterize the features that the user is interested in. Then, through the knowledge graph, it recalls a list of related music videos with similar semantics for the user to achieve preliminary screening. Then, it comprehensively characterizes the user's long-term and short-term preferences according to different historical spans, and based on the comprehensive characterization, it predicts associated music videos based on semantics to further refine and screen out a candidate music video list from the initially screened associated music video list. Finally, it calculates the similarity of the deep semantic vectors representing the content information of each music video in the candidate music video list, and de-duplicates the music videos according to the similarity values to finally obtain a recommended music video list. The whole process goes from preliminary recall to refined screening and then to refined de-duplication, going deeper layer by layer, achieving the goal of accurately obtaining a recommended music video list that the user is interested in.

[0046] Secondly, in each execution link of the present application, multiple links optimize the recommendation of music videos layer by layer based on the semantic information implied in user behavior data. With the assistance of the user's personal preference information represented by user behavior data and the knowledge graph, the interest boundary of the user can be reasonably explored, so as to effectively obtain the target music videos that match the user's personal preferences. The recommended audio and video obtained has a higher accuracy rate. On the one hand, it can probe or expand the user's interests and distribute music videos related to the user's interest boundary, enabling the user to get rid of the bondage of the information cocoon; on the other hand, for newly added music videos, it is also easier to achieve cold start.

[0047] Furthermore, the present application not only performs basic sorting based on the recall data to obtain a list of candidate music videos, but further performs refined deduplication based on the content similarity between music videos on the basis of the list of candidate music videos, so that music videos with similar content in the recommended music video list will not be piled up and presented to the user, and the recommended music video list obtained by the user is more concise.

[0048] In addition, the technical solution of the present application is suitable for being deployed in an online music platform to serve a large number of platform users. Since it can perform recall, rough sorting, re-sorting, etc. for each user, it greatly compresses the output of the recommended music video list for each user, taking into account the two aspects of recall and precision in the field of recommended sorting. Therefore, it can improve the platform service efficiency and even comprehensively improve the user experience of the online music platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0050] Figure 1 is a schematic flowchart of a typical embodiment of the music video recommendation method of the present application;

[0051] Figure 2 is a schematic flowchart of the knowledge graph creation process in the embodiment of the present application;

[0052] Figure 3 is a schematic diagram of the music video recall process in the embodiment of the present application;

[0053] Figure 4 is a schematic flowchart of the working process of the neural network model in the embodiment of the present application;

[0054] Figure 5 is a schematic flowchart of the process of performing recommended sorting on the recalled relevant music video list to obtain a list of candidate music videos in the embodiment of the present application;

[0055] Figure 6It is a schematic flowchart of the process of pushing a recommended music video list to a user in an embodiment of the present application;

[0056] Figure 7 It is a principle block diagram of the music video recommendation device of the present application;

[0057] Figure 8 It is a schematic structural diagram of a computer device adopted by the present application. Detailed implementation manners

[0058] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as a limitation of the present application.

[0059] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0060] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0061] Those skilled in the art can understand that the "client", "terminal", and "terminal device" used herein include both devices with wireless signal receivers that only have the ability to receive and no ability to transmit, and devices with receiving and transmitting hardware that have the receiving and transmitting hardware capable of two-way communication on a two-way communication link. Such devices can include: cellular or other communication devices such as personal computers, tablet computers, etc., which have a single-line display or a multi-line display or a cellular or other communication device without a multi-line display; PCS (Personal Communications Service), which can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant), which can include a radio frequency receiver, a pager, Internet / intranet access, a web browser, a notepad, a calendar, and / or a GPS (Global Positioning System) receiver; conventional laptop and / or palm computers or other devices, which are conventional laptop and / or palm computers or other devices with and / or including a radio frequency receiver. The "client", "terminal", and "terminal device" used herein can be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to run locally, and / or run in a distributed form at any other location on the earth and / or in space. The "client", "terminal", and "terminal device" used herein can also be a communication terminal, an Internet access terminal, a music / video playback terminal, such as a PDA, a MID (Mobile Internet Device), and / or a mobile phone with music / video playback function, or can also be a smart TV, a set-top box, etc.

[0062] The hardware referred to by names such as "server", "client", and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer, and is a hardware device with the necessary components disclosed by the von Neumann principle, including a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit loads the program stored in the external memory into the memory for execution, executes the instructions in the program, and interacts with the input / output devices to complete specific functions.

[0063] It should be noted that the concept of "server" as referred to in this application can similarly be extended to the case applicable to a server cluster. According to the network deployment principles understood by those skilled in the art, the various servers should be logically divided. Physically, these servers can either be independent of each other but can be invoked through interfaces, or can be integrated into a single physical computer or a set of computer clusters. Those skilled in the art should understand this flexibility and should not be restricted by this in the implementation manner of the network deployment method of this application.

[0064] One or several technical features of this application, unless explicitly specified, can either be deployed on a server for implementation and accessed by a client remotely invoking the online service interface provided by the server, or can be directly deployed and run on the client for implementation and access.

[0065] The neural network models cited or possibly cited in this application, unless explicitly specified, can either be deployed on a remote server and remotely invoked on the client, or can be directly invoked on the client with sufficient device capabilities. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operation resources and avoid over-occupying the client's hardware operation resources.

[0066] All kinds of data involved in this application, unless explicitly specified, can either be remotely stored on a server or stored on a local terminal device, as long as it is suitable for being invoked by the technical solution of this application.

[0067] Those skilled in the art should be aware of this: Although the various methods of this application are described based on the same concept and thus show commonality with each other, unless otherwise specified, these methods can all be executed independently. Similarly, for each of the embodiments disclosed in this application, they are all proposed based on the same inventive concept. Therefore, for concepts with the same expression, as well as concepts that are only appropriately transformed for convenience although the concept expressions are different, they should be equivalently understood.

[0068] For each of the embodiments to be disclosed in this application, unless explicitly pointed out that there is a mutually exclusive relationship between them, otherwise, the relevant technical features involved in each embodiment can be cross-combined to flexibly construct new embodiments, as long as this combination does not deviate from the creative spirit of this application and can meet the requirements in the prior art or solve certain deficiencies in the prior art. Those skilled in the art should be aware of this flexibility.

[0069] A music video recommendation method of this application can be programmed into a computer program product and deployed to run in a server. Thereby, a client can access the interface opened after the computer program product runs in the form of a web program or an application program, and achieve human-computer interaction with the process of the computer program product through a graphical user interface.

[0070] Please refer to Figure 1 , in a typical embodiment of the music video recommendation method of the present application, it includes the following steps:

[0071] Step S1100: In response to a user's music video recommendation request, determine the user's historical access music video list according to the user's historical behavior data:

[0072] One of the application scenarios of the present application is an online music service platform. By running the product implemented by the present application, using the relevant data of the music library and user database of the platform, providing services in the background, and responding to the user's music video recommendation request to recommend music videos that the user is interested in.

[0073] The music video is a type of music data object that can be played in video format on the client. It can be in the form of a video file containing the audio content of the music, or an audio file carrying cover pictures, cover titles, introductions, etc. used to form a video playback form on the client. No matter what form the music video is, it can be uniformly processed in the present application. The characteristics of the music video referred to in the present application, because it contains cover information such as pictures, audio information that is the song, text information such as introductions, titles, etc., therefore, the amount of graphic and text information is relatively rich, and it is easier to show various characteristics, so that rich semantic information can be obtained from its pictures and texts.

[0074] When the user needs to obtain the music videos that he / she is interested in, he / she can trigger a music video recommendation request on his / her client device, and the server implementing the technical solution of the present application responds. The server of the present application responds to this music video recommendation request and starts to execute each step of the present application for this user to obtain the music videos that he / she is interested in.

[0075] The user mentioned above is a platform user. His / her historical access behavior on the platform will generate corresponding historical behavior data stored in the user database. The historical behavior data usually stores the mapping relationship data corresponding to the access time, access behavior, and access music video (or its list) when the user accesses a certain song (music video) or playlist (music video list) in a certain behavior type at a certain moment. Since there are usually multiple portrait tags corresponding to the music videos in the music library, these portrait tags actually also represent the user's interest types. The portrait tags mentioned above can usually also be regarded as type tags divided from different dimensions of the music video, such as type tags divided according to dimensions such as era, singer, style, language, etc.

[0076] In this step, according to the above mapping relationship data existing in the historical behavior data of the user, the full set of music videos generated by his historical visits can be determined, and thus a list of historical visited music videos can be determined. The uniqueness feature information of specific music videos in the music library can be used to represent each corresponding music video in this list.

[0077] Step S1200: Obtain a user interest vector according to the list of historical visited music videos, and recall a list of relevant music videos that match the user interest vector from the music library knowledge graph:

[0078] In this application, a knowledge graph is constructed for the music library corresponding to the platform. This knowledge graph is implemented in a directed graph structure and has multiple entity nodes. Each song (music video) in the music library can be stored as one of the entity nodes, and various portrait labels used to describe the music videos in the music library can also be stored as the entity nodes. Thus, a knowledge propagation path in the form of "music video 1 - portrait label 1 - music video 2 - portrait label 2..." in the music library can be realized in the knowledge graph. The process of establishing the knowledge propagation path can be established through the pre-annotated portrait labels of each music video, that is, the connection relationship of each entity node in the corresponding knowledge propagation path is marked according to different music videos pointing to the same portrait label.

[0079] More specifically, the portrait label is a label suitable for describing the characteristics of a music video, mainly including various specific labels included in any one or any combination of style type, era type, language type, emotion type, rhythm type, singer name, singer characteristics, lyric characteristics, evaluation characteristics, song relationship, singer relationship, user behavior characteristics of the song, etc. For example, the era type can be divided into specific labels such as "post-80s" and "post-90s". Such portrait labels can be used to describe the characteristics of each music video and realize the semantic representation of the common characteristics of many music videos.

[0080] So far, it can be understood that through the knowledge graph, different music videos associated with the same portrait label can be conveniently found, and at the same time, the full set of portrait labels required to form the same music video can also be found. By calling the knowledge graph, the recall of music videos can be realized.

[0081] For users who trigger music video recommendation requests, based on the historical access music video list obtained from their historical behavior data, this application can recall music videos corresponding to these portrait tags from the knowledge graph according to the full portrait tags included in all music videos in the list, and construct these music videos into a relevant music video list. When implementing this recall operation, the vector representation of the full portrait tags in the user's historical access music video list can be first used to form a user interest vector, and then vector matching can be performed in the knowledge graph according to the user interest vector to obtain the relevant music video list. It is not difficult to understand that the result obtained by semantic matching based on the user interest vector can be implemented with the help of a convolutional neural network model. Since semantic relevance is fully considered, the exploration and expansion of the user interest boundary are realized, which helps to break the user's information cocoon and also helps the cold start of new music videos, making them more likely to be recommended to users. For the more detailed operation of obtaining a candidate music video list according to the user interest vector, the technical personnel of this application can implement it by themselves according to the principles disclosed here, or a subsequent embodiment will also give an example with better actual measurement results, which will not be elaborated here for the time being.

[0082] Step S1300: Construct a comprehensive coding vector according to the historical access music video lists corresponding to multiple historical spans of the user, and screen out a candidate music video list from the relevant music video list according to the deep semantic information of the comprehensive coding vector.

[0083] The user's interest may change over time. Therefore, if the full historical access music video list is generally used to characterize the user's long-term interest characteristics, the dynamic changes of the user's interest cannot be reflected. On the contrary, if only the local data of the historical access music video list is used to characterize the user's short-term interest characteristics, it is easy to overgeneralize. Therefore, in this application, by pre-determining multiple historical spans and constructing historical access music video lists corresponding to each historical span from the historical access music video list, a multi-faceted characterization of the user's interest characteristics is realized, so as to assist the user in further screening out a candidate music video list that more matches the user's interest from the relevant music video list according to such characteristics.

[0084] In this application, this purpose can be achieved by means of a convolutional neural network model. To this end, the historical access audio and video lists corresponding to each historical span can be vectorized first to obtain their corresponding encoded vectors, and then these encoded vectors can be concatenated into a comprehensive encoded vector as the input of the model. Then, the model performs representation learning on this comprehensive encoded vector to obtain its deep semantic information. On this basis, classification mapping is performed according to this deep semantic information, and according to the mapping result, some music videos that match this deep semantic information are screened out from the relevant audio and video lists to form the candidate music video list mentioned above. The model here can be implemented using a two-tower model, and the matching between the user branch (User) and the music video branch (Item) is achieved through the two-tower model. Thus, the corresponding candidate music video list can be obtained according to the comprehensive encoded vector of the user.

[0085] Step S1400: Calculate the similarity of the deep semantic vectors of pairwise candidate music videos in the candidate music video list, and perform duplicate removal on the candidate music videos that are similar to each other to obtain the recommended music video list for answering the recommendation request:

[0086] After obtaining the candidate music video list mentioned above, considering that there may be different singing versions of a music video in the music library, and users tend to obtain high-quality versions among them, and in practice, it is not necessary to recommend multiple singing versions based on the same song. Therefore, in this step, the candidate music video list can be further optimized based on this principle.

[0087] In view of this, based on the candidate music video list, calculate the similarity between the deep semantic vectors of pairwise candidate music videos in it to obtain the corresponding similarity values, which can be represented as a two-dimensional storage matrix. Then, based on these similarity values, examine whether there is a situation where the similarity values between other candidate music videos that are similar to the same candidate music video reach a certain degree of closeness. If there is such a situation, perform duplicate removal on the candidate music videos that reach this degree of closeness, keep the one with the highest similarity value in the candidate music video list, and delete the others with lower similarity values. Finally, after the candidate music video list is purified, it is used as the recommended music video list and pushed to the user to complete the response to the user's music video recommendation request.

[0088] It is not difficult to understand that after the above processing, from the relevant music video list retrieved roughly, to the candidate music video list obtained through preliminary screening, and then to the recommended music video list obtained through duplicate removal, in the whole process, the music videos that the user may be interested in are filtered and optimized layer by layer according to the semantic information contained in the user's historical behavior data. Therefore, the obtained recommended music video list should be able to accurately match the user's interests.

[0089] Through the introduction of this exemplary embodiment, it can be understood that the implementation of this application has rich positive meanings, including but not limited to the following aspects:

[0090] First, based on the user behavior data, this application determines the user's historical access audio and video list. First, it determines the user interest vector according to the historical access audio and video list to characterize the features that the user is interested in. Then, through the knowledge graph, it recalls a list of related music videos with similar semantics for this user to achieve preliminary screening. Then, it comprehensively characterizes the long-term and short-term preferences of the user according to different historical spans. Based on the comprehensive characterization, it predicts associated music videos based on semantics to further refine and screen out a candidate music video list from the initially screened associated music video list. Finally, it calculates the similarity of the deep semantic vectors representing the content information of each music video in the candidate music video list, and performs duplicate removal processing on the music videos according to the similarity values. Finally, a recommended music video list is obtained. The whole process goes from preliminary recall to refined screening and then to refined duplicate removal, going deeper layer by layer to achieve the goal of accurately obtaining a recommended music video list that the user is interested in.

[0091] Second, in each execution link of this application, multiple links optimize the recommendation of music videos based on the semantic information implied in the user behavior data. With the assistance of the user personal preference information represented by the user behavior data and the knowledge graph, it can reasonably explore the user's interest boundary, so as to effectively obtain target music videos that match the user's personal preferences. The obtained recommended audio and video has a higher accuracy. On the one hand, it can explore or expand the user's interests and distribute music videos related to the user's interest boundary, enabling the user to get rid of the bondage of the information cocoon. On the other hand, for newly added music videos, it is also easier to achieve cold start.

[0092] Furthermore, this application not only performs basic sorting based on the recalled data to obtain a candidate music video list, but further performs refined duplicate removal processing on the candidate music video list according to the content similarity between music videos, so that music videos with similar content in the recommended music video list will not be presented to the user in a pile. The recommended music video list obtained by the user is more concise.

[0093] In addition, the technical solution of this application is suitable for being deployed in an online music platform to serve a large number of platform users. Since it can perform recall, rough ranking, re-ranking, etc. for each user, it greatly compresses the output of the recommended music video list for each user, taking into account the two goals of recall and precision in the field of recommendation ranking. Therefore, it can improve the platform service efficiency and even comprehensively improve the user experience of the online music platform.

[0094] Please refer to Figure 2, in an extended embodiment, before the step S1100, the step of responding to the user's music video recommendation request, the following steps are included:

[0095] Step S0100, create a knowledge graph corresponding to the music library:

[0096] The online music service platform can initialize and create the knowledge graph as needed. The knowledge graph can be organized using a database and can be flexibly implemented by those skilled in the art.

[0097] Step S0200, extract the knowledge propagation paths between the historical music videos accessed by each user and the interest tags of each user from the historical behavior data of all users. The interest tags are portrait tags used to label the music videos:

[0098] In this embodiment, it is recommended to use the EGES model to implement graph embedding. This model is suitable for obtaining relevant music video sequences based on the historical music video sequences extracted from the historical behavior data of users, so that the portrait tags of the music videos in these music video sequences can be determined as the interest tags of users. Then, according to the mapping relationship between the interest tags and each music video, the knowledge propagation paths from music videos to interest tags and then to music videos can be constructed, and the knowledge graph can be further completed based on these knowledge propagation paths.

[0099] Step S0300, represent the music videos and the interest tags as entity nodes in the knowledge graph, and establish the association relationships between the entity nodes according to the knowledge propagation paths:

[0100] In the knowledge graph, each of the interest tags and music videos can be stored as entity nodes in the knowledge graph. Further, based on the knowledge propagation paths, a many-to-many association between each music video and each interest tag is established, thereby establishing the association relationship architecture of the knowledge graph.

[0101] Step S0400, store the content information of the music videos as attribute items in their corresponding entity nodes. The content information includes the access link of the corresponding music video and the deep semantic vector of the music video. The deep semantic vector is a comprehensive feature representation of the cover information, audio information, and text information of the music video:

[0102] The content information of the music video, including but not limited to the cover information, music information, text information, access link, deep semantic vector, etc. of the music video, can be stored in the corresponding entity node of the music video and stored as an attribute item therein for convenient subsequent invocation. Among them, the deep semantic vector is a comprehensive feature representation of the cover information, audio information, and text information of the music video, and can be obtained by means of a convolutional neural network model pre-trained to a convergent state. In this regard, those skilled in the art can implement it as needed.

[0103] This embodiment provides an implementation method of the knowledge graph. It can be seen that the knowledge graph realizes the representation of the semantic association relationship between the portrait tags of the music video. When the music video is data determined according to user behavior data, it can better reflect the characteristics of the association relationship between the content of interest to the user.

[0104] Please refer to Figure 3 , in a further embodiment, the step S1200 of obtaining a user interest vector according to the historical access music video list and recalling a relevant music video list matching the user interest vector from the music library knowledge graph includes the following steps:

[0105] Step S1210: Use a multi-interest recall model pre-trained to a convergent state to perform representation learning according to the encoded information corresponding to the historical access music video list to obtain corresponding deep semantic information, and perform multi-classification mapping according to the deep semantic information to obtain multiple interest tags:

[0106] In this embodiment, it is recommended to use the MIND model for implementation. The MIND model is a multi-interest network model of the dynamic routing algorithm applying the capsule network by the Alibaba team in 2019. With the help of this model, deep semantic information can be extracted from the encoded information corresponding to the historical access music video list, and then multi-classification mapping is performed on the basis of the deep semantic information. Correspondingly, multiple interest tags are obtained. As mentioned above, the interest tags are also portrait tags used to label music videos.

[0107] Specifically, the historical access music video list represents a list that the corresponding user is relatively interested in, so it represents the user interest characteristics. After splitting and encoding it, it is input into the capsule network structure of the MIND model for processing, and then a multi-interest vector of the user, that is, the user interest vector, is calculated to obtain corresponding multiple interest tags, representing the change direction of the user interest granularity and enhancing the interest diversity. Based on this, the target music video can be recalled in the knowledge graph.

[0108] Step S1220: According to the interest tags, recall all target music videos carrying any one of the multiple interest tags from the music library knowledge graph:

[0109] After determining the user interest vector, that is, obtaining multiple interest tags representing the user's interests, accordingly, the target music videos matching the interest tags can be recalled from the knowledge graph of the music library to obtain all the target music videos.

[0110] Step S1230: Construct a list of relevant music videos from all the target music videos:

[0111] Finally, add all the target music videos that match all the interest tags to the list of relevant music videos, that is, complete all the work of recall according to the user interest vector.

[0112] In this embodiment, a multi-interest recall model is used to determine and expand the interest tags that match the user's interests according to the user's historical access music video list, realizing the expansion of the user's interests according to the semantics of the user's historical behavior data. Accordingly, the target music videos are recalled to ensure the recall rate of the recall for the user's interests, laying an important foundation for the subsequent work.

[0113] Please refer to Figure 4 , in the in-depth embodiment, in step S1300, a comprehensive coding vector is constructed according to the historical access music video lists corresponding to multiple historical spans of the user, and according to the deep semantic information of the comprehensive coding vector, candidate music video lists are screened out from the list of relevant music videos, including the following steps:

[0114] Step S1310: Obtain the long-term historical access music video list and the short-term historical access music video list of the user according to two historical spans representing long-term and short-term respectively, where the long-term historical span covers the short-term historical span:

[0115] In this step, two historical spans can be preset, for example, one month before the current day and one year before the current day. According to these two historical spans, a corresponding short-term historical access music video list and a long-term historical access music video list are respectively extracted from the user's historical access music video list. Through these two lists, the short-term preferences and long-term preferences of the user can be respectively represented. Generally speaking, the historical span corresponding to the long-term preference is suitable to cover the historical span corresponding to the short-term preference in order to describe the long-term interest characteristics and short-term interest characteristics within a relatively long time range.

[0116] Step S1320: Vectorize the long-term historical access music video list and the short-term historical access music video list respectively and splice them into a high-dimensional vector to obtain a comprehensive coding vector:

[0117] The long-term historical access music video list and the short-term historical music video list mentioned above are both represented in the form of the unique feature (ID) of the music video, which is essentially an ID sequence. Encoding according to this ID sequence and realizing vectorization can obtain the corresponding encoding vector. On this basis, by splicing the two encoding vectors before and after, a comprehensive encoding vector can be obtained. This comprehensive encoding vector should be processed into a high-dimensional vector form, where the former part is the vector representation of one historical span, and the other part is the vector representation of another historical span.

[0118] Step S1330: Use a recommendation ranking model pre-trained to the convergence state to perform representation learning on the comprehensive encoding vector to obtain its deep semantic information, and perform multi-classification mapping according to this deep semantic information to obtain a list of candidate music videos corresponding to multiple preset targets. Each list of candidate music videos is sorted respectively according to the click access statistical quantity, effective play statistical quantity, and complete play statistical quantity of the candidate music videos:

[0119] In this embodiment, it is recommended to use the Song2vec model pre-trained to the convergence state as the recommendation ranking model of the present application to implement the optimization of the list of candidate music videos according to the comprehensive encoding vector, so as to realize the preliminary recommendation ranking. Song2 is a multi-objective classification model, which is improved from item2vec improved by Word2vec and is well-known to those skilled in the art. After being trained, this model can extract deep semantic information according to the input comprehensive encoding vector, and then according to this deep semantic information and multiple preset targets, obtain the music videos in the relevant music video list that match the deep semantic information, and determine the corresponding multiple lists of candidate music videos according to different targets. The above-mentioned targets can be set during the model training stage so that the model can acquire the corresponding multi-objective classification ability. By setting the above-mentioned various targets, the determined lists of candidate music videos can be organized and sorted according to different targets. For example, according to targets such as the click access statistical quantity, effective play statistical quantity, and complete play statistical quantity of the candidate music videos, the corresponding lists of candidate music videos are determined.

[0120] In this embodiment, with the help of the recommendation ranking model, according to the deep semantic information of the comprehensive encoding vector representing the user's interest, and combined with multiple preset targets, the relevant music video list is refined, so that the music videos in one or more lists of candidate music videos obtained by refinement can better match the user's interest, and the precision rate in the recommendation ranking process is improved.

[0121] Please refer to Figure 5, in a refined embodiment, the step S1400 of calculating the similarity of the deep semantic vectors of pairwise candidate music videos in the candidate music video list, removing duplicates from the candidate music videos that are similar to each other, and obtaining the recommended music video list for the recommendation request to answer the recommendation request includes the following steps:

[0122] Step S1410: According to the candidate music video list, obtain the deep semantic vectors in the entity nodes corresponding to each candidate music video from the knowledge graph. The deep semantic vector is a comprehensive feature representation of the cover information, audio information, and text information of the corresponding music video:

[0123] As mentioned above, in the knowledge graph of the music library, deep semantic vectors obtained by extracting features from the content information of music videos with a neural network model are pre-stored to realize the feature representation of music videos. The content information includes any one or any combination of cover information, audio information, and text information. Accordingly, for each of the candidate music video lists, the corresponding deep semantic vectors of the music videos therein can be obtained one by one from the corresponding entity nodes of the knowledge graph.

[0124] Step S1420: For each candidate music video list, calculate the similarity of the deep semantic vectors of pairwise candidate music videos therein to obtain the similarity values between pairwise candidate music videos:

[0125] For each candidate music video list, a preset similarity calculation formula can be used to calculate the data distance between the deep semantic vectors of pairwise candidate music videos therein, determine the normalized similarity, and obtain the similarity values between pairwise candidate music videos. The similarity calculation formula can be implemented using a variety of data distance algorithms, including but not limited to any one of the cosine similarity algorithm, Pearson correlation coefficient algorithm, Jaccard similarity algorithm, Euclidean distance algorithm, etc. Those skilled in the art can implement it flexibly.

[0126] Step S1430: For each candidate music video list, filter the similarity values using the maximum greedy algorithm to obtain the recommended music video list:

[0127] Furthermore, duplicate removal processing can be performed on each candidate music video list. Specifically, the square matrix is transformed into matrix operations through the DPP algorithm, that is, the determinant point process algorithm, the square matrix is decomposed, and the maximum greedy algorithm is used to search for a subset with high relevance between music videos and users and low similarity between music videos to realize the optimization of the candidate music video list, thereby determining the recommended music video list corresponding to each candidate music video list.

[0128] Step S1440: After associating the access links of the music videos in the recommended music video list, push the recommended music video list to the user:

[0129] Finally, for each of the recommended music video lists, the access links corresponding to the respective music videos can be obtained from the knowledge graph, associated with the corresponding music videos and stored in the recommended music video list, and then the recommended music video list can be pushed to the user who initiated the music video recommendation request, thus completing the response process.

[0130] In this embodiment, according to the similarity between the music videos in each candidate music video list, further deduplication processing is performed on each candidate music video list, and multiple candidate music videos with similar content are streamlined, so as to obtain a recommended music video list, realizing refined selection of recommended music videos, so that the recommended music videos obtained by the user not only match their interest fields but also are relatively concise.

[0131] Please refer to Figure 6 , in a specific embodiment, the step S1440: After associating the access links of the music videos in the recommended music video list, push the recommended music video list to the user, includes the following steps:

[0132] Step S1441: Obtain the unique feature information corresponding to each music video in the recommended music video list from the knowledge graph:

[0133] In this embodiment, if the knowledge graph does not store rich various content information of music videos but only stores the unique feature information of each music video, then after obtaining the unique feature information corresponding to the music video, various specific content information of each music video can also be called from the music library according to the unique feature information.

[0134] Step S1442: Obtain the content information corresponding to each music video from the music library according to the unique feature information, and the content information includes the access link, cover information, audio information and text information of the corresponding music video:

[0135] According to the unique feature information of each music video, access the music library of the online music service platform, which pre-stores rich various content information of each music video. Accordingly, various content information of its corresponding music video can be obtained according to the unique feature information, including but not limited to the access link, cover information, audio information and text information of the music video, etc.

[0136] Step S1443: Format the content information and associate it with the corresponding music video in the recommended music video list:

[0137] Format the content information of each music video according to the display requirements of the client device, and then associate each music video with its formatted content information and store them in the recommended music video list, making the recommended music video list more suitable for being parsed and displayed by the client device.

[0138] Step S1444: Push the recommended music video list to the user.

[0139] After formatting the recommended music video list, the list can be pushed to the client device of the user who triggered the music video recommendation request for parsing and display. Based on this, the user can obtain music videos that are precisely matched to their fields of interest.

[0140] This embodiment meets the display requirements of the client device, improves the recommended music video list, enriches the content information of each music video therein, and can enhance the user experience.

[0141] Please refer to Figure 7 , a music video recommendation device provided by this application is functionally deployed according to the music video recommendation method of this application, including: a request response module 1100, an interest recall module 1200, a data pre-screening module 1300, and a re-ranking processing module 1400. Among them, the request response module 1100 is used to respond to the user's music video recommendation request and determine the user's historical accessed music video list according to the user's historical behavior data; the interest recall module 1200 is used to obtain the user interest vector according to the historical accessed music video list and recall the relevant music video list matching the user interest vector from the music library knowledge graph; the data pre-screening module 1300 is used to construct a comprehensive coding vector according to the historical accessed music video lists corresponding to multiple historical spans of the user, and screen out the candidate music video list from the relevant music video list according to the deep semantic information of the comprehensive coding vector; the re-ranking processing module 1400 is used to calculate the similarity of the deep semantic vectors of two-by-two candidate music videos in the candidate music video list, perform duplicate removal processing on the candidate music videos with similar compositions, and obtain the recommended music video list of the recommendation request to answer the recommendation request.

[0142] In an extended embodiment, the music video recommendation device of the present application further includes: a knowledge graph creation module for creating a knowledge graph corresponding to a music library; a path sorting module for extracting, from the historical behavior data of all users, the knowledge propagation paths between the historical music videos accessed by each user and the interest tags of each user, where the interest tags are portrait tags used to label the music videos; a relationship representation module for representing the music videos and the interest tags as entity nodes in the knowledge graph, and establishing association relationships between the entity nodes according to the knowledge propagation paths; an information association module for storing the content information of the music videos as attribute items in their corresponding entity nodes, where the content information includes the access link of the corresponding music video and the deep semantic vector of the music video, and the deep semantic vector is a comprehensive feature representation of the cover information, audio information, and text information of the music video.

[0143] In a refined embodiment, the interest recall module 1200 includes: an interest classification sub-module for performing representation learning on the encoded information corresponding to the historical access music video list using a multi-interest recall model pre-trained to a convergent state to obtain corresponding deep semantic information, and performing multi-classification mapping according to the deep semantic information to obtain multiple interest tags; a data recall sub-module for recalling, from the music library knowledge graph, all target music videos carrying any one of the multiple interest tags according to the interest tags; a list construction sub-module for constructing all the target music videos into a relevant music video list.

[0144] In a refined embodiment, the data pre-screening module 1300 includes: a duration division sub-module for obtaining, according to two historical spans representing long-term and short-term, the long-term historical access music video list and the short-term historical access music video list of the user respectively, where the long-term historical span covers the short-term historical span; a vector encoding sub-module for vectorizing and splicing the long-term historical access music video list and the short-term historical access music video list respectively into a high-dimensional vector to obtain a comprehensive encoded vector; a recommendation sorting sub-module for using a recommendation sorting model pre-trained to a convergent state to perform representation learning on the comprehensive encoded vector to obtain its deep semantic information, and performing multi-classification mapping according to the deep semantic information to obtain candidate music video lists corresponding to multiple preset targets, and each candidate music video list is sorted according to the click access statistical quantity, effective play statistical quantity, and complete play statistical quantity of the candidate music videos.

[0145] In a refined embodiment, the rearrangement processing module 1400 includes: a candidate call sub-module, configured to obtain, according to the candidate music video list, deep semantic vectors in the entity nodes corresponding to the respective candidate music videos from the knowledge graph, where the deep semantic vector is a comprehensive feature representation of the cover information, audio information, and text information of the corresponding music video; a similarity calculation sub-module, configured to calculate, for each candidate music video list, the similarity between the deep semantic vectors of two candidate music videos in the list to obtain the similarity value between the two candidate music videos; a filtering processing sub-module, configured to perform filtering on the similarity values for each candidate music video list by adapting the maximum greedy algorithm to obtain a recommended music video list; and an associated recommendation sub-module, configured to push the recommended music video list to the user after associating the access links of the music videos in the recommended music video list.

[0146] In a specific embodiment, the associated recommendation sub-module includes: a feature determination unit, configured to obtain the unique feature information corresponding to each music video in the recommended music video list from the knowledge graph; a content call unit, configured to obtain the content information corresponding to each music video from the music library according to the unique feature information, where the content information includes the access link, cover information, audio information, and text information of the corresponding music video; a format processing unit, configured to format the content information and associate it with the corresponding music video in the recommended music video list; and a push execution unit, configured to push the recommended music video list to the user.

[0147] To solve the above technical problems, an embodiment of the present application further provides a computer device. As Figure 8 shown, it is a schematic internal structure diagram of the computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected through a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The control information sequence can be stored in the database. When the computer-readable instructions are executed by the processor, the processor can implement a music video recommendation method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The computer-readable instructions can be stored in the memory of the computer device. When the computer-readable instructions are executed by the processor, the processor can execute the music video recommendation method of the present application. The network interface of the computer device is used to connect and communicate with the terminal. Those skilled in the art can understand that Figure 8 the structure shown in merely represents a block diagram of some structures related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0148] In this embodiment, the processor is used to execute Figure 7 the specific functions of each module and its sub-modules in []. The memory stores the program codes and various types of data required to execute the above modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. The memory in this embodiment stores the program codes and data required to execute all modules / sub-modules in the music video recommendation device of the present application, and the server can call the program codes and data of the server to execute the functions of all sub-modules.

[0149] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the music video recommendation method according to any embodiment of the present application.

[0150] The present application also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by one or more processors, the steps of the method according to any embodiment of the present application are implemented.

[0151] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments of the present application can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disc, a Read-Only Memory (ROM), or a Random Access Memory (RAM), etc.

[0152] In summary, the present application performs multi-level screening processing on the relevant music videos recalled from the knowledge graph and matching the user's historical behavior data, breaks the constraints of the information cocoon, and accurately obtains a list of recommended music videos matching the user's behavior data, which is suitable for providing standardized music recommendation services in an online music platform.

[0153] Those skilled in the art of the present technology can understand that the steps, measures, and solutions in the various operations, methods, and processes discussed in the present application can be alternated, changed, combined, or deleted. Further, the other steps, measures, and solutions in the various operations, methods, and processes discussed in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and solutions in the prior art that are the same as those disclosed in the various operations, methods, and processes in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted.

[0154] The above are only some embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A music video recommendation method, characterized in that, it includes the following steps: Respond to the user's music video recommendation request, and determine the user's historical accessed music video list according to the user's historical behavior data; Obtain the user interest vector according to the historical accessed music video list, and recall the relevant music video list that matches the user interest vector from the music library knowledge graph; Construct a comprehensive coding vector according to the historical accessed music video lists corresponding to multiple historical spans of the user, and screen out the candidate music video list from the relevant music video list according to the deep semantic information of the comprehensive coding vector, including: obtaining the user's long-term historical accessed music video list and short-term historical accessed music video list according to two historical spans representing long-term and short-term respectively, where the long-term historical span covers the short-term historical span; respectively vectorize the long-term historical accessed music video list and the short-term historical accessed music video list and splice them into a high-dimensional vector to obtain the comprehensive coding vector; use the recommendation ranking model pre-trained to the convergence state to perform representation learning on the comprehensive coding vector to obtain its deep semantic information, and perform multi-classification mapping according to the deep semantic information to obtain the candidate music video list corresponding to multiple preset targets, and each candidate music video list is sorted according to the click access statistical quantity, effective play statistical quantity, and complete play statistical quantity of the candidate music video; Calculate the similarity of the deep semantic vectors of two candidate music videos in the candidate music video list, and perform deduplication processing on the similar candidate music videos to obtain the recommended music video list of the recommendation request to respond to the recommendation request.

2. The music video recommendation method according to claim 1, characterized in that, before the step of responding to the user's music video recommendation request, it includes the following steps: Create a knowledge graph corresponding to the music library; Extract the knowledge propagation path between the historical music videos accessed by each user and the interest tags of each user from the historical behavior data of all users, and the interest tags are portrait tags used to label the music videos; Represent the music video and the interest tag as entity nodes in the knowledge graph, and establish the association relationship between each entity node according to the knowledge propagation path; Store the content information of the music video as an attribute item in its corresponding entity node, and the content information includes the access link of the corresponding music video and the deep semantic vector of the music video, and the deep semantic vector is the comprehensive feature representation of the cover information, audio information, and text information of the music video.

3. The music video recommendation method according to claim 1, characterized in that, obtaining the user interest vector according to the historical accessed music video list and recalling the relevant music video list that matches the user interest vector from the music library knowledge graph includes the following steps: The multi-interest recall model pre-trained to the convergence state performs representation learning based on the encoded information corresponding to the historical visited music video list to obtain corresponding deep semantic information, and performs multi-classification mapping based on the deep semantic information to obtain multiple interest tags; According to the interest tags, recall all target music videos carrying any one of the multiple interest tags from the music library knowledge graph; Construct the full amount of the target music videos into a relevant music video list.

4. The music video recommendation method according to claim 1, wherein, Calculating the similarity of the deep semantic vectors of pairwise candidate music videos in the candidate music video list, and performing deduplication processing on the similar candidate music videos to obtain the recommended music video list for answering the recommendation request, including the following steps: According to the candidate music video list, obtain the deep semantic vectors in the entity nodes corresponding to each candidate music video from the knowledge graph, and the deep semantic vector is a comprehensive feature representation of the cover information, audio information, and text information of the corresponding music video; For each candidate music video list, calculate the similarity of the deep semantic vectors of pairwise candidate music videos therein to obtain the similarity values between pairwise candidate music videos; For each candidate music video list, filter the similarity values by adapting the maximum greedy algorithm to obtain the recommended music video list; After associating the access link of the music video in the recommended music video list, push the recommended music video list to the user.

5. The music video recommendation method according to claim 4, wherein, After associating the access link of the music video in the recommended music video list, push the recommended music video list to the user, including the following steps: Obtain the unique feature information corresponding to each music video in the recommended music video list from the knowledge graph; Obtain the content information corresponding to each music video from the music library according to the unique feature information, and the content information includes the access link, cover information, audio information, and text information of the corresponding music video; Format the content information and associate it with the corresponding music video in the recommended music video list; Push the recommended music video list to the user.

6. A music video recommendation device, wherein, comprising: A request response module, configured to respond to a user's music video recommendation request and determine the user's historical visited music video list according to the user's historical behavior data; An interest recall module, configured to obtain a user interest vector according to the historical visited music video list and recall a relevant music video list matching the user interest vector from the music library knowledge graph; A data preliminary screening module, which is used to construct a comprehensive coding vector according to the historical access music video lists corresponding to multiple historical spans of a user, and screen out a candidate music video list from the relevant music video list according to the deep semantic information of the comprehensive coding vector, including: obtaining the long-term historical access music video list and the short-term historical access music video list of the user respectively according to two historical spans representing long-term and short-term, wherein the long-term historical span covers the short-term historical span; respectively vectorizing the long-term historical access music video list and the short-term historical access music video list and splicing them into a high-dimensional vector to obtain a comprehensive coding vector; using a recommendation ranking model pre-trained to a convergent state to perform representation learning on the comprehensive coding vector to obtain its deep semantic information, and performing multi-classification mapping according to the deep semantic information to obtain a candidate music video list corresponding to multiple preset targets, and each candidate music video list is sorted according to the click access statistical quantity, effective play statistical quantity, and complete play statistical quantity of the candidate music video; A rearrangement processing module, which is used to calculate the similarity of the deep semantic vectors of two candidate music videos in the candidate music video list, perform duplicate removal processing on the candidate music videos with similar compositions, and obtain the recommended music video list of the recommendation request to answer the recommendation request.

7. A computer device, including a central processing unit and a memory, characterized in that, the central processing unit is used to call and run a computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, it stores a computer program implemented according to the method according to any one of claims 1 to 5 in the form of computer-readable instructions, and when the computer program is called and run by a computer, it executes the steps included in the corresponding method.

9. A computer program product, including a computer program / instructions, characterized in that, when the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Project information processing method and device based on project recommendation model

    CN111026858A

  • Feature extraction network training method, information recommendation method, device and equipment

    CN113672820A