Listening sequence music recommendation method based on combined time convolutional network and KAN
By combining temporal convolutional networks and KAN, a music recommendation model is constructed, which solves the shortcomings of existing recommendation systems in user interest modeling and dynamic perception, and achieves more efficient user interest extraction and personalized recommendations.
Patent Information
- Application Number
- CN202511001866.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-04
AI Technical Summary
Existing recommendation systems suffer from insufficient recommendation accuracy, incomplete utilization of time-series information, and lack of dynamic perception capabilities in terms of user behavior prediction and interest modeling, making it difficult to meet users' diverse and personalized needs and real-time changing recommendation requirements.
By combining temporal convolutional networks and Kolmogorov-Arnold networks (KANs), and through conditional density functions and attention mechanisms, a music recommendation model is constructed to extract key features from user item interaction sequences, thereby achieving joint modeling of long-term and short-term interests.
It significantly improves the utilization rate of user information and the effectiveness of recommendations, accurately depicts users' dynamic interests, and provides more precise personalized music recommendations.
Smart Images

Figure CN120892598A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the fields of data mining and recommendation, and particularly relates to a listening sequence music recommendation method based on a combined time convolution network and KAN. BACKGROUND
[0002] At present, the Internet content information grows exponentially, and a recommendation system becomes a core technology for solving the information screening dilemma by screening interesting items through predicting user behaviors, and user behavior prediction is a key to accurate modeling.
[0003] At present, the traditional recommendation method has the following technical defects: first, the recommendation accuracy is insufficient, and it is difficult to meet the increasingly diversified personalized needs of users; second, the mining depth of user-item interaction data is limited, and the time sequence information, behavior patterns and other key features contained in the historical interaction records cannot be fully utilized; third, the dynamic perception ability to the real-time demand change of users is poor, and it is difficult to provide effective recommendation services when the user behavior changes immediately.
[0004] Although the existing next item recommendation algorithm combines the traditional strategy and the interaction sequence data, the accuracy is improved, but there are still bottlenecks: the value of the sequence data is not fully utilized, and the user dynamic interest evolution process reflected in the sequence cannot be accurately modeled; the long-term and short-term interest collaborative modeling is insufficient, and it is difficult to realize the comprehensive and dynamic description of the user interest.
[0005] Therefore, how to efficiently analyze the user-item interaction sequence, accurately extract the key features of the items, and realize the joint modeling of the long-term and short-term dynamic interest of the user has become a core technical difficulty for breaking through the bottleneck of the existing technology and improving the performance of the recommendation system. SUMMARY
[0006] The application aims to overcome the deficiencies of the prior art and provides a listening sequence music recommendation method based on a combined time convolution network and KAN, which can significantly improve the user information utilization rate and improve the recommendation effect.
[0007] In order to achieve the above purpose, the technical scheme adopted by the application is as follows:
[0008] A listening sequence music recommendation method based on a combined time convolution network and KAN comprises the following steps:
[0009] Step 1, collecting all the music listening sequence data of users, and taking the music listening sequence of a target user as an ordered set of music listened by the user;
[0010] Step 2: Based on the target user's music listening sequence, use a temporal convolutional network to establish a conditional density function for the target user, the target user's historical listening sequence, and the target music, and then calculate the probability that the user is interested in the music through the conditional density function.
[0011] Step 3: Based on the probability of users being interested in music in the music listening sequence data of all users, construct a logarithmic objective function and solve for the conditional density function, thereby maximizing the long-term interest and short-term interest through the conditional density function.
[0012] Step 4: Calculate the user's interest value for each piece of music in all ordered sets based on long-term and short-term interests. Sort by interest value in descending order;
[0013] Step 5: Recommend the top K music tracks with the highest interest scores to the user.
[0014] Preferably, the conditional density function is:
[0015]
[0016] in: These are the target users u i Interest vectors and music m j m k eigenvectors, This refers to the historical listening sequence, F KAN (·) is the Kolmogorov-Arnold network, for It provides a learnable activation function that adaptively adjusts the non-linear mapping for each dimension of the data distribution, F. TCN (·) represents a temporal convolutional network. Indicates user u i For music m j The probability of being interested This refers to user u i Long-term interest in music j Feature similarity, This refers to the impact of short-term historical behavior (L) on user (u). j For music m j The degree of interest.
[0017] Preferably, the temporal convolutional network F TCN (·) is defined as four layers:
[0018]
[0019] Where: f = (f1, f2, ..., f k-1) is a filter, x is an input sequence, k and w are filter size and convolution network layer number respectively.
[0020] As preferred, in the step (2), the long-term interest of the user u i and the feature similarity of the music m j are defined by a cosine similarity function as follows:
[0021]
[0022] where: is the feature vector representation of the music m j , is the interest vector representation of the user u i .
[0023] As preferred, in the step (2), the historical behavior L influences the degree of the user u j interest in the music m j is defined as:
[0024]
[0025] where: is the feature vector representation of the historical music sequence L, is the interest vector representation of the user u i , is the attention mechanism weight perceived by the user and the historical behavior.
[0026] As preferred, the logarithmic objective function expression is as follows:
[0027]
[0028] where: represents the historical music listening sequence of the given user u i , the probability of the user u i being interested in the music m j .
[0029] As preferred, in the step (4), given the historical interaction record of the user u i , the interest value of the user u i in the music m j is defined as:
[0030]
[0031] where: f KAN (·) is a Kolmogorov-Arnold network, for A learnable activation function is provided, and the nonlinear mapping mode is adaptively adjusted according to the data distribution of each dimension, μm, u representing the long-term interest of the user, representing the short-term interest of the user, and t is the time threshold of the long and short term sequence.
[0032] As preferred, the ranking calculation formula defined in step (5) is as follows:
[0033]
[0034] Wherein: u represents a target user; m and m' are music in the database.
[0035] The present application has the following characteristics and beneficial effects:
[0036] The present application firstly utilizes the attention mechanism combining the convolutional network and the Kolmogorov-Arnold network (KAN) and the multi-dimensional Hoeksema process to establish a music recommendation model, obtains a key feature vector of an item from an item interaction sequence of a user, and provides a feasible method for solving the difficulty in item feature extraction; the present application accurately obtains the dynamic interest preference of a user according to the feature vector of an item in the interaction sequence of the user, and provides a reliable method for the extraction and modeling of the interest preference of the user; by comprehensively utilizing the key feature vector of the item and the dynamic interest of the user, the present application can improve the recommendation effect. BRIEF DESCRIPTION OF DRAWINGS
[0037] Figure 1 The figure is a schematic diagram of the system architecture of the recommendation method of the present application.
[0038] Figure 2 The figure is a schematic diagram of the user preference prediction process in the recommendation method of the present application. DETAILED DESCRIPTION
[0039] The present application will be described in detail below in combination with specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any form. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0040] A listening sequence music recommendation method based on the combined time convolutional network and KAN, as shown in Figure 1 and Figure 2 includes the following steps:
[0041] Step 1, collecting all user music listening sequence data, taking the music listening sequence of a target user as an ordered set of music listened by the user.
[0042] In this embodiment, all user music listening sequence data is collected User u i The music listening sequence is an ordered set of music listened to by the user. The user set and the music set are defined as U and M, respectively.
[0043] Step 2, based on user u i Music listening sequence Using temporal convolutional networks to connect user u i Historical listening sequence and target music m j We establish a conditional density function for (j≤n), and then calculate the probability that the user is interested in music using the conditional density function:
[0044]
[0045] in: These are the target users u i Interest vectors and music m j m k eigenvectors, This refers to the historical listening sequence, F KAN (·) is the Kolmogorov-Arnold network, for It provides a learnable activation function that adaptively adjusts the non-linear mapping for each dimension of the data distribution, F. TCN (·) represents a temporal convolutional network. Indicates user u i For music m j The probability of being interested This refers to user u i Long-term interest in music j Feature similarity, This refers to the impact of short-term historical behavior (L) on user (u). j For music m j The degree of interest.
[0046] Furthermore, the aforementioned user u i For the target music m j long-term interest The cosine similarity function is defined as follows:
[0047]
[0048] in: It's music m j The feature vector representation from music m j The feature tags, such as duration and music genre, are obtained after standardization. User u iis the feature vector representation of the interest vector of user u
[0049] The historical behavior L influences the user u i The degree of interest of the target music m j The degree of interest of the target music m is defined as:
[0050]
[0051] wherein: is the feature vector representation of the historical music sequence L, is the feature vector representation of the interest vector of user u i is the feature vector representation of the interest vector of user u is the attention mechanism weight perceived by the user and the historical behavior.
[0052] Step 3, based on the probability of the user's interest in the music in the music listening sequence data of all users, a logarithmic objective function is constructed and maximized to solve, and then the long-term interest and short-term interest are calculated by applying the conditional density function after maximization.
[0053] Specifically, given the music interaction sequence data of all users The objective function in logarithmic form can be defined as:
[0054]
[0055] wherein: is the feature vector representation of the music m in the music set M, is the music interaction sequence of the given user u i before time t The probability of the user u i being interested in the music m.
[0056] The objective function O is maximized by the gradient ascent optimization method.
[0057] After maximization, the optimal attention mechanism weight perceived by the user and the historical behavior is obtained, and then the degree of interest of the historical behavior L to the target music m i of the user u j is calculated. The optimal short-term interest of the user is obtained
[0058] Step 4, according to the long-term and short-term interests of the user, the interest value of each music in all ordered sets is calculated Sort by interest value in descending order.
[0059] Specifically, given user u i Historical interaction records, user u i Interest in music m is defined as:
[0060]
[0061] Where: f KAN (·) represents the KAN process, which is... It provides a learnable activation function that adaptively adjusts the nonlinear mapping for each dimension of the data distribution, μ m,u This indicates that it represents the user's long-term interests. This represents the user's short-term interests, and t is the time threshold of the long-short-term series.
[0062] The above exponential kernel function κ(tt) L ) is defined as:
[0063] κ(tt L )=exp(-δ u (tt L )),
[0064] Where: δ u These are user-related parameters used to represent the impact of historical behavior (h) on the target music (m) for different users. j The effects are different.
[0065] Sort all music in the database from highest to lowest based on the user's interest score, using the following formula:
[0066]
[0067] Where: u represents the target user; m∈M and m′∈M are the music in the database.
[0068] Step 5: Recommend the top K music tracks with the highest interest scores to the user.
[0069] Figure 1 This paper illustrates the architecture of the music recommendation method based on temporal convolutional networks combined with KAN and listening sequences in this implementation. The recommendation method consists of two main modules: a preprocessing module and a prediction module. In the preprocessing module, the music interaction sequences of all users are first obtained; then, a multidimensional Hawkes process model and an attention mechanism are used to learn the music feature vectors and the users' long-term interest vectors from the music sequences and corresponding temporal information. In the prediction module, the short-term dynamic interest preferences of the target users are first obtained from their music interaction sequences; then, suitable music is recommended to the users based on their interests and music feature vectors. Figure 2Detailed steps of the user preference prediction are shown, which first acquire the user's interactive behavior record, and extract the user's dynamic preference therefrom, and then calculate the target user u's preference for music by using the user's preference and the feature vector of music.
[0070] The above shows and describes the basic principles, main features and advantages of the present application. Those skilled in the art should understand that the present application is not limited to the above-mentioned embodiments, and the above-mentioned embodiments and descriptions in the specification are only preferred examples of the present application and are not intended to limit the present application. Various changes and improvements can be made to the present application without departing from the spirit and scope of the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection of the present application is defined by the appended claims and their equivalents.
Claims
1. A method for recommending music based on a combined temporal convolutional network and KAN, characterized in that, Includes the following steps: Step 1: Collect music listening sequence data from all users, and use the music listening sequence of the target user as an ordered set of music listened to by that user; Step 2: Based on the target user's music listening sequence, use a temporal convolutional network to establish a conditional density function for the target user, the target user's historical listening sequence, and the target music, and then calculate the probability that the user is interested in the music through the conditional density function. Step 3: Based on the probability of users being interested in music in the music listening sequence data of all users, construct a logarithmic objective function and maximize the conditional density function, and then apply the maximized conditional density function to calculate long-term interest and short-term interest. Step 4: Calculate the user's interest value for each piece of music in all ordered sets based on long-term and short-term interests. Sort by interest value in descending order; Step 5: Recommend the top K music tracks with the highest interest scores to the user.
2. The method according to claim 1, characterized in that, The conditional density function is: in: These are the target users u i Interest vectors and music m j m k eigenvectors, This refers to the historical listening sequence, F KAN (·) is the Kolmogorov-Arnold network, for It provides a learnable activation function that adaptively adjusts the non-linear mapping for each dimension of the data distribution, F. TCN (·) represents a temporal convolutional network. Indicates user u i For music m j The probability of being interested This refers to user u i For music m j Long-term interest This refers to the impact of short-term historical behavior (L) on user (u). j For music m j The degree of interest.
3. The method according to claim 1, characterized in that, The temporal convolutional network F TCN (·) is defined as four layers: Where: f = (f1, f2, ..., f k-1 ) is the filter, x is the input sequence, and k and w are the filter size and the number of convolutional network layers, respectively.
4. The method according to claim 2, characterized in that, In step (2), user u i For music m j Long-term interest is defined using the cosine similarity function as: in: It's music m j The eigenvector representation, User u i Interest vector representation.
5. The method according to claim 4, characterized in that, In step (2), historical behavior L influences user u j For music m j The degree of interest is defined as: in: It is the feature vector representation of the historical music sequence L. User u i Interest vector representation, It is the weight of the attention mechanism for user and historical behavior perception.
6. The method according to claim 1, characterized in that, The logarithmic objective function is expressed as follows: in: Indicates a given user u i Historical music listening sequence In the middle, user u i For music m j The probability of being interested.
7. The method according to claim 5, characterized in that, In step (4), given user u i Historical interaction records, user u i For music m j Interest value is defined as: Where: f KAN (·) is the Kolmogorov-Arnold network, for It provides a learnable activation function that adaptively adjusts the nonlinear mapping for each dimension of the data distribution, μm. u This indicates that it represents the user's long-term interests. This represents the user's short-term interests, and t is the time threshold of the long-short-term series.
8. The method according to claim 1, characterized in that: The sorting calculation formula in step (5) is defined as follows: Where: u represents the target user; m∈M and m′∈M are the music in the database.