Recommendation system optimization method and device, and computer storage medium
By constructing a user preference model and predicting the actual preference needs of users, the problem of existing recommendation systems ignoring the differences in individual preferences of users and failing to capture changes in user interests is solved, and the personalization and accuracy of the recommendation system are improved.
Patent Information
- Application Number
- CN202411898758.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-13
AI Technical Summary
When generating recommendation results, the existing recommendation system ignores the preference differences between individual users and fails to effectively capture the dynamic changes in user interests and needs, resulting in deviations in recommendation results.
By obtaining the recommended content list of users in each preset period, calculating the similarity of the recommended content list, obtaining individual deviation data; sorting the deviation data to construct an individual deviation time series; extracting statistical features to construct a user preference model; using the model to predict the user's actual preferences, and generating a recommended content list that meets actual preferences.
It improves the personalization and accuracy of the recommendation system, can more effectively capture the changing trends of user preferences, and provide recommendation results that are more in line with user needs.
Smart Images

Figure CN119988722A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an optimization method, device and computer storage medium for a recommendation system. Background Art
[0002] With the development of Internet technology, recommendation systems are widely used in various Internet platforms. Recommendation systems mainly provide users with goods, content or services of interest by analyzing users' historical behavior data, such as browsing and evaluation. In the process of generating recommendation results, most current recommendation systems rely on analyzing the behavioral similarities between user groups to infer the preferences of target users.
[0003] However, current recommendation systems often ignore the differences in preferences between individual users while relying on the similarity of group behavior. In particular, as time goes by, users' interests and needs change dynamically, and the recommendation system fails to effectively capture the changing trend of such preferences, resulting in deviations in generating recommendation results. Summary of the invention
[0004] The main purpose of this application is to provide a recommendation system optimization method, device and computer storage medium, aiming to solve the technical problem of how to improve the accuracy of the recommendation system.
[0005] To achieve the above purpose, an embodiment of the present application provides a method for optimizing a recommendation system, and the method for optimizing a recommendation system includes: Obtaining a recommended content list of a user in each preset period, calculating the similarity of the recommended content lists corresponding to the user in different preset periods, and obtaining individual deviation data; The individual deviation data are sequentially sorted according to the time sequence of the preset period to construct an individual deviation time series; Extracting statistical features from the individual deviation sequence, and constructing a user preference model based on the statistical features; The user preference model is used to predict the actual preferences of the user, and a recommended content list that meets the actual preferences is generated.
[0006] In one embodiment, the step of obtaining a recommended content list of a user in each preset period, calculating the similarity of the recommended content lists corresponding to the user in different preset periods, and obtaining individual deviation data includes: Obtaining a list of the recommended contents received by the user in each of the preset periods, and recording the interactive behavior of the user on each of the recommended contents; quantifying the interaction behavior; The individual deviation data is calculated by comparing the deviation between the actual interaction behavior response of the user and the expected response.
[0007] In one embodiment, the step of quantifying the interaction behavior includes: Converting the interaction behavior into numerical data; A weight is assigned to each of the interactive behaviors according to the behavior type of the interactive behaviors.
[0008] In one embodiment, before the step of obtaining the recommended content lists of the user in each preset period, calculating the similarity of the recommended content lists corresponding to the user in different preset periods, and obtaining individual deviation data, the step includes: Determining a preference tag of the user according to the user's historical feedback behavior on the recommended content list; When user behavior or system events are detected to meet the preset trigger conditions, the current interaction log is recorded; Based on the interaction log, the recommended content list received by the user in each preset period is determined.
[0009] In one embodiment, the step of extracting statistical features from the individual deviation sequence further includes: Performing descriptive statistical analysis on the individual deviation time series, and calculating the central tendency and dispersion index of the individual deviation time series; performing an autocorrelation analysis on the individual deviation time series to determine the correlation between each time point in the individual deviation time series; Perform frequency domain analysis on the individual deviation time series to extract the periodic characteristics of the individual deviation time series.
[0010] In one embodiment, the step of extracting statistical features from the individual deviation sequence and constructing a user preference model based on the statistical features includes: Calculating the correlation coefficient between each of the statistical features and the user's preference tag; Determine the statistical features whose absolute values of the correlation coefficients are greater than a preset threshold as a key feature subset; Dividing the data including the key feature subset and the corresponding preference labels into a training set and a test set; The data of the training set is input into the model for iterative training until the performance index of the model on the test set reaches a preset standard or reaches a preset number of iterations, thereby obtaining a user preference model.
[0011] In one embodiment, the step of using the user preference model to predict the user's actual preferences and generating a list of recommended content that meets the actual preferences includes: Acquire the latest interaction data of the user, convert the latest interaction data into statistical features, and input the statistical features into the user preference model to obtain the prediction result of the user's current actual preference; Based on the prediction results, screening and matching are performed in the recommended content library to generate the recommended content list.
[0012] In one embodiment, the step of screening and matching in the recommended content library based on the prediction result to generate the recommended content list further includes: Obtaining a category label of the recommended content; Calculating the matching degree between the user's preference label and the category label; The order of the recommended contents in the recommended content list is adjusted according to the order of the matching scores from high to low.
[0013] An embodiment of the present application also provides an optimization device for a recommendation system, the optimization device for the recommendation system comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the optimization method for the recommendation system as described above.
[0014] An embodiment of the present application further provides a computer storage medium, which is a computer-readable storage medium. A computer program is stored on the computer storage medium. When the computer program is executed by a processor, the steps of the optimization method of the recommendation system described above are implemented.
[0015] The embodiment of the present application discloses a method for optimizing a recommendation system, which obtains the recommended content list of a user in each preset period, calculates the similarity of the recommended content list corresponding to the user in different preset periods, and obtains individual deviation data; sorts the individual deviation data in the time sequence of the preset period to construct an individual deviation time series; extracts statistical features from the individual deviation series, and constructs a user preference model based on the statistical features; uses the user preference model to predict the actual preference of the user, and generates a recommended content list that meets the actual preference. The present application improves the personalization and accuracy of the recommendation system by constructing a user preference model to predict the actual preference needs of the user. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a flowchart of a first embodiment of an optimization method for a recommendation system involved in an embodiment of the present application; Figure 2This is a flow chart of a second embodiment of an optimization method for a recommendation system according to an embodiment of the present application; Figure 3 This is a flow chart of a third embodiment of an optimization method for a recommendation system according to an embodiment of the present application; Figure 4 This is a flow chart of a fourth embodiment of an optimization method for a recommendation system according to an embodiment of the present application; Figure 5 This is a flow chart of a fifth embodiment of an optimization method for a recommendation system according to an embodiment of the present application; Figure 6 Schematic diagram of the structure of the optimization equipment recommended for this application.
[0017] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0018] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0019] With the development of Internet technology, recommendation systems are widely used in various Internet platforms. Recommendation systems mainly provide users with goods, content or services of interest by analyzing users' historical behavior data, such as browsing and evaluation. In the process of generating recommendation results, most current recommendation systems rely on analyzing the behavioral similarities between user groups to infer the preferences of target users.
[0020] However, current recommendation systems often ignore the differences in preferences between individual users while relying on the similarity of group behavior. In particular, as time goes by, users' interests and needs change dynamically, and the recommendation system fails to effectively capture the changing trend of such preferences, resulting in deviations in generating recommendation results.
[0021] In order to solve the above defects existing in the related art, the embodiment of the present application proposes a method for optimizing the recommendation system, which obtains the recommended content list of the user in each preset period, calculates the similarity of the recommended content list corresponding to the user in different preset periods, and obtains individual deviation data; sorts the individual deviation data in the time sequence of the preset period to construct an individual deviation time series; extracts statistical features in the individual deviation series, and constructs a user preference model based on the statistical features; uses the user preference model to predict the actual preference of the user, and generates a recommended content list that meets the actual preference. The present application improves the personalization and accuracy of the recommendation system by constructing a user preference model to predict the actual preference needs of the user.
[0022] It should be noted that the execution subject of this embodiment may be an optimization system of a recommendation system, or a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an optimization device of a recommendation system capable of realizing the above functions, etc. The following takes the optimization system of a recommendation system as an example to illustrate this embodiment and the following embodiments.
[0023] Please refer to the optimization method of the recommendation system of the first embodiment proposed in this application. Figure 1 The method comprises steps S10 to S40: Step S10: obtaining a recommended content list of a user in each preset period, calculating the similarity of the recommended content lists corresponding to the user in different preset periods, and obtaining individual deviation data.
[0024] In the recommendation system, the content recommended to users changes over time and under the influence of various factors. In order to better optimize the recommendation system so that it can make recommendations more accurately, stably and in line with the real interests of users, it is necessary to analyze the correlation between the recommended content received by users in different time periods.
[0025] It should be noted that the preset cycle refers to a pre-set time interval, such as every hour, every day, etc. The recommendation situation at different stages is divided by a fixed time period, which facilitates comparative analysis of the recommended content recommended to users at different stages.
[0026] The recommended content list refers to the collection of content that the recommendation system pushes to users within a certain preset period, which depends on the scenario in which the recommendation system is used. For example, the recommended content list of an e-commerce platform is a collection of product information, and the recommended content list of a video platform is a collection of video resources.
[0027] It should be clear that before performing any operation to obtain the recommended content list of any user, the embodiment of the present application will inform the user of information such as the data collection action, content, and purpose of use through pop-up windows, agreement confirmation, etc.
[0028] With the user's explicit authorization, the recommended content list is obtained. For user privacy data that may be involved, it is also used only after obtaining the user's authorization to protect user rights and data security.
[0029] In an optional embodiment, step S10 may include steps S11 to S13: Step S11: obtaining the list of recommended contents received by the user in each preset period, and recording the interactive behavior of the user on each recommended content.
[0030] Interactive behavior refers to various feedback from users on recommended content, such as clicks, views, likes, etc., which reflects the user's interest in and acceptance of the recommended content.
[0031] Step S12: quantifying the interaction behavior.
[0032] Quantification can express users' interactive behaviors in numerical form for mathematical calculations and analysis. For example, a click behavior can be recorded as 1, a purchase behavior can be recorded as 5, and so on. By quantifying the interactive behaviors, the behaviors can be made measurable, making it easier to analyze the differences in users' attention to different recommended content under a unified standard.
[0033] Further, step S12 includes steps S121-S122: Step S121: converting the interaction behavior into numerical data.
[0034] When obtaining the recommended content list, it also includes matching the corresponding assignment rules according to the behavior type corresponding to each recommended content in the recommended content list, such as click, watch, like, share, etc., so as to quantify the interactive behavior and form a data form in which the numerical value reflects the user's attention to each recommended content. For example, the user's click behavior can be mapped to the response value to the recommended content.
[0035] The distribution of values can be linear, such as click=1, view=2, like=3, purchase=4, or non-linear, such as click=1, view=3, like=5, purchase=10.
[0036] For example, set click = 1, view = 2, like = 3, purchase = 4. Within a preset period, a user clicks 5 times, views 2 times, and likes 1 time for a certain video A01. Then the click, view, and like scores of the user's video A01 are 5, 4, and 3 respectively.
[0037] Step S122: assigning a weight to each of the interactive behaviors according to the behavior type of the interactive behaviors.
[0038] Different interactive behaviors can reflect different levels of user interest and engagement. For example, purchase behavior is more indicative of user interest than click behavior. Therefore, a weight value can be assigned to each interactive behavior type to reflect the relative importance of the behavior type in the judgment of user preference.
[0039] For example, the weight of click behavior is 0.1, the weight of viewing is 0.2, the weight of like behavior is 0.3, and the weight of purchase behavior is 0.4. When the click, viewing, and like scores of a certain video A01 of a user are 5, 4, and 3 respectively, when calculating the comprehensive interaction score of the user for the recommended content, the weighted score can be calculated by applying the weight, that is, the comprehensive interaction score is: (5×0.1) based on the behavior type of the interaction behavior + based on the behavior type of the interaction behavior (4×0.2) based on the behavior type of the interaction behavior + based on the behavior type of the interaction behavior (3×0.3) == 2.2.
[0040] In this embodiment, after the specific interactive behaviors in the user's recommended content list are quantified, the quantified values are used as components of the feature vector of the recommended content list, and their size and distribution reflect the degree of user interaction with different recommended contents. The similarity is calculated based on these feature vectors to evaluate the consistency of the recommended content list at different time points.
[0041] Specifically, after quantifying the user's interactive behavior into numerical values, each recommended content can be represented as a multi-dimensional feature vector, where each dimension corresponds to the quantitative value of a type of interactive behavior. For example, if a user clicks, watches, and likes a video, the quantitative values of these behaviors will constitute the feature vector of the video. In this way, the entire list of recommended content is converted into a set of points in a high-dimensional space, and each point represents a recommended content.
[0042] Next, the similarity between these feature vectors is calculated using methods such as the Jaccard coefficient or cosine similarity. The Jaccard coefficient determines the similarity of the recommended content lists at different time points by comparing the ratio of the intersection size to the union size between the sets of two recommended content lists. The cosine similarity calculates the cosine value of the angle between the two feature vectors to obtain the calculation result of the similarity of the two recommended content lists.
[0043] Step S13: Calculate the individual deviation data by comparing the deviation between the actual interaction behavior response of the user and the expected response.
[0044] In another optional implementation, users are divided into multiple user groups according to their different attribute characteristics (such as age, gender, region, occupation, etc.). For each user group, a recommended content list corresponding to the user group within a preset period is obtained, and then a statistical method such as a chi-square test or KL divergence is used to calculate the distribution difference of the recommended content lists between different user groups, thereby obtaining group deviation data.
[0045] Then, based on the group bias data, we analyze the behavior patterns and preference characteristics of different user groups, explore the factors that lead to group bias, and then adjust and optimize the recommendation strategy of the recommendation system for different user groups.
[0046] In another optional implementation, based on the category labels of the recommended content list received by the user who recommended the content within a preset period, an evaluation method such as entropy or Gini coefficient (Gini coefficient / index) is used to calculate the diversity of the types of recommended content in the recommended content list, such as the distribution ratio of the category labels of the recommended content.
[0047] Then, the diversity feature data is integrated into the set of statistical features to train the user preference model. During the training of the user preference model, a penalty term of the diversity target can be added to the loss function. When the category label of the recommended content in the output prediction result is lower than the diversity threshold, the parameters of the user preference model are adjusted.
[0048] In this way, the user preference model can output prediction results containing multiple category labels when predicting user preferences, thereby ensuring that the recommendation system maintains the diversity of recommended content when generating a list of recommended content.
[0049] Furthermore, the proportion of users covered by the recommendation system to the total number of users may be calculated based on the coverage of the recommendation system to users, that is, the user coverage rate of the recommendation system may be calculated when the recommendation content is generated.
[0050] Specifically, user coverage is first defined, and if a user receives at least one recommended content within a specific period, the user is considered to be covered. Then, the ratio between the number of users covered by the recommendation system and the total number of users within the preset period is calculated.
[0051] It should be noted that the total number of users refers to the number of active users in the current time period.
[0052] Next, the feature data of user coverage is also integrated into the set of statistical features to train the user preference model. By introducing user coverage as the optimization target, the user preference model can cover multiple user groups when outputting prediction results, especially users in niche or marginal groups.
[0053] Step S20: sorting the individual deviation data in sequence according to the time sequence of the preset period to construct an individual deviation time series.
[0054] After the individual deviation data is obtained, since the individual deviation data is calculated corresponding to different preset periods, it reflects the deviation presented when the user receives the recommended content in each time period.
[0055] The individual deviation data is sorted in the time sequence of the preset period, with the aim of orderly integrating the deviation information scattered at different time points and constructing an individual deviation time series. In this way, the trend of individual deviations changing over time can be intuitively displayed, such as whether the deviation is gradually increasing, gradually decreasing, or showing periodic fluctuations or trends.
[0056] Determine the specific time label of the preset period corresponding to each individual deviation data, such as a specific date, hour period, etc., and then arrange the individual deviation data one by one in chronological order to form an orderly time series.
[0057] Exemplarily, the preset period is in hours, and the individual deviation data corresponding to the user at 18:00, 19:00, 20:00, etc. on a certain day are obtained, then the individual deviation data at 18:00 is placed at the front, and the individual deviation data is arranged in sequence to form an individual deviation time series.
[0058] In an optional implementation, a visualization tool is used to display the individual deviation time series. For example, a line chart or other chart format is used, with time as the horizontal axis and the value of the individual deviation data as the vertical axis, so as to clearly and intuitively display the fluctuation of the individual deviation over time through visualization.
[0059] Step S30: extracting statistical features from the individual deviation sequence, and constructing a user preference model based on the statistical features.
[0060] In this embodiment, the individual deviation time series includes individual deviation data arranged in chronological order, and the statistical feature is to use statistical methods to mine the information characteristics in the individual deviation data.
[0061] For example, the mean is calculated to understand the average level of individual deviation data within a preset time period, reflecting the roughly stable state of user preferences for recommended content; the variance is calculated to determine the degree of discreteness of individual deviation data and observe the fluctuation of user preferences; the maximum and minimum values are counted to grasp the extremes of individual deviations, and thus understand the maximum and minimum degrees to which user preferences deviate from the norm.
[0062] By extracting statistical features, the overall characteristics of individual deviation data can be obtained from different angles.
[0063] After obtaining the above statistical features, a user preference model is constructed based on these features. The user preference model aims to use the extracted statistical features to simulate and present the user's preferences for recommended content in different time periods. For example, based on the statistical features, it can be determined which types of recommended content the user prefers in a certain time period, how stable the interest in different recommended content is, etc., so that the recommendation system can adjust the recommendation strategy for the user in different time periods according to the user preference model.
[0064] In an optional implementation scheme, when constructing the user preference model, supervised machine learning algorithms such as decision trees and random forests can be used to construct it.
[0065] Specifically, the extracted statistical features are first converted into a feature matrix, where each row represents a statistical feature vector corresponding to a preset period, and each column corresponds to a specific statistical feature, such as a mean column, a variance column, etc. At the same time, the corresponding user's real feedback data on the recommended content is obtained as label data. For example, if the user has purchased the recommended content within a preset period, the label is set to 1, otherwise it is set to 0. Or if the user is more inclined to watch photography-related videos within a preset period, the statistical feature vector within the preset period will be associated with the user's preference for photography videos.
[0066] Then, a machine learning algorithm, such as a logistic regression model, is used. The feature matrix is used as input, the label data is used as output, and the data is divided into a training set and a test set. The logistic regression model is trained using the training set, and the parameters of the model are adjusted through an optimization algorithm (such as the gradient descent method) to minimize the loss function (such as the cross entropy loss) of the model on the training set.
[0067] Then use the test set to evaluate the trained model and calculate evaluation indicators such as accuracy, recall, and F1 value. If the evaluation indicators indicate that the model performance is not ideal, the model can be optimized by adjusting the model's hyperparameters, such as the regularization coefficient, or increasing the amount of data and performing feature engineering, such as feature combination and feature scaling. For example, grid search can be used to find the optimal hyperparameters.
[0068] When the model performance meets the requirements, it can be applied to new individual deviation data. The statistical features of the new data are input into the model, and the model will output the prediction results of the user's preference for recommended content. The recommendation system adjusts the recommendation strategy based on these results. For example, if the prediction result is that in the T1 time period, the user tends to watch photography-related content, the recommendation system can increase the weight or priority of the recommended content under the photography category.
[0069] In another optional implementation, a probability model can be used to construct a user preference model. Taking the Bayesian network as an example, the network structure is first determined based on the relationship between user preferences and statistical characteristics. For example, if it is found that the user's age characteristics are directly related to the preference for a certain type of recommended content, and the regional characteristics indirectly affect the preference by affecting consumption habits, then the corresponding nodes and connections are established in the network structure.
[0070] Then, we use historical data containing statistical features and user preference results to learn the conditional probability table in the Bayesian network. For example, we can calculate the probability that users within a certain age range have a preference for specific recommended content, as well as the preference probability under different regions and age combinations.
[0071] When there is new individual deviation data, the new statistical feature value is input into the trained Bayesian network, and the user's preference probability for different recommended content is calculated through probabilistic reasoning. For example, given the age and region of a new user, the network can be used to infer the probability distribution of their preference for various types of recommended content. The recommendation system then screens and sorts the recommended content based on the probability distribution of preference, giving priority to content with a higher probability.
[0072] Step S40: using the user preference model to predict the user's actual preference, and generating a list of recommended content that meets the actual preference.
[0073] Based on the user preference model constructed above, the corresponding input data is provided to the user preference model, so that the user preference model can output the prediction results of user preferences according to the rules and characteristics it has learned.
[0074] For example, a logistic regression model based on a machine learning algorithm inputs the statistical features corresponding to the new individual deviation data, and the model outputs the classification results of the user's preference for recommended content, such as whether they prefer or not a certain type of content. For a Bayesian network model, the probability distribution of preferences for different recommended content is output, so as to determine which recommended content the user may currently prefer, providing a basis for the subsequent generation of a recommendation list.
[0075] Based on the actual user preferences predicted by the model, the recommendation system will filter out recommended content that meets the preferences from the recommended content library, sort and integrate them according to the rules, and finally generate a recommended content list to present to the user.
[0076] Recommended content can be sorted according to popularity (such as number of views, number of likes, etc.), release time (latest first), relevance to past user interactions (such as similar content that has been clicked on in the past is ranked higher), and other rules, ultimately forming a recommended content list that is pushed to the user.
[0077] Please refer to Figure 2The optimization method of the recommendation system of the second embodiment proposed in this application includes steps S110 to S130 before step S10: Step S110: determining the user's preference tag according to the user's historical feedback behavior on the recommended content list.
[0078] Preference labels are marks of user interest preferences. By analyzing users' past feedback on recommended content lists, such as which recommended content users clicked to view, browsed for a long time, liked, and purchased, or which content users directly ignored or quickly swiped past, we can get labels that represent user interest tendencies.
[0079] In this embodiment, the user's preference label may be determined based on the interaction frequency of the user's historical feedback behavior with respect to the category label of each recommended content.
[0080] Step S120: When it is detected that the user behavior or system event meets the preset trigger condition, the current interaction log is recorded.
[0081] Step S130: Based on the interaction log, determine the recommended content list received by the user in each preset period.
[0082] The preset trigger conditions usually include the following three situations: user behavior trigger, system time trigger and timer trigger. The interaction log includes user behavior log and recommended content log, which includes interaction information such as user ID, behavior type, behavior timestamp, recommended content ID and recommendation timestamp.
[0083] User behavior triggering refers to the recording operation of the current user behavior log being triggered when the user performs a specific behavior, such as searching, clicking, or watching.
[0084] System event triggering means that when the recommendation system generates a list of recommended content, the recommended content log is automatically recorded, regardless of whether the user interacts with the recommended content.
[0085] Scheduled triggering is achieved by setting a scheduled task, which will periodically trigger and record user behavior logs and recommended content logs.
[0086] According to the preset period, the user behavior log and the recommendation system's recommendation content log are traversed, and the interaction records within the preset period are filtered out according to the recommendation timestamp, and then the interaction records are classified and divided according to different preset periods.
[0087] For each set of interaction records within a preset period, extract the user ID, recommended content ID, and recommendation timestamp data. Since users may have the same behavior on the same recommended content multiple times during the interaction process, such as viewing the same video multiple times, the extracted recommended content IDs may be repeated. Therefore, the extracted data needs to be preprocessed, including data cleaning and deduplication.
[0088] For records with the same recommended content ID and the same behavior type, the recommendation system will retain the user's interaction frequency information for the same recommended content and the same behavior type when performing deduplication operations, so as to assign weights to different recommended content in subsequent analysis, and thus better analyze the user's actual preferences and behavior patterns.
[0089] For the same user ID, based on the deduplicated data records, the corresponding complete recommended content information is queried from the recommended content library of the recommendation system. Finally, the recommended content list is generated through data structures such as arrays and lists. The recommended content list includes the recommended content ID, user ID, corresponding complete recommended content information, recommended content category identifier, recommendation timestamp, behavior type, and behavior frequency.
[0090] For example, the following records are recorded in the interaction log within a preset period: for the video of recommended content A, user B has 5 viewing behavior records and 1 like behavior record. When constructing the recommended content list within the preset period of user B, recommended content A retains two records, namely "recommended content ID: A, user ID: B, content information: C; category identification: sports, recommendation timestamp: 2024-10-05 14:25:30, behavior type: click, behavior frequency 5, ...", and "recommended content ID: A, user ID: B, content information: C; category identification: sports, recommendation timestamp: 2024-10-05 14:26:06, behavior type: like, behavior frequency 1, ...".
[0091] Please refer to Figure 3 , the optimization method of the recommendation system of the third embodiment proposed in this application, step S30 includes steps S31~S33: Step S31: performing descriptive statistical analysis on the individual deviation time series, and calculating the central tendency and dispersion index of the individual deviation time series.
[0092] Descriptive statistical analysis is to describe the basic characteristics of individual deviation time series in a general way through statistical indicators. Among them, the central tendency indicator is mainly used to reflect the typical value or central position in the data set, such as the mean, median and mode.
[0093] The mean is the average value of all individual deviation data; the median is the value in the middle after the data is arranged in order of size. It is insensitive to extreme values and can more robustly reflect the intermediate state of the data; the mode is the data value that appears the most times and can reflect the most common deviation in the data.
[0094] The dispersion index indicates the dispersion of data relative to the central position, such as variance, standard deviation, etc. Variance is the average of the sum of the squares of the differences between each data and the mean, and the standard deviation is the square root of the variance. The dispersion degree can show the magnitude of fluctuations in individual deviation data and help determine the stability of user preferences.
[0095] Step S32: performing autocorrelation analysis on the individual deviation time series to determine the correlation between each time point in the individual deviation time series.
[0096] Autocorrelation analysis is to determine the correlation between data at different time points in the individual deviation time series. In simple terms, it is to determine whether the individual deviation data at a certain time point is correlated with the individual deviation data at other time points and to what extent. For example, whether the preference deviation of users for recommended content today is correlated with yesterday or a historical time point.
[0097] This correlation is quantified by calculating the autocorrelation coefficient. The autocorrelation coefficient usually ranges from -1 to 1. A value close to 1 indicates a strong positive correlation, which means that the deviations at different time points show a trend of changing in the same direction; a value close to -1 indicates a strong negative correlation, that is, the deviations show a trend of changing in the opposite direction; a value close to 0 indicates a weak correlation, and the deviations at each time point are relatively independent.
[0098] Based on the correlation analysis of individual deviation data at each time point, the model can learn the possible trends in user preferences, and thus more accurately predict user preferences in different time periods based on the laws of this trend.
[0099] Step S33: performing frequency domain analysis on the individual deviation time series to extract periodic features of the individual deviation time series.
[0100] Frequency domain analysis is to convert the individual deviation time series from the time domain (time dimension) to the frequency domain (frequency dimension) for analysis, with the aim of extracting the periodic characteristics of the individual deviation time series.
[0101] In the time domain, the user preference model can learn how individual deviation data changes over time; in the frequency domain, the user preference model will learn the strength of different frequency components, that is, whether there are periodic fluctuations in individual deviation data and the length of the period. For example, through frequency domain analysis, it is found that the user's preference deviation for recommended content will show a similar change pattern every week or month, which indicates that there is a corresponding periodic pattern.
[0102] After extracting the periodic features, the user preference model can predict user preferences based on the periodic rules, and then optimize the recommendation strategy of the recommendation system. For example, in the time period when the periodic preference deviation is large, the type or order of recommended content can be adjusted.
[0103] Please refer to Figure 4 , the optimization method of the recommendation system of the fourth embodiment proposed in this application, step S30 also includes steps S34~S37: Step S34: Calculate the correlation coefficient between each of the statistical features and the user's preference tag.
[0104] The purpose of calculating the correlation coefficient is to clarify the degree of correlation between the statistical features and the labels representing the actual preferences of users. For example, the relationship between the mean and variance of the individual deviation data in the statistical features and the preferences represented by labels, such as whether the user purchased the recommended product or whether he prefers to watch a certain type of video.
[0105] The correlation coefficient can be calculated by Pearson correlation coefficient, Spearman rank correlation coefficient, etc. The correlation coefficient is used to determine to what extent the statistical feature is related to the user preference, whether it is positively correlated (coefficient greater than 0) or negatively correlated (coefficient less than 0).
[0106] Step S35: Determine the statistical features whose absolute values of the correlation coefficients are greater than a preset threshold as a key feature subset.
[0107] In this embodiment, since there are statistical features that are weakly correlated or irrelevant to user preferences, these features will interfere with the prediction of the model. Therefore, after obtaining the correlation coefficient between each statistical feature and the user preference label, the key features are screened according to the preset threshold of the correlation coefficient.
[0108] When the key features are strongly correlated with user preferences, the model can learn the user preference patterns more accurately when building the model, which in turn helps improve the accuracy of the model's user preference predictions.
[0109] Step S36: Divide the data including the key feature subset and the corresponding preference labels into a training set and a test set.
[0110] Divide the data containing the key feature subsets and the corresponding preference labels into training sets and test sets.
[0111] The training set is used to allow the model to learn the logical relationship between key features and user preference labels, and user preferences continuously adjust their parameters based on the data in the training set.
[0112] The test set is used to test the generalization ability of the model after the model training is completed. The evaluation results of the test set can be used to determine whether the model is overfitting (i.e., it performs well on the training set but poorly on the test set) or underfitting (i.e., it performs poorly on both the training set and the test set), and then optimize the parameters or structure of the model.
[0113] When dividing the training set and the test set, a random number generation tool is used, and random sampling is performed according to a preset ratio to complete the division.
[0114] When using group-biased data for model training, a stratified sampling method is used to divide the training set and the test set.
[0115] Step S37: input the data of the training set into the model for iterative training until the performance index of the model on the test set reaches a preset standard or reaches a preset number of iterations, thereby obtaining a user preference model.
[0116] Iterative training allows the model to continuously adjust its parameters based on the data in the training set, so that the model's prediction results for the training data are as consistent as possible with the actual user preference labels.
[0117] During the iterative training process, the quality of the model prediction is measured through a loss function, such as the cross entropy loss function, and then the optimization algorithm, such as the gradient descent method, is used to update the model parameters and continuously reduce the value of the loss function, thereby improving the performance of the model. The model training is completed until the performance indicators of the model on the test set, such as accuracy and recall, reach the preset standards, or the preset number of iterations is reached, and the user preference model for predicting user preferences is obtained.
[0118] Please refer to Figure 5 In the optimization method of the recommendation system of the fifth embodiment proposed in the present application, step S40 includes steps S41-S42: Step S41: Acquire the latest interaction data of the user, convert the latest interaction data into statistical features, and input them into the user preference model to obtain the prediction result of the user's current actual preference.
[0119] When the user has the latest interaction data, the latest interaction data is converted into statistical features and input into the user preference model. The user preference model uses the rules and features learned to predict the user's actual preferences in the current time period. The user's latest interaction data contains various feedback behavior information made by the user in recent times for the recommended content, such as clicks, likes, purchases, etc. These feedback behavior information reflects the user's latest interest changes.
[0120] Step S42: Based on the prediction results, screening and matching are performed in the recommended content library to generate the recommended content list.
[0121] Based on the prediction results obtained in the previous steps, the recommended content library is screened and matched to generate a recommended content list, and then reasonably sorted according to the rules to finally form a recommended content list presented to the user. For example, if the prediction results show that the user currently prefers photography content, photography-related content is screened out from the recommended content library, and then the screened content is further sorted considering factors such as popularity, relevance, and release time, so that the recommended list is more attractive.
[0122] In an optional implementation, the relevance between the recommended content and the user's photography preference is determined by text keyword matching or the like. For example, natural language processing technology is used to extract keywords from the recommended content, and the frequency and weight of keywords related to photography are counted to screen out recommended content with a higher relevance score.
[0123] Further, step S42 includes steps S421 to S423: Step S421: Obtain the category label of the recommended content.
[0124] Step S422: Calculate the matching degree between the user's preference tag and the category tag.
[0125] Step S423: adjusting the order of the recommended contents in the recommended content list according to the order of the matching degree scores from high to low.
[0126] Get the category label of the recommended content, calculate the matching degree between the user's preference label and the category label, and determine whether the recommended content meets the user's preference.
[0127] In an optional implementation, when the preference tag and the category tag are both clear and fixed categories, such as photography, sports, music, etc., an exact matching method can be used. When the user's preference tag and the category tag of the recommended content are completely consistent, the matching score is set to 1; otherwise, it is set to 0.
[0128] In another optional implementation, text vector similarity calculation is used to convert text into vector representation, for example, a word vector model is used to map the label text into a vector, and then the matching score is determined by calculating the cosine similarity between the vectors.
[0129] Finally, based on the calculated matching score, the original recommended content list is rearranged so that the recommended content with a high matching degree (i.e., more in line with user preferences) is placed at the front and displayed to users first, thereby increasing the probability of users seeing content of interest and improving the accuracy of the recommendation system.
[0130] In another optional implementation, the system can also use post-processing technology to adjust the recommendation list to ensure the diversity and coverage of the recommendation results. Specifically, the recommendation system implements a re-ranking algorithm to promote content of a specific category to a higher position in the recommendation list according to preset diversity rules and user preferences, thereby ensuring that users can be exposed to a wider range of topics and categories when browsing recommended content.
[0131] In addition, the recommendation system ensures that the recommended content list covers multiple categories and topics through the diverse feature data of the recommended content, meeting the diverse information needs of users.
[0132] In order to further improve the user experience of the recommendation system, after the recommendation system generates a list of recommended content, it provides explanations for the recommended content. Specifically, the recommendation system uses rule-based methods or natural language generation technology to generate explanations for the recommendation results. These explanations will detail the source of the recommended content, such as the user's past behavior, interest preferences, or the influence of social networks.
[0133] After generating the explanation, the recommendation system will display the explanation on the user interface to help users understand why they received specific recommended content, such as through an information prompt "Because you often browse technology articles, the latest technology news is recommended to you."
[0134] The present application provides an optimization device for a recommendation system, the optimization device for a recommendation system comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the optimization method for the recommendation system in the above-mentioned embodiment 1.
[0135] Reference below Figure 6 , which shows a schematic diagram of the structure of an optimization device suitable for implementing the recommendation system of the embodiment of the present application. The optimization device of the recommendation system in the embodiment of the present application may include various hardware and software components for implementing the optimization method of the recommendation system. Figure 6The optimization device of the recommendation system shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0136] like Figure 6 As shown, the optimization device of the recommendation system may include a processing device 1001 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the optimization device of the recommendation system are also stored. The processing device 1001, ROM1002 and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the optimization device of the recommendation system to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows the optimization device of the recommendation system with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.
[0137] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0138] The optimization device for the recommendation system provided by the present application adopts the optimization device method for the recommendation system in the above embodiment, which can solve the technical problem of how to improve the accuracy of the recommendation system. Compared with the prior art, the beneficial effects of the optimization device for the recommendation system provided by the present application are the same as the beneficial effects of the optimization method for the recommendation system provided by the above embodiment, and the other technical features in the optimization device for the recommendation system are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0139] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0140] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0141] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the optimization method of the recommendation system in the above-mentioned embodiment.
[0142] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM: Random Access Memory), a read-only memory (ROM: Read Only Memory), an erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency: Radio Frequency), etc., or any suitable combination of the above.
[0143] The computer-readable storage medium may be included in the optimization device of the recommendation system; or may exist independently without being assembled into the optimization device of the recommendation system.
[0144] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the optimization device of the recommendation system, the optimization device of the recommendation system: obtains the recommended content list of the user in each preset period, calculates the similarity of the recommended content list corresponding to the user in different preset periods, and obtains individual deviation data; sorts the individual deviation data in sequence according to the time sequence of the preset period to construct an individual deviation time series; extracts statistical features in the individual deviation series, and constructs a user preference model based on the statistical features; uses the user preference model to predict the actual preferences of the user, and generates a recommended content list that meets the actual preferences.
[0145] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0146] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0147] The modules involved in the embodiments described in the present application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.
[0148] The readable storage medium provided in the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the optimization method of the recommendation system described above, and can solve the technical problem of how to improve the accuracy of the recommendation system. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the present application are the same as the beneficial effects of the optimization method of the recommendation system provided in the above embodiment, and will not be elaborated here.
[0149] An embodiment of the present application provides a computer program product, including a computer program, which implements the steps of the optimization method of the recommendation system as described above when the computer program is executed by a processor.
[0150] The computer program product provided in this application can solve the technical problem of how to improve the accuracy of the recommendation system. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiment of this application are the same as the beneficial effects of the optimization method of the recommendation system provided in the above embodiment, which will not be repeated here.
[0151] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent processing scope of the present application.
[0152] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.
[0153] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.
[0154] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for optimizing a recommendation system, characterized in that: The optimization method of the recommendation system includes: Obtaining a recommended content list of a user in each preset period, calculating the similarity of the recommended content lists corresponding to the user in different preset periods, and obtaining individual deviation data; The individual deviation data are sequentially sorted according to the time sequence of the preset period to construct an individual deviation time series; Extracting statistical features from the individual deviation sequence, and constructing a user preference model based on the statistical features; The user preference model is used to predict the actual preferences of the user, and a recommended content list that meets the actual preferences is generated.
2. The optimization method of the recommendation system according to claim 1, characterized in that: The step of obtaining a recommended content list of a user in each preset period, calculating the similarity of the recommended content list corresponding to the user in different preset periods, and obtaining individual deviation data comprises: Obtaining a list of the recommended contents received by the user in each of the preset periods, and recording the interactive behavior of the user on each of the recommended contents; quantifying the interaction behavior; The individual deviation data is calculated by comparing the deviation between the actual interaction behavior response of the user and the expected response.
3. The optimization method of the recommendation system according to claim 2, characterized in that: The step of quantifying the interactive behavior comprises: Converting the interaction behavior into numerical data; A weight is assigned to each of the interactive behaviors according to the behavior type of the interactive behaviors.
4. The method for optimizing a recommendation system according to claim 1, wherein: Before the step of obtaining the recommended content lists of the user in each preset period, calculating the similarity of the recommended content lists corresponding to the user in different preset periods, and obtaining individual deviation data, the method includes: Determining a preference tag of the user according to the user's historical feedback behavior on the recommended content list; When user behavior or system events are detected to meet the preset trigger conditions, the current interaction log is recorded; Based on the interaction log, the recommended content list received by the user in each preset period is determined.
5. The method for optimizing a recommendation system according to claim 1, wherein: The step of extracting statistical features from the individual deviation sequence further includes: Performing descriptive statistical analysis on the individual deviation time series, and calculating the central tendency and dispersion index of the individual deviation time series; performing an autocorrelation analysis on the individual deviation time series to determine the correlation between each time point in the individual deviation time series; Perform frequency domain analysis on the individual deviation time series to extract the periodic characteristics of the individual deviation time series.
6. The method for optimizing a recommendation system according to claim 1, wherein: The step of extracting statistical features from the individual deviation sequence and constructing a user preference model based on the statistical features includes: Calculating the correlation coefficient between each of the statistical features and the user's preference tag; Determine the statistical features whose absolute values of the correlation coefficients are greater than a preset threshold as a key feature subset; Dividing the data including the key feature subset and the corresponding preference labels into a training set and a test set; The data of the training set is input into the model for iterative training until the performance index of the model on the test set reaches a preset standard or reaches a preset number of iterations, thereby obtaining a user preference model.
7. The optimization method of the recommendation system according to claim 1, characterized in that: The step of using the user preference model to predict the user's actual preference and generating a recommended content list that meets the actual preference includes: Acquire the latest interaction data of the user, convert the latest interaction data into statistical features, and input the statistical features into the user preference model to obtain the prediction result of the user's current actual preference; Based on the prediction results, screening and matching are performed in the recommended content library to generate the recommended content list.
8. The method for optimizing a recommendation system according to claim 7, wherein: The step of screening and matching in the recommended content library based on the prediction result to generate the recommended content list further includes: Obtaining a category label of the recommended content; Calculating the matching degree between the user's preference label and the category label; The order of the recommended contents in the recommended content list is adjusted according to the order of the matching scores from high to low.
9. An optimization device for a recommendation system, characterized in that: The optimization device of the recommendation system comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the optimization method of the recommendation system according to any one of claims 1 to 8.
10. A computer storage medium, characterized in that: The computer storage medium is a computer-readable storage medium, and a computer program is stored on the computer storage medium. When the computer program is executed by a processor, the steps of the optimization method of the recommendation system according to any one of claims 1 to 8 are implemented.