A deep learning recommendation system that fuses item audience features

CN116452293BActive Publication Date: 2026-09-29CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310416927.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-18
Publication Date
2026-09-29
Estimated Expiration
2043-04-18

AI Technical Summary

Technical Problem

然而,随着神经网络层数的加深,在网络前向传播过程中可能存在信息丢失的问题

Benefits of technology

[0054]本发明的有益效果在于:本发明利用注意力机制从用户与物品的历史交互记录中自适应地计算出物品的受众特征,将其作为反映目标用户偏好的重要补充信息,从而增强应对数据稀疏性问题的能力。提出了特征交互层,能够显式地控制各种特征的交叉,丰富特征交叉的方式,并通过残差连接来保证信息在神经网络前向传播的完整性。设计了低阶特征提取模块和高阶特征模块,同时学习数据中的低阶特征和高阶特征,以提高推荐性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116452293B_ABST
    Figure CN116452293B_ABST
Patent Text Reader

Abstract

The application provides a kind of fusion article audience feature deep learning recommendation system, belong to recommended technical field. Including data collection and data cleaning;Convert score into implicit feedback matrix;Get the historical interactive user list of article;Adaptive calculation of personalized audience features of article by using attention mechanism;Learn low-order features in data through linear regression and vector inner product;Design feature interaction layer to explicitly cross features;Further learn high-order features using multi-layer fully connected neural network;Fuse low-order and high-order feature information to output the predicted value of target user to article;Sort the prediction set and perform top-N recommendation;Through multiple experimental verification, the application can fully exploit the potential value in historical interaction information, improve the recommendation quality, and show good application potential.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This invention belongs to the field of recommendation technology, specifically a deep learning recommendation system that integrates the characteristics of the target audience for items. Background Technology

[0002] In the era of information overload, how to quickly filter personalized products and services suitable for users from massive amounts of information to improve user experience is a primary problem that internet digital platforms need to solve. Recommendation models, as important decision support tools, can predict the likelihood of a user accepting recommended products based on their historical behavior (such as ratings, clicks, or browsing of products), thus effectively alleviating information overload. Currently, recommendation systems are widely used in e-commerce, online marketing, news platforms, and many other fields.

[0003] While the introduction of deep learning methods has effectively enhanced the model's ability to perform non-linear modeling of data, these models struggle to effectively capture the intrinsic value information within the data. This is because they simply map user and item features into low-dimensional dense vectors, failing to emphasize the underlying mechanisms guiding recommendation applications and the effective mining of historical user-item interactions. Many models use user-item interaction information only for optimizing model parameters, ignoring potential collaborative information in the data that could reveal behavioral similarities between users, severely hindering further performance improvements. Secondly, existing deep learning-based recommendation models primarily utilize vector concatenation combined with neural networks to extract high-order features. However, as the number of neural network layers increases, information loss may occur during forward propagation. Furthermore, neural networks typically perform feature crossing implicitly, a process that is difficult to control, potentially leading to situations where they cannot effectively learn multiple feature combinations. Summary of the Invention

[0004] To address the aforementioned problems, this invention provides a deep learning recommendation system that integrates the characteristics of the target audience for items. The system includes a data acquisition and processing module, a model training module, a model prediction module, and a recommendation module, all housed in a server connected to multiple databases, which in turn connect to multiple smart terminals.

[0005] The smart terminal is used to collect user behavior data and store it in a database;

[0006] The database is used to store users' historical behavior information and is then processed by the server.

[0007] The data acquisition and processing module is used to process the data information in the database to obtain standardized data that meets the needs of the system.

[0008] The model training module is used to model and train standardized data until the model converges.

[0009] The model prediction module is used to predict users' personalized preferences using the trained model.

[0010] The recommendation module is used to sort items according to their predicted ratings and generate a list of recommended items to improve user satisfaction.

[0011] The deep learning recommendation system of this invention includes a data processing module that converts user feedback information on items into an implicit feedback matrix with values ​​of 0 or 1. A value of 1 indicates that the user has interacted with the item, while a value of 0 indicates that the user has not interacted with the item. Indicates the number of users. This represents the number of items. It collects the historical user interactions for each item and transforms the item and user information into a low-dimensional, dense feature vector. The main processing steps are as follows:

[0012] S1: Retrieve all data information: user ID, item ID, rating information, and the set of users who have interacted with the item in the past. ;

[0013] S2: Generate an implicit feedback matrix based on the scoring information. ;

[0014] S3: Convert user ID to a one-hot vector ;

[0015] S4: Convert project IDs to one-hot vectors ;

[0016] S5: Set up user groups Each user in the list is converted into their own one-hot vector;

[0017] S6: Randomly initialize the mapping matrix of users and items that follows a Gaussian distribution.

[0018] S7: Convert the user's onehot vector into a low-dimensional dense feature vector. The method is as follows:

[0019]

[0020] The deep learning recommendation system of the present invention comprises a model training module consisting of an item audience feature aggregation module, a low-order feature extraction module, a feature interaction module, a fully connected neural network module, and a model parameter update module.

[0021] The item audience feature aggregation module includes the following steps:

[0022] S1: Calculate the correlation coefficient between the target user and the historical users who interacted with the item using a multilayer perceptron-based attention mechanism, as shown below:

[0023]

[0024] in, These are the weight matrix, bias vector, regression coefficients, and activation function, respectively. For target users eigenvectors, The number of neurons in the hidden layer, and the activation function. ReLU was chosen to enhance nonlinear expressive power. This is a function for vector concatenation. This represents the element-wise product of vectors;

[0025] S2: Obtain the attention distribution by normalizing the correlation coefficients using the softmax function. The method is as follows:

[0026]

[0027] S3: Based on attention distribution The audience feature vector of an item is obtained by weighted summation, as shown below:

[0028]

[0029] in, For items Historical user interaction set For items The number of users who interacted. For users eigenvectors.

[0030] The low-order feature extraction module is used to learn low-order feature information in the data through linear regression and vector dot product, and its expression is as follows:

[0031]

[0032] in, This is the vector obtained by concatenating the one-hot vectors of the target user and the item. Feature vectors for the target user and the item, respectively. For the feature vector of the object's audience, For regression coefficients, This represents the vector dot product operation.

[0033] The feature interaction module is used to perform explicit feature crossing to obtain higher-order feature information and to maintain the integrity of information during the forward propagation of the neural network through residual connections. Specifically, it includes the following steps:

[0034] S1: Target user's feature vector Feature vectors of items and the audience feature vector of the item And stack all input vectors to form a specific matrix. ;

[0035] S2: Explicit feature crossing is performed by calculating the element-wise product between any two eigenvectors in the matrix, and the output matrix of the previous layer is used as the input of the next layer.

[0036] S3: Use residual connections to fuse the output matrices of each layer into a feature combination information vector. :

[0037]

[0038] in, The number of feature interaction layers. For the first The output matrix of the layer, Represents the first of the matrix OK.

[0039] The fully connected neural network module is used to combine feature information vectors. Based on this, higher-order features in the data are further learned through a multi-layer nonlinear neural network, the expression of which is as follows:

[0040]

[0041] in, This represents the number of layers in a fully connected neural network. The output of the feature interaction layer, Representing the first The layer's weight matrix, bias vector, and activation function are used. Finally, the predicted value is output by combining low-order and high-order feature information.

[0042] The model parameter update module is used to update the model parameters according to the loss value until convergence, and specifically includes the following steps:

[0043] S1: Construct the loss function, whose expression is as follows:

[0044]

[0045] in, The weight parameters are defined for the model embedding layer, the item audience feature aggregation module, the high-order feature extraction module, and the low-order feature extraction module, respectively. The feature interaction module and the fully connected neural network module are collectively referred to as the high-order feature extraction module. These represent the sets of positive and negative samples used in model training. For the true value, For predicted values, The coefficient of the regularization term;

[0046] S2: Based on the predicted value and the true value Calculate the loss value;

[0047] S3: Update model parameters using the Adam optimizer;

[0048] S4: Iterate through the set of positive and negative samples until the loss function value no longer decreases.

[0049] The deep learning recommendation system of this invention includes a model prediction module, which is used to obtain low-order and high-order feature information using a trained and fused model, fuse them, and output a predicted value, as expressed below:

[0050]

[0051] in, This is the output of the low-order feature extraction module. The output of the higher-order feature extraction module, activation function The sigmoid function has the following expression:

[0052]

[0053] The deep learning recommendation system of the present invention includes a recommendation module that sorts and recommends items based on the predicted ratings of each user for different items, and forms an item recommendation list, thereby improving the recommendation quality.

[0054] The beneficial effects of this invention are as follows: This invention utilizes an attention mechanism to adaptively calculate the audience characteristics of items from the historical interaction records of users and items, using this as important supplementary information reflecting the preferences of target users, thereby enhancing the ability to cope with data sparsity problems. A feature interaction layer is proposed, which can explicitly control the interaction of various features, enriching the ways in which features interact, and using residual connections to ensure the integrity of information during the forward propagation of the neural network. A low-order feature extraction module and a high-order feature module are designed to simultaneously learn low-order and high-order features from the data to improve recommendation performance. Attached Figure Description

[0055] Figure 1 This is a system block diagram of the present invention;

[0056] Figure 2 This is a model framework diagram of the present invention.

[0057] Figure 3 This is an example diagram of the feature interaction layer in this invention;

[0058] Figure 4 This invention demonstrates the variation of precision values ​​under different recommendation list lengths across three datasets.

[0059] Figure 5 This invention demonstrates the variation of recall values ​​under different recommendation list lengths across three datasets.

[0060] Figure 6 This invention demonstrates the variation of the f1 value under different recommendation list lengths across three datasets.

[0061] Figure 7 This invention demonstrates the variation of ndcg values ​​under different recommendation list lengths across three datasets. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of the present invention clearer, the specific embodiments and working principles of the present invention will be further described in detail below with reference to the accompanying drawings.

[0063] from Figure 1 As can be seen, a deep learning recommendation system integrating the characteristics of the target audience for items includes a data acquisition and processing module, a model training module, a model prediction module, and a recommendation module. These modules are located on a server connected to multiple databases, which in turn connect to multiple smart terminals. Users generate behavioral information on these smart terminals. The databases are responsible for collecting and storing behavioral data from different smart terminals. The data acquisition and processing module processes the data from the databases to obtain standardized data that meets the system's requirements. The model training module models and trains the standardized data until the model converges. The model prediction module uses the trained model to predict users' personalized preferences. The recommendation module sorts the items according to their predicted ratings and generates a list of recommended items to improve user satisfaction.

[0064] from Figure 2As can be seen, the deep learning recommendation system described in this invention comprises an item audience feature aggregation module, a low-order feature extraction module, a feature interaction module, a fully connected neural network module, and a model parameter update module. First, it obtains the historical user interaction set for each item and converts user IDs and item IDs into one-hot vectors. The one-hot vectors are then embedded to convert them into low-dimensional, dense feature vectors. An attention mechanism is used to calculate the correlation between the target user and users in the historical interaction set of the item, and personalized audience features for the item are adaptively obtained through weighted summation. In the low-order feature extraction module, linear regression and eigenvector inner product are used to learn low-order features in the data. A feature interaction layer is proposed to extract multi-order feature information combinations through element-wise product of corresponding feature vectors and residual connections. A multi-layer fully connected neural network is built to further learn potential high-order feature information in the data. Finally, the low-order and high-order feature information are fused to output predicted values.

[0065] Figure 3 This is an example diagram of the feature interaction layer in this invention. From... Figure 2 As can be seen, the feature interaction layer designed in this invention first stacks the input feature vectors into a specific matrix; it completes feature crossing by multiplying corresponding elements between feature vectors, and uses the output of the previous layer as the input of the next layer. Therefore, the order of the cross features increases as the feature interaction layer deepens; finally, it fuses the cross features captured by each layer through residual connections. This not only obtains multi-order feature combination information, but also maintains the integrity of information in the forward propagation of the neural network, laying a good foundation for the subsequent learning of higher-order features by the neural network.

[0066] Furthermore, the following example illustrates this further:

[0067] Assume there is individual users and Project This is used to convert user feedback on items into an implicit feedback matrix with values ​​of 0 or 1. A value of 1 indicates that the user has interacted with the item, while a value of 0 indicates that the user has not interacted with the item. Indicates the number of users. Indicates the number of items. Items Historical user interaction collection .

[0068] First, standardized data information that meets the system's needs is obtained through data collection and cleaning. The specific implementation steps of the proposed recommendation system are as follows:

[0069] S1: Retrieve all data information: user ID, item ID, rating information, and the set of users who have interacted with the item in the past. ;

[0070] S2: Generate an implicit feedback matrix based on the scoring information. Convert user IDs to one-hot vectors. Convert project IDs to one-hot vectors. ; Set up user groups Each user in the list is converted into their own one-hot vector;

[0071] S3: Randomly initialize the mapping matrix of users and items that follows a Gaussian distribution. Transform a one-hot vector into a low-dimensional, dense feature vector. The method is as follows:

[0072]

[0073] S4: Calculate the correlation coefficient between the target user and the historical users who interacted with the item using a multilayer perceptron-based attention mechanism, as shown below:

[0074]

[0075] in, These are the weight matrix, bias vector, regression coefficients, and activation function, respectively. For target users eigenvectors, The number of neurons in the hidden layer, and the activation function. ReLU was chosen to enhance nonlinear expressive power. This is a function for vector concatenation. The element-wise product of the representative vectors is used; the attention distribution is obtained by normalizing the correlation coefficients using the softmax function. The method is as follows:

[0076]

[0077] According to attention distribution The audience feature vector of an item is obtained by weighted summation, as shown below:

[0078]

[0079] S5: In the low-order feature extraction module, low-order feature information in the data is learned through linear regression and vector dot product, and its expression is as follows:

[0080]

[0081] S6: Target user's feature vector Feature vectors of items and the audience feature vector of the item And stack all input vectors to form a specific matrix. Explicit feature crossing is performed by calculating the element-wise product of any two eigenvectors in the matrix, and the output matrix of the previous layer is used as the input of the next layer; residual connections are used to fuse the output matrices of each layer into a feature combination information vector. :

[0082]

[0083] S7: In the fully connected neural network module, in the feature combination information vector Based on this, higher-order features in the data are further learned through a multi-layer nonlinear neural network, the expression of which is as follows:

[0084]

[0085] in, This represents the number of layers in a fully connected neural network. The output of the feature interaction layer, Representing the first Layer weight matrix, bias vector, activation function.

[0086] S7: The predicted value is output by fusing low-order and high-order feature information, and its expression is as follows:

[0087]

[0088] in, This is the output of the low-order feature extraction module. The output of the higher-order feature extraction module, activation function The sigmoid function has the following expression:

[0089]

[0090] S8: Train parameters and complete the recommendation. Construct a loss function based on cross-entropy error and use... Regularization prevents overfitting, and its expression is as follows:

[0091]

[0092] in, These are the weight parameters for the embedding layer, aggregation layer, high-order feature extraction module, and low-order feature extraction module, respectively. These represent the sets of positive and negative samples used in the training process, respectively. For the true value, For predicted values, The coefficients are regularization terms; the Adam optimizer is used to update the parameters; the optimized parameters are used to output the predicted user preferences for items, and a list of recommended items is formed according to the predicted rating values.

[0093] Figure 4-7 The model's performance was measured in four metrics, namely recommendation accuracy ( ), recall rate ), Value and normalized loss cumulative gain ( Its definition is as follows:

[0094]

[0095] in, To recommend to users The collection of items For users The set of items that have actually been interacted with. This is used to test the ratio of the number of items in the recommended item set that the user has actually interacted with to the total number of items in the recommended item set.

[0096]

[0097] The F1 score is used to test the ratio of the number of items in the recommended item set that the user has actually interacted with to the total number of items the user has actually interacted with. The F1 score considers both precision and recall.

[0098]

[0099]

[0100]

[0101] Used to measure the quality of recommendation ranking. Among them, Indicates to users recommend The cumulative gain from the loss of an item, Indicates position The relevance of the recommendation results; if the item is in the user's actual purchase set, then... The value is 1 if it is 1, otherwise it is 0. Is it pressed The ideal calculated on the descending recommended list This indicates the length of the recommended list.

[0102] It should be noted that the above description of the embodiments is not intended to limit the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions, alterations, or substitutions made by those skilled in the art within the scope of the present invention should be covered by the claims of the present invention.

Claims

1. A deep learning recommendation system that integrates the characteristics of the target audience for items, characterized in that, It includes a data acquisition and processing module, a model training module, a model prediction module, and a recommendation module, which are located in a server. The server is connected to multiple databases, and the databases are connected to multiple smart terminals. The smart terminal is used to collect user behavior data and store it in a database; The database is used to store users' historical behavior information and is then processed by the server. The data acquisition and processing module is used to process the data information in the database to obtain standardized data that meets the needs of the system. The model training module is used to model and train standardized data until the model converges. It includes the following sub-modules: an item audience feature aggregation module, a low-order feature extraction module, a feature interaction module, a fully connected neural network module, and a model parameter update module. The feature interaction module is used to perform explicit feature crossing to obtain high-order feature information and to maintain the integrity of information during the forward propagation of the neural network through residual connections. Its execution steps are as follows: S1: Target user's feature vector Feature vectors of items and the audience feature vector of the item Stack them to form a specific matrix. ; S2: Explicit feature crossing is performed by calculating the element-wise product between any two eigenvectors in the matrix, and the output matrix of the previous layer is used as the input of the next layer. S3: Use residual connections to fuse the output matrices of each layer into a feature combination information vector. : , in, The number of feature interaction layers. Indicates the number of users. For the first The output matrix of the layer, Represents the first of the matrix OK; The model parameter update module is used to update the model parameters according to the loss value until convergence. Its execution steps are as follows: S1: Construct the loss function, whose expression is as follows: , in, These are the weight parameters for the model embedding layer, the item audience feature aggregation module, the high-order feature extraction module, and the low-order feature extraction module, respectively; the feature interaction module and the fully connected neural network module are collectively referred to as the high-order feature extraction module; These represent the sets of positive and negative samples used in model training, respectively. For users In items The actual value on For users In items The predicted value; The coefficient of the regularization term; S2: Based on the predicted value and the true value Calculate the loss value; S3: Update model parameters using the Adam optimizer; S4: Iterate through the set of positive and negative samples until the value of the loss function no longer decreases; The model prediction module is used to predict users' personalized preferences using the trained model. The recommendation module is used to sort items according to their predicted ratings and generate a list of recommended items to improve user satisfaction.

2. The deep learning recommendation system that integrates the characteristics of the audience of items according to claim 1, characterized in that, The data acquisition and processing module is used to convert user feedback information about items into an implicit feedback matrix with values ​​of 0 or 1. A value of 1 indicates that the user has interacted with the item, while a value of 0 indicates that the user has not interacted with the item. Indicates the number of users. Represents the number of items; collects the historical user interactions for each item, and converts the item and user information into a low-dimensional, dense feature vector. The processing steps are as follows: S1: Retrieve all data information: user ID, item ID, rating information, and the set of users who have interacted with the item in the past. ; S2: Generate an implicit feedback matrix based on the scoring information. ; S3: Convert user ID to a one-hot vector ; S4: Convert project IDs to one-hot vectors ; S5: Set up user groups Each user in the list is converted into their own one-hot vector; S6: Randomly initialize the mapping matrix of users and items that follows a Gaussian distribution. ; S7: Will the user Transforming a one-hot vector into a low-dimensional dense feature vector The method is as follows: 。 3. The deep learning recommendation system that integrates the characteristics of the audience of items according to claim 1, characterized in that, The model training module consists of an item audience feature aggregation module, a low-order feature extraction module, a feature interaction module, a fully connected neural network module, and a model parameter update module; wherein, the execution steps of the item audience feature aggregation module are as follows: S1: Calculate the correlation coefficient of the target user's historical interaction records with items using an attention mechanism based on a multilayer perceptron. The method is as follows: , in, These are the weight matrix, bias vector, regression coefficients, and activation function, respectively. The number of neurons in the hidden layer, and the activation function. ReLU was chosen to enhance nonlinear expressive power; For target users eigenvectors, For users eigenvectors; This is a function for vector concatenation. This represents the element-wise product of vectors; S2: Normalize the correlation coefficients using the softmax function to obtain the attention distribution. The method is as follows: , in, For items The historical user interaction set; S3: Based on attention distribution The audience feature vector of the item is obtained by weighted summation. The method is as follows: , in, For items Interactive users; The low-order feature extraction module is used to learn low-order feature information in the data through linear regression and vector dot product, and its expression is as follows: , in, This is the vector obtained by concatenating the one-hot vectors of the target user and the item; For the feature vector of the target item, The feature vector of the object's audience; For regression coefficients, This represents the vector dot product operation; The fully connected neural network module is used to combine feature information vectors. Based on this, higher-order features in the data are learned through a multi-layer nonlinear neural network, the expression of which is as follows: , in, This represents the number of layers in a fully connected neural network. This is the output of the feature interaction layer; Representing the first The layer's weight matrix, bias vector, and activation function are used; finally, the predicted value is output by combining low-order and high-order feature information.

4. The deep learning recommendation system that integrates the characteristics of the audience of items according to claim 1, characterized in that, The model prediction module is used to obtain low-order and high-order feature information from the trained and fused model, fuse them, and output a predicted value, as shown in the following expression: , in, This is the output of the low-order feature extraction module. The output of the higher-order feature extraction module, activation function This is the sigmoid function.