User behavior analysis-based data filtering and intelligent scheduling method
By building client-side and server-side models of user purchasing behavior, combined with ridge regression and FedRec algorithm, the problem of collaborative filtering recommendation algorithm leaking user scoring behavior when building recommendation models is solved, achieving user privacy protection, reduction of computing costs and improvement of model accuracy.
Patent Information
- Application Number
- PCT/CN2024/123327
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-02
- Filing Date
- 2024-10-08
- Publication Date
- 2025-05-08
AI Technical Summary
The existing collaborative filtering recommendation algorithms are prone to leak user rating behavior when building recommendation models, resulting in user privacy leakage, high computing costs and model training deviations.
By building a client model of user purchasing behavior and a server-side data decomposition model, the ridge regression algorithm is used for data filtering, and the data is scored and scheduled using the user average method and mixed filling method in the FedRec algorithm, and high-rating data are transferred first and threshold adjustment calculations are performed.
It effectively protects users' scoring behavior, reduces communication and computing costs, and improves the accuracy and efficiency of the recommendation model.
Smart Images

Figure CN2024123327_08052025_PF_FP_ABST
Abstract
Description
A method of data filtering and intelligent scheduling based on user behavior analysis Technical Field
[0001] The present invention belongs to the technical field of user behavior analysis, and specifically relates to a method for data filtering and intelligent scheduling based on user behavior analysis. Background Art
[0002] In recent years, with the rise of big data, research on consumer behavior analysis has flourished. Scholars from a wide range of fields, including databases and data mining, information systems and information management, image processing and computer vision, social network analysis, and e-commerce, have joined the consumer behavior research community. This research field has also attracted significant attention from businesses operating in digital economies, such as e-commerce and social networks. Consumer behavior analysis is considered an effective means for businesses in this digital economy to understand their customers and conduct marketing activities. Within these emerging fields, consumer behavior research is referred to as consumer profiling, and it also holds a prominent position in research areas such as social computing.
[0003] However, during the calculation process, the existing collaborative filtering recommendation algorithm for consumer portraits may over-rate special items or have inaccurate problems in obtaining data, thereby leaking user rating behavior when building the recommendation model.
[0004] Summary of the Invention
[0005] In order to solve the technical problem that the traditional collaborative filtering recommendation algorithm leaks user rating behavior when building a recommendation model, the present invention provides a method for data filtering and intelligent scheduling based on user behavior analysis.
[0006] The present invention provides a method for data filtering and intelligent scheduling based on user behavior analysis, comprising:
[0007] S101: Build a customer service model and server data analysis model for user purchasing behavior;
[0008] S102: Obtain user consumption data, input the consumption data into the customer service model and the server data analysis model, and obtain valid training data and invalid training data;
[0009] S103: Scoring the valid training data and the invalid data using a user scoring algorithm to obtain high-scoring data and low-scoring data;
[0010] S104: Constructing a thread pool scheduling model, inputting the high-scoring data and the low-scoring data into the thread pool scheduling model, so that the high-scoring data is preferentially transmitted, and the low-scoring data is subjected to threshold adjustment calculation.
[0011] Compared with the prior art, the present invention has at least the following beneficial technical effects:
[0012] In this invention, we first construct a client-side model of user purchasing behavior and a server-side data decomposition model, and employ a ridge regression algorithm to filter data on both ends, prioritizing the distribution of valid training data. Then, by re-evaluating valid and invalid training data, we address the privacy, accuracy, and efficiency issues of the FedRec algorithm through the user averaging and hybrid filling methods used in the FedRec algorithm to protect user rating behavior. The user privacy issue stems from the fact that the new scoring algorithm can calculate the gradients of the feature vectors of a user's rated items and the gradients of the feature vectors of their unrated items. Uploading these gradients together to the server prevents the server from identifying the user's rated item set from the item IDs contained in the gradients, thereby protecting the user's rating behavior. The efficiency issue stems from the fact that the gradients of all rated items and a small portion of unrated items for each user u are uploaded to the server, rather than uploading the gradients of both rated and unrated items for each user u as in traditional algorithms. This avoids the high communication and computational costs associated with a large number of items. The accuracy problem means that the gradient of unrated items calculated using the average rating or predicted rating contains less noise than the gradient of items calculated using a value of 0 using traditional algorithms, thus avoiding large deviations in model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The preferred embodiments will be described below in a clear and understandable manner with reference to the accompanying drawings to further illustrate the above-mentioned characteristics, technical features, advantages and implementation methods of the present invention.
[0014] FIG1 is a flow chart of a method for data filtering and intelligent scheduling based on user behavior analysis provided by the present invention;
[0015] FIG2 is a schematic structural diagram of a data filtering and intelligent scheduling system based on user behavior analysis provided by the present invention. DETAILED DESCRIPTION
[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the specific embodiments of the present invention will be described below with reference to the accompanying drawings. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings and other embodiments can be obtained based on these drawings without inventive work.
[0017] To simplify the drawings, only portions relevant to the invention are schematically depicted in each figure; they do not represent the actual structure of the product. Furthermore, to simplify the drawings and facilitate understanding, in some figures, only one component with the same structure or function is schematically depicted or labeled. In this document, "one" not only means "only one" but also "more than one."
[0018] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0019] It should be noted that, unless otherwise specified or limited, the terms "mounted," "connected," and "connected" should be understood broadly. For example, they can refer to fixed connections, removable connections, or integral connections. They can also refer to mechanical connections or electrical connections. They can also refer to direct connections or indirect connections through an intermediary, or to internal communication between two components. Those skilled in the art will understand the specific meanings of these terms in the present invention based on the specific circumstances.
[0020] In addition, in the description of the present invention, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0021] Example 1
[0022] In one embodiment, referring to FIG1 of the specification, a flow chart of a method for data filtering and intelligent scheduling based on user behavior analysis provided by the present invention is shown. Referring to FIG2 of the specification, a structural chart of a method for data filtering and intelligent scheduling based on user behavior analysis provided by the present invention is shown.
[0023] The present invention provides a method for data filtering and intelligent scheduling based on user behavior analysis, comprising:
[0024] S101: Build a customer service model and server data analysis model for user purchasing behavior.
[0025] Alternatively, a preference context-based matrix factorization algorithm (SVD++) is used. This algorithm introduces preference context data based on the conventional matrix factorization algorithm. This assumes that user u's predicted rating for item i is not only related to item i, but also to other items rated by user u. Therefore, the customer service model uses the formula:
[0026] Among them, u is the user, i is the item, τ u The set of items rated by user u, Wi′. is the potential feature vector of item i′.
[0027] The server-side model is the implicit feedback federated collaborative filtering algorithm FCF, that is, the gradient calculation formula of the potential feature vector of item i at client u is:
[0028] Among them, y ui Indicates whether user u has rated item i in the training set, and y ui =1 means there is a rating record of user u for item i, y ui =0 means there is no rating record of user u for item i, 1+λy ui Represents the confidence weight, λ>0.
[0029] On client u, the gradient calculation formula of the potential feature vector of item i is:
[0030] S102: Obtain user consumption data, input the consumption data into the customer service model and the server-side data analysis model, and obtain valid training data and invalid training data.
[0031] Optionally, the S102 specifically includes:
[0032] S1021: Obtain user consumption data, input the consumption data into the customer service model and the server-side data analysis model, and obtain first data and second data.
[0033] S1022: Filter the first data and the second data through a ridge regression algorithm to obtain the valid training data and the invalid training data.
[0034] Optionally, the S1022 specifically includes:
[0035] S10221: Input the first data and the second data into the ridge regression algorithm.
[0036] S10222: Using the ridge regression algorithm, perform an overfitting prevention operation on the first data and the second data to obtain third data and fourth data.
[0037] Optionally, a ridge regression method is used to prevent overfitting of the client and server models. Traditional least squares methods lack stability and reliability. To solve this problem, it is necessary to transform the ill-posed problem into a well-posed problem. Therefore, a regularization term is added to the above loss function. The ridge regression algorithm uses the formula: ||Xθ-y|| 2 +||Γθ|| 2; Γ=aI; θ(a)=(X T X+aI) -1 X T y;
[0038] Among them, X is the input, y is the output, r is the objective training result, aI is the fitting value, θ(a)=(X T X+aI) -1 X T y is an operation to prevent overfitting, I is the identity matrix, θ is the fitting hyperparameter, T is the weight constant, a is the weight of the identity matrix, and θ(a) represents the calculation of θ when a is determined.
[0039] S10223: Compare the third data and the fourth data by performing a difference comparison. If the difference is smaller than a preset difference, the difference is used as the valid training data; if the difference is larger than the preset difference, the difference is used as the invalid training data.
[0040] Optionally, the preset difference is 10%.
[0041] That is, the client-side fitted values are compared with the server-side fitted values. If the difference is less than or equal to 10%, it is considered normal and valid training data, and priority scheduling and data transmission are performed in the thread pool. If the difference is greater than 10%, it is considered invalid training data, and threshold adjustment calculations are performed in the thread pool.
[0042] S103: Scoring the valid training data and the invalid data using a user scoring algorithm to obtain high-scoring data and low-scoring data.
[0043] That is to say, the invalid training data with a large difference between the valid training data and the fitting value is orderly solved through the user averaging method and mixed filling method used in the FedRec algorithm to protect the user rating behavior to solve the privacy, accuracy and efficiency problems of the FedRec algorithm.
[0044] Optionally, the S103 specifically includes:
[0045] S1031: Scoring the valid training data and the invalid data using a user scoring algorithm.
[0046] Optionally, the user scoring algorithm, that is, the FedRec algorithm, adopts the formula: R′ v ∪R u ;
[0047] in, is the average rating of all rated items by user u, For any unrated item i′∈τ′ by user u u The prediction score, R′ v ∪R u is a new set of rating records,
[0048] S1032: Solve user privacy issues, accuracy issues and efficiency issues through the scoring.
[0049] Optionally, a new rating record set R′ for each user u is generated v ∪R u , can solve three problems:
[0050] First of all, there is the issue of user privacy. v ∪R u The gradients of the feature vectors of the user's rated items and the gradients of the feature vectors of the user's unrated items can be calculated. Uploading these gradients together to the server prevents the server from identifying the set of items rated by the user from the item IDs contained in the gradients, thereby protecting the user's rating behavior.
[0051] Secondly, there is the issue of efficiency. By uploading the gradients of all rated items and a small number of unrated items for each user u to the server, instead of uploading the gradients of all rated items and all unrated items for each user u as in traditional algorithms, we can avoid the high communication and computing costs caused by too many items.
[0052] Finally, there is the issue of accuracy. The gradient of unrated items calculated using the average rating or predicted rating contains less noise than the gradient of items calculated using a value of 0 using traditional algorithms, thus avoiding large deviations in model training.
[0053] S104: Constructing a thread pool scheduling model, inputting the high-scoring data and the low-scoring data into the thread pool scheduling model, so that the high-scoring data is preferentially transmitted, and the low-scoring data is subjected to threshold adjustment calculation.
[0054] Optionally, the thread pool scheduling model adopts the formula:
[0055] Among them, ω is the thread pool load index, N is the number of working threads in the thread pool when it is running, and N max The maximum number of threads to set. To describe the saturation of the working thread; T cur is the number of tasks in the current acquisition time window, T pre is the number of tasks in the previous acquisition time window, Q is the size of the task buffer queue, To describe the current task saturation, is the growth rate of the task buffer queue; ξ is the weight coefficient.
[0056] Optionally, the load degree is converted from data such as the number of working threads, the maximum number of threads, and the size of the task buffer queue when the thread pool is running, and a percentage value is calculated through different weighted proportions.
[0057] Optionally, the S104 specifically includes:
[0058] S1041: Build a thread pool scheduling model, input the high-scoring data and the low-scoring data into the thread pool scheduling model, and obtain a thread pool load;
[0059] S1042: If the thread pool load exceeds the preset thread pool load, threshold adjustment calculation is performed; if the thread pool load is less than the preset thread pool load, the current collection time window is skipped and priority data transmission is performed.
[0060] Compared with the prior art, the present invention has at least the following beneficial technical effects:
[0061] In this invention, we first construct a client-side model of user purchasing behavior and a server-side data decomposition model, and employ a ridge regression algorithm to filter data on both ends, prioritizing the distribution of valid training data. Then, by re-evaluating valid and invalid training data, we address the privacy, accuracy, and efficiency issues of the FedRec algorithm through the user averaging and hybrid filling methods used in the FedRec algorithm to protect user rating behavior. The user privacy issue stems from the fact that the new scoring algorithm can calculate the gradients of the feature vectors of a user's rated items and the gradients of the feature vectors of their unrated items. Uploading these gradients together to the server prevents the server from identifying the user's rated item set from the item IDs contained in the gradients, thereby protecting the user's rating behavior. The efficiency issue stems from the fact that the gradients of all rated items and a small portion of unrated items for each user u are uploaded to the server, rather than uploading the gradients of both rated and unrated items for each user u as in traditional algorithms. This avoids the high communication and computational costs associated with a large number of items. The accuracy problem means that the gradient of unrated items calculated using the average rating or predicted rating contains less noise than the gradient of items calculated using a value of 0 using traditional algorithms, thus avoiding large deviations in model training.
[0062] Example 2
[0063] In one embodiment, the present invention provides a system for data filtering and intelligent scheduling based on user behavior analysis, which is used to execute the method for data filtering and intelligent scheduling based on user behavior analysis in Example 1.
[0064] A data filtering and intelligent scheduling system 30 based on user behavior analysis, comprising:
[0065] The first construction module 301 is used to construct a customer service model and a server data analysis model of user purchase behavior;
[0066] An acquisition module 302 is used to acquire user consumption data, input the consumption data into the customer service model and the server data analysis model, and obtain valid training data and invalid training data;
[0067] Scoring module 303, configured to score the valid training data and the invalid data using a user scoring algorithm to obtain high-scoring data and low-scoring data;
[0068] The second building module is used to build a thread pool scheduling model, and input the high-scoring data and the low-scoring data into the thread pool scheduling model, so that the high-scoring data is preferentially transmitted and the low-scoring data is threshold-adjusted and calculated.
[0069] The system for data filtering and intelligent scheduling based on user behavior analysis provided by the present invention can implement the method for data filtering and intelligent scheduling based on user behavior analysis in the above-mentioned embodiment 1. To avoid repetition, the present invention will not go into details.
[0070] In this invention, we first construct a client-side model of user purchasing behavior and a server-side data decomposition model, and employ a ridge regression algorithm to filter data on both ends, prioritizing the distribution of valid training data. Then, by re-evaluating valid and invalid training data, we address the privacy, accuracy, and efficiency issues of the FedRec algorithm through the user averaging and hybrid filling methods used in the FedRec algorithm to protect user rating behavior. The user privacy issue stems from the fact that the new scoring algorithm can calculate the gradients of the feature vectors of a user's rated items and the gradients of the feature vectors of their unrated items. Uploading these gradients together to the server prevents the server from identifying the user's rated item set from the item IDs contained in the gradients, thereby protecting the user's rating behavior. The efficiency issue stems from the fact that the gradients of all rated items and a small portion of unrated items for each user u are uploaded to the server, rather than uploading the gradients of both rated and unrated items for each user u as in traditional algorithms. This avoids the high communication and computational costs associated with a large number of items. The accuracy problem means that the gradient of unrated items calculated using the average rating or predicted rating contains less noise than the gradient of items calculated using a value of 0 using traditional algorithms, thus avoiding large deviations in model training.
[0071] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0072] The above embodiments merely illustrate several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.
Claims
1. A method for data filtering and intelligent scheduling based on user behavior analysis, characterized in that: include: S101: Build a customer service model and server data analysis model for user purchasing behavior; S102: Obtain user consumption data, input the consumption data into the customer service model and the server data analysis model, and obtain valid training data and invalid training data; S103: Scoring the valid training data and the invalid data through a user scoring algorithm to obtain high-scoring data and low-scoring data; S104: constructing a thread pool scheduling model, and inputting the high-scoring data and the low-scoring data into the thread pool scheduling model, so that the high-scoring data is preferentially transmitted, and the low-scoring data is threshold-adjusted and calculated.
2. The method for data filtering and intelligent scheduling based on user behavior analysis according to claim 1 is characterized in that: The customer service model uses the formula: Among them, u is the user, i is the item, τ u The set of items rated by user u, W i′ . is the potential feature vector of item i′.; The server-side model uses the formula: Among them, y ui Indicates whether user u has rated item i in the training set, and y ui =1 means there is a rating record of user u for item i, y ui =0 means there is no rating record of user u for item i, 1+λy ui represents the confidence weight, λ>0.
3. The method for data filtering and intelligent scheduling based on user behavior analysis according to claim 2 is characterized in that: The S102 specifically includes: S1021: Obtain user consumption data and input the consumption data into the customer service model and The server-side data analysis model obtains the first data and the second data; S1022: Filter the first data and the second data through a ridge regression algorithm to obtain the valid training data and the invalid training data.
4. The method for data filtering and intelligent scheduling based on user behavior analysis according to claim 3 is characterized in that: The S1022 specifically includes: S10221: inputting the first data and the second data into the ridge regression algorithm; S10222: performing an overfitting prevention operation on the first data and the second data by using the ridge regression algorithm, and obtaining third data and fourth data; S10223: Compare the third data and the fourth data for difference. If the difference is smaller than a preset difference, use the difference as the valid training data; if the difference is larger than the preset difference, use the difference as the invalid training data.
5. The method for data filtering and intelligent scheduling based on user behavior analysis according to claim 4 is characterized in that: The ridge regression algorithm uses the formula: ||Xθ-y|| 2 +||Γθ|| 2 ; Γ=aI; θ(a)=(X T X+aI) -1 X T y; Among them, X is the input, y is the output, Γ is the objective training result, aI is the fitting value, θ(a)=(X T X+aI) -1 X T y is an operation to prevent overfitting, I is the identity matrix, θ is the fitting hyperparameter, T is the weight constant, a is the weight of the identity matrix, and θ(a) represents the calculation of θ when a is determined.
6. The method for data filtering and intelligent scheduling based on user behavior analysis according to claim 4 is characterized in that: The preset difference is 10%.
7. The method for data filtering and intelligent scheduling based on user behavior analysis according to claim 1, characterized in that: The S103 specifically includes: S1031: Scoring the valid training data and the invalid data using a user scoring algorithm; S1032: Solve user privacy issues, accuracy issues and efficiency issues through the scoring.
8. The method for data filtering and intelligent scheduling based on user behavior analysis according to claim 7, characterized in that: The user rating algorithm uses the formula: R′ v ∪R u ; in, is the average rating of all rated items by user u, For any unrated item i′∈τ′ by user u u The prediction score, R′ v ∪R u , is the new scoring record set, 9. The method for data filtering and intelligent scheduling based on user behavior analysis according to claim 1, characterized in that: The thread pool scheduling model adopts the formula: Among them, ω is the thread pool load index, N is the number of working threads when the thread pool is running, and N max To set the maximum number of threads, To describe the saturation of the working thread; T cur is the number of tasks in the current acquisition time window, T pre is the number of tasks in the previous acquisition time window, Q is the size of the task buffer queue, To describe the current task saturation, is to describe the growth rate of the task buffer queue; ξ is the weight coefficient.
10. The method for data filtering and intelligent scheduling based on user behavior analysis according to claim 9, characterized in that: The S104 specifically includes: S1041: construct the thread pool scheduling model, input the high-scoring data and the low-scoring data into the thread pool scheduling model, and obtain the thread pool load degree; S1042: If the thread pool load exceeds the preset thread pool load, threshold adjustment calculation is performed; if the thread pool load is less than the preset thread pool load, priority data transmission is performed.
Citation Information
Patent Citations
Social collaborative filtering recommendation method based on federal learning
CN114510652A
Data filtering and intelligent scheduling method based on user behavior analysis
CN117474573A
Recommender Systems and Methods
US20120030159A1