Personalized recommendation method and device based on real-time user behaviors and storage medium
By adopting streaming processing and distributed computing architecture in a large-scale real-time personalized recommendation method, combined with Flink-SQL and fault-tolerant recovery mechanism, efficient, accurate and real-time personalized recommendations are achieved, solving the complex problems of processing massive user behavior data, and improving the performance and reliability of recommendation services.
Patent Information
- Application Number
- CN202510437104.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-09
AI Technical Summary
In the large-scale real-time personalized recommendation method, how to achieve efficient, accurate and real-time recommendations under massive user behavior data and complex business scenarios, and make full use of user behavior characteristic data while ensuring data security and user privacy.
Using a personalized recommendation method based on real-time user behavior, real-time extraction and optimization of user behavior characteristics is achieved through streaming processing and distributed computing architecture. Use Flink-SQL for real-time statistics and aggregation, combining fault-tolerant recovery mechanism and dynamic resource scheduling to ensure the high availability and scalability of recommended services. Continuously collect user feedback and optimize recommendation results in real time, cache them to a distributed cache cluster and adopt load balancing strategies.
It realizes user behavior analysis and recommendation within millisecond delay, improves the processing efficiency, reliability and access speed of recommendations, and ensures high-quality personalized recommendation services for a large number of users.
Smart Images

Figure CN119961526A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a personalized recommendation method, device and storage medium based on real-time user behavior. Background Art
[0002] In large-scale real-time personalized recommendation methods, the core technical challenge is how to achieve efficient, accurate, and real-time recommendations under massive user behavior data and complex business scenarios. With the rapid growth of user scale and data volume, traditional recommendation methods face multiple technical bottlenecks. First, there is a contradiction between real-time and accuracy. The method needs to complete complex calculations such as user behavior analysis, feature extraction, and interest modeling within milliseconds, while also ensuring the accuracy of the recommendation results. Second, there is a conflict between the scalability of the method and the efficiency of resource utilization. In the face of sudden traffic and data surges, it is necessary to be able to flexibly expand to maintain service quality while avoiding resource waste. Furthermore, there is a tension between the continuous optimization and stability of the recommendation algorithm. Frequent model updates may lead to unstable operation, while fixed algorithms are difficult to adapt to changes in user interests. In addition, how to make full use of user behavior feature data for personalized recommendations while ensuring data security and user privacy is also a thorny issue. These technical difficulties are interrelated and mutually influential, and constitute the core challenges faced by large-scale real-time personalized recommendation methods. How to find a balance among these contradictions and create a high-performance, highly available, and sustainably optimized recommendation method is a key issue that needs to be solved urgently. Summary of the invention
[0003] In order to achieve the purpose of the present invention, in a first aspect, the present invention provides a personalized recommendation method based on real-time user behavior, which mainly includes: Collect real-time user behavior data, perform streaming processing on the real-time user behavior data, smooth traffic fluctuations through message queue buffering, and obtain user behavior feature data; input the user behavior feature data into the distributed computing architecture, process the user behavior feature data through shard storage and parallel computing tasks, and obtain optimized user behavior feature data; use Flink-SQL to perform real-time statistics and aggregation on the optimized user behavior feature data, and implement task scheduling of multiple computing nodes through distributed coordination services, so as to determine popular items and user interest preferences, and generate recommendation results through a personalized recommendation engine; in view of the risks of failure and data loss during the operation of the personalized recommendation method, adopt a fault-tolerant recovery mechanism and data backup and disaster recovery, implement load balancing through dynamic resource scheduling, and judge individual The system monitors the running status of the personalized recommendation method and generates an availability evaluation report; obtains the user scale growth data and analyzes the user scale growth rate. If the user scale growth rate is greater than the preset growth rate threshold, it is determined that the user scale has increased significantly, and dynamically increases the distributed storage and computing nodes. The principle of data proximity computing is used to optimize the data transmission path, reduce network transmission overhead, and adjust the resource allocation results to cope with the user scale growth of the personalized recommendation method; continuously collects user feedback data on the recommendation results, optimizes the recommendation results according to the feedback data, and obtains the optimized personalized recommendation results; caches the optimized personalized recommendation results in the distributed cache cluster, adopts a multi-copy storage mechanism, and routes user requests to the nearest cache node through load balancing to improve the access speed of the optimized personalized recommendation results.
[0004] In a second aspect, the present invention further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.
[0005] In a third aspect, a computer-readable storage medium stores a computer program, wherein the computer program implements the steps of any of the above methods when executed by a processor.
[0006] The technical solution provided by the embodiment of the present invention may have the following beneficial effects: The present invention discloses a personalized recommendation method, device and storage medium based on real-time user behavior. Aiming at the business scenarios of large-scale user real-time behavior data processing and personalized recommendation, the present invention adopts a personalized recommendation engine and a streaming processing architecture to realize the extraction of user behavior features with millisecond delay. Through distributed storage and parallel computing, combined with Flink-SQL real-time statistical aggregation, popular items and user interest preferences are efficiently determined. The present invention introduces a fault-tolerant recovery mechanism and dynamic resource scheduling to ensure the high availability and scalability of the recommendation service. At the same time, user feedback is continuously collected and processed in real time to form a closed-loop recommendation algorithm optimization process. Finally, the present invention caches the personalized recommendation results to a distributed cluster and adopts a load balancing strategy, which significantly improves the processing efficiency, reliability and access speed of the recommendation method, and provides high-quality personalized recommendation services for massive users. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 The figure is a flow chart of the personalized recommendation method based on real-time user behavior of the present invention.
[0008] Figure 2 The figure is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0009] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0010] like Figure 1 The personalized recommendation method based on real-time user behavior in this embodiment may specifically include: S101. Collect real-time user behavior data, perform streaming processing on the real-time user behavior data, smooth traffic fluctuations through message queue buffering, and obtain user behavior feature data.
[0011] Real-time user behavior data is obtained from the data collection layer to obtain the original behavior data stream. Flink is used to establish a real-time data stream processing pipeline. The original behavior data stream is buffered by Kafka to smooth the instantaneous traffic fluctuations, and the buffered data stream is input into the real-time data stream processing pipeline. In Flink, a window function is used to perform real-time aggregation processing on the input buffered data stream to obtain aggregated user behavior data. The aggregated user behavior data is obtained, and the K nearest neighbor classification algorithm is used to classify it to obtain classification result data that characterizes the user behavior pattern. The classification result data is obtained, and the classification result data is subjected to feature screening by the random forest algorithm through the feature importance screening method to obtain user behavior feature data that reflects the user behavior characteristics, wherein the feature importance screening method refers to calculating the feature importance score through the random forest algorithm to screen out features that have a significant impact on the target variable.
[0012] Specifically, obtaining real-time user behavior data from the data collection layer is the starting point of the entire process, and the core is to capture the user's dynamic operations. For example, in an e-commerce platform, the user's behavior data such as browsing products, adding to shopping carts, placing orders and paying can be collected in real time through the tracking technology to form an original behavior data stream. These behavior data usually contain fields such as timestamps, user IDs, and behavior types, reflecting the user's real-time activities on the platform. The original behavior data stream is buffered by Kafka to smooth traffic fluctuations. When using Flink to establish a real-time data stream processing pipeline, the focus is on its low latency and high throughput characteristics. Exemplarily, the data stream after the original behavior data stream is buffered by Kafka is input into the real-time data stream processing pipeline established by Flink, and the accuracy of subsequent processing is ensured by defining data processing logic, such as filtering out invalid clicks or repeated operations. Specifically, Flink can process tens of thousands of user behavior records per second, which is suitable for scenarios with surges in traffic during peak periods. In one possible implementation, assuming that the e-commerce platform generates 100,000 behavior data per second during a promotion, Kafka can store it in shards and set the buffer size to 1 minute of data, that is, 6 million records. In this way, even if the traffic surges instantly, it will not directly impact Flink, ensuring the stability of the system. When using window functions for real-time aggregation processing in Flink, it can be understood as dividing the buffered data stream by time windows. For example, a window is created every 5 seconds to count the number of views of a certain product or the purchase frequency of a user, and obtain aggregated user behavior data. This method can effectively reduce data redundancy and improve the efficiency of subsequent analysis. After obtaining the aggregated data, the K nearest neighbor classification algorithm is used for classification processing to identify user behavior patterns. Preferably, users can be divided into categories such as "high-frequency buyers" and "main browsers" based on their browsing and purchase records. In one embodiment, assuming that a user browses 10 times and places orders 2 times within 5 minutes, he is classified as a "potential high-value user" by calculating the distance from historical data, providing a basis for personalized recommendations. Subsequently, the random forest algorithm is used to perform feature screening on the classification result data through the feature importance screening method to further explore deep features. For example, from the classification result data, random forests screen out high-weight features such as "historical low-priced product click-through rate", "brand repurchase times", and "price sensitivity (discount participation times)" by calculating feature importance; based on these high-weight features, user behavior features are constructed in combination with business logic (such as "brand repurchase times / total purchase times" to quantify "brand preference", and "discount participation times / total purchase times" to reflect "low price preference") to obtain user behavior feature data that reflects user behavior characteristics. The advantage of this method is that it can integrate multi-dimensional information and improve the ability to express features. Furthermore, this application stores user behavior feature data in Redis and establishes a real-time query interface to achieve efficient access.Exemplarily, Redis can store user behavior feature data for each user, and the query interface supports processing 10,000 requests per second, meeting the needs of real-time recommendation systems. It should be noted that the memory storage feature of Redis greatly reduces latency. This application also uses Kubernetes to dynamically adjust Flink computing resources. In one embodiment, the access frequency threshold is set to 5,000 times per second, and the number of Flink task slots is automatically increased when it exceeds the threshold, for example, from 10 to 15; when it is less than 2,000 times, it is reduced to 8. This dynamic adjustment can optimize resource utilization and reduce costs. Obtaining user behavior feature data through the above process can not only support real-time business decisions, but also improve user experience. For example, e-commerce platforms can optimize recommendation algorithms based on this, increase user retention and conversion rates, and reflect the efficiency and practicality of the technology.
[0013] S102: Input the user behavior feature data into a distributed computing architecture, process the user behavior feature data through shard storage and parallel computing tasks, and obtain optimized user behavior feature data.
[0014] Input the user behavior feature data into the streaming computing architecture, and use the distributed storage architecture to shard the user behavior feature data according to the preset sharding strategy, and evenly distribute the user behavior feature data to at least one storage node to obtain the sharded data. Start at least one parallel computing task for the sharded data, and use the MapReduce computing framework to process the sharded data to obtain the MapReduce processed data. Obtain the MapReduce processed data, use the KMeans algorithm to perform cluster analysis on the MapReduce processed data, and obtain the user behavior grouping result data. Use the principal component analysis algorithm to reduce the dimension of the user behavior grouping result data to obtain the optimized user behavior feature data.
[0015] Specifically, when inputting user behavior feature data into a distributed computing architecture, it is first necessary to clarify the data sharding strategy to ensure that the data can be evenly distributed to multiple storage nodes. The sharding strategy may be based on user ID, geographic location, or timestamp. For example, an e-commerce platform can shard users by province to ensure that user data in the same region is stored in similar nodes to improve query efficiency. Assuming there are 10 million users distributed in 30 provinces, each shard has an average of about 330,000 user data. The MapReduce framework processes massive data in a distributed environment. The Map phase may convert user behavior feature data into key-value pairs, such as {user ID: [browse product ID list]}. The Reduce phase may aggregate user behavior feature data for the same user. This parallel processing method significantly improves the processing speed of large-scale data. The KMeans clustering algorithm is used to discover user groups. In e-commerce scenarios, users may be divided into groups such as "price-sensitive", "brand loyal", and "new product seekers". The algorithm iteratively calculates the centers of each cluster until convergence. Assuming that 5 clusters are set, a stable result is obtained after 50 iterations, and each cluster contains about 2 million users. The principal component analysis algorithm (PCA) reduces the data dimension and retains key features. For example, for user A, 10 main indicator data are extracted from the original 100 user behaviors, such as user A's browsing, clicking and purchasing behavior data, to obtain optimized user behavior feature data. This not only reduces the storage space, but also improves the efficiency of subsequent analysis. Furthermore, the optimized user behavior feature data in this application is stored in the HBase database, and a query interface based on the main feature data of user behavior is established. As a big data storage system, the HBase database is suitable for storing and quickly retrieving massive user feature data. The data table design may use user ID as a row and various features as columns. This structure supports efficient random read and write and range scanning, and provides data support for applications such as personalized recommendations. This application will also determine whether the access volume of the query interface exceeds the preset access volume threshold. If it exceeds, the number of computing nodes of the Hadoop cluster will be dynamically expanded. If it is lower, the number of computing nodes will be reduced. Dynamically adjusting the number of computing nodes in the Hadoop cluster is the key to ensuring system performance. In one embodiment, assuming that the original cluster has 40 nodes, when the access volume exceeds the limit, the system automatically expands 20 additional nodes to join the cluster through a script. After the node expansion, the total processing capacity is increased from 50,000 behavioral data per second to 80,000, ensuring the smooth operation of the system. On the contrary, if the access volume drops to 3,000 times per second, which is lower than the access volume threshold, the system can be reduced to 30 nodes to release excess resources. This dynamic adjustment can also optimize costs, such as reducing nodes during low-peak hours at night to reduce costs. The advantage of this architecture lies in its scalability and flexibility. Through distributed storage and computing, the method can easily cope with the growth of data volume.Dynamic resource adjustment ensures performance stability under different loads. For example, an e-commerce platform may process 10,000 user behavior feature data per second on a daily basis, and this number may surge to 100,000 during a big promotion. Through the above architecture, the method can smoothly expand processing capabilities to ensure that user experience is not affected. In general, this big data processing architecture optimizes the processing of user behavior feature data through distributed storage, parallel computing, intelligent algorithms, and dynamic resource scheduling, achieving efficient processing and utilization of massive user behavior feature data, and providing enterprises with strong data support and decision-making basis.
[0016] S103. Use Flink-SQL to perform real-time statistics and aggregation on the optimized user behavior feature data, and implement task scheduling of multiple computing nodes through distributed coordination services, so as to determine popular items and user interest preferences, and generate recommendation results through a personalized recommendation engine.
[0017] The optimized user behavior feature data is obtained, and Flink-SQL is used to perform real-time statistical processing on the optimized user behavior feature data to obtain statistical values; real-time aggregation calculations are performed on the statistical values to obtain scheduling values; distributed computing processing is performed on multiple computing nodes according to the scheduling values to obtain calculation values; for the calculation values, popular items are analyzed by a preset popularity analysis algorithm to obtain item values; a preset user interest analysis model is used to perform user interest analysis according to the item values to obtain interest values; the interest values are input into a preset user preference analysis model for processing to obtain user preference values; if the preference value is greater than a preset preference threshold, user preference portrait data is generated according to the preference value; the user preference portrait data is used to generate targeted recommendation results in real time through a personalized recommendation engine.
[0018] Specifically, real-time statistical processing is the key to efficiently analyzing user behavior. Flink-SQL (a stream-batch integrated SQL query engine based on Apache Flink) can process massive data streams within milliseconds due to its low latency and high throughput. In one possible implementation, when the optimized user behavior feature data is processed in real time by Flink-SQL, the data stream can be segmented based on the time window. The following is an example of the whole process based on the above steps. For example, the e-commerce platform extracts the browsing, clicking, and purchasing behavior data of user A in the past 10 days from the optimized user behavior feature data and summarizes them to count the number of views and purchase conversion rate of user A for the product "sports shoes". Assuming that "sports shoes" is viewed 50 times and purchased 2 times in 10 days, the statistical value is calculated according to the following formula: "sports shoes" statistical value = number of purchases / number of views × 100% = 2 / 50 × 100% = 4%. Perform real-time aggregation calculations on the statistical values, group and aggregate them according to purchase time periods, etc., to obtain scheduling values. Exemplarily, the scheduling value is calculated as follows: scheduling value = click-through rate × statistical value × time period weight (preset weight 1.2 for 8-10 pm), calculate the scheduling value of user A for "sports shoes", extract the click-through rate of user A for "sports shoes" as 40% from the optimized user behavior feature data, and the statistical value is 4%. User A's scheduling value for "sports shoes" = 40% × 4% × 1.2 = 1.92%. When performing distributed computing processing on multiple computing nodes according to the scheduling value, task slicing can be used. Assuming there are 10 computing nodes, commodity categories with higher scheduling values will be allocated more node resources, and the corresponding calculation values will be higher. For example, the scheduling values of all commodities are obtained, among which "sports shoes" have the highest scheduling value (1.92%), and are allocated to 4 nodes for parallel processing, while "books" are only allocated 1 node due to the low scheduling value, and the calculation value of "sports shoes" is 4, and the calculation value of "books" is 1. According to the preset popularity analysis algorithm, such as the weighted scoring algorithm based on time decay, the calculated value is combined with real-time data such as click volume, collection volume, purchase volume, sharing volume, time decay factor, etc., and the calculated value is analyzed for popular items to obtain the item value. For example, the item value of sports shoes is calculated, and the item value of "sports shoes" = calculated value × {α × (click volume) + β × (collection volume) + γ × (purchase volume) + δ × (sharing volume) - θ}, where the weight coefficients (α = 0.4, β = 0.3, γ = 0.2, δ = 0.1) can be obtained in advance through AB testing, and θ (time decay factor) = log 10 (Current time - item listing time + 1). Input sample data: "Sports shoes" calculated value = 4, clicks = 1000, favorites = 50, purchases = 20, shares = 10, where the calculated value is calculated based on the optimized user A behavior feature data. Listing time: 7 days ago, corresponding to θ = log 10(7×24×3600+1)≈5.8, and the item value of "sports shoes" is 4×[0.4×1000+0.3×50+0.2×20+0.1×10-5.8]=4×[400+15+4+1-5.8]=4×414.2=1656.8. When analyzing the interest value using the preset user interest analysis model, the interest point can be inferred by combining the item value and the user's historical behavior. For example, interest value = total behavior score × (item value / maximum item value of similar items), where the total behavior score = Σ (number of behaviors × behavior weight × time decay factor); behavior weight: browse = 1, favorite = 3, purchase = 5. Input sample data: The item value of "sports shoes" is 1656.8. User A's interest in "sports shoes": browse 5 times (score after each decay = 1×0.9=0.9) → 5×0.9=4.5; collect 2 times (score after each decay = 3×0.8=2.4) → 2×2.4=4.8; purchase 1 time (score after decay = 5×0.7=3.5) → 1×3.5=3.5; total behavior score = 4.5+4.8+3.5=12.8; maximum value of similar items: item value = 2000 ("sports shoes" category), where the item value is calculated based on the optimized user A behavior feature data, and user A's related data (browsing, collection, purchase) is extracted from the optimized user behavior feature data. The interest value of user A in "sports shoes" is obtained = 12.8×(1656.8 / 2000)≈12.8×0.828≈10.60. When user A's interest value in "sports shoes" is input into the user preference analysis model, the preference value can be determined through a weighted scoring mechanism, where preference value = Σ (behavior type weight × interest value) × category coefficient. For example, the behavior type weights may be as follows: browse 0.3 (must last > 30 seconds), collect 0.5 (add to favorites / wishlist), add to cart 0.6 (add to cart but not purchased), purchase 0.7 (actual payment successful), review 0.4 (extra +0.1 for reviews with pictures), category coefficient (sports shoes): 1.2. Input sample data: user A's interest value in "sports shoes" = 10.60. Behavior type: browse (0.3) collect (0.5) purchase (0.7), where the interest value is calculated based on the optimized user A behavior feature data. The obtained preference value of user A for "sports shoes" = (0.3+0.5+0.7)×10.60×1.2=1.5×10.60×1.2=19.08. If the preference value of user A for "sports shoes" is greater than the preset preference threshold of 15, it is determined that user A has a strong preference for "sports shoes" and generates preference portrait data, such as "sports enthusiast". This portrait clearly depicts the user's tendency. When using a personalized recommendation engine to generate recommendation results based on the preference portrait, the possible calculation basis is obtained by using the user preference portrait data through calculation based on the collaborative filtering method and machine learning model, including the similarity between users and items, predicted click-through rate, and real-time weight data of items.Specific actual calculation process example: By analyzing the user's historical behavior data (such as browsing, purchasing, rating, etc.) and the attribute characteristics of the item, the collaborative filtering method is used to calculate the similarity between the user and the item; based on the user's behavior data and item characteristics, the machine learning model (such as logistic regression, decision tree, neural network, etc.) is used to predict the probability of whether the user will click on a certain item, thereby estimating the predicted click-through rate; the real-time weight data of the item is pre-set according to the user's usage time, usage platform and other data (such as 7 pm, using a mobile phone). The similarity between the user and the item, the predicted click-through rate, and the real-time weight data of the item are obtained and the score is calculated. The basic score = the similarity between the user and the item (0.8) × predicted click-through rate (0.9) × the real-time weight of the item (1.2), the adjustment score = heat (0.3) + diversity (same category penalty 0.7) + clearance bonus (0.2), the final score = basic score + adjustment score, and the final output result: the top 10 ranked items (such as running socks, sports water bottles, cushioning running shoes, etc.). The qualified items (recommendation results) are screened in real time and pushed to the user. This approach enhances the pertinence of recommendations. Understandably, this architecture transforms user behavior into precise recommendations through multi-level analysis and real-time processing. The distributed design and algorithm optimization of each link ensures efficient connection from data input to result output, providing strong support for e-commerce platforms.
[0019] S104. To address the risks of failure and data loss during the operation of the personalized recommendation method, a fault-tolerant recovery mechanism and data backup and disaster recovery are adopted, load balancing is achieved through dynamic resource scheduling, the operation status of the personalized recommendation method is determined, and an availability evaluation report is generated.
[0020] The operation log data of the personalized recommendation method is obtained, and the operation log data refers to a collection of events and data automatically recorded when the personalized recommendation method is running. The operation log data is analyzed by using a pre-established fault-tolerant recovery mechanism, the operation status of the personalized recommendation method is detected, and an operation status detection result is obtained. According to a preset data backup strategy, the optimized user behavior feature data in the personalized recommendation method is backed up regularly in full. In combination with a preset disaster recovery strategy, the backup data of the full backup is deployed in a remote data center. According to the operation status detection result, if at least one fault point is detected, the preset data backup disaster recovery mechanism is triggered, the user request is switched to the backup data, and the lost data caused by the at least one fault point is restored to obtain the restored data. According to the restored data, the load condition of the personalized recommendation method in operation is analyzed by using a pre-established dynamic resource scheduling algorithm to obtain a current load value. According to the current load value, the computing resources of the personalized recommendation method are dynamically allocated to obtain a computing resource allocation result. According to the computing resource allocation result, the pre-established load balancing algorithm is adopted to dynamically adjust the various resources of the personalized recommendation method to obtain a resource balancing result. According to the resource balancing result, the availability index of the personalized recommendation method is calculated to obtain a method availability value. It is determined whether the method availability value is greater than a preset availability threshold. If the method availability value is greater than the preset availability threshold, it is determined that the personalized recommendation method is currently in a high availability state, and an availability evaluation report is generated.
[0021] Specifically, the high availability of personalized recommendation methods is the key to ensuring user experience and business continuity. By obtaining the operation log data of the personalized recommendation method, various indicators of the personalized recommendation engine can be monitored in real time. For example, in an e-commerce platform, the operation log data can record request and response data, performance indicators, user behavior feedback, system status, errors and alarms, etc., to evaluate whether the personalized recommendation method runs smoothly. The establishment of a fault-tolerant recovery mechanism can improve the stability of the personalized recommendation method. For example, in the recommendation system of an e-commerce platform, a multi-level fault detection threshold can be set. When the error rate of a recommendation module exceeds the preset value, the backup algorithm is automatically triggered to ensure that the service is not interrupted. This mechanism can detect and respond to anomalies within milliseconds, greatly reducing the risk of method anomalies. Regular full backup is the cornerstone of data security. For the recommendation system of an e-commerce platform, incremental backup can be performed every morning and full backup can be performed once a week. This not only protects data such as user historical purchase information, but also provides a guarantee for emergency recovery. The deployment of off-site data centers further enhances the risk resistance of the method. For example, setting up the main data center in Beijing and setting up a mirror center in Guangzhou can effectively deal with regional disasters. When a failure is detected, the data backup disaster recovery mechanism will play a key role. In the e-commerce platform recommendation system, if the main database (failure point) is unavailable due to hardware failure, the system can switch to the backup data source within seconds and restore the lost data through the backup data to ensure that the user's investment decision is not affected. Or if the operation status detection results show that at least one remote data center (failure point) has failed, the system can also switch to other backup data sources or primary data sources within seconds and restore the lost data through the backup data to ensure that the user's investment decision is not affected. Dynamic resource scheduling algorithms can optimize method performance. When dynamically allocating computing resources based on the current load value, computing resources can be added to nodes with higher loads. For example, an additional 2 CPU cores are allocated to a node with a load of 1200 times / minute, and the computing resource allocation result is that the resources of each node are balanced to a processing capacity of about 1000 times / minute. When using the load balancing algorithm to adjust resources, the various resources of the personalized recommendation method that can be adjusted include computing resources (such as CPU, GPU), memory resources, service instances (number of containers and virtual machines), storage resources (such as database access), network bandwidth, task parallelism (number of shards of distributed tasks), Kafka partitions, etc.For example, the recommendation system faces a surge in load during the evening peak hours: computing resources: expanded from 50 to 100 to process real-time click streams; memory resources: 3 new nodes are added to the Redis cluster to cache the embedding vectors of popular products; service instances: Kubernetes expands the Pods of the recommendation API from 20 to 50; network bandwidth resources: enable dedicated network channels for cross-availability zone communication to reduce latency; load balancing: adjust the Nginx weight to direct 70% of the traffic to the newly expanded Pods. Result: The recommendation response time is reduced from 800ms to 200ms, and the system throughput is increased by 3 times. Availability indicators are a collection of indicators such as request response time, fault tolerance and fault recovery time, and request success rate (SLA achievement rate). The calculation of availability indicators is an important means of evaluating service quality. For example, the availability percentage can be calculated by "(total service time - failure time) / total service time". For critical businesses, an availability of more than 99.99% is usually required, that is, the failure time throughout the year does not exceed 52.56 minutes. Through this series of measures, the personalized recommendation method can maintain high availability in the face of various challenges and provide users with continuous, stable, and high-quality services. The availability evaluation report not only reflects the current operating status, but also provides direction for future optimization, so that the recommendation service can continue to evolve and meet the growing personalized needs of users.
[0022] S105. Obtain user scale growth data and analyze the user scale growth rate. If the user scale growth rate is greater than the preset growth rate threshold, it is determined that the user scale has increased significantly. Dynamically increase distributed storage and computing nodes, use the principle of data locality computing, optimize data transmission paths, reduce network transmission overhead, adjust resource allocation results, and respond to the user scale growth of personalized recommendation methods.
[0023] Obtain user scale growth data from the log database, and use linear regression to analyze the user scale growth rate. If the growth rate is less than or equal to the preset growth rate threshold, it is determined that the user scale is stable and no processing is performed. If the growth rate is greater than the preset growth rate threshold, it is determined that the user scale has increased significantly. If it is determined that the user scale has increased significantly, the storage capacity is dynamically increased according to the result of the user scale growth rate. Deploy computing nodes in areas with dense user distribution, allocate computing tasks nearby according to user geographic location information, use the shortest path algorithm to optimize the data transmission path, reduce the data transmission distance, and reduce network transmission overhead. Obtain load data from the preset monitoring system, use the dynamic resource scheduling algorithm to allocate idle computing resources, and improve computing efficiency. According to the load data, use the polling algorithm to adjust the resource allocation results to ensure the stability and efficiency of the personalized recommendation method.
[0024] Specifically, the user scale growth analysis of the personalized recommendation method is the key starting point of the optimization method. By extracting data such as user activity and registration volume from the log database, a linear regression model can be used to analyze the user scale growth rate. For example, an e-commerce platform analyzed the user data of the past 12 months and found that the number of users showed a stable month-on-month growth. The analysis showed that the number of users will increase by 15% in the next quarter. This growth rate directly affects the storage demand. In order to cope with the surge in data caused by user growth, the number of storage nodes needs to be adjusted dynamically. When the number of users is analyzed to increase significantly, the capacity can be expanded in advance. For example, based on the user growth analysis, an e-commerce platform decided to add 20 nodes to the original 100 storage nodes, increasing the total storage capacity from 1PB to 1.2PB to ensure the data storage needs for the next 6 months. User geographic distribution information is crucial to the performance of the optimization method. By analyzing user IP addresses or GPS data, user-dense areas can be determined. For example, an e-commerce platform found that 80% of its active users are concentrated in 10 large cities. Based on this finding, the platform deployed additional computing nodes in these cities, so that most users' requests can be processed within a range of 100 kilometers, significantly reducing network latency. The shortest path algorithm plays an important role in optimizing data transmission. Taking the product recommendation method as an example, when a user requests to view a product, the system calculates the optimal path from the nearest storage node to the user's device. This not only reduces the data transmission distance, but also reduces the possibility of network congestion. Practice has shown that this optimization can shorten the loading time of product details by more than 20%, greatly improving the user experience. Load data is an important basis for resource scheduling. The monitoring system collects indicators such as CPU usage, memory usage, and network throughput in real time to provide a decision-making basis for dynamic resource scheduling. For example, when an e-commerce platform detects a surge in load between 7 and 9 p.m., it automatically allocates idle computing resources to the recommendation engine to cope with the peak period of concentrated user access. This elastic scheduling strategy makes the method's operating response time during peak periods only 5% higher than usual, ensuring the consistency of user experience. Polling algorithm is an effective means to ensure service stability. In the personalized recommendation method, the polling algorithm can evenly distribute user requests to multiple servers. For example, during the Double 11 period, the recommendation system of an e-commerce platform adopted a weighted polling strategy to dynamically adjust the weight according to the processing capacity of each server to ensure load balancing. This not only improves the concurrent processing capability of the method, but also reduces the risk of single point failure, so that even at traffic peaks, the average response time of the method is controlled within 200 milliseconds. Through the comprehensive application of these technologies, the personalized recommendation method can maintain efficient and stable operation under the rapid growth of the user scale, provide users with a smooth and personalized service experience, and cope with the growth of the user scale of the personalized recommendation method.
[0025] S106: Continuously collect user feedback data on recommendation results, optimize the recommendation results according to the feedback data, and obtain optimized personalized recommendation results.
[0026] The click-through rate, dwell time and conversion rate data of the recommendation results are extracted from the user behavior log, and the data is cleaned and feature extracted through the streaming computing architecture. The user behavior log records the detailed operation and interaction data of the user. The click-through rate, dwell time and conversion rate data of the recommendation results are aggregated by the sliding window mechanism to obtain the time series characteristics of the user behavior. If the characteristic value of the time series characteristic exceeds the preset characteristic value threshold, the engine update mechanism is triggered, and the time series characteristics exceeding the characteristic value threshold are input into the personalized recommendation engine. The personalized recommendation engine parameters are optimized by the gradient descent algorithm, and the user preference portrait data in the personalized recommendation engine is updated to obtain the updated user preference portrait data. The updated user preference portrait data is used by the personalized recommendation engine to generate new recommendation results. The collaborative filtering algorithm is used to calculate the similarity between the updated user preference portrait data and other user preference portrait data, and the target user preference portrait data with a similarity greater than the preset similarity threshold is selected, and the target recommendation result corresponding to the target user preference portrait data is obtained. The recommendation result is updated by combining the new recommendation result and the target recommendation result to obtain the optimized personalized recommendation result.
[0027] Specifically, the personalized recommendation method optimizes the recommendation effect by analyzing the user behavior log. The user behavior log is the detailed operation and interaction data of the user in the application (such as click, stay, purchase, etc.) recorded by the system, which contains information such as time, object, context, etc., and is used to calculate key indicators such as click rate, stay time and conversion rate in real time. Taking an e-commerce platform as an example, the click rate, stay time and conversion rate data are extracted from the user behavior log. The click rate reflects the user's initial interest in the recommended content, the stay time reflects the attractiveness of the content, and the conversion rate shows whether the user has completed the expected behavior, such as adding to the shopping cart or purchasing. The streaming computing architecture can process massive data in real time. The platform uses Apache Flink for data cleaning and feature extraction. During the cleaning process, the method will filter out abnormal data, such as records with too short stay time. Feature extraction includes calculating the user's viewing preferences, active time periods, etc. The sliding window mechanism is used to capture the dynamic changes of user behavior. Assume that the platform sets a sliding window of 5 minutes and updates it every 1 minute. Within this window, the method will calculate the user's average stay time, click rate and other time series features. If a user's stay time suddenly increases from an average of 10 minutes to 30 minutes, exceeding the preset 20-minute threshold, the update mechanism of the recommendation engine will be triggered, and the time series features exceeding the feature value threshold will be input into the personalized recommendation engine, and the updated user preference portrait data can be obtained. Specifically, when optimizing the parameters of the personalized recommendation engine through the gradient descent algorithm, the weight of the user preference portrait data in the personalized recommendation engine can be adjusted based on the newly input time series features. For example, the system analysis found that the proportion of users' browsing time for fitness equipment products increased from 40% to 60%, so the gradient descent algorithm will gradually adjust the weight of the user preference portrait data for fitness equipment preferences from 0.5 to 0.7. The core of this user preference portrait data optimization method is to use the difference between historical data and new data to gradually approach the optimal solution, ensuring that the user preference portrait data can reflect the latest points of interest in a timely manner. Preferably, this method can also avoid the instability of recommendation results caused by too fast parameter adjustment. After obtaining the updated user preference portrait data, the updated user preference portrait data is used by the personalized recommendation engine to generate new recommendation results. Exemplarily, a collaborative filtering algorithm is used to calculate the similarity between the updated user preference portrait data and other user preference portrait data using cosine similarity or Pearson correlation coefficient based on the interaction data of different users on items (such as click-through rate, dwell time and conversion rate data, etc.). Other user preference portrait data with a similarity higher than a preset similarity threshold can be used as target user preference portrait data, and then the target recommendation result corresponding to the target user preference portrait data is obtained, and finally the target recommendation result and the new recommendation result are used to update the recommendation result to obtain an optimized personalized recommendation result.For example, the preset similarity threshold is 80%. After a user's preference profile data is updated, it shows a strong interest in sports shoes and fitness equipment. The system finds the target user preference profile data with a similarity greater than 80%, and finds that these users often browse yoga mat recommendations, that is, yoga mats exist in the target recommendation results. Combining these target recommendation results and new recommendation results, the system merges the target recommendation results, new recommendation results and recommendation results, such as removing duplicate data, etc., to obtain personalized recommendation results. For example, in the above yoga mat example, if there is no recommendation for yoga mats in the recommendation results and target recommendation results, yoga mats will be added to the optimized personalized recommendation results. This method can explore the user's potential interests and improve the diversity of recommendations. It is understandable that after the personalized recommendation results are input into the streaming computing architecture, the personalized recommendation results can be continuously optimized. For example, when a user browses a page, the system updates the recommendation results every 5 minutes through the streaming computing architecture based on the recommendation list generated in the previous steps. Specifically, when a user clicks on a recommended pair of sneakers, the system will immediately feed this behavior back to the engine, triggering a new round of feature extraction and recommendation list generation, and updating personalized recommendation results. The core of this closed-loop iterative process is to quickly respond to changes in user behavior and ensure that the recommendation results are always close to the user's current needs. Through the comprehensive application of this series of technologies, the system can provide users with more accurate and personalized content recommendations. This not only improves the user experience, but also increases the user stickiness of the platform. At the same time, this data-driven approach also provides valuable insights into the platform's content creation and operation strategies.
[0028] S107, cache the optimized personalized recommendation results to the distributed cache cluster, adopt a multi-copy storage mechanism, and route user requests to the nearest cache node through load balancing to improve the access speed of the optimized personalized recommendation results.
[0029] The distributed cache cluster obtains the optimized personalized recommendation results and uses a multi-copy storage mechanism to store the recommendation results in multiple cache nodes. The source of user requests is analyzed through a load balancing algorithm, and the optimal cache node is determined based on the geographic location and network delay information. The user request is routed to the optimal cache node, and the optimized personalized recommendation results are obtained from the optimal cache node, thereby improving the access speed of the optimized personalized recommendation results.
[0030] Specifically, the distributed cache cluster is a key technology to improve the performance of the personalized recommendation method. The multi-copy storage mechanism ensures the high availability and fault tolerance of the data. For example, an e-commerce platform stores the personalized recommendation results of users in multiple cache nodes distributed across the country. Each recommendation result has at least three copies, distributed in different geographical locations to cope with possible node failures. The load balancing algorithm selects the optimal cache node by analyzing the source of user requests. For example, user requests from Beijing may be routed to cache nodes in North China, while users from Guangzhou are directed to nodes in South China. This geographical location-based routing strategy can significantly reduce network latency and improve user experience. Furthermore, in this application, the access pressure of the cache node is monitored in real time by adopting a sliding window mechanism. If the node access pressure exceeds the preset access volume threshold, dynamic load balancing adjustment is triggered. The cache hit rate is analyzed according to the user access log. If the hit rate is lower than the preset hit rate value, the cache strategy is updated and the copy distribution of the optimized personalized recommendation result is adjusted. The user access log records the detailed data of the user's request frequency, access time, resource path and cache hit status of the cache node. The load balancing parameters are optimized by the gradient descent algorithm, and the routing strategy of user requests is dynamically adjusted. The collaborative filtering algorithm is used to calculate the similarity of user preferences, and the recommendation results of similar users are stored in the nearest cache node. The sliding window mechanism is used to monitor the access pressure of the cache node in real time. Assume that an e-commerce platform sets a 5-minute sliding window and updates it every 1 minute. If within this window, the number of visits to a node suddenly increases from an average of 1,000 times per minute to 3,000 times, exceeding the preset threshold of 2,000 times, the method will trigger dynamic load balancing adjustment. This may involve rerouting some user requests to other relatively idle nodes, or adding new cache nodes to share the pressure. The user access log is a detailed data that records the user's request frequency, access time, resource path and cache hit status (such as hit / miss) of the cache node, which is used to analyze node pressure in real time, trigger load balancing strategies and optimize cache distribution to improve the performance of the recommendation system. Cache hit rate is an important indicator for evaluating the effectiveness of cache strategies. Combined with the data in the user access log, if the cache hit rate of a certain type of recommendation result is found to be lower than the preset threshold of 80%, the cache strategy update will be triggered. This may include increasing the cache time of popular content, or adjusting the distribution of recommendation results among different nodes. For example, the recommendation results that have been frequently accessed in the last 24 hours are copied to more cache nodes to improve the overall hit rate. The gradient descent algorithm is used to optimize the load balancing parameters. The method continuously adjusts the routing weight of user requests based on historical data. For example, if it is found that the number of visits in East China has surged within a certain period of time, the method will automatically increase the weight of the cache nodes in the region so that more requests are allocated to these nodes. This dynamic adjustment can better cope with traffic fluctuations and ensure the stability of the method.Collaborative filtering algorithms also play an important role in cache optimization. The method calculates the preference similarity between users and stores the recommendation results of similar users in similar cache nodes. For example, user A and user B who like technology products are identified as similar users, and the method stores their recommendation results in the same or adjacent cache nodes. This strategy not only improves the cache hit rate, but also reduces the replication of data between nodes and improves the overall efficiency of the method. Through the comprehensive application of these technologies, the personalized recommendation method can significantly improve the response speed and stability of the method while ensuring the quality of the recommendation. This not only improves the user experience, but also brings higher user retention and economic benefits to the platform. At the same time, this data-driven approach also provides reliable technical support for the continuous optimization of the method.
[0031] Reference Figure 2 In an embodiment of the present application, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Figure 2 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as user behavior characteristics and recommendation results. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the personalized recommendation method based on real-time user behavior of any of the above embodiments is implemented.
[0032] Those skilled in the art will understand that Figure 2 The structure shown in is merely a block diagram of a portion of the structure related to the present application solution and does not constitute a limitation on the computer device to which the present application solution is applied.
[0033] The embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the personalized recommendation method based on real-time user behavior of any of the above embodiments is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0034] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Any reference to memory, storage, database or other media provided in this application and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0035] Although the present invention has been described in detail above by general description and specific embodiments, it is obvious to those skilled in the art that some modifications or improvements can be made to the present invention. Therefore, these modifications or improvements made without departing from the spirit of the present invention all belong to the scope of protection claimed by the present invention.
Claims
1. A personalized recommendation method based on real-time user behavior, characterized in that: The method comprises: Collect real-time user behavior data, perform streaming processing on the real-time user behavior data, smooth traffic fluctuations through message queue buffering, and obtain user behavior feature data; input the user behavior feature data into the distributed computing architecture, process the user behavior feature data through shard storage and parallel computing tasks, and obtain optimized user behavior feature data; use Flink-SQL to perform real-time statistics and aggregation on the optimized user behavior feature data, and implement task scheduling of multiple computing nodes through distributed coordination services, so as to determine popular items and user interest preferences, and generate recommendation results through a personalized recommendation engine; in view of the risks of failure and data loss during the operation of the personalized recommendation method, adopt a fault-tolerant recovery mechanism and data backup and disaster recovery, implement load balancing through dynamic resource scheduling, and judge individual The system monitors the running status of the personalized recommendation method and generates an availability evaluation report; obtains the user scale growth data and analyzes the user scale growth rate. If the user scale growth rate is greater than the preset growth rate threshold, it is determined that the user scale has increased significantly, and dynamically increases the distributed storage and computing nodes. The principle of data proximity computing is used to optimize the data transmission path, reduce network transmission overhead, and adjust the resource allocation results to cope with the user scale growth of the personalized recommendation method; continuously collects user feedback data on the recommendation results, optimizes the recommendation results according to the feedback data, and obtains the optimized personalized recommendation results; caches the optimized personalized recommendation results in the distributed cache cluster, adopts a multi-copy storage mechanism, and routes user requests to the nearest cache node through load balancing to improve the access speed of the optimized personalized recommendation results.
2. The method according to claim 1, characterized in that The real-time user behavior data is collected, the real-time user behavior data is stream processed, and the flow fluctuation is smoothed by message queue buffering to obtain user behavior feature data, including: Obtain user real-time behavior data from the data collection layer to obtain the original behavior data stream; Use Flink to build a real-time data stream processing pipeline; Buffering the original behavior data stream through Kafka to smooth instantaneous traffic fluctuations, and inputting the buffered data stream into the real-time data stream processing pipeline; In Flink, window functions are used to perform real-time aggregation processing on the buffered data stream to obtain aggregated user behavior data; Obtain aggregated user behavior data, classify it using the K nearest neighbor classification algorithm, and obtain classification result data that characterizes the user behavior pattern; The classification result data is obtained, and the classification result data is subjected to feature screening by a feature importance screening method using a random forest algorithm to obtain user behavior feature data reflecting user behavior characteristics, wherein the feature importance screening method refers to calculating feature importance scores by a random forest algorithm to screen out features that have a significant impact on the target variable.
3. The method according to claim 1, characterized in that The step of inputting the user behavior characteristic data into a distributed computing architecture and processing the user behavior characteristic data through shard storage and parallel computing tasks to obtain optimized user behavior characteristic data includes: Input the user behavior feature data into the distributed computing architecture, and according to the preset sharding strategy, use the distributed storage architecture to shard the user behavior feature data, evenly distribute the user behavior feature data to at least one storage node, and obtain sharded data; For the sharded data, at least one parallel computing task is started, and the sharded data is processed using the MapReduce computing framework to obtain the data processed by MapReduce. Obtain the data processed by MapReduce, use the KMeans algorithm to perform cluster analysis on the data processed by MapReduce, and obtain the grouping result data of user behavior; The principal component analysis algorithm is used to reduce the dimension of the grouping result data of user behavior to obtain optimized user behavior feature data.
4. The method according to claim 1, characterized in that The method uses Flink-SQL to perform real-time statistics and aggregation on the optimized user behavior feature data, and implements task scheduling of multiple computing nodes through distributed coordination services, thereby determining popular items and user interest preferences, and generating recommendation results through a personalized recommendation engine, including: Obtain the optimized user behavior feature data, and use Flink-SQL to perform real-time statistical processing on the optimized user behavior feature data to obtain statistical values; Performing real-time aggregation calculation on the statistical value to obtain a scheduling value; Performing distributed computing processing on multiple computing nodes according to the scheduling value to obtain a computing value; For the calculated value, a hot item analysis is performed using a preset popularity analysis algorithm to obtain an item value; Using a preset user interest analysis model, perform user interest analysis according to the item value to obtain an interest value; Inputting the interest value into a preset user preference analysis model for processing to obtain the user's preference value; If the preference value is greater than a preset preference threshold, generating user preference portrait data according to the preference value; By using this user preference portrait data, targeted recommendation results are generated in real time through a personalized recommendation engine.
5. The method according to claim 1, characterized in that: In view of the risks of failure and data loss during the operation of the personalized recommendation method, a fault-tolerant recovery mechanism and data backup and disaster recovery are adopted, load balancing is achieved through dynamic resource scheduling, the operation status of the personalized recommendation method is judged, and an availability evaluation report is generated, including: Acquire operation log data of the personalized recommendation method, wherein the operation log data refers to a collection of events and data automatically recorded when the personalized recommendation method is running; Adopting a pre-established fault-tolerant recovery mechanism, analyzing the operation log data, detecting the operation status of the personalized recommendation method, and obtaining an operation status detection result; According to the preset data backup strategy, the optimized user behavior feature data in the personalized recommendation method is backed up regularly in full; Combined with the preset disaster recovery strategy, full backup data is deployed in the off-site data center; According to the running status detection result, if at least one fault point is detected, a preset data backup disaster recovery mechanism is triggered to switch the user request to the backup data, restore the lost data caused by the at least one fault point, and obtain the restored data; According to the restored data, a pre-established dynamic resource scheduling algorithm is used to analyze the load condition of the personalized recommendation method in operation to obtain a current load value; According to the current load value, dynamically allocate computing resources of the personalized recommendation method to obtain a computing resource allocation result; Adopting a pre-established load balancing algorithm, dynamically adjusting various resources of the personalized recommendation method according to the computing resource allocation result, and obtaining a resource balancing result; According to the resource balancing result, calculating the availability index of the personalized recommendation method to obtain a method availability value; It is determined whether the method availability value is greater than a preset availability threshold. If the method availability value is greater than the preset availability threshold, it is determined that the personalized recommendation method is currently in a high availability state, and an availability evaluation report is generated.
6. The method according to claim 1, characterized in that The method of obtaining user scale growth data and analyzing the user scale growth rate is as follows: if the user scale growth rate is greater than a preset growth rate threshold, it is determined that the user scale has grown significantly, and distributed storage and computing nodes are dynamically added. The method uses the principle of data proximity computing to optimize the data transmission path, reduce network transmission overhead, and adjust resource allocation results to cope with the user scale growth of the personalized recommendation method, including: Obtain user scale growth data from the log database, and use linear regression to analyze the user scale growth rate. If the growth rate is less than or equal to the preset growth rate threshold, it is determined that the user scale is stable and no processing is performed. If the growth rate is greater than the preset growth rate threshold, it is determined that the user scale has increased significantly. If it is determined that the user scale has increased significantly, the storage capacity will be dynamically increased based on the result of the user scale growth rate; Deploy computing nodes in areas with dense user distribution, assign computing tasks to nearby users based on their geographic location information, and use the shortest path algorithm to optimize data transmission paths, reduce data transmission distances, and reduce network transmission overhead; Obtain load data from the preset monitoring system and use dynamic resource scheduling algorithms to allocate idle computing resources to improve computing efficiency; According to the load data, a polling algorithm is used to adjust the resource allocation results to ensure the stability and efficiency of the personalized recommendation method.
7. The method according to claim 1, characterized in that The continuously collecting the user's feedback data on the recommendation results, optimizing the recommendation results according to the feedback data, and obtaining the optimized personalized recommendation results include: Extract the click-through rate, dwell time, and conversion rate data of the recommendation results from the user behavior log, and clean and extract features of the data through the streaming computing architecture. The user behavior log records the user's detailed operations and interaction data; A sliding window mechanism is used to aggregate the click-through rate, dwell time, and conversion rate data of the recommendation results to obtain the temporal characteristics of user behavior; If the feature value of the time series feature exceeds the preset feature value threshold, the engine update mechanism is triggered to input the time series feature exceeding the feature value threshold into the personalized recommendation engine; Optimize the personalized recommendation engine parameters through the gradient descent algorithm, update the user preference portrait data in the personalized recommendation engine, and obtain the updated user preference portrait data; Generate new recommendation results through the personalized recommendation engine using updated user preference profile data; A collaborative filtering algorithm is used to calculate the similarity between the updated user preference portrait data and other user preference portrait data, and the target user preference portrait data with a similarity greater than a preset similarity threshold is screened, and the target recommendation result corresponding to the target user preference portrait data is obtained; The recommendation results are updated by combining the new recommendation results with the target recommendation results to obtain optimized personalized recommendation results.
8. The method according to claim 1, characterized in that: The optimized personalized recommendation results are cached in a distributed cache cluster, a multi-copy storage mechanism is adopted, and user requests are routed to the nearest cache node through load balancing to improve the access speed of the optimized personalized recommendation results, including: The distributed cache cluster obtains the optimized personalized recommendation results and stores the recommendation results in multiple cache nodes using a multi-copy storage mechanism; Analyze the source of user requests through load balancing algorithms and determine the optimal cache node based on geographic location and network latency information; The user request is routed to the optimal cache node, and the optimized personalized recommendation results are obtained from the optimal cache node, thereby improving the access speed of the optimized personalized recommendation results.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Cloud computing based real-time mass user behavior analyzing method and system
CN103793465A
Position-based edge cloud resource scheduling method and system
CN110266744A
News recommendation method and system based on user behavior detection and computer equipment
CN110489652A
Article score prediction method, device and system and storage medium
CN115759381A
Cross-server cluster grouping theory and processing method
CN117411887A
Cited By
Machine learning-based time sequence personalized recommendation system and method
CN120670669A
Marketing method and device based on user behavior analysis, equipment and medium
CN120952924A
Intelligent ticket business recommendation method and system based on dynamic flight data
CN121190162A
Intelligent Ticketing Recommendation Method and System Based on Dynamic Flight Data
CN121190162B