Content caching and replacing method and system based on context awareness

By constructing context feature vectors and XGBoost models to predict future access popularity, combined with NSGA-II algorithm to optimize the cache strategy, the problems of low cache hit rate and low resource utilization rate in the edge cache method are solved, and efficient and intelligent cache management is achieved in complex network environments.

CN120354972AActive Publication Date: 2025-07-22JIANGXI NORMAL UNIV

Patent Information

Application Number
CN202510838313.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-22
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

Existing edge caching methods lack the ability to perceive and dynamic adjustment of multi-dimensional context information, resulting in low cache hit rate, low resource utilization rate, and traditional strategies are difficult to achieve the intelligence of multi-objective optimization and content replacement in complex network environments.

Method used

By constructing context feature vectors, using XGBoost regression model to predict future content access popularity, and combining NSGA-II multi-objective evolution algorithm to optimize the cache strategy, design cache scoring functions and greedy replacement mechanisms to achieve multi-dimensional perception and dynamic adjustment of user behavior, node state and content characteristics.

Benefits of technology

It significantly improves the cache hit rate, reduces transmission delay and resource consumption, enhances content caching efficiency and resource utilization efficiency in edge computing environments, and realizes an intelligent and dynamic cache management strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354972A_ABST
    Figure CN120354972A_ABST
Patent Text Reader

Abstract

The invention discloses a content caching and replacing method and system based on context awareness, and relates to the field of content caching, and the method comprises the steps: obtaining context information, constructing a context feature vector, constructing a training sample, training an XGBoost regression model, and carrying out the prediction, thereby obtaining the future content access popularity; a multi-objective optimization function and constraint conditions are designed, an NSGA-II multi-objective evolutionary algorithm is used for solving, a comprehensive scoring function is defined, and a solution with the maximum comprehensive scoring function value is selected from an obtained solution set to serve as a final caching strategy; and designing a cache scoring function, calculating to obtain a cache score, and replacing the cache content by adopting a greedy strategy. According to the content caching and replacing method, the caching hit rate and the resource utilization efficiency can be improved, the service capability and the response efficiency of the edge node in a complex environment can be further enhanced, and therefore a more intelligent and more efficient content caching and replacing strategy is achieved in a dynamic network scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of content caching, and in particular, to a context-aware content caching and replacement method and system. Background Art

[0002] With the rapid development of mobile Internet, Internet of Things, and edge computing technologies, end users have increasingly higher requirements for data access latency, service availability, and transmission efficiency. To alleviate the transmission pressure on the core network and improve the user experience, edge caching technology is widely used to pre-store content in edge nodes close to the user side, so as to achieve fast response. However, most of the current mainstream edge caching methods are based on traditional content access statistical information, such as access frequency or recent access time and other strategies. These methods rely on the historical access patterns of content for caching decisions and cannot effectively predict future access trends, and are prone to problems such as low hit rate and poor caching efficiency in practical applications.

[0003] In addition, when making caching decisions, existing methods generally lack the perception and utilization of multi-dimensional context information. In real edge computing scenarios, user behavior characteristics, the operating status of network nodes, and the attributes of the requested content will all have a significant impact on the access probability of content. Ignoring these context factors closely related to content requests will make the caching strategy lack environmental adaptability and prediction accuracy, resulting in unreasonable allocation and inefficient use of system resources.

[0004] On the other hand, edge caching systems usually face multiple optimization goals, and many current strategies are only optimized for a single metric, lacking the ability to model the balance relationship between multiple goals, and it is difficult to meet the diverse performance requirements in complex network environments. At the same time, in the content replacement process, existing methods mostly adopt fixed or static replacement strategies, without considering the future access possibility of content, resource consumption cost, and the current system state. The replacement decision lacks intelligence and dynamics, reducing the utilization efficiency of the cache space.

[0005] CN111629217A, titled: VOD Service Cache Optimization Method Based on XGBoost Algorithm in Edge Network Environment. The method includes: collecting video data; using the XGBoost algorithm to perform regression modeling with the average access volume as the prediction target to obtain a prediction model; using the prediction model to predict the average access volume; establishing a cache optimization model according to the prediction results; using the knapsack algorithm to solve the optimization model to obtain the final cache scheme. This method takes into account that edge servers need to process a large amount of video information and the excellent data analysis ability of machine learning in big data processing, so that edge servers can minimize the service access delay to the greatest extent and improve the cache efficiency of edge servers. However, this method lacks the ability to perceive environmental changes in real time and make dynamic adjustments. It only uses the average access volume of videos as the prediction target and lacks a comprehensive model of multi-dimensional context factors, resulting in the prediction results may not be able to dynamically adapt to complex and changing network and user environments. At the same time, it also lacks fine-grained evaluation and replacement decisions for the cached content, so the cache efficiency is limited in high-frequency update scenarios.

[0006] CN114513514A, titled: An Edge Network Content Caching and Prefetching Method for Vehicle Users. The method includes: according to constraints such as content popularity, vehicle speed, request initiation location, and edge server cache capacity, adopting the method of content segmentation, and through vertical collaboration between the cloud server and the edge server, completing the caching and prefetching configuration of the content requested by the vehicle, maximizing the resource utilization rate of the edge server, reducing the content download delay of vehicle users, and reducing the content update cost of the edge server. By combining factors such as content popularity, vehicle speed, and location, and adopting the method of content segmentation and edge-cloud collaboration, this method effectively improves the content acquisition efficiency of vehicle users in high-speed mobile scenarios. However, it lacks consideration for replacing cached content, and the optimization goal is relatively single, making it difficult to achieve a more efficient and refined caching strategy in a complex network environment. Summary of the Invention

[0007] In view of the above problems, the present invention proposes a context-aware content caching and replacement method and system. This method can comprehensively perceive multi-dimensional context information such as user behavior, node status, and content characteristics. By constructing context feature vectors and introducing an XGBoost (Extreme Gradient Boosting) regression model, it accurately predicts the future content access popularity. On this basis, the NSGA-II (Non-dominated Sorting Genetic Algorithm II) is used to optimize the caching strategy, achieving global balanced optimization of cache hit rate, transmission delay, and bandwidth occupancy. This method not only significantly improves the cache hit rate but also effectively reduces system latency and resource consumption, thus realizing an efficient, intelligent, and sustainable caching management strategy in a dynamic and complex edge network environment.

[0008] One technical solution of the present invention is: A context-aware content caching and replacement method, the method comprising: Obtain context information at different times and construct context feature vectors; Based on the context feature vectors, construct training samples, train an XGBoost regression model, and obtain a trained XGBoost prediction model; Construct a cache decision candidate set, collect the context feature vectors, and use the trained XGBoost prediction model for prediction to obtain the future content access popularity; Based on the cache strategy requirements, design a multi-objective optimization function and constraint conditions; Based on the multi-objective optimization function and constraint conditions, use the NSGA-II multi-objective evolutionary algorithm to solve and obtain a Pareto solution set; Define a comprehensive scoring function, select the individual with the maximum comprehensive scoring function value from the Pareto solution set as the final caching strategy, use the corresponding content set as the final cached content set, and deploy the cache to the edge node; Design a cache scoring function for a single content, obtain a cache score based on the cache scoring function, and use a greedy strategy to replace the cached content based on the cache score; The content refers to the content requested by the user for access.

[0009] Further, the obtaining context information at different times and constructing context feature vectors includes: Perform preprocessing operations on the context information to obtain preprocessed context information; Construct the context feature vectors according to the preprocessed context information.

[0010] Furthermore, the context information includes: user behavior information, node status information, and request content information; The user behavior information includes: user access frequency, user location, user speed, and user activity; The node status information includes: node bandwidth, node transmission delay, and node load; The request content information includes: the total number of times the content is accessed within a sliding time window, and the content size; The preprocessing operations include: outlier removal processing, missing value filling processing, and normalization processing; The context feature vector has the following expression: ; Among them, refers to the context feature vector, refers to the content, refers to the user, refers to the edge node, refers to the current moment, refers to the user 's user access frequency to the content within the sliding time window, refers to the user's location at the current moment , refers to the user's speed at the current moment , refers to the user's activity at the current moment, that is, the total number of user requests, , refers to the node bandwidth at the current moment , refers to the node transmission delay at the current moment , refers to the node load at the current moment , refers to the content 's total number of accesses within the sliding time window, refers to the content size, that is, the size of the content .

[0011] Furthermore, constructing training samples based on the context feature vector, training an XGBoost regression model, and obtaining a trained XGBoost prediction model includes: Based on the context feature vector, collecting the actual access popularity within the time span after its corresponding time point as label data; Forming a training sample set by combining the context feature vectors at all historical time points with their corresponding label data; Set the input of the XGBoost regression model as the context feature vector and the output as the future content access popularity. Define a loss function and, through iterative training, continuously optimize the parameters of the XGBoost regression model using gradient boosting trees to minimize the loss function, and finally obtain the trained XGBoost prediction model.

[0012] Further, the labeled data refers to the actual access popularity of the content at at time, and its expression is: ; The time span is expressed as: ; The training sample set is expressed as: ; where refers to the training sample set, refers to the context feature vector, refers to the labeled data, refers to the time span, refers to the number of samples; The future content access popularity is expressed as: ; The parameters are expressed as: ; The loss function is expressed as: ; where refers to the loss function, refers to the model parameters, refers to the number of samples, refers to the index variable, refers to the true content access popularity, refers to the model-predicted content access popularity, refers to the model complexity regularization term, which represents the model complexity penalty term, including the number and depth of regression trees, and is used to prevent overfitting, is the regularization parameter.

[0013] Further, the construction of the cache decision candidate set, collecting the context feature vector, and using the trained XGBoost prediction model for prediction to obtain the future content access popularity includes: At the edge node, count and generate the set of all user-requested content, denoted as the candidate content set; For each piece of content in the candidate content set, collect its context feature vector, and use the trained XGBoost prediction model to make a prediction to obtain the predicted future content access popularity. The candidate content set, whose expression is: ; where refers to the candidate content set, refers to each corresponding piece of content.

[0014] Furthermore, the cache policy requirement refers to maximizing the cache hit rate, minimizing the transmission delay, and minimizing the bandwidth occupancy. The multi-objective optimization function includes a cache hit rate maximization objective function, a transmission delay minimization objective function, and a bandwidth occupancy minimization objective function. The cache hit rate maximization objective function, whose expression is: ; where refers to the cache hit rate maximization objective function, refers to the set of actually selected cached content, refers to the candidate content set, refers to the future content access popularity, refers to the time span; The transmission delay minimization objective function, whose expression is: ; where refers to the transmission delay minimization objective function, refers to the set of actually selected cached content, refers to the current moment of the node transmission delay; The bandwidth occupancy minimization objective function, whose expression is: ; where refers to the bandwidth occupancy minimization objective function, refers to the set of actually selected cached content, refers to the content size, refers to the future content access popularity, refers to the time span, refers to the current moment of the node bandwidth; The constraint condition refers to the cache capacity constraint condition, whose expression is: ; Among them, refers to the actually selected cache content set, refers to the content size, refers to the total cache capacity of the edge nodes.

[0015] Furthermore, based on the multi-objective optimization function and constraint conditions, the NSGA-II multi-objective evolutionary algorithm is used to solve, and a Pareto solution set is obtained, including: Encode each individual, use a binary vector to represent the cache selection scheme of the content, and judge whether to cache the content to obtain a cache selection vector; Randomly generate several legal individuals to form an initial population, set the population size, and initialize the iteration counter; Calculate the three objective functions of each individual to obtain the objective function values; Perform an evolutionary operation on the current population to generate the next generation population; Merge the current population and the next generation population into a combined population, re-perform non-dominated sorting and crowding degree evaluation, select the individuals with the number of the previous population size to form a new population, update the iteration counter, and stop the evolutionary process until the set maximum iteration number is reached, and finally obtain the Pareto solution set.

[0016] Furthermore, the individual refers to a complete content cache selection scheme, which is represented in the form of a binary vector. Each bit in the vector corresponds to the cache decision of the content in a candidate content set, so as to form the cache selection vector; The cache selection vector has the following expression: ; Among them, refers to the cache selection vector, refers to the binary vector, indicating whether to select the content in the th candidate content set to be cached on the edge node. Its value of 1 means caching, and its value of 0 means not caching, refers to the total number of candidate content sets; The several legal individuals need to satisfy the cache capacity constraint conditions; The cache capacity constraint conditions have the following expression: ; Among them, refers to the actually selected cache content set, refers to the content size, refers to the total cache capacity of the edge nodes; The initial population is expressed as: ; The population size is expressed as: ; The initialized iteration counter is expressed as: ; The current population is expressed as: ; The three objective functions refer to: the objective function of maximizing the cache hit rate, the objective function of minimizing the transmission delay, and the objective function of minimizing the bandwidth occupation; The evolutionary operations include: non - dominated sorting, crowding degree calculation, selection operation, crossover operation, and mutation operation; The new population is expressed as: ; The maximum number of iterations is expressed as: ; The Pareto solution set is expressed as: .

[0017] Furthermore, defining a comprehensive scoring function, selecting the individual with the maximum comprehensive scoring function value from the Pareto solution set as the final caching strategy, and taking the corresponding content set as the final cached content set and deploying the cache to the edge node includes: Defining a comprehensive scoring function for each individual in the Pareto solution set; Among all the solutions in the Pareto solution set, selecting the individual that makes the comprehensive scoring function reach the maximum value, denoted as the optimal solution, and obtaining the optimal cache selection vector; Based on the value of each bit in the binary of the optimal cache selection vector, determining whether to cache the corresponding content, obtaining the final cached content set, and caching the content in the final cached content set to the edge node.

[0018] Furthermore, each individual refers to each cache selection vector; The comprehensive scoring function has the following expression: ; Where, refers to the comprehensive scoring function, refers to a caching strategy individual in the Pareto solution set, , refers to the caching strategy 's predicted hit rate value, refers to the caching strategy 's average transmission delay at the current edge node, refers to the caching strategy 's total bandwidth consumption at the current edge node, , and Refers to the weight parameter, which simultaneously satisfies ; The optimal solution is expressed as: ; The optimal cache selection vector is expressed as: ; The final cache content set is expressed as: .

[0019] Furthermore, the designed cache scoring function for a single content, based on which the cache score is obtained, and based on the cache score, the cache content is replaced using a greedy strategy, including; Design the cache scoring function, and calculate the cache scores of the contents in all the candidate content sets and the cached content at the current edge node according to the future content access popularity, the content size, and the node load; Sort the contents in all the candidate content sets and the cached content according to the cache score from high to low, and use a greedy strategy to replace the cache content.

[0020] Furthermore, the cache scoring function has the following expression: ; Where refers to the cache scoring function, refers to the future content access popularity, refers to the time span, refers to the content size, refers to the current moment node load, which is used as the load penalty term in the cache scoring function to penalize the cache selection of high-load nodes and reflect the current cache pressure; The greedy strategy includes: Sort the contents in all the candidate content sets according to the cache score from high to low; Traverse the sorted list and perform cache replacement judgment; The cache replacement judgment is specifically: If the content in the candidate content set is not cached and the cache space is sufficient, directly add it to the cache until the cache space is exhausted; If the cache space is insufficient, compare the cache score of the content in the candidate content set with the content with the lowest cache score among the cached content; If the cache score of the content in the candidate content set is higher, replace the content with the lowest cache score among the cached content, otherwise skip the content in the candidate content set.

[0021] Based on the above context-aware content caching and replacement method, the present invention also provides a context-aware content caching and replacement system, which includes: A data processing module, configured to obtain context information, perform preprocessing operations on the context information to obtain data after preprocessing operations, and construct a context feature vector; A popularity prediction module, configured to train an XGBoost prediction model based on the context feature vector obtained by the data processing module, and use the XGBoost prediction model to predict the future content access popularity; A content caching module, configured to design a multi-objective optimization function based on cache policy requirements and solve it using the NSGA-II algorithm to obtain a Pareto solution set, define a comprehensive scoring function, select the individual with the maximum comprehensive scoring function value from the Pareto solution set as the final caching policy, use the corresponding content set as the final cached content set, and deploy the cache to the edge node; A cache replacement module, configured to design a cache scoring function for a single content based on the future content access popularity predicted by the popularity prediction module, obtain a cache score based on the cache scoring function, and replace the cached content of the content caching module using a greedy strategy based on the cache score.

[0022] A context-aware content caching and replacement method and system provided by an embodiment of the present invention accurately predict the future content access popularity by collecting and analyzing multi-dimensional context feature information, combine with the XGBoost prediction model, optimize the cache policy using a multi-objective evolutionary algorithm, and finally adopt a greedy replacement mechanism based on cache scoring, achieving a significant improvement in cache hit rate, effectively reducing transmission delay and bandwidth occupancy, and significantly enhancing the efficiency and resource utilization efficiency of content caching in the edge computing environment.

[0023] The above invention content is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically describes the specific embodiments of the present invention. Brief Description of the Drawings

[0024] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those skilled in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0025] Figure 1Shows a specific flowchart of a context - aware content caching and replacement method provided by an embodiment of the present invention.

[0026] Figure 2 Shows a framework diagram of a context - aware content caching and replacement system provided by an embodiment of the present invention. Detailed implementation manners

[0027] Hereinafter, exemplary embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.

[0028] With the rapid development of services such as mobile Internet, Internet of Things, and high - definition video, users have put forward higher requirements for the real - time and reliability of data access. As an important means to relieve the pressure on the central server and reduce access latency, edge computing has been widely applied to content distribution systems. However, most of the existing content caching strategies rely on static rules or simple content access statistics, making it difficult to accurately capture dynamic context information such as user behavior, node status, and content characteristics, resulting in problems such as low cache hit rate, low resource utilization, and untimely content replacement. At the same time, traditional strategies lack flexibility and global optimality in the face of complex multi - objective optimization requirements and are difficult to achieve the optimal distribution efficiency with limited cache resources. How to achieve more intelligent and efficient content caching and replacement at edge nodes has become an urgent problem to be solved.

[0029] Based on the above problems, the present invention proposes a context - aware content caching and replacement method and system, which fully excavates multi - source context information such as user behavior, node status, and content characteristics, uses the XGBoost model to accurately predict the future content access popularity, combines multi - objective optimization algorithms to generate efficient caching strategies, and dynamically adjusts the cached content by designing a cache scoring function and a greedy replacement mechanism to improve the cache hit rate and optimize the system performance. This method effectively addresses the deficiencies in cache accuracy, dynamics, and resource scheduling in existing solutions and has strong practicality and promotion value.

[0030] The specific implementation includes the following content:

[0031] Embodiment 1: Refer to Figure 1 , a context - aware content caching and replacement method provided in this embodiment, the method includes: Step S1: Obtain context information at different times and construct a context feature vector; Step S2: Based on the context feature vector, construct training samples to train an XGBoost regression model, and obtain a trained XGBoost prediction model; Step S3: Construct a cache decision candidate set, collect the context feature vector, and use the trained XGBoost prediction model for prediction to obtain the future content access popularity; Step S4: Design a multi-objective optimization function and constraint conditions based on the cache policy requirements; Step S5: Based on the multi-objective optimization function and constraint conditions, use the NSGA-II multi-objective evolutionary algorithm to solve and obtain a Pareto solution set; Step S6: Define a comprehensive scoring function, select the individual with the maximum comprehensive scoring function value from the Pareto solution set as the final caching strategy, use the corresponding content set as the final cached content set, and deploy the cache to the edge node; Step S7: Design a cache scoring function for a single content, obtain a cache score based on the cache scoring function, and use a greedy strategy to replace the cached content based on the cache score.

[0032] In step S1, the obtaining of context information at different times and constructing of the context feature vector include: performing a preprocessing operation on the context information to obtain preprocessed context information; constructing the context feature vector according to the preprocessed context information.

[0033] Specifically, the preprocessing operation includes data alignment and denoising: unifying the time base of all context data (aligning according to the sliding window sampling points), outlier processing (including negative node load and position jump), and missing value filling (interpolation, mean filling or deletion); feature normalization: using Min-Max normalization for real-valued features to ensure that all feature values are within a reasonable range, which is beneficial to model training.

[0034] Adopting context information can significantly improve the intelligent level of the content caching strategy. By introducing multi-dimensional information such as user behavior, node status, and content features, it helps the system more accurately predict the future access popularity of content, and realizes the joint modeling of people, content, and environment. This not only enhances the adaptability of the model to complex network environments and user dynamic behaviors, but also makes the cache decision more refined and personalized, thereby improving the cache hit rate while effectively reducing transmission latency and bandwidth occupancy, and improving the utilization efficiency of edge resources.

[0035] In the above steps, the content refers to the content requested by the user for access.

[0036] The context information includes: user behavior information, node status information, and requested content information.

[0037] The user behavior information includes: user access frequency, user location, user speed, and user activity.

[0038] The node status information includes: node bandwidth, node transmission delay, and node load.

[0039] The request content information includes: the total number of times the content is accessed within a sliding time window, and the content size.

[0040] The preprocessing operations include: outlier removal processing, missing value filling processing, and normalization processing.

[0041] The context feature vector has the following expression: ; Where refers to the context feature vector, refers to the content, refers to the user, refers to the edge node, refers to the current moment, refers to the user 's access frequency to the content within the sliding time window, refers to the user's location at the current moment, refers to the user's speed at the current moment, refers to the user's activity at the current moment, that is, the total number of user requests, refers to the node bandwidth at the current moment, refers to the node transmission delay at the current moment, refers to the node load at the current moment, refers to the total number of times the content is accessed within the sliding time window, refers to the content size, that is, the size of the content .

[0042] In specific implementation, consider the following scenario: The user requests to watch short video content through a mobile device in different regions, and the system deploys edge nodes in multiple geographical regions to be responsible for responding to the content requests of surrounding users.

[0043] In step S2, constructing a training sample based on the context feature vector, training an XGBoost regression model, and obtaining a trained XGBoost prediction model includes: collecting the actual access popularity within a time span after the corresponding time point of the context feature vector as label data; forming a training sample set with the context feature vectors at all historical time points and their corresponding label data; setting the input of the XGBoost regression model as the context feature vector and the output as the future content access popularity; defining a loss function, and through iterative training, continuously optimizing the parameters of the XGBoost regression model using gradient boosting trees to minimize the loss function, and finally obtaining the trained XGBoost prediction model.

[0044] Using the XGBoost prediction model for prediction can achieve accurate estimation of the content access popularity within a specific future time period. By inputting the context feature vector, the XGBoost model can comprehensively consider user behavior characteristics, edge node states, and content characteristics, and learn the non-linear relationship of the content access pattern from historical data. The prediction result is the access popularity value of each candidate content within a certain future time span, providing key support for subsequent caching decisions. Compared with traditional static or rule-driven prediction methods, XGBoost has stronger fitting ability and generalization performance, can improve the cache hit rate, reduce unnecessary content replacement and data transmission, and thus effectively optimize the overall performance of the edge caching system.

[0045] In the above steps, the label data refers to the actual access popularity of the content at the moment under the condition that all historical time points are known, and its expression is: .

[0046] The time span is expressed as: .

[0047] The training sample set is expressed as: ; where refers to the training sample set, refers to the context feature vector, refers to the label data, refers to the time span, refers to the number of samples.

[0048] The future content access popularity is expressed as: .

[0049] The parameters are expressed as: .

[0050] The loss function has the following expression: ; where denotes the loss function, denotes the model parameters, denotes the number of samples, denotes the index variable, denotes the actual access popularity of the content, denotes the predicted access popularity of the content by the model, denotes the regularization term of the model complexity, representing the model complexity penalty term, including the number and depth of regression trees, used to prevent overfitting, is the regularization parameter.

[0051] In step S3, to construct the cache decision candidate set, collect the context feature vectors, and use the trained XGBoost prediction model for prediction to obtain the future content access popularity, the following steps are included: At the edge node, count and generate the set of all content requested by users, denoted as the candidate content set; for each piece of content in the candidate content set, collect its context feature vectors and use the trained XGBoost prediction model for prediction to obtain the predicted future content access popularity.

[0052] In this step, the candidate content set has the following expression: ; where denotes the candidate content set, denotes each corresponding piece of content.

[0053] In step S4, the cache policy requirements refer to maximizing the cache hit rate, minimizing the transmission delay, and minimizing the bandwidth occupancy.

[0054] The multi-objective optimization function includes the objective function for maximizing the cache hit rate, the objective function for minimizing the transmission delay, and the objective function for minimizing the bandwidth occupancy.

[0055] The objective function for maximizing the cache hit rate has the following expression: ; where denotes the objective function for maximizing the cache hit rate, denotes the set of actually selected cached content, denotes the candidate content set, denotes the future content access popularity, denotes the time span.

[0056] Minimize the transmission delay objective function, and its expression is: ; where, refers to the minimize transmission delay objective function, refers to the actually selected cache content set, refers to the current time of the node transmission delay.

[0057] Minimize the bandwidth occupancy objective function, and its expression is: ; where, refers to the minimize bandwidth occupancy objective function, refers to the actually selected cache content set, refers to the content size of, refers to the future content access popularity, refers to the time span, refers to the current time of the node bandwidth.

[0058] The constraint condition refers to the cache capacity constraint condition, and its expression is: ; where, refers to the actually selected cache content set, refers to the content size of, refers to the total cache capacity of the edge node.

[0059] Specifically, the design of the multi-objective optimization function and the cache capacity constraint condition aims to comprehensively balance the key indicators of cache performance under the condition of limited edge node resources. By maximizing the cache hit rate objective function, preferentially select the content with high predicted access popularity for caching, increase the probability of user requests hitting the local cache, and reduce the frequency of backhaul requests; by minimizing the transmission delay objective function, encourage the cache policy to tend to select the content with lower node delay at the current time, improve the data response efficiency, and improve the user experience; while the minimize bandwidth occupancy objective function guides the system to preferentially cache the content with lower unit bandwidth consumption and higher cost performance, and improve the network resource utilization rate. The above three are parallel optimization objectives. Under the overall limited cache capacity constraint, they are solved by the NSGA-II multi-objective evolutionary algorithm to obtain a set of Pareto optimal cache policy sets, enabling the system to weigh performance indicators from multiple perspectives and ultimately achieve refined cache content decision-making and optimal resource allocation.

[0060] In step S5, based on the multi-objective optimization function and the constraint conditions, the NSGA-II multi-objective evolutionary algorithm is used to solve and obtain the Pareto solution set, including: encoding each individual, representing the content caching selection scheme with a binary vector, and judging whether to cache the content to obtain the caching selection vector; randomly generating a number of legal individuals to form an initial population, setting the population size, and initializing the iteration counter; calculating the three objective functions of each individual to obtain the objective function values; performing an evolutionary operation on the current population to generate the next generation population; merging the current population and the next generation population into a combined population, performing non-dominated sorting and crowding degree evaluation again, selecting the individuals with the number of the population size to form a new population, updating the iteration counter, until the set maximum iteration number is reached, stopping the evolutionary process, and finally obtaining the Pareto solution set.

[0061] Specifically, the NSGA-II multi-objective evolutionary algorithm simulates the natural selection and evolution mechanism, and gradually optimizes the content caching selection strategy based on the multi-objective optimization function and the cache capacity constraint. In the initial stage, the algorithm randomly generates a number of individuals that meet the cache capacity constraint. Each individual is encoded with a binary vector, and each bit in the vector corresponds to the candidate content. 1 indicates that the content is cached, and 0 indicates that it is not cached. Subsequently, the three objective function values of cache hit rate, transmission delay, and bandwidth occupancy are calculated for each individual in the population. The individuals are ranked according to the non-dominated sorting principle and divided into different non-dominated levels. The higher the level, the better the solution. The crowding degree distance is calculated to maintain the diversity of the population. During the iteration process, genetic operations such as selection, crossover, and mutation are used to generate the next generation population. After merging with the current population, new sorting and screening are performed, continuously promoting the population to converge to a better Pareto front. Through the set maximum iteration number limit, a set of Pareto optimal caching strategy sets covering different trade-off relationships are finally output, providing multiple feasible solutions for subsequent selection of specific caching schemes.

[0062] Furthermore, the evolutionary operation of the population includes: non-dominated sorting, crowding degree calculation, selection operation, crossover operation, and mutation operation.

[0063] Furthermore, the non-dominated sorting refers to sorting according to the objective function values, divided into different non-dominated levels. The higher the level, the better the solution.

[0064] The crowding degree calculation refers to calculating the density of the objective space within the same level to maintain the diversity of the solutions. The individuals with a larger crowding degree are more preferentially retained.

[0065] The selection operation refers to adopting the tournament selection method and selecting the parents by combining the non-dominated level and the crowding degree.

[0066] The crossover operation refers to generating new individuals using single-point crossover or uniform crossover methods, ensuring that the new individuals meet the cache space constraint; otherwise, prune (delete the content with the lowest popularity).

[0067] The mutation operation refers to flipping the bit (0 ↔ 1) at a certain position with a probability to generate mutant individuals, and the capacity limit also needs to be ensured.

[0068] In the above steps, the individual refers to a complete content caching selection scheme, represented in the form of a binary vector. Each bit in the vector corresponds to the caching decision of the content in a candidate content set, thus forming the caching selection vector.

[0069] The caching selection vector has the following expression: ; where refers to the caching selection vector, refers to the binary vector, indicating whether to cache the content in the -th candidate content set on the edge node. The value of 1 means caching, and the value of 0 means not caching. refers to the total number of candidate content sets.

[0070] The several legal individuals need to meet the cache capacity constraint condition.

[0071] The cache capacity constraint condition has the following expression: ; where refers to the actually selected cache content set, refers to the content size, refers to the total cache capacity of the edge node.

[0072] The initial population is expressed as: .

[0073] The population size is expressed as: .

[0074] The initialized iteration counter is expressed as: .

[0075] The current population is expressed as: .

[0076] The three objective functions refer to: the objective function of maximizing the cache hit rate, the objective function of minimizing the transmission delay, and the objective function of minimizing the bandwidth occupancy.

[0077] The evolutionary operations include: non-dominated sorting, crowding degree calculation, selection operation, crossover operation, and mutation operation.

[0078] The new population is expressed as: .

[0079] The maximum number of iterations is expressed as: .

[0080] The Pareto solution set is expressed as: .

[0081] In step S6, defining the comprehensive scoring function, selecting the individual with the maximum comprehensive scoring function value from the Pareto solution set as the final caching strategy, taking the corresponding content set as the final cached content set, and deploying the cache to the edge node includes: defining the comprehensive scoring function for each individual in the Pareto solution set; selecting the individual that makes the comprehensive scoring function obtain the maximum value from all solutions in the Pareto solution set, denoted as the optimal solution, and obtaining the optimal cache selection vector; determining whether to cache the corresponding content based on the value of each bit in the binary of the optimal cache selection vector to obtain the final cached content set, and caching the content in the final cached content set to the edge node.

[0082] Specifically, the comprehensive scoring function is used to perform quantitative evaluation and unified comparison among multi-objective solutions with their respective advantages and disadvantages in the Pareto solution set to assist in selecting the final caching strategy that best meets the overall optimization requirements of the system. This scoring function can assign weights to different objectives according to the actual scenario to dynamically adjust to meet different needs.

[0083] In the above steps, each individual refers to each cache selection vector.

[0084] The expression of the comprehensive scoring function is: ; where refers to the comprehensive scoring function, refers to an individual caching strategy in the Pareto solution set, , refers to the caching strategy 's predicted hit rate value, refers to the caching strategy 's average transmission delay at the current edge node, refers to the caching strategy 's total bandwidth consumption at the current edge node, , and refer to the weight parameters, and at the same time satisfy .

[0085] The optimal solution is expressed as: .

[0086] The optimal cache selection vector is expressed as: .

[0087] The final cache content set is expressed as: .

[0088] In step S7, designing the cache scoring function for a single content, obtaining the cache score based on the cache scoring function, and replacing the cache content using the greedy strategy based on the cache score, including: designing the cache scoring function, and calculating the cache scores of the contents in all the candidate content sets and the cached contents at the current edge node according to the future content access popularity, the content size, and the node load; sorting the contents in all the candidate content sets and the cached contents from high to low according to the cache score, and replacing the cache content using the greedy strategy.

[0089] Specifically, the cache scoring function is defined as an evaluation index for comprehensively measuring the cache value of content in the current network environment. Its calculation method takes into account the predicted future content access popularity, the content size, and the load level of the current edge node to achieve the efficient use of cache resources. This method has dynamic adaptability and high efficiency, and can quickly replace the cache content under the changing content popularity and network status, improve the overall cache hit rate, and reduce the access latency and bandwidth occupancy.

[0090] In the above steps, the cache scoring function has the following expression: ; where denotes the cache scoring function, denotes the future content access popularity, denotes the time span, denotes the content size, denotes the node load at the current moment , which is used as the load penalty term in the cache scoring function to penalize the cache selection of high-load nodes and reflect the current cache pressure.

[0091] Furthermore, the greedy strategy includes: sorting the contents in all the candidate content sets from high to low according to the cache score; traversing the sorted list to perform cache replacement judgment.

[0092] Further, the cache replacement determination is specifically as follows: If the content in the candidate content set is not cached and there is enough cache space, it is directly added to the cache until the cache space is exhausted; if the cache space is insufficient, the cache score of the content in the candidate content set is compared with the content with the lowest cache score among the cached content; if the cache score of the content in the candidate content set is higher, the content with the lowest cache score among the cached content is replaced, otherwise the content in the candidate content set is skipped.

[0093] Embodiment 2: Reference Figure 2 , based on the context-aware content caching and replacement method in the above Embodiment 1, the present invention further provides a context-aware content caching and replacement system, which includes: A data processing module, configured to obtain context information, perform preprocessing operations on the context information to obtain preprocessed data, and construct a context feature vector; A popularity prediction module, configured to train an XGBoost prediction model based on the context feature vector obtained by the data processing module, and use the XGBoost prediction model to predict the future content access popularity; A content caching module, configured to design a multi-objective optimization function based on cache policy requirements and solve it using the NSGA-II algorithm to obtain a Pareto solution set, define a comprehensive scoring function, select the individual with the maximum comprehensive scoring function value from the Pareto solution set as the final caching policy, use the corresponding content set as the final cached content set, and deploy the cache to the edge node; A cache replacement module, configured to design a cache scoring function for a single content based on the future content access popularity predicted by the popularity prediction module, obtain a cache score based on the cache scoring function, and replace the cached content of the content caching module using a greedy strategy based on the cache score.

[0094] The specific implementation method of this embodiment is the same as that of Embodiment 1 and will not be elaborated here. For details, refer to the description of Embodiment 1.

[0095] Adopting the technical solution of the above embodiment, in a context-aware content caching and replacement method, by introducing a context-aware mechanism, making full use of multi-dimensional dynamic information such as user behavior, node status, and requested content characteristics, using the XGBoost regression model to accurately predict the future content access popularity, and combining with the multi-objective optimization algorithm NSGA-II to simultaneously optimize the cache hit rate, transmission delay, and bandwidth occupancy, realizing intelligent caching decision-making for multi-scenarios; at the same time, designing a comprehensive scoring function to select the optimal caching strategy from the Pareto solution set, further improving the decision-making quality, and introducing a scoring function based on access popularity, content size, and node load in the cache replacement stage, using a greedy strategy to dynamically replace the cached content, thereby effectively improving the cache space utilization efficiency and the overall system performance, significantly superior to traditional static or single-objective caching methods, with stronger environmental adaptability, prediction accuracy, and real-time response ability. This method can not only improve the cache hit rate and resource utilization efficiency, but also further enhance the service ability and response efficiency of edge nodes in complex environments, thus realizing a more intelligent and efficient content caching and replacement strategy in dynamic network scenarios.

[0096] Those skilled in the art can understand that although some embodiments herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means that it is within the scope of the present invention and forms different embodiments.

[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A context-aware content caching and replacement method, characterized in that The method includes: Obtaining context information at different times and constructing a context feature vector; Based on the context feature vector, constructing training samples, training an XGBoost regression model, and obtaining a trained XGBoost prediction model; Constructing a cache decision candidate set, collecting the context feature vector, and using the trained XGBoost prediction model for prediction to obtain the future content access popularity; Based on the cache policy requirements, designing a multi-objective optimization function and constraint conditions; Based on the multi-objective optimization function and constraint conditions, using the NSGA-II multi-objective evolutionary algorithm to solve and obtain a Pareto solution set; Defining a comprehensive scoring function, selecting the individual with the maximum comprehensive scoring function value from the Pareto solution set as the final caching policy, taking the corresponding content set as the final cached content set, and deploying the cache to the edge node; Designing a cache scoring function for a single content, obtaining a cache score based on the cache scoring function, and replacing the cached content using a greedy strategy based on the cache score; The content refers to the content requested by the user for access.

2. The method for content caching and replacement based on context awareness according to claim 1, wherein: The obtaining context information at different times and constructing a context feature vector includes: Performing preprocessing operations on the context information to obtain preprocessed context information; Constructing the context feature vector according to the preprocessed context information.

3. The method for content caching and replacement based on context awareness according to claim 2, wherein: The context information includes: user behavior information, node status information, and requested content information; The user behavior information includes: user access frequency, user location, user speed, and user activity; The node status information includes: node bandwidth, node transmission delay, and node load; The requested content information includes: the total number of times the content is accessed within a sliding time window and the content size; The preprocessing operations include: outlier removal processing, missing value filling processing, and normalization processing; The expression of the context feature vector is: ; Among them, refers to the context feature vector, refers to the content, refers to the user, refers to the edge node, refers to the current moment, refers to the user 's access frequency to the content within the sliding time window, refers to the current moment of the user's location, refers to the current moment of the user's speed, refers to the current moment of the user's activity, that is, the total number of user requests, refers to the current moment of the node bandwidth, refers to the current moment of the node transmission delay, refers to the current moment of the node load, refers to the content 's total number of accesses within the sliding time window, refers to the content size, that is, the size of the content itself.

4. The method for content caching and replacement based on context awareness according to claim 3, wherein: The constructing training samples based on the context feature vector, training an XGBoost regression model, and obtaining a trained XGBoost prediction model includes: Based on the context feature vector, collecting the actual access popularity within the time span after its corresponding time point as label data; Forming a training sample set with the context feature vectors at all historical time points and their corresponding label data; Setting the input of the XGBoost regression model as the context feature vector and the output as the future content access popularity; Defining a loss function, and through iterative training, continuously optimizing the parameters of the XGBoost regression model using gradient boosting trees to minimize the loss function, and finally obtaining the trained XGBoost prediction model.

5. A context-aware content caching and replacement method according to claim 4, characterized in that: The label data refers to the content under the condition known at all the historical time points at the actual access popularity at the moment, and its expression is: ; The time span is expressed as: ; The training sample set, its expression is: ; Among them, refers to the training sample set, refers to the context feature vector, refers to the label data, refers to the time span, refers to the number of samples; The future content access popularity is expressed as: ; The parameter is expressed as: ; The loss function, its expression is: ; Among them, refers to the loss function, refers to the model parameters, refers to the number of samples, refers to the index variable, refers to the access popularity of the real content, refers to the access popularity of the content predicted by the model, refers to the regularization term of the model complexity, representing the model complexity penalty term, including the number and depth of regression trees, used to prevent overfitting, refers to the regularization parameter.

6. A context-aware content caching and replacement method according to claim 5, characterized in that: To construct a cache decision candidate set, collect the context feature vectors, and use the trained XGBoost prediction model for prediction to obtain the future content access popularity, including: At the edge node, count and generate the set of all user-requested content, denoted as the candidate content set; For each content in the candidate content set, collect its context feature vector, and use the trained XGBoost prediction model for prediction to obtain the predicted future content access popularity; The candidate content set, its expression is: ; Among them, refers to the candidate content set, refers to each corresponding content.

7. A context-aware content caching and replacement method according to claim 6, characterized in that: The cache policy requirements refer to maximizing the cache hit rate, minimizing the transmission delay, and minimizing the bandwidth occupancy; The multi-objective optimization function includes a cache hit rate maximization objective function, a transmission delay minimization objective function, and a bandwidth occupancy minimization objective function; The cache hit rate maximization objective function, its expression is: ; Among them, refers to the objective function of maximizing the cache hit rate, refers to the set of actually selected cached content, refers to the candidate content set, refers to the future content access popularity, refers to the time span; The transmission delay minimization objective function, its expression is: ; Among them, refers to the objective function of minimizing the transmission delay, refers to the actually selected cache content set, refers to the current time of the node transmission delay; The bandwidth occupancy minimization objective function, its expression is: ; Among them, refers to the objective function of minimizing bandwidth occupancy, refers to the actually selected cache content set, refers to the content size, refers to the future content access popularity, refers to the time span, refers to the current moment node bandwidth; The constraint condition refers to the cache capacity constraint condition, and its expression is: ; Among them, refers to the actually selected cache content set, refers to the content size of, refers to the total cache capacity of the edge nodes.

8. A context-aware content caching and replacement method according to claim 7, characterized in that: Based on the multi-objective optimization function and the constraint condition, use the NSGA-II multi-objective evolutionary algorithm to solve and obtain the Pareto solution set, including: Encode each individual, use a binary vector to represent the content caching selection scheme, and determine whether to cache the content to obtain the caching selection vector; Randomly generate several legal individuals to form an initial population, set the population size, and initialize the iteration counter; Calculate the three objective functions for each individual to obtain the objective function values; Perform an evolutionary operation on the current population to generate the next generation population; Merge the current population and the next generation population into a combined population, re-perform non-dominated sorting and crowding degree evaluation, select the top population size number of individuals to form a new population, update the iteration counter, and stop the evolutionary process until the set maximum iteration number is reached, and finally obtain the Pareto solution set.

9. A context-aware content caching and replacement method according to claim 8, characterized in that: The individual refers to a complete content caching selection scheme, which is represented in the form of a binary vector, and each bit in the vector corresponds to the caching decision of a content in the candidate content set, thus forming the caching selection vector; The caching selection vector, its expression is: ; Among them, refers to the cache selection vector, is a binary vector indicating whether to select the content in the th candidate content set to be cached on the edge node. A value of 1 means caching, and a value of 0 means not caching, refers to the total number of candidate content sets; The several legal individuals need to satisfy the cache capacity constraint condition; The cache capacity constraint condition, its expression is: ; Among them, refers to the actually selected cached content set, refers to the content size of, refers to the total cache capacity of the edge nodes; The initial population is expressed as: ; The population size, expressed as: ; The initialized iteration counter is expressed as: ; The current population, denoted as: ; The three objective functions refer to: the objective function of maximizing the cache hit rate, the objective function of minimizing the transmission delay, and the objective function of minimizing the bandwidth occupancy; The evolutionary operations include: non-dominated sorting, crowding degree calculation, selection operation, crossover operation, and mutation operation; The new population, expressed as: ; The maximum number of iterations, expressed as: ; The Pareto solution set, expressed as: .

10. A context-aware content caching and replacement method according to claim 9, wherein: Defining a comprehensive scoring function, selecting the individual with the maximum comprehensive scoring function value from the Pareto solution set as the final caching strategy, taking the corresponding content set as the final cached content set, and deploying the cache to the edge node, including: Defining a comprehensive scoring function for each individual in the Pareto solution set; Among all the solutions in the Pareto solution set, selecting the individual that makes the comprehensive scoring function reach the maximum value, denoted as the optimal solution, and obtaining the optimal cache selection vector; Based on the value of each bit in the binary of the optimal cache selection vector, determine whether to cache the corresponding content, obtain the final cached content set, and cache the content in the final cached content set to the edge node.

11. A context-aware content caching and replacement method according to claim 10, wherein: Each individual refers to each cache selection vector; The expression of the comprehensive scoring function is: ; Among them, refers to the comprehensive scoring function, refers to a caching strategy individual in the Pareto solution set, , refers to the caching strategy 's predicted hit rate value, refers to the caching strategy 's average transmission delay at the current edge node, refers to the caching strategy 's total bandwidth consumption at the current edge node, , and refer to the weight parameters, and at the same time satisfy ; The optimal solution is expressed as: ; The optimal cache selection vector is expressed as: ; The final cached content set, expressed as: .

12. A context-aware content caching and replacement method according to claim 11, wherein: Designing a cache scoring function for a single content, obtaining a cache score based on the cache scoring function, and replacing the cached content using a greedy strategy based on the cache score, including; Designing the cache scoring function, and calculating the cache scores of the content in all candidate content sets and the cached content at the current edge node according to the future content access popularity, the content size, and the node load; Sorting the content in all candidate content sets and the cached content according to the cache score from high to low, and replacing the cached content using a greedy strategy.

13. A context-aware content caching and replacement method according to claim 12, wherein: The expression of the cache scoring function is: ; Among them, refers to the cache scoring function, refers to the future content access popularity, refers to the time span, refers to the content size, refers to the node load at the current moment which, as the load penalty term in the cache scoring function, is used to penalize the cache selection of high-load nodes and reflects the current cache pressure; The greedy strategy includes: Sorting the content in all candidate content sets according to the cache score from high to low; Traversing the sorted list to perform cache replacement judgment; The cache replacement judgment is specifically: If the content in the candidate content set is not cached and the cache space is sufficient, directly add it to the cache until the cache space is exhausted; If the cache space is insufficient, compare the cache score of the content in the candidate content set with the content with the lowest cache score among the cached content; If the cache score of the content in the candidate content set is higher, replace the content with the lowest cache score among the cached content, otherwise skip the content in the candidate content set.

14. A context-aware content caching and replacement system, characterized in that, Implementing a context-aware content caching and replacement method as described in any one of claims 1-13, including: A data processing module, which is used to obtain context information, perform preprocessing operations on the context information to obtain data after preprocessing operations, and construct a context feature vector; A popularity prediction module, which is used to train an XGBoost prediction model based on the context feature vector obtained by the data processing module, and use the XGBoost prediction model to predict the future content access popularity; A content caching module, which is used to design a multi-objective optimization function based on cache policy requirements and solve it using the NSGA-II algorithm to obtain a Pareto solution set, define a comprehensive scoring function, select the individual with the largest comprehensive scoring function value from the Pareto solution set as the final caching policy, use the corresponding content set as the final cached content set, and deploy the cache to the edge node; A cache replacement module, which is used to design a cache scoring function for a single content based on the future content access popularity predicted by the popularity prediction module, obtain a cache score based on the cache scoring function, and replace the cached content of the content caching module using a greedy strategy based on the cache score.

Citation Information

Patent Citations

  • Optimal design method and system of wind power plant, electronic equipment and storage medium

    CN111339713A

  • VOD service cache optimization method based on XGBoost algorithm in edge network environment

    CN111629217A

  • Content updating method based on deep reinforcement learning

    CN113064907A

  • Vehicle user-oriented edge network content caching and pre-caching method

    CN114513514A

  • Dynamic caching method and device based on deep learning, equipment and medium

    CN116684491A

Cited By

  • Containerized GPU training-oriented cache acceleration method, apparatus and device, and medium

    CN121092310A

  • File caching performance improving method and system based on distributed storage

    CN121722325A

  • A method and system for improving file caching performance based on distributed storage

    CN121722325B