Video Cooperative Caching Method for Multiple Unit Trains Based on Multi-Agent Reinforcement Learning
By adopting the video collaborative caching method of multi-agent reinforcement learning in the EMU, the problems of network instability and dynamic resource changes in the EMU video service are solved, and the accurate and rapid video services and efficient utilization of cache space are achieved.
Patent Information
- Application Number
- CN202510187523.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-20
AI Technical Summary
In the EMU scenario, video services face challenges such as instability in the network communication environment, large passenger mobility and dynamic resource variation, which makes it difficult for existing cache algorithms to effectively learn resource request rules, and the cache strategy effect is average.
The video collaborative cache method of EMU based on multi-agent reinforcement learning is adopted. By treating each car as an agent, using a neural network to fit the actions of the agent, and setting up an evaluation network for collaborative optimization, the video content cache strategy for each period is determined.
It realizes the accuracy and speed of cabin video services, balances server load, improves cache space utilization, and adapts to the differences and dynamics in EMU scenarios.
Smart Images

Figure CN119676476B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video data caching, and specifically relates to a collaborative video caching method for multiple intelligent agents based on reinforcement learning for EMUs. Background Art
[0002] With the rapid development of the high-speed railway network, EMUs have become an important means of transportation for people to travel. In recent years, videos have become rapidly popular due to their direct expression and easy consumption characteristics, and have become the main component of Internet traffic. People are more inclined to obtain information through videos in their daily lives. Therefore, video services in EMU carriages have become increasingly important. However, the sharp increase in traffic brought by videos poses higher requirements for network bandwidth and latency. Especially in a mobile environment such as an EMU, the network connection may be more unstable, and the uncertainty of the quality of the user experience has a greater impact. Video services in the EMU scenario face a series of challenges such as unstable network communication environments, large passenger mobility, and high resource dynamic variability.
[0003] By caching video content at the edge of the carriage, the viewing experience of passengers can be significantly improved, and problems such as buffering and freezing caused by network latency and bandwidth limitations can be reduced. By deploying cache servers at the edge of the carriage, video content can be pre-loaded closer to passengers, thereby reducing the distance and time of data transmission and accelerating the video loading speed.
[0004] However, previous studies often considered optimizing the EMU network communication environment in terms of channels or base stations to accelerate the rate at which users obtain video content, and rarely considered how to provide more stable and reliable video services for users from the perspective of edge caching in EMUs. The main reasons for this phenomenon are as follows:
[0005] First, there are significant differences in the user attributes within EMU carriages, which are often manifested in the seat class, identity, occupation of users, and cultural characteristics that vary with regions, resulting in large differences in content preferences among users in different carriages. Secondly, the mobility of users in EMU carriages is greater than that of users around traditional servers, resulting in a large degree of resource dynamic variability, and there are differences in the degree of dynamic variability between different carriages. Algorithms such as LRU and LFU, which essentially try to learn the request patterns of resources through training of historical data, often find it difficult to apply the learned historical patterns to the future when facing scenarios with large differences and dynamics such as EMUs, and the effects of prediction-based caching algorithms are generally average.
[0006] Second, the current EMU carriages do not have much design for the edge caching architecture, and the single-carriage service architecture design often has problems such as large load pressure and inability to dynamically adjust.
[0007] For the above reasons, although caching video content at the edge of the carriage can significantly improve the viewing experience of passengers, there are still many difficulties in implementing this technical means. Summary of the Invention
[0008] The present invention aims to solve at least one of the technical problems existing in the related art to a certain extent.
[0009] An object of the present invention is to provide a video collaborative caching method for multiple units based on multi-agent reinforcement learning, which considers the differences between different carriages and the problem of linear networking, and performs multi-carriage collaborative caching of videos, so that the carriage video service is more accurate and fast, while balancing the load of each server and fully improving the utilization rate of the caching space of the multiple units.
[0010] In order to achieve the above object, on the one hand, the present invention provides a video collaborative caching method for multiple units based on multi-agent reinforcement learning, including the following steps:
[0011] S100. Construct a video collaborative caching service system for multiple carriages of the multiple units, including defining a set of carriages of the multiple units and a set of video content, and dividing the running time of the multiple units into time periods according to the driving time between stations;
[0012] S200. Define the request response time cost according to the cache hit situation of the user request.
[0013] S300. Define the resource replacement time cost according to the difference in cached video content between adjacent time periods.
[0014] S400. Formalize the problem of minimizing the video service time cost of the multiple units.
[0015] S500. Call the multi-carriage collaborative caching algorithm based on multi-agent reinforcement learning to determine the video content caching strategy for each time period, where each carriage is regarded as an agent, the actions of the agent are fitted by a neural network, and two evaluation networks are set for collaborative optimization; the data of all agents in the time period are stored in the sample pool in the form of a five-tuple, including the global state, global action, global state of the next time period, reward set and end flag; if the number of samples reaches the minimum batch processing threshold, the neural network is trained to update the target network parameters.
[0016] A further preferred technical solution of the present invention is that the construction of the video collaborative caching service system for multiple carriages of the multiple units in step S100 includes:
[0017] S110. Define a set of carriages of the multiple units and a set of video content, and number the carriages in sequence from the front of the train to the rear of the train , let represent the set of carriages of the multiple units, Represents a set of video content;
[0018] S120. Divide the running time of the multiple unit train according to stations, and each time period represents the running time between the current station and the next station;
[0019] S130. Constructing the multiple unit train video service includes:
[0020] S131. Set up a car buffer server on each car as a reverse proxy, and define the upper limit of the buffer capacity of each car buffer server as , ; The car buffer servers are connected by wire or wireless to form a multiple unit train local area network. Each car buffer server communicates with an external base station through the roof antenna of the car where it is located; Users connect to the multiple unit train local area network through the wireless access points set on each car; The multiple unit train is provided with a central server for querying the deployment situation of the multiple unit train video content;
[0021] S132. When the users in the car access the network of this car and request video content, check the cache hit situation of the user request, and select to send the user-requested video content to the user from this car, other cars or the external base station;
[0022] S133. Let the number of requests for the video content by all users in the car during the time period be , and define a binary indicator variable , indicating whether the requests for the video content by all users on the car during the time period are responded to by the target . If it is responded to by the target , then , otherwise, , where , is the multiple unit train car buffer server or the external base station.
[0023] As an optimization, when the users in the car access the network of this car and request video content as described in step S132, check the cache hit situation of the user request, and select to send the user-requested video content to the user from this car, other cars or the external base station; The specific method is:
[0024] S132-a. Judge whether the video content requested by the user is cached in the car where the user is located. If so, execute step S132-e; otherwise, execute step S132-b;
[0025] S132-b. Obtain the deployment situation of the user-requested video content on the multiple unit train from the central server. If there is a cache of the user-requested video content on the multiple unit train, execute step S132-c; otherwise, execute step S132-d.
[0026] S132-c. Select the carriage that is closest to the carriage where the request is sent and has cached the user-requested video content to join the candidate response carriage set. Check whether there are two candidate carriages with the same distance. If not, directly select the carriage that is closest to the carriage where the request is sent and has cached the user-requested video content as the response carriage. If so, select the carriage with a smaller load among the two candidate carriages as the response carriage. The carriage where the user is located acts as an agent to request the video content from the response carriage, and after receiving the response, execute step S132-e.
[0027] S132-d. The carriage where the user is located acts as an agent to request the required video content from the external base station, and after receiving the response, execute step S132-e.
[0028] S132-e. The carriage where the user is located sends the user-requested video content to the users in this carriage.
[0029] Preferably, according to the cache hit situation of the user request, define the request response time cost in step S200; specifically including:
[0030] According to the cache hit situation of the user request, there are three different responses in total, namely in-carriage response, inter-carriage response, and external base station response, corresponding to three different request response time costs;
[0031] Define a binary indicator variable , indicating whether the request between carriage and carriage passes through the link section connected by carriage and carriage . If it passes through, then , otherwise , that is:
[0032] (1);
[0033] Let the total link bandwidth be , and define the request response time cost from carriage to the target as:
[0034] (2);
[0035] Among them, is the response time cost for the carriage to request video content from the external base station, which is set as a fixed value; Indicates the time period Inside the carriage The number of requests for video content by all users ; Indicates that in Time period carriage Whether the requests for video content by all users are responded to by the carriage If it is responded to by the carriage then , otherwise ; Indicates the carriage and the carriage Whether the requests between are passed through the carriage and the target The link interval connected when it is the carriage;
[0036] In the time period Meet the carriage The request response time cost of all users is:
[0037] (3).
[0038] Preferably, according to the difference in cached video content between adjacent time periods in step S300, define the resource replacement time cost; specifically include:
[0039] S310. Define a binary indicator variable , indicating Whether the video content in the time period is cached on the carriage . If the video content in the time period is cached on the carriage then , otherwise ;
[0040] S320. At the beginning of each time period , if the video content to be cached in the current time period is different from the video content cached in the previous time period , then each carriage cache server obtains the video content required for the current time period from other carriage cache servers or external base stations to replace the video content cached in the previous time period . Define the resource replacement time cost of the carriage at the time period as:
[0041] (4).
[0042] Preferably, the formalized problem of minimizing the video service time cost of the multiple unit train in step S400 is specifically as follows:
[0043] Let the video service time cost of the multiple unit train be the weighted sum of the request response time cost and the resource replacement time cost of all carriages in all time periods. The formalized problem of minimizing the video service time cost of the multiple unit train is expressed as:
[0044] (5);
[0045] Constraint conditions:
[0046] (6);
[0047] (7);
[0048] (8);
[0049] (9);
[0050] (10);
[0051] Among them, and are weight factors used to adjust the weights between the request response time cost and the resource replacement time cost. represents the upper limit of the cache capacity of the cache server of carriage ; represents whether all users' requests for video content in carriage during time period are responded to by carriage . If it is responded to by carriage , then , otherwise, .
[0052] Preferably, step S500 includes:
[0053] S510. Regard each carriage as an agent, and use the carriage number to represent the agent corresponding to the carriage; use a neural network to fit the action of agent , denoted as the training action network , where is the training action network parameter of agent , and the corresponding target action network is denoted as ; for each agent , set two training evaluation networks, denoted as and , respectively, where For the first training evaluation network parameters of the agent , the corresponding target network is , and For the second training evaluation network parameters of the agent , the corresponding target network is ;
[0054] S520. Define the state of each agent during the period, including the resource caching situation of the carriage itself during the period, the user request situation in the carriage during the period, the carriage attributes, and the location attributes during the period;
[0055] Denote the set of carriage attributes as , and the set of location attributes as . Then the state of the agent during the period is represented as , where represents the carriage attribute where the agent is located, represents the location attribute where this agent is currently located, , indicating the caching situation of the cache server of carriage during the period, indicating the request situation of the users in carriage for video content during the period;
[0056] S530. According to the state of the agent obtained during the period, train the action network , and output the caching probability of each video content ; sort the caching probabilities of each video content from large to small, and select the top video contents as the final caching actions;
[0057] Define the action of the agent as , where, if the agent caches the content during the period, then , otherwise, ; and ;
[0058] S540. Each agent executes the action , and the state changes from Enter the next state ; The environment gives a reward , defined as Time period agent The negative value of the weighted sum of the request response time cost and the resource replacement time cost, expressed as:
[0059] (11);
[0060] S550. Remove the time attribute from the data of all agents in the time period and store it in the sample pool in the form of a five-tuple , where , represents the global state composed of the local observation states of all agents; , represents the global action composed of the local actions of all agents; , represents the global state of the next time period; , represents the reward set; is a binary indicator variable. If the EMU arrives at the terminal station, then , otherwise, ;
[0061] S560. If the number of samples reaches the minimum batch processing threshold , train the neural network and update the target network parameters.
[0062] Preferably, the training of the neural network described in step S560 is specifically as follows:
[0063] S561. First, calculate the two temporal difference errors of the evaluation network of the agent :
[0064] (12);
[0065] (13);
[0066] Among them, is the discount factor; is the reward of the agent ; represents the relative action made by the target action network for the next time period; Select the Huber error as the final error , where is the evaluation network number, expressed as:
[0067] (14);
[0068] S562. Each time for Batch process a random sample, and update the training evaluation network parameters in the way of gradient descent, expressed as:
[0069] (15);
[0070] Among them, is the step size;
[0071] S563. For the update of the policy, use gradient ascent; for the training action network and all target networks, delay the update, that is, every time the training evaluation network is updated, then update the training action network and all target networks, expressed as: After updating the training evaluation network times, update the training action network and all target networks, expressed as:
[0072] (16);
[0073] Among them, is the step size; is the local observation state of the agent of.
[0074] S564. Adopt the soft update method to update the parameters of the target network, expressed as:
[0075] (17);
[0076] (18);
[0077] Among them, is a coefficient between 0 and 1, used to control the update step; soft update slowly follows the parameter changes of the training network by gradually adjusting the parameters of the target network.
[0078] Beneficial effects: By caching video content on the EMU carriages, the present invention can greatly reduce the time for users to obtain video content.
[0079] The present invention considers the problems that the content preferences of users in different carriages in the EMU scenario vary greatly, and the dynamic change differences of user requests in different carriages are relatively large, and designs a multi-carriage collaborative caching system, which can make the carriage video service more accurate and fast, while balancing the loads of each server and fully improving the utilization rate of the EMU caching space.
[0080] The present invention analyzes the linear networking of EMUs, making the defined cross-carriage hit response time cost more in line with reality. At the same time, the present invention considers the problem of poor communication conditions along the way of EMUs and defines the resource replacement time cost, so that the deployment of resources not only considers the current situation, but also takes into account the possible future user request needs, leaving bandwidth for the resources that really need to be downloaded and cached. Through the constraint and balance between the request response time cost and the resource replacement time cost, the resource deployment becomes more reasonable.
[0081] Due to the large differences and dynamics in the EMU scenario, it is usually difficult to apply the learning of historical laws to the future, so that the effect of the prediction-based caching algorithm is average. Starting from the perspective of decision-making, the present invention uses the technology of multi-agent reinforcement learning to decide the caching strategy of resources, which can better cope with the dynamics and differences of environmental changes.
[0082] When the present invention adopts the multi-agent reinforcement learning technology, special indicators for the EMU scenario are also added to make the action decision more accurate and efficient. And a series of optimization methods are adopted to ensure the smoothness and accuracy of the training process and the handling of the overestimation problem. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Figure 1 is the overall flowchart of the EMU video collaborative caching method based on multi-agent reinforcement learning of the present invention.
[0084] Figure 2 is a schematic diagram of the EMU multi-carriage collaborative caching video service system constructed by the present invention.
[0085] Figure 3 is the flowchart of the user request response processing of the EMU multi-carriage collaborative caching video service system constructed by the present invention.
[0086] Figure 4 is the flowchart of the multi-carriage collaborative caching algorithm based on multi-agent reinforcement learning of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0087] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention, and they should not be construed as limitations on the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. In the description of the present invention, it should be understood that the terms used are only for the purpose of description and cannot be construed as indicating or implying relative importance.
[0088] Embodiment: The following is combined withFigures 1 - 4 Specifically describe the video collaborative caching method for multiple units based on multi-agent reinforcement learning provided in this embodiment.
[0089] The video collaborative caching method for multiple units based on multi-agent reinforcement learning in this embodiment is as Figure 1 shown, and the main steps include:
[0090] S100. Construct a video service system for collaborative caching of multiple carriages of multiple units;
[0091] S200. Define the request response time cost according to the cache hit situation of user requests;
[0092] S300. Define the resource replacement time cost according to the difference in cached video content in adjacent time periods;
[0093] S400. Formalize the problem of minimizing the video service time cost of multiple units;
[0094] S500. Invoke the multi-carriage collaborative caching algorithm based on multi-agent reinforcement learning to determine the video content caching policy for each time period.
[0095] Among them, S100 constructs a video service system for collaborative caching of multiple carriages of multiple units, as Figure 2 shown, including:
[0096] S110. Define the set of carriages of multiple units and the set of video content, and number the carriages in sequence from the head to the tail of the train as , let represent the set of carriages of multiple units, represent the set of video content.
[0097] S120. Divide the running time of multiple units according to stations, and each time period represents the running time between the current station and the next station.
[0098] S130. Construct the video service for multiple units, including:
[0099] S131. Set a carriage cache server on each carriage as a reverse proxy, and define the upper limit of the cache capacity of each carriage cache server as , ; it can be used to respond to requests from users on multiple units for video content; the carriage cache servers are connected by wire to form a local area network for multiple units, and each carriage cache server communicates with an external base station through the roof antenna of the carriage where it is located; users connect to the local area network for multiple units through the wireless access points set on each carriage; a central server is set on multiple units, which is used to query the deployment situation of video content for multiple units and perform other services during the operation of multiple units.
[0100] S132. When a user in the carriage accesses the network of this carriage and requests video content, check the cache hit situation of the user request, and select to send the video content requested by the user to the user from this carriage, other carriages or the external base station; as Figure 3 shown, it includes:
[0101] S132-a. Determine whether the video content requested by the user is cached in the carriage where the user is located. If so, execute step S132-e; otherwise, execute step S132-b.
[0102] S132-b. Obtain the deployment situation of the video content requested by the user in the multiple unit train from the central server. If there is a cache of the video content requested by the user in the multiple unit train, execute step S132-c; otherwise, execute step S132-d.
[0103] S132-c. Select the carriage that is the closest to the carriage where the request is sent and has cached the video content requested by the user to join the candidate response carriage set. Check whether there are two candidate carriages with the same distance. If not, directly select the carriage that is the closest to the carriage where the request is sent and has cached the video content requested by the user as the response carriage. If so, select the carriage with a smaller load among the two candidate carriages as the response carriage. The carriage where the user is located acts as an agent to request the video content from the response carriage. After receiving the response, execute step S132-e;
[0104] S132-d. The carriage where the user is located acts as an agent to request the video content required by the user from the external base station. After receiving the response, execute step S132-e.
[0105] S132-e. The carriage where the user is located sends the video content requested by the user to the user in this carriage.
[0106] S133. Let the number of requests for the video content by all users in the carriage during the time period be . Define the binary indicator variable , which indicates whether all users' requests for the video content during the time period in the carriage are responded to by the target . If it is responded to by the target , then , otherwise, . Among them, , . can be the cache server of the multiple unit train carriage or the external base station.
[0107] Step S200 defines the request response time cost according to the cache hit situation of the user request; specifically including:
[0108] According to the cache hit situation of the user request, there are three different responses in total, namely in-car response, inter-car response, and external base station response, corresponding to three different request response time costs;
[0109] Define a binary indicator variable , indicating the carriage and the carriage whether the request between passes through the carriage and the carriage connected link section, if passing through then , otherwise , that is:
[0110] (1);
[0111] Let the total link bandwidth be , define the request response time cost from carriage to the target as:
[0112] (2);
[0113] Among them, is the response time cost for the carriage to request video content from the external base station; represents the time period inside the carriage all users' requests for the video content quantity; represents at time period carriage all users' requests for the video content whether it is responded by the carriage , if responded by the carriage , then , otherwise, ; represents the carriage and the carriage whether the request between passes through the carriage and the target when it is a carriage connected link section;
[0114] Therefore, in the time period to meet the request response time cost of all users in the carriage is:
[0115] (3).
[0116] Step S300 defines the time cost of resource replacement according to the difference in cached video content in adjacent time periods, specifically including:
[0117] S310. Define a binary indicator variable , representing whether the video content in the time period is cached in the carriage . If the video content in the time period is cached in the carriage , then , otherwise ;
[0118] S320. At the beginning of each time period , if the video content to be cached in the current time period is different from the video content cached in the previous time period, each carriage cache server obtains the video content required in the current time period from other carriage cache servers or external base stations to replace the video content cached in the previous time period. Define the resource replacement time cost of carriage at time period as:
[0119] (4).
[0120] Step S400 formalizes the problem of minimizing the time cost of the EMU video service, specifically as follows:
[0121] Let the time cost of the EMU video service be the weighted sum of the request response time cost and the resource replacement time cost of all carriages in all time periods. Formalize the problem of minimizing the time cost of the EMU video service, expressed as:
[0122] (5);
[0123] Constraints:
[0124] (6);
[0125] (7);
[0126] (8);
[0127] (9);
[0128] (10);
[0129] Among them, and is a weight factor used to adjust the weight between the request response time cost and the resource replacement time cost; represents a carriage the upper limit of the cache capacity of the cache server; represents at time period, the carriage all users' requests for video content whether the request is responded to by the carriage if it is responded to by the carriage, then otherwise, ;
[0130] Equation (6) ensures that each request is responded to by only one carriage cache server or external base station. Equation (7) indicates that the cache upper limit of each carriage cache server should not exceed its own maximum cache capacity limit. Equation (8) ensures that a user request can only be responded to by a carriage server or external base station that has cached the requested video content. Equations (9) and (10) limit the indication quantity to a binary indication quantity.
[0131] Step S500 calls the multi-carriage collaborative caching algorithm based on multi-agent reinforcement learning to determine the video resource caching policy for each time period, as Figure 4 shown, including:
[0132] S510. Consider each carriage as an agent. For convenience, use the symbol to represent the agent corresponding to the carriage ; Use a neural network to fit the actions of the agent , denoted as the training action network , where is the training action network parameter of the agent , and the corresponding target action network is denoted as ; For each agent , set two training evaluation networks, denoted as and respectively, where is the first training evaluation network parameter of the agent , and the corresponding target network is , while is the second training evaluation network parameter of the agent , and the corresponding target network is ;
[0133] S520. Define the state of each agent at time period, including the resource caching situation of the carriage itself at time period, the user request situation in the carriage at time period, the carriage attributes, and at Location attribute of the time period;
[0134] Denote the set of carriage attributes as , and the set of location attributes as . Then the state of the agent at time period is represented as , where represents the carriage attribute where the agent is located, represents the location attribute where this agent is currently located, , indicating the cache situation of carriage cache server at time period, indicating the request situation of carriage users for video content at time period;
[0135] S530. According to the state of the agent obtained at time period, train the action network , and output the cache probability of each video content . Since the cached video content cannot exceed the maximum capacity of carriage which is 30, by sorting the cache probabilities of each video content from large to small, select the top 30 video contents as the final cache actions;
[0136] Define the action of the agent as , where, if the agent caches the content at time period, then , otherwise, ; and ;
[0137] S540. Each agent executes the action , and the state changes from to the next state ; The environment gives a reward , defined as the negative value of the weighted sum of the request response time cost and the resource replacement time cost of the agent at
[0138] (11);
[0139] S550. Remove the time attribute from the data of all agents at time period and store it in the sample pool in the form of five-tuples , where represents the global state composed of the local observation states of all agents; represents the global action composed of the local actions of all agents; represents the global state at the next time step; represents the reward set; is a binary indicator variable. If the multiple unit train arrives at the terminal station, then , otherwise, ;
[0140] S560. If the number of samples reaches the minimum batch processing threshold , train the neural network and update the target network parameters.
[0141] The specific steps for training the neural network are as follows:
[0142] S561. First, calculate the two temporal difference errors of the evaluation network of agent :
[0143] (12);
[0144] (13);
[0145] Among them, is the discount factor, is the environmental reward of agent ; represents the action relative to at the next time step made by the target action network; The Huber error is selected as the final error , where is the evaluation network number, expressed as:
[0146] (14);
[0147] S562. Each time, perform batch processing on random samples, and update the training evaluation network parameters in the way of gradient descent, expressed as:
[0148] (15);
[0149] Among them, is the step size;
[0150] S563. For the update of the policy, use gradient ascent; To prevent the policy from being updated in the incorrect direction due to incorrect action evaluation, the training action network and all target networks are updated with a delay, that is, every time Evaluate the training network again, and then update the training action network and all target networks, expressed as:
[0151] (16);
[0152] where is the step size; is the local observation state of the agent .
[0153] S564. Update the parameters of the target network in a soft update manner, expressed as:
[0154] (17);
[0155] (18);
[0156] where , used to control the update step; soft update gradually adjusts the parameters of the target network to slowly follow the parameter changes of the training network.
[0157] Example 2: This example provides a non-transitory computer-readable storage medium, on which computer instructions are stored. The computer instructions cause the computer to execute the multi-agent video collaborative caching method for multiple units based on multi-agent reinforcement learning. The method includes the following steps:
[0158] S100. Construct a multi-carriage collaborative caching video service system for multiple units;
[0159] S200. Define the request response time cost according to the cache hit situation of user requests;
[0160] S300. Define the resource replacement time cost according to the difference in cached video content in adjacent time periods;
[0161] S400. Formalize the problem of minimizing the time cost of the multiple-unit video service;
[0162] S500. Call the multi-carriage collaborative caching algorithm based on multi-agent reinforcement learning to determine the video content caching policy for each time period.
[0163] Example 3: This example provides an electronic device, which may include: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory communicate with each other through the communication bus. The processor can call the logical instructions in the memory to execute the multi-agent video collaborative caching method for multiple units based on multi-agent reinforcement learning. The method includes the following steps:
[0164] S100. Build a collaborative caching video service system for multiple carriages of EMUs;
[0165] S200. Define the request response time cost according to the cache hit situation of user requests;
[0166] S300. Define the resource replacement time cost according to the difference in cached video content in adjacent time periods;
[0167] S400. Formalize the problem of minimizing the time cost of EMU video services;
[0168] S500. Invoke the multi-carriage collaborative caching algorithm based on multi-agent reinforcement learning to determine the video content caching strategy for each time period.
[0169] In addition, when the logical instructions in the above-mentioned memory can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0170] Embodiment 4: This embodiment provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the EMU video collaborative caching method based on multi-agent reinforcement learning. The method includes the following steps:
[0171] S100. Build a collaborative caching video service system for multiple carriages of EMUs;
[0172] S200. Define the request response time cost according to the cache hit situation of user requests;
[0173] S300. Define the resource replacement time cost according to the difference in cached video content in adjacent time periods;
[0174] S400. Formalize the problem of minimizing the time cost of EMU video services;
[0175] S500. Call the multi-carriage collaborative caching algorithm based on multi-agent reinforcement learning to determine the video content caching policy for each time period.
[0176] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.
[0177] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0178] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for collaborative caching of EMU videos based on multi-agent reinforcement learning, characterized in that: The following steps are involved: S100. Constructing a multi-carriage collaborative caching video service system for EMUs, including defining an EMU carriage set and a video content set, and dividing the EMU running time into time periods according to the travel time between stations; S200. According to the cache hit of the user request, define the request response time cost, including the response within the car, the response between cars and the external base station response; S300. Define the resource replacement time cost according to the difference in cached video content in adjacent time periods; S400.Formalize the time cost minimization problem of EMU video service; S500. Calling a multi-carriage collaborative caching algorithm based on multi-agent reinforcement learning to determine the video content caching strategy for each period, where each car is regarded as an agent, a neural network is used to fit the action of the agent, and two evaluation networks are set for collaborative optimization; The data of all agents in the period are stored in the sample pool in the form of five-tuples, including global state, global action, global state of the next period, reward set and end mark; If the number of samples reaches the minimum batch processing threshold, the neural network is trained and the target network parameters are updated; specifically: S510, each carriage is regarded as an intelligent agent, and the carriage number i represents the intelligent agent corresponding to the carriage; a neural network is used to fit the action of intelligent agent i, which is recorded as the training action network where θ i is the training action network parameter of agent i, and the corresponding target action network is recorded as For each agent i, two training and evaluation networks are set up, denoted as and where w i,1 is the first training evaluation network parameter of agent i, and the corresponding target network is And w i,2 is the second training evaluation network parameter of agent i, and the corresponding target network is S520, defining the state of each agent i in time period t, including the resource cache status of the carriage itself in time period t, the user request status in the carriage in time period t, the carriage attributes and the location attributes in time period t; Let the car attribute set be U and the location attribute set be V, then the state of agent i in time period t is expressed as where u i ∈U represents the attributes of the compartment where agent i is located, Represents the current location attribute of the agent i, represents the cache status of the cache server in carriage i during period t, Indicates the video content requests of the users in carriage i during time period t; S530: train the action network according to the state of the agent in period t Output the cache probability of each video content f; sort the cache probabilities of each video content f from large to small, and select the top c i The video content is used as the final cache action; Define the action of the agent as If agent i caches content f during period t, then otherwise, and S540, each agent i performs an action Status by Enter the next state The environment rewards It is defined as the negative value of the weighted sum of the time cost of agent i’s request response and the time cost of resource replacement in period t, expressed as: S550, after removing the time attribute from the data of all agents in time period t, store them in the sample pool D in the form of a five-tuple (o, a, o′, r, d), where: Represents the global state composed of the local observation states of all agents; Represents the global action composed of all the local actions of the agents; Indicates the global state of the next period; represents the reward set; d∈{0,1} is a binary indicator variable, if the EMU arrives at the terminal station, d=1, otherwise, d=0; S560: If the number of samples reaches the minimum batch processing threshold G, the neural network is trained and the target network parameters are updated.
2. The EMU video collaborative caching method based on multi-agent reinforcement learning according to claim 1 is characterized in that: The step S100 of constructing a multi-carriage collaborative caching video service system for a train set includes: S110, define a set of train carriages and a set of video contents, number the carriages 1, 2, ..., m from the front to the rear, let N = {1, 2, ..., m} represent the set of train carriages, and F = {1, 2, ..., f} represent the set of video contents; S120, dividing the EMU running time according to the stations, each time period t∈T={1,2,…,h} represents the travel time from the current station to the next station; S130, constructing EMU video service includes: S131. Set up a carriage cache server as a reverse proxy in each carriage, and define the upper limit of the cache capacity of each carriage cache server as c i , i∈N; the cache servers in each carriage form a train LAN through wired or wireless connections, and each carriage cache server communicates with the external base station C through the roof antenna of the carriage where it is located; users connect to the train LAN through the wireless access point set in each carriage; the train sets a central server to query the deployment of the video content of the train; S132, when a user in compartment i accesses the compartment network and requests video content, the video content requested by the user is sent to the user from the compartment, other compartments or an external base station according to the cache hit of the user request; S133: Let the number of requests for video content f by all users in carriage i in time period t be Define binary indicator variables Indicates whether the request for video content f by all users in carriage i during period t is responded by target j′. If it is responded by target j′, then otherwise, Where j′∈N∪{C}, j′ is the EMU carriage cache server or external base station.
3. The EMU video collaborative caching method based on multi-agent reinforcement learning according to claim 2 is characterized in that: In step S132, when a user in compartment i accesses the compartment network and requests video content, the video content requested by the user is sent to the user from the compartment, other compartments or an external base station according to the cache hit of the user request; the specific method is: S132-a, determining whether the video content requested by the user is cached in the car where the user is located, if so, executing step S132-e, otherwise executing step S132-b; S132-b, obtaining the deployment status of the video content requested by the user in the EMU from the central server, if the EMU has a cache of the video content requested by the user, executing step S132-c, otherwise executing step S132-d; S132-c, select the car that is closest to the car that issued the request and has cached the video content requested by the user to add to the candidate response car set, check whether there are two optional candidate cars with the same distance, if not, directly select the car that is closest to the car that issued the request and has cached the video content requested by the user as the response car, if so, select the car with the smaller load among the two candidate cars as the response car, and the car where the user is located acts as a proxy to request the video content from the response car, and execute step S132-e after receiving the response; S132-d, the carriage where the user is located acts as an agent to request the video content required by the user from the external base station, and after receiving the response, executes step S132-e; S132-e. The carriage where the user is located sends the video content requested by the user to the users in the same carriage.
4. The EMU video collaborative caching method based on multi-agent reinforcement learning according to claim 3 is characterized in that: Step S200 defines the request response time cost according to the cache hit situation of the user request; specifically includes: According to the cache hit situation of the user request, there are three different responses, namely, in-car response, inter-car response and external base station response, corresponding to three different request response time costs; Define a binary indicator variable z i,j,b,e ∈{0,1}, indicating whether the request between carriage b and carriage e passes through the link section connecting carriage i and carriage j. If so, z i,j,b,e =1, otherwise z i,j,b,e =0, that is: Assume the total link bandwidth is B, and define the request response time cost from carriage i to target j′ as: Among them, l C The response time cost of the car requesting video content from the external base station is set to a fixed value; represents the number of requests for video content f by all users in carriage b during time period t; Indicates whether the requests for video content f by all users in carriage b during period t are responded by carriage e. If so, then otherwise, z i,j',b,e represents whether the request between car b and car e passes through the link interval connected when car i and target j′ are cars; The time cost of satisfying the request response of all users in carriage i during time period t is:
5. The method for collaborative caching of train group videos based on multi-agent reinforcement learning according to claim 4 is characterized in that: Step S300 defines the resource replacement time cost according to the difference in cached video content in adjacent time periods; specifically includes: S310. Define binary indicator variables Indicates whether the video content f during period t is cached in carriage j. If the video content during period t is cached in carriage j, then otherwise, S320. At the beginning of each time period t, if the video content to be cached in the current time period is different from the video content cached in the previous time period t-1, each carriage cache server obtains the video content required in the current time period from other carriage cache servers or external base stations to replace the video content cached in the previous time period t-1. The resource replacement time cost of carriage i in time period t is defined as:
6. The EMU video collaborative caching method based on multi-agent reinforcement learning according to claim 5 is characterized in that: Step S400 formalizes the problem of minimizing the time cost of EMU video service, specifically: Let the EMU video service time cost be the weighted sum of the request response time cost and resource replacement time cost of all carriages in all time periods, and formalize the EMU video service time cost minimization problem as follows: Constraints: Among them, μ and ξ are weight factors, which are used to adjust the weight between request response time cost and resource replacement time cost; c j Indicates the upper limit of cache capacity of cache server in carriage j; Indicates whether the requests for video content f by all users in carriage i during period t are responded by carriage j. If so, then otherwise, 7. The EMU video collaborative caching method based on multi-agent reinforcement learning according to claim 1 is characterized in that: Step S560 is to train the neural network, and the specific steps are as follows: S561, first calculate the temporal difference error of the two evaluation networks of agent i: Where γ is the discount factor; r i is the reward of agent i; a′ represents the action of the target action network relative to a in the next period; Huber error is used as the final error L(δ i,k ), where k is the evaluation network number, expressed as: S562, batch processing is performed on G random samples each time, and the training evaluation network parameters are updated by gradient descent, which is expressed as: Among them, α is the step size; S563, for the update of the strategy, gradient ascent is used; for the training action network and all target networks, the update is delayed, that is, each update Train the evaluation network again, and then update the training action network and all target networks, expressed as: Among them, β is the step size; o i is the local observation state of agent i; S564, using a soft update method to update the parameters of the target network, expressed as: w i,k,targ ←τw i,k,targ +(1-τ)w i,k ,k=1,2 (17); i i,targ ←tth i,targ +(1-τ)θ i (18); Among them, τ is a coefficient between 0 and 1, which is used to control the step of the update; the soft update slowly follows the parameter changes of the training network by gradually adjusting the parameters of the target network.
Citation Information
Patent Citations
Automatically augmenting user resources dedicated to serving content to a content delivery network
US10743036B1
KR20240072735A