Wireless network edge caching method and system based on D2D, block chain and social perception

By introducing D2D, blockchain and social perception technologies into wireless networks, combined with deep reinforcement learning algorithm PPO, dynamically adjusting content placement and cache strategies, the problems of network congestion and uneven resource allocation are solved, and efficient and reliable wireless communication is achieved.

CN120238965APending Publication Date: 2025-07-01XIAN UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510267794.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the face of problems of network congestion, uneven allocation of cache resources, complex environment dynamics and high-dimensional decision-making space, the prior art is difficult to effectively improve the efficiency and quality of wireless communications.

Method used

Using the wireless network edge caching method based on D2D, blockchain and social perception, through the interaction between the agent and the environment, the content placement, D2D UE pairing and blockchain cache strategies are dynamically adjusted, and the PPO algorithm is used to optimize the strategy to improve network performance.

Benefits of technology

It significantly improves the network transmission efficiency and service quality, reduces latency and energy consumption, enhances the stability and reliability of the system, and realizes intelligent management of network resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238965A_ABST
    Figure CN120238965A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of wireless communication, and discloses a wireless network edge caching method based on device-to-device (D2D), a block chain and social perception, which realizes content caching on user equipment (UE) and allows the UE to cache the content from the UE, obtain the content from other UE through a D2D link or obtain the required content from a content server. The intimacy between users is evaluated through social perception cooperation, and a real D2D communication link is reconstructed, so that a caching strategy is optimized. Considering the problem that terminals from different operation institutions cannot mutually trust to share contents, the block chain technology is combined to supervise cache transactions, UE transaction records are continuously updated, and a safe and credible transaction platform is constructed. In a D2D edge cache system, the invention aims to optimize content deployment and access strategies so as to reduce D2D communication delay and energy consumption. In order to cope with the dynamic change of the environment, the invention provides a deep reinforcement learning (DRL) framework to solve the problem.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a wireless network edge caching method and system based on D2D, blockchain, and social awareness, specifically involving content placement, D2D UE pairing, and blockchain caching strategies, and belongs to the field of wireless communication technology. Background Art

[0002] With the rapid development of mobile Internet and wireless communication technologies, the surging information traffic has posed unprecedented challenges to network transmission efficiency, service quality, and quality of experience. Especially in densely populated areas, network congestion and latency problems have become more prominent. To address this challenge, edge caching technology deploys caching resources at the network edge, significantly reducing the load on the core network and access latency by storing and distributing content closer to users. Related research shows that edge caching technology can effectively improve the quality of wireless communication experience and provide an important solution for alleviating the pressure on the core network.

[0003] Device-to-Device (D2D) communication enables direct information exchange between devices, thereby improving data transmission efficiency and enhancing network flexibility and reliability. Combining D2D communication, the effectiveness of edge caching is further enhanced. Especially in 5G and future networks, the combination of D2D communication and edge caching brings new development opportunities for wireless communication.

[0004] Deep Reinforcement Learning (DRL), as a cutting-edge machine learning technology, has been successfully applied to edge caching and D2D network optimization in wireless communication. Through the interaction between the agent and the environment, DRL can find the optimal strategy in high-dimensional input and complex decision spaces, thereby effectively improving network performance. In edge caching and D2D communication, deep reinforcement learning can dynamically adjust cached content, optimize resource allocation, and improve data transmission efficiency. However, there are still many challenges in practical applications:

[0005] (1) Network congestion and latency

[0006] With the increase in the number of users, especially in high-density areas, the traditional communication network architecture faces great pressure, resulting in particularly serious network congestion and latency problems. Especially during peak data traffic periods, the core network's bearing capacity is insufficient to handle a large number of data requests in a timely manner, leading to a decline in communication quality. To address this challenge, edge caching technology can effectively reduce the load on the core network by caching content at the network edge, reducing cross-network data transmission, and thus reducing access latency.

[0007] (2) Uneven distribution of caching resources

[0008] The core of edge caching technology lies in storing content closer to users, thereby improving data transmission efficiency and user experience. However, in practical applications, the storage capacities, demand volumes, and cache resource allocation requirements of different devices vary. How to effectively allocate cache resources according to real-time demand remains a challenge. If the resource allocation is unreasonable, some devices may not be able to obtain the required content, resulting in a low cache hit rate and thus affecting the performance of the system.

[0009] (3) Complex environmental dynamics

[0010] The wireless communication environment itself has a high degree of dynamics, and user demands, network environments, device states, etc. are all constantly changing. For example, factors such as user location, signal interference, and device movement will affect the quality and reliability of D2D communication links, while changes in user demands directly affect the selection of cached content and resource allocation. Therefore, how to quickly respond to these changes and timely adjust cache strategies and communication links is a huge challenge.

[0011] (4) High-dimensional decision space

[0012] In the D2D communication and edge caching system, there are numerous decision variables involved, such as the number of user devices, cache capacity, communication channel status, etc. These variables interact with each other, constituting a high-dimensional decision space. In such a complex decision space, how to efficiently find the optimal strategy to improve network performance has become a technical problem. Traditional optimization methods often have difficulty finding effective solutions in high-dimensional spaces due to high computational costs and slow convergence speeds. Summary of the invention

[0013] Aiming at the problems existing in the prior art, the present invention provides a wireless network edge caching method based on D2D, blockchain, and social awareness. The network consists of two parts, namely an edge caching subsystem and a blockchain subsystem. The edge caching subsystem is used for caching and sharing content files, and it includes a remote content server, a BS, and I users (User Equipment, UE) randomly distributed within the coverage area of the base station. The blockchain subsystem is used to build a secure and reliable data sharing information trading platform. In addition, each UE also acts as a blockchain node. The method includes the following steps:

[0014] S101. Initialize the parameters θ and ω of the policy actor network and the critic network of the agent, and set the relevant parameters θ′ and ω′ of the sampling policy actor network and the critic network. At the same time, determine the number of episodes, the learning rates μ and σ of the actor network and the critic network, and the discount factor γ. In addition, initialize the experience pool and configure network layout parameters, such as the number of UEs I, the number of contents F, etc.;

[0015] After initializing the state of the agent in S102, the agent interacts with the environment, and the policy network generates corresponding actions according to the current policy;

[0016] In S103, the agent executes the generated actions, obtains immediate rewards, and causes the environmental state to transition to the next state;

[0017] In S104, according to the parameters of the sampling policy actor network, samples are taken from the complete segments in the environment, and the trajectories are saved in the memory;

[0018] In S105, discounted rewards are calculated;

[0019] In S106, the advantage function is calculated, and a clipping factor constraint is added to limit the update rate, while the objective function is solved;

[0020] In S107, the parameters of the policy actor and critic networks are updated;

[0021] In S108, according to the updated parameters of the policy actor and critic networks, the parameters of the sampling policy actor and critic networks are adjusted;

[0022] In S109, through repeated iterative training, the optimal actions are selected in each state to obtain the maximum benefit, and finally the optimal content placement strategy, D2D UE pairing strategy, and blockchain caching strategy are obtained.

[0023] Furthermore, in S102, the agent interacts with the environment. At the beginning of each episode, the system state s(t) is initialized, specifically expressed as

[0024]

[0025] where

[0026] v(t - 1) = {v i (t - 1)}, where v i (t - 1) is the speed of the UE i at time slot t - 1;

[0027] θ(t - 1) = {θ i (t - 1)}, where θ i (t - 1) is the direction of the UE i at time slot t - 1;

[0028] where q i (t - 1) = {x i (t - 1), y i (t - 1)} is the coordinate of the UE i at the end of time slot t - 1;

[0029] is the UE i and the UE j the distance at time slot t;

[0030] is the UE i the distance from the base station at time slot t;

[0031] where k i,f,l (t) ∈ {0, 1}, which represents the UE i whether to request the l-th quality of content f. Note that at time slot t, a UE can only request one quality of one content at a time;

[0032] represents the UE i and the UE j the social trust level at the start of time slot t;

[0033] represents the UE i and the UE j the social association strength at time slot t;

[0034] represents the UE i and the UE j the Rayleigh fast fading coefficient at time slot t (for D2D communication links);

[0035] represents the UE i the Rayleigh fast fading coefficient with the BS at time slot t;

[0036] represents the UE i and the UE j the D2D communication link bandwidth at time t;

[0037] represents the UE i the bandwidth with the BS at time slot t;

[0038] represents the UE i the transaction size at time slot t;

[0039] χ(t - 1) represents the block size at the end of time slot t - 1.

[0040] In each time slot t, the agent determines its action, denoted as where:

[0041] represents the content placement strategy;

[0042] Represents the D2D UE pairing strategy;

[0043] Represents the blockchain caching strategy.

[0044] Furthermore, after the agent executes the generated action in S103, it obtains an immediate reward and transfers the environmental state to the next state. The calculation formula for the reward is as follows:

[0045]

[0046] In the above formula, O(t) represents the objective function of the present invention.

[0047] Furthermore, in S106: Calculate the advantage function:

[0048]

[0049] At this time, there is:

[0050] δ t = r(t) + γV(s t+1 ; w) - V(s t ; w).

[0051] To improve performance, PPO further optimizes the actor-based objective function. By introducing a clip factor to limit the update rate, the actor of PPO can be updated by maximizing the objective function. The specific formula is:

[0052]

[0053] Where ∈ is a hyperparameter that controls the clip function, which limits the value within a specific range. This method ensures that after minimizing the clip function, the two distributions remain relatively close and large differences are avoided.

[0054] Furthermore, in S107: Update the parameters of the policy actor and critic networks; the parameters of the actor are updated by the following formula:

[0055] θ ← θ + η·▽ θ L clip (θ).

[0056] This algorithm considers the mean squared error function of value estimation and gives the loss function of the critic network:

[0057] L critic (w) = [R(t) - V w (s t )] 2 ,

[0058] Meanwhile, it can be updated through this formula:

[0059] w←w-α·▽ w L critic (w).

[0060] Furthermore, the system further includes:

[0061] System initialization module: Responsible for initializing all necessary parameters, including parameters of experience memory, policy actor network, and critic network, as well as parameters of sampling policy actor and critic network.

[0062] Configuration module: Used to set parameters for specific applications, such as parameters related to the number of users, the number of contents, etc.

[0063] Agent module: At the beginning of each cycle, generate actions according to the current network state; collect data samples, calculate the advantage function, and update the policy and value network during the policy evaluation process.

[0064] Action execution module: Responsible for executing content placement, D2D UE pairing, and blockchain caching strategy.

[0065] Reward acquisition module: Calculate the immediate reward according to whether the system constraint conditions are met. Obtain the reward when all constraint conditions are met, otherwise be punished.

[0066] State transition module: Used to transfer the system from the current state to the next state.

[0067] Experience replay module: Store the experience tuples of system state, action, reward, and next state.

[0068] Data sampling module: Extract a certain segment from the stored experience for learning.

[0069] Network update module: Update the policy actor network and critic network according to the data provided by the experience replay module.

[0070] Parameter update module: Responsible for updating the parameters of the policy actor network and critic network, and adjusting the parameters of the sampling policy actor and critic network.

[0071] Combined with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by the present invention are:

[0072] First, the present invention has important benefits in the integration of D2D, blockchain, and social awareness-based wireless networks, and the specific analysis is as follows:

[0073] First is efficiency and security. The present invention contemplates a cellular network architecture empowered by blockchain and edge caching technologies and assisted by D2D communication. The network consists of two main subsystems: an edge caching subsystem and a blockchain subsystem. The edge caching subsystem is responsible for caching and sharing of content files, and includes a remote content server, a BS, and UEs randomly distributed within the coverage area of the base station. The blockchain subsystem provides a secure and reliable data sharing platform for content sharing transactions. Each UE acts as a blockchain node at the same time, and the blockchain subsystem processes content sharing transactions from the edge caching subsystem in a secure manner to ensure the traceability of content sharing. All cache-related transactions and reputation update transactions will be packaged into blocks and recorded in the blockchain after being verified by a consensus mechanism.

[0074] Considering the dynamic nature of the network environment, the conditional volatility, and the random behavior of nodes during the consensus process, the present invention adopts a DRL method based on PPO to learn the environmental state and provide an optimal joint strategy for content placement, D2D UE pairing, and blockchain caching decisions. Through this strategy, intelligent management of various resources in the network can be achieved, thereby improving the efficiency and security of the overall system.

[0075] Second, the significant technical progress specifically achieved by the present invention lies in realizing a wireless network edge caching method based on D2D, blockchain, and social awareness, which has made significant progress in the following key aspects:

[0076] 1) Precise content placement and blockchain caching strategy

[0077] The present invention significantly improves the hit rate of edge caching by introducing content placement and blockchain caching strategies. The content placement strategy dynamically adjusts the cached content according to user requirements and network status, ensuring that popular or soon-to-be-accessed content is cached first, thereby improving the cache hit rate and reducing latency. At the same time, the blockchain provides a decentralized and immutable data sharing platform to ensure the security and traceability of the cached content. Through smart contracts and reputation mechanisms, the system can incentivize good caching behaviors, ensure network fairness, and optimize cache resource allocation.

[0078] 2) Efficient D2D UE pairing strategy

[0079] The present invention introduces a D2D UE pairing strategy, which can directly transmit data between UEs without passing through the BS, thus effectively reducing communication latency and network load. This strategy dynamically adjusts the pairing strategy to optimize the data transmission path by intelligently selecting the pairing relationship between devices, based on factors such as signal strength, geographical location, communication requirements, and social relationships of the devices. Through D2D pairing, the network can not only improve data transmission efficiency but also ensure the reliability and stability of data transmission in case of network congestion. This direct device-to-device communication method significantly improves system resource utilization, reduces dependence on the base station, and enhances the performance of the entire network.

[0080] 3) Maximize economic cost

[0081] The reward mechanism of this method focuses on the economic rewards of each UE, including D2D content sharing and blockchain caching rewards. In D2D sharing, UEs obtain rewards by directly transmitting data, reducing the burden on the base station and lowering latency; in blockchain caching, UEs obtain rewards by sharing cached content.

[0082] 4) Minimize content acquisition latency

[0083] This method aims to minimize content acquisition latency by optimizing content placement and D2D UE pairing strategies. By caching popular content at the network edge and dynamically adjusting the location of cached content according to user needs, the access distance when users obtain content is reduced. At the same time, D2D communication enables UEs to directly exchange data with other devices, avoiding transmission latency of intermediate nodes. Combining blockchain to ensure the security and credibility of data sharing, these strategies effectively reduce data transmission latency and improve the speed of content acquisition and system response ability.

[0084] 5) Minimize system energy consumption

[0085] This method minimizes system energy consumption by optimizing D2D UE pairing and edge caching strategies. By selecting the optimal pairing path and caching content close to users, the energy consumption of redundant transmission and remote access is reduced. At the same time, the decentralized nature of blockchain improves the efficient utilization of resources and further reduces energy consumption.

[0086] 6) Autonomous learning and optimization ability of the network

[0087] This method enables the system to continuously optimize the decision-making process and thus improve the overall performance through continuous iterative training and experience-based network updates. Using the DRL algorithm, the network can autonomously learn and optimize strategies, dynamically adjust content placement, D2D pairing, and caching strategies, and adapt to environmental changes in real time, thereby improving the overall performance of the system.

[0088] 7) Joint optimization problem

[0089] This method jointly optimizes content placement, D2D UE pairing, and blockchain caching strategies, coordinates the decisions of each subsystem, and comprehensively improves network performance.

[0090] 8) Complexity and dynamics

[0091] The complexity of the network and the dynamics of the environment make it difficult for traditional optimization methods to effectively address these challenges. To address this issue, this method uses the PPO algorithm to adapt to changing network conditions and user requirements.

[0092] 9) Improvement of system stability and reliability

[0093] This method introduces blockchain technology to ensure the security, transparency, and traceability of data sharing. At the same time, combined with deep reinforcement learning optimization strategies, it dynamically adjusts system behavior, enhances network stability and reliability, effectively prevents data loss and malicious attacks, and improves the fault tolerance and continuous operation capabilities of the system in complex environments.

[0094] Third, the core of the wireless network edge caching method based on D2D, blockchain, and social awareness provided by the present invention lies in using mathematical models to guide the system's behavior and learning process. The technical effects brought by them can be explored based on the characteristics of these mathematical models:

[0095] 1) Calculation of even rewards

[0096] The calculation consensus of even rewards focuses on local and D2D cache hit ratios, content acquisition latency, economic returns, and energy consumption.

[0097] Local and D2D cache hit ratios: Defined as the number of sub-files obtained through local hits and D2D hits divided by the total number of all sub-files.

[0098] Content acquisition latency: Content acquisition latency is a performance metric. When its value is small, the quality of experience of UEs is better.

[0099] Economic returns: The economic rewards for each UE come from two aspects, namely the rewards obtained through D2D content sharing and the rewards obtained through blockchain caching.

[0100] Energy consumption: This method only considers the energy consumption of the sender in D2D data sharing and hopes that this energy consumption is as small as possible.

[0101] 2) Update of actor and critic networks

[0102] By randomly sampling small batches of empirical data and using the gradient ascent method, the parameters of the actor and critic networks are updated. In terms of policy optimization, by continuously adjusting the policy network, the system can gradually learn and adopt more effective decision-making strategies. Through the use of small batches of data, the network can quickly respond to environmental changes, enhancing the system's dynamic adaptability.

[0103] 3) Parameter update formula

[0104] By introducing a clipping function, the system gradually adjusts the parameters of the policy actor and critic networks, making the current network gradually approach the target network and avoiding performance fluctuations caused by rapid changes. This parameter update mechanism supports the continuous learning and adaptation of the system, ensuring that the system can smoothly adjust and maintain optimized performance in the face of long-term environmental changes.

[0105] Fourth, as an auxiliary evidence of the inventiveness of the claims of the present invention, it is also reflected in the expected benefits and commercial value after the transformation of the technical solution as follows:

[0106] The present invention proposes a wireless network edge caching method based on D2D, blockchain, and social awareness by introducing the PPO algorithm and combining blockchain technology with social awareness mechanisms, effectively improving the network transmission efficiency and service quality. This method can not only dynamically adjust content placement, caching strategies, and D2D pairing, but also optimize system performance while reducing latency and energy consumption. Especially in scenarios with high requirements for large-scale data processing, real-time communication, and security, the present invention demonstrates strong application potential, bringing significant technological breakthroughs and economic value to related fields. With its high efficiency and security, the present invention has broad market prospects and commercial value in fields such as smart cities, the Internet of Things, and 5G. Brief Description of the Drawings

[0107] To more clearly show the technical solutions of the embodiments of the present invention, the required drawings are briefly introduced below. Obviously, the described drawings only show some embodiments of the present invention, and those skilled in the art can derive other implementation manners from these drawings without creative efforts.

[0108] Figure 1 is a flowchart of the wireless network edge caching method based on D2D, blockchain, and social awareness provided by an embodiment of the present invention;

[0109] Figure 2 is a scenario diagram that can be applied provided by an embodiment of the present invention.

[0110] Figure 3 is an implementation flowchart of the wireless network edge caching method based on D2D, blockchain, and social awareness provided by an embodiment of the present invention.

[0111] Figure 4 It is a comprehensive convergence performance comparison diagram provided by an embodiment of the present invention. Specific embodiments

[0112] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined purpose, the specific embodiments, structural features, and corresponding effects of the present invention are described in detail below in conjunction with the drawings and embodiments.

[0113] Embodiment 1

[0114] Aiming at the problems existing in the prior art, this embodiment provides a wireless network edge caching method based on D2D, blockchain, and social awareness. Figure 2 It is a scenario diagram to which the method of the present invention can be applied. First, it includes two systems: an edge caching subsystem and a blockchain subsystem. The edge caching subsystem is used for caching and sharing of content files, and includes a remote content server, a BS, and UEs randomly distributed within the coverage area of the base station. The blockchain subsystem is used to build a secure and reliable data sharing information trading platform. In addition, each UE also acts as a blockchain node.

[0115] The base station is connected to the content server through a wired optical fiber link. At the same time, all UEs can communicate directly with the base station via a wireless 4G or 5G wireless link. The set of user equipment is represented by , and the index of each user equipment is denoted as i. Macro base stations near scenarios with a large number of users often experience overload situations and thus cannot meet the traffic demands of users during peak hours. For example, this situation is very common during football games. To prevent potential traffic congestion caused by repeated content download requests, UEs can act as caching nodes to provide caching services and share cached content with other UEs through D2D communication.

[0116] The content server stores F content items. The set of video files is represented by . Each video file is encoded into layered files, and the size of each layered file is bits. The encoded layered files are represented by It is represented that the first element corresponds to the base layer file, and all other elements in the set are enhancement layer files. The base is expected to provide the most basic perception quality, and users who can access the enhanced content can obtain higher video quality. Since each video file is encoded into L layered files, each video file accordingly has L quality levels. Hereinafter, the present invention will refer to each layer as a sub-file. In this framework, the content server stores all the layers of all F files. At the same time, the UE with caching function can cache the copies of the base and enhancement layer sub-files of popular video files. Assume that the system of the present invention operates in time slots, and the time slots are indexed by t. Each time slot has a fixed length Δt, and the set of time slots is represented by It is represented.

[0117] The blockchain subsystem mainly processes the content sharing transactions originating from the edge caching subsystem in a secure manner, aiming to ensure the traceability of content sharing. All cache-related transactions and reputation update transactions will be packaged into blocks and then recorded in the blockchain after being verified through the consensus mechanism process.

[0118] Embodiment 2

[0119] Given the dynamics of the network and the uncertainty of information acquisition, the problem becomes complex. To effectively address this challenge, the present invention transforms the problem into a Markov Decision Process (MDP). Since the variables exhibit discontinuity, the present invention selects the PPO algorithm, which is a deep reinforcement learning algorithm designed for the continuous action space reinforcement learning problem and can support real-time online decision-making.

[0120] As Figure 1 shown, the embodiment of the present invention provides a wireless network edge caching method based on D2D, blockchain, and social awareness, specifically including the following steps:

[0121] S101. Initialize the parameters θ and ω of the policy actor network and the critic network of the agent, and set the relevant parameters θ′ and ω′ of the sampling policy actor network and the critic network. At the same time, determine the number of episodes, the learning rates μ and σ of the actor network and the critic network, and the discount factor γ. In addition, initialize the experience pool and configure the network layout parameters, such as the number I of UEs, the number F of contents, etc.;

[0122] S102. After initializing the state of the agent, the agent interacts with the environment, and the policy network generates corresponding actions according to the current policy;

[0123] S103. The agent executes the generated actions, obtains the immediate reward, and makes the environmental state transition to the next state;

[0124] S104. Sample the complete segments in the environment according to the parameters of the sampling policy actor network, and save the trajectories in memory;

[0125] S105. Calculate the discounted reward;

[0126] S106. Calculate the advantage function, add a clipping factor constraint to limit the update rate, and solve the objective function;

[0127] S107. Update the parameters of the policy actor and critic networks;

[0128] S108. Adjust the parameters of the sampling policy actor and critic networks according to the updated parameters of the policy actor and critic networks;

[0129] S109. Through repeated iterative training, select the optimal actions in each state to obtain the maximum benefit, and finally obtain the optimal content placement policy, D2D UE pairing policy, and blockchain caching policy.

[0130] Further, in S102, the agent interacts with the environment. At the beginning of each episode, the system state s(t) is initialized, specifically represented as

[0131]

[0132] where

[0133] v(t - 1) = {v i (t - 1)}, where v i (t - 1) is the speed of the UE i at time slot t - 1;

[0134] θ(t - 1) = {θ i (t - 1)}, where θ i (t - 1) is the direction of the UE i at time slot t - 1;

[0135] where q i (t - 1) = {x i (t - 1), y i (t - 1)} is the coordinate of the UE i at the end of time slot t - 1;

[0136] is the distance between the UE i and the UE j at time slot t;

[0137] is the UE iDistance from the base station at time slot t;

[0138] where k i,f,l (t) ∈ {0, 1}, which represents whether the UE i requests the l-th quality of content f. Note that at time slot t, a UE can only request one quality of one content at a time;

[0139] denotes the UE i and the UE j 's social trust level at the start of time slot t;

[0140] denotes the UE i and the UE j 's social association strength at time slot t;

[0141] denotes the UE i and the UE j 's Rayleigh fast fading coefficient at time slot t (for D2D communication links);

[0142] denotes the UE i and the BS's Rayleigh fast fading coefficient at time slot t;

[0143] denotes the UE i and the UE j 's D2D communication link bandwidth at time t;

[0144] denotes the UE i and the BS's bandwidth at time slot t;

[0145] denotes the UE i 's transaction size at time slot t;

[0146] χ(t - 1) represents the block size at the end of time slot t - 1.

[0147] In each time slot t, the agent determines its action, denoted as where:

[0148] represents the content placement strategy;

[0149] represents the D2D UE pairing strategy;

[0150] represents the blockchain caching strategy.

[0151] Further, in S103: After the agent executes the generated action, it obtains an immediate reward and transfers the environmental state to the next state. The calculation formula for the reward is as follows:

[0152]

[0153] In the above formula, O(t) represents the objective function of the present invention.

[0154] Further, in S106: Calculate the advantage function:

[0155]

[0156] At this time, there is:

[0157] δ t =r(t)+γV(s t+1 ;w)-V(s t ;w).

[0158] To improve performance, PPO further optimizes the actor-based objective function. By introducing a clip factor to limit the update rate, the actor of PPO can be updated by maximizing the objective function. The specific formula is:

[0159]

[0160] Where ∈ is a hyperparameter that controls the clip function, which limits the value within a specific range. This method ensures that after minimizing the clip function, the two distributions remain relatively close and avoid large differences.

[0161] Further, in S107: Update the parameters of the policy actor and critic networks; the parameters of the actor are updated by the following formula:

[0162] θ←θ+η·▽ θ L clip (θ).

[0163] This algorithm considers the mean square error function of value estimation and gives the loss function of the critic network:

[0164] L critic (w)=[R(t)-V w (s t )] 2 ,

[0165] Meanwhile, it can be updated by this formula:

[0166] w←w-α·▽ w L critic (w).

[0167] To elaborate in detail on the wireless network edge caching method based on D2D, blockchain, and social awareness, the present invention provides two specific application embodiments, including the key details of the implementation scheme.

[0168] Application Embodiment 1: Data Sharing and Content Distribution in Smart Cities

[0169] 1) System Initialization

[0170] Initialize the edge caching system, including the BS, remote content server, and UE.

[0171] Configure the blockchain subsystem to ensure the security and traceability of data sharing, and initialize the policy actor network and critic network parameters of the agent.

[0172] Set the initial state of the environment, including the network topology within the urban area, the caching status of each UE, and the communication links.

[0173] 2) Interaction between the Agent and the Environment

[0174] The agent generates actions based on the current network state (such as the caching status of user devices, available bandwidth, etc.), including selecting a content placement strategy, D2D UE pairing, and a caching content update strategy. Based on the interaction between the current policy and the environment, the agent executes the actions and undergoes a state transition.

[0175] 3) Action Execution and Reward Obtaining

[0176] Execute the content placement operation to update the caching content of user devices to appropriate locations. Optimize the data exchange between user devices according to the social awareness mechanism and D2D communication strategy to ensure that data can be shared among local devices and avoid excessive core network requests.

[0177] 4) Obtain Rewards

[0178] Calculate the immediate reward according to the goal of the task.

[0179] If the data is successfully shared without delay, a positive reward is given; if there is a transmission delay or network congestion, a penalty is given.

[0180] 5) Learning and Optimization

[0181] Use the collected experience data to update the policy actor and critic networks of the agent, and optimize the policy through the PPO algorithm. Adjust the policy according to the reward signal so that the agent can make better content placement and D2D communication decisions under different network states. Repeatedly iterate and optimize until the policy converges to achieve the optimal caching content placement and data sharing strategy.

[0182] Application Example 2: Real-time Data Transmission and Content Update in Telemedicine

[0183] 1) System Initialization

[0184] Initialize the agents in the hospital network, including setting up the policy actor and critic networks, defining the initial state of the environment, including the states of the devices (doctors' and patients' devices) in the hospital, the cache distribution of medical data, and the bandwidth and cache capacity of each UE device.

[0185] Configure the blockchain subsystem to ensure the security and privacy of medical data, and to ensure the transparency and credibility of data sharing.

[0186] 2) Interaction between Agent and Environment

[0187] Based on the current medical data requirements, device states, and network conditions, the agent generates actions, decides on the placement location of content, the pairing of D2D devices, and the content update strategy. After executing the corresponding actions, the agent observes the feedback from the environment and further adjusts and optimizes the strategy.

[0188] 3) Action Execution and Reward Obtaining

[0189] During the medical data transmission process, the agent performs content placement operations to ensure that the data is updated to the patient device or doctor device in a timely manner for quick response. Based on the D2D communication strategy, the best path is selected for medical data transmission to reduce network latency and core network load.

[0190] 4) Obtaining Rewards

[0191] Calculate the rewards based on factors such as the latency, success rate, and system energy efficiency of medical data transmission. If the data is transmitted in a timely manner and the energy efficiency is high, the agent receives a positive reward; if the latency is too long or the transmission fails, a penalty is given.

[0192] 5) Learning and Optimization

[0193] Use the collected experience data to update the policy network, apply the PPO algorithm to optimize the policy, and ensure that in a dynamically changing medical network environment, the agent can dynamically adjust decisions such as caching, content placement, and D2D pairing. Through repeated iterative training, gradually improve the optimization ability of the system, and finally achieve efficient transmission, content update, and cache policy optimization of medical data. Through the above steps, the present invention provides an efficient and intelligent solution for ensuring real-time data transmission and content update in telemedicine.

[0194] In these two embodiments, the wireless network edge caching method based on D2D, blockchain, and social awareness provides an efficient, reliable, and energy-saving solution applicable to different application scenarios, from smart cities to remote medical systems, demonstrating its broad application potential and technical advantages.

[0195] To comprehensively evaluate the performance of the present invention, the present invention compares it with multiple representative benchmark algorithms. These benchmark algorithms each have their own characteristics and advantages and can effectively handle similar problems. By comparing the performance of these algorithms under the same conditions, the present invention can more clearly understand the superiority of the present invention and potential improvement space.

[0196] 1) Fixed blockchain algorithm: In each time slot, each user caches a fixed proportion of the blockchain ledger, which means that the size of the blockchain ledger cached by each user remains constant. By caching ledgers of different sizes, the impact of the blockchain ledger cache size within the framework proposed by the present invention is evaluated.

[0197] 2) Fixed pairing algorithm: During the entire simulation process, the D2D communication pairing between users remains unchanged. That is to say, each user is restricted to accessing and caching local content only from a single fixed counterpart. This algorithm is adopted to evaluate the impact of user pairing within the system proposed by the present invention.

[0198] 3) The present invention: It represents the proposed joint content placement, D2D UE pairing, and blockchain caching strategy based on PPO.

[0199] In Figure 4 , the present invention comparatively evaluates the convergence performance of the three algorithms with default parameter settings, mainly considering two aspects: convergence speed and the final reward obtained. From the curve trend, all three algorithms reach a stable stage at around the 200th round, indicating that they all have relatively reliable convergence capabilities. In contrast, the algorithm proposed by the present invention shows a faster reward accumulation speed during the training process and can further improve the obtained reward after stabilization.

[0200] By comparing the final performance of the three curves, it can be seen that the proposed algorithm significantly exceeds the other two algorithms in terms of the reward value, indicating that it can achieve better strategy selection in decision-making scenarios such as resource scheduling and content placement. This result verifies the effectiveness and superiority of the algorithm of the present invention and provides a feasible basis for subsequent extended applications in complex network environments.

[0201] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modification, equivalent replacement or improvement made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A wireless network edge caching method based on D2D, blockchain and social perception, comprising the following steps: 1) Initialize the parameters of the first and second networks of the strategy, set the parameters of the sampling strategy actor network and critic network, and the hyperparameters required for training. Based on the user social graph, historical access records, and D2D reachability between devices, build a user-content relationship model to assist in the decision-making of the caching strategy; 2) In each round, an action is generated according to the environment state, the environment state is transferred to the next state, and the instant information corresponding to the action is recorded; 3) Based on the parameters of the sampling strategy actor network, the complete segment is sampled and the sampled trajectory data is stored in the memory; 4) performing discount processing on the sampled trajectory data, and calculating a corresponding difference value based on the discount result; 5) Within the target scope including parameter update constraints, the parameters of the first network and the second network are iterated, and the cache replacement strategy of the D2D device is adaptively adjusted through the dynamic exploration mechanism of reinforcement learning, and the blockchain transaction model is optimized to reduce cache redundancy and on-chain computing overhead; 6) According to the updated first network and second network parameters, the first network and second network parameters of the sampling strategy are adjusted and the training process is repeated until the training is completed to obtain strategies for content placement, D2D pairing, and blockchain caching.

2. The wireless network edge caching method based on D2D, blockchain and social perception as claimed in claim 1, characterized in that: When initializing the system state in step 2), the speed, direction, coordinates, distance, social trust, social association, fast fading coefficient, bandwidth and size information related to transactions or blocks of each user device in different time periods are represented; the actions of the user device consist of content placement instructions, D2D pairing instructions and blockchain cache instructions.

3. The wireless network edge caching method based on D2D, blockchain and social perception as claimed in claim 1, characterized in that: The process of acquiring the instant information between step 2) and step 3) includes recording corresponding instant feedback data based on the results of the specific indicators generated in the environment according to content placement, D2D pairing and blockchain operations, and combining the feedback data with the environment state transformation to obtain a new environment state.

4. The wireless network edge caching method based on D2D, blockchain and social perception as claimed in claim 1, characterized in that: When discounting the sampled trajectory data in step 4), first weight it according to the feedback data generated at each time step, then aggregate the data of all time steps, calculate the corresponding difference to characterize the deviation of the actual performance from the reference benchmark, and add a constraint factor in the target range to limit the update amplitude.

5. The wireless network edge caching method based on D2D, blockchain and social perception as claimed in claim 1, characterized in that: In step 5), the parameters of the first network (actor network) are adjusted using an iterative method based on the negative gradient direction. The iterative method gradually corrects the existing parameters in the target range according to the difference amount and a preset learning rate, so that the new parameters are close to the target distribution within a certain range.

6. The wireless network edge caching method based on D2D, blockchain and social perception as claimed in claim 1, characterized in that: In step 5), the second network corrects the value estimate based on the mean square error, uses the analysis results of the difference between the reference benchmark and the existing estimate, updates the parameters in the negative gradient direction at a preset learning rate, and generates a new value estimate distribution.

7. A wireless network edge caching system based on D2D, blockchain and social perception, comprising: at least one processor; A storage device, used to store parameters of the first network and the second network and related hyperparameters used for training; The data acquisition module is used to obtain the status information of the user device, the user request instructions and the block data related to the blockchain from the external environment; A trajectory storage module is used to store trajectory data obtained after sampling a complete segment of the external environment; A parameter updating module, configured to iterate parameters of the first network and the second network based on the trajectory data; Communication interface, used to exchange data with distributed user equipment or base stations.

8. The system according to claim 7, characterized in that The first network stored in the storage device is a policy network, the second network stored in the storage device is a value evaluation network, and a set of constraint parameters for representing a clipping factor is also stored.

9. The system according to claim 7, characterized in that The track storage module includes a set of storage units arranged in chronological order, which are used to sequentially store the environment state, execution instructions and instant information returned by the external environment, and index and manage them in a data structure manner.

10. The system according to claim 7, characterized in that A short-range wireless communication module and an access gateway are used between the communication interface and the base station, user equipment or other nodes, and a group of bandwidth configuration units and user identity recognition units are set in the access gateway to identify access requests and allocate bandwidth resources during data exchange.