Double-stage edge computing task unloading method suitable for human body action recognition

Through the dual-stage edge computing task unloading method and the DPC-GCN model, the problem of limited processing capabilities of terminal equipment is solved, efficient and accurate human movement recognition is achieved, and recognition accuracy and robustness are improved.

CN119992292APending Publication Date: 2025-05-13SHENYANG INST OF ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510049838.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing human action recognition technology has insufficient computing resources and energy consumption, which makes it difficult for terminal devices to process complex image and video data efficiently and in real time, and behavioral identification lacks fine-grained classification and rating.

Method used

The dual-stage edge computing task unloading method is adopted, and the task unloading strategy is designed through the DQN agent and the PPO agent, and the calculation tasks are reasonably allocated to the cloud and edge ends, and the dynamic point cloud graph convolution network DPC-GCN model is used for human body movement recognition.

Benefits of technology

It effectively reduces the calculation delay and energy consumption of terminal equipment, significantly improves the accuracy and robustness of human movement recognition, and achieves efficient and accurate human movement recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992292A_ABST
    Figure CN119992292A_ABST
Patent Text Reader

Abstract

The invention discloses a two-stage edge computing task unloading method suitable for human body action recognition, which comprises the following steps of: capturing a video by adopting terminal monitoring equipment, and unloading a human body action recognition task through a two-stage human body action recognition task unloading strategy in an edge computing environment, wherein the global task intelligent scheduling is responsible for selecting an optimal server, and the single task layering method is responsible for determining the unloading percentage of each task. Afterwards, a dynamic part clustering graph convolution action recognition network is adopted, interaction among body parts is captured, and action recognition is completed with low calculation cost. According to the dual-stage edge computing task unloading strategy suitable for human body action recognition, the problems of computing unloading and resource optimization of human body action recognition in an edge computing environment are solved, and the accuracy and efficiency of judging dangerous and abnormal behaviors in safety monitoring are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and edge computing, and in particular provides a two-stage edge computing task offloading method suitable for human motion recognition. Background Art

[0002] In recent years, with the rise of deep learning, human action recognition methods based on deep neural networks have made great progress and achieved some results, but there are still the following shortcomings: 1. The accuracy of human action recognition has been greatly improved, but it requires a lot of computing resources and ignores the problem of energy consumption; 2. At present, most human action recognition calculations mainly rely on terminal devices. However, these monitoring terminal devices generally have limited computing power. They not only face the challenge of processing huge video data streams, but are also limited by their own load and high energy consumption, making it difficult to process the recognized content efficiently and in real time. Furthermore, due to the identification of these restricted behaviors, there is no emphasis on more fine-grained division and rating of actions.

[0003] Therefore, proposing a new method for human action recognition, efficient computational offloading and resource optimization, and improving recognition rate and accuracy have become urgent issues to be solved. Summary of the invention

[0004] In view of this, the purpose of the present invention is to provide a two-stage edge computing task offloading method suitable for human motion recognition. Through the two-stage offloading strategy, the computing tasks are reasonably allocated between the cloud and the edge according to the complexity of motion recognition and the computing resource requirements, so as to solve the energy consumption and efficiency problems, and the lightweight deep learning model and dynamic component clustering graph convolutional network are used to improve the accuracy of motion recognition.

[0005] The technical solution provided by the present invention is: a two-stage edge computing task offloading strategy suitable for human motion recognition, comprising the following steps:

[0006] S1: Use the terminal to collect video images to generate human behavior recognition tasks. Instead of sending all computing tasks to the edge server or cloud for processing, the terminal will first complete a part of the computing tasks locally based on its own computing capabilities, and use the DQN agent to select the server to perform the remaining task calculations;

[0007] S2: In the first stage of offloading, the DQN agent is used to collect offloading environment information, and the environment data is standardized and input into the DQN agent to obtain the offloading behavior; then, according to the dynamic adjustment of the server load and range, it is decided to offload to the selected server or the cloud. When the task is processed on the server side, the DQN agent receives a positive reward, and when it is processed on the cloud or locally, the DQN agent receives a negative reward;

[0008] S3: After the first stage of unloading is completed, the second stage of unloading is carried out, and the cumulative reward is calculated. Next, the identified action tasks are rated for danger and grouped accordingly. In each group, the task with the longest calculation time in the previous round is selected as the representative task. A PPO agent different from the DQN agent is used to collect the status of the representative tasks of each group and calculate the unloading behavior, that is, to determine the unloading percentage of the task;

[0009] S4: After completing the second stage of unloading, the action recognition task is distributed according to the selected server and the unloading percentage. The action recognition task is completed through the dynamic point cloud graph convolutional network DPC-GCN model to obtain the required features and action recognition results.

[0010] Furthermore, S2 specifically includes the following steps:

[0011] S21: Based on the received environmental data, in the first round of iteration, the server closest to the terminal is selected as the initial offloading target for offloading. In subsequent rounds, the DQN agent makes autonomous selections, and tasks that exceed the server range or are larger than the load are offloaded to the cloud.

[0012] S22: The reward obtained by the DQN agent is based on the Markov decision process, which calculates the reward value by unloading information in the environment;

[0013]

[0014] The system state S1 includes the total task computation delay of task i The transmission rate between the terminal and the server is r i,j , computing power is better than g i and the energy consumption of the i-th task of the terminal e i , M represents the number of servers, and N represents the number of tasks;

[0015] S23: Train the DQN agent, guide the DQN agent training through the objective function, and obtain the reward R1, where The total delay of task i in round t is calculated, z and b represent hyperparameters;

[0016]

[0017] Among them, R1 represents the reward function, and z1 represents the total calculation delay of the adjustment task The weight parameter of the influence on the reward function, z2 is used to adjust the weight parameter of the influence of the terminal energy consumption on the reward function, e iIncluding communication energy consumption and computing energy consumption, b1b2 is a bias term used to adjust the baseline value of the reward function to reflect different offloading strategies or environmental conditions;

[0018] S24: The goal of the DQN agent is to maximize the cumulative reward, and the reward is calculated through the designed reward function R1, where the energy consumption e is set to a negative value, indicating that the reward obtained is the largest when the task delay is low and the energy consumption is small.

[0019] Furthermore, S3 specifically includes the following steps:

[0020] S31: Use a lightweight high-risk action prediction network to group the risk levels. In each group, select the task with the longest calculation time in the previous round as the representative task, and use the PPO agent to collect the status of the representative task. After round 0, calculate the target workload based on the output action of the PPO agent, and select the layer closest to the target workload as the decomposition layer;

[0021] S32: Use a lightweight high-risk action prediction network HDN to predict the danger level D of each action, and divide all action categories into three levels according to the probability of danger. The high-risk action prediction network HDN adopts an efficient network implementation and calculates the task time of transmitting the danger level D to server j. Where l(D i ) represents the amount of data of the danger level, and the total delay of executing the HDN is composed of the calculation delay of the high-risk action prediction network HDN and the transmission time of the danger level;

[0022]

[0023] in represents the computational delay of the high-risk action prediction network HDN, that is, the time required to perform the risk level prediction on the terminal. Perform high-risk actions to predict the total delay of the network HDN;

[0024] S33: The total computational delay and local computational load percentage of the representative task are input into the PPO agent as states, and the agent is trained to generate actions based on these states. The preprocessed actions are then assigned to each group as the next round of offloading decisions, and modeled based on Markov decision making:

[0025]

[0026] Where S2 is the computational latency of the representative task and the percentage of local computational work, is the previous round of unloading decision, G is the number of groups;

[0027] S34: Each group of tasks executes a consistent offloading decision logic. During the execution of each task, when it reaches a specific decomposition level on the terminal, the intermediate results generated will be transmitted to the corresponding servers. These servers will be responsible for continuing to calculate the remaining levels of the task and feeding the results back to the system after the calculation is completed;

[0028] S35: The server receives a reward R2 based on the delay in processing the task, which includes The local calculation delay of task i in round t and The total delay composition is calculated for task i in round t, providing feedback to the PPO agent to adjust future unloading decisions;

[0029]

[0030] in It refers to the computation delay of the i-th task completed locally on the terminal.

[0031] Furthermore, S4 specifically includes the following steps:

[0032] S41: In the human action recognition task, in order to perform action recognition and classification, the data that needs to be transmitted and processed, that is, the serialized data of the transmitted skeleton, is ensured to ensure the integrity and security of the data through the first layer multi-input branch of the designed dynamic point cloud convolutional network;

[0033] S42: Next, the transmitted feature pair is input with feature f in Perform simple graph convolution SGC feature extraction, ns q s represents the spatial features of node q in a certain state, (q, s) represents a pair of nodes, where q and s are two nodes in the graph,

[0034] Get the enhanced feature ns of each node;

[0035] S43: Use the attention mechanism to calculate the weight distribution m between nodes q , based on the features of adjacent nodes, calculate the weight α between nodes i,j , focus on important nodes, according to the weight α i,j , get the weighted node feature X q ;

[0036]

[0037] Where W * represents the weight matrix, which is used for linear transformation, b is the bias term, α q,s It is represented as the attention weight between node q and its neighbor node s, Represented as the weight vector in the attention mechanism, σ is the activation function, and δ is the conversion function;

[0038] S44: After obtaining the node features, the node clustering is performed. Xq is input into the evaluation function ADconv, and the evaluation score φ is calculated for each node. The top C nodes with the highest evaluation scores are selected as cluster centers. i,j And the natural connection between nodes, assign the nodes to the corresponding clusters, and get the clustering result P, where W′, W″, W″′ are the attention weights;

[0039] S45: The clustering result P is fed into the adaptive spatiotemporal attention module to perform spatial attention calculation, using the clustering result P and node feature f in , the spatial attention distribution f is obtained through attention calculation p , distribute the spatial attention of all clustering results P to f p Connect them together to get the overall spatial attention representation f' in ;

[0040]

[0041] f i ' n =concat({f p |p=1,2,…,P})

[0042] S46: By calculating the spatial attention, the spatiotemporal joint attention extraction is performed, and the overall spatial attention f i ' n and the original node feature f in Perform spatiotemporal joint attention extraction to obtain the attention score of the entire skeleton sequence:

[0043] f inner =θ * ([pool t (f i ' n )||pool v (f in )]·W2)

[0044]

[0045] S47: Output the results obtained in S45 and S46 and use them as input for subsequent layers to perform further feature extraction and classification, and finally output the final feature map or classification result;

[0046] S48: The received joint position, joint velocity and skeleton information are taken as input. After batch normalization, features are extracted and enhanced through a series of feature extraction layers, global graph convolutional network GCN layers and adaptive spatial-temporal attention ASTA modules. The behavior recognition results are output through feature connection, global average pooling and fully connected layers.

[0047] The present invention proposes a two-stage edge computing task offloading method suitable for human motion recognition, aiming to solve the problem that the terminal device has limited processing power and is difficult to directly process complex image and video data. The method first uses edge computing technology to intelligently allocate computing tasks to terminal devices and edge servers through a carefully designed two-stage offloading strategy. In the first stage, the deep reinforcement learning algorithm DQN (Deep Q-Network) is used to perform preliminary evaluation and selection of tasks, and low global task intelligent scheduling is realized to select the optimal server to reduce the network burden. Subsequently, in the second stage, the policy optimization algorithm PPO (Proximal Policy Optimization) is used to further determine the offloading percentage of each task and offload tasks with high computing requirements to the edge server to ensure efficient execution of the task. After the task is offloaded, the present invention introduces the DPC-GCN model for human motion recognition. The DPC-GCN model can dynamically capture changes in human posture, effectively integrate the spatial relationship and time series information between joints through a graph convolutional network, and achieve high-precision recognition of complex actions. This method not only makes full use of the distributed processing advantages of edge computing, but also improves the recognition accuracy through advanced deep learning technology.

[0048] In summary, the dual-stage offloading strategy of the present invention combined with the DPC-GCN model effectively solves the problem of limited processing power of terminal devices and realizes efficient and accurate human motion recognition. Experimental results show that this method significantly improves the accuracy and robustness of motion recognition while reducing computational latency, providing an innovative solution for human motion recognition applications in the field of edge computing. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The following is a further detailed description of this patent in conjunction with the accompanying drawings and implementation methods:

[0050] Figure 1 The overall flow chart of the method provided by the present invention;

[0051] Figure 2 The model framework diagram of DPC-GCN of the two-stage edge computing task offloading strategy suitable for human motion recognition provided by the present invention. DETAILED DESCRIPTION

[0052] The technical solution of the present invention is further described in detail below.

[0053] See also Figure 1 As shown, the technical solution provided by the present invention is: a two-stage edge computing task offloading method based on human motion recognition, comprising the following steps:

[0054] S1: Use the terminal to collect video images to generate human behavior recognition tasks, and do not send all computing tasks to the edge server or cloud for processing. The terminal device will first complete a part of the computing tasks locally based on its own computing capabilities, and use the DQN agent to select a suitable server to perform the remaining task calculations;

[0055] S2: In the first stage of offloading, the DQN agent is used to collect offloading environment information, and the environment data is standardized and input into the DQN agent to obtain the offloading behavior; then, according to the dynamic adjustment of the server load and range, it is decided to offload to the selected server or the cloud. When the task is processed on the server side, the DQN agent receives positive rewards, and when it is processed on the cloud or locally, the DQN agent can receive negative rewards;

[0056] S2 specifically includes the following steps:

[0057] S21: Based on the received environmental data, in the first round of iteration, the server closest to the terminal is selected as the initial offloading target for offloading. In subsequent rounds, the DQN agent makes autonomous selections, and tasks that exceed the server range or are larger than the load are offloaded to the cloud.

[0058] S22: The reward obtained by the DQN agent is based on the Markov decision process, which calculates the reward value by unloading information in the environment;

[0059]

[0060] The system state S1 includes the total task computation delay of task i The transmission rate between the terminal and the server is r i,j , computing power is better than g i and the energy consumption of the terminal's i-th task e i , while M represents the number of servers and N represents the number of tasks;

[0061] S23: Train the DQN agent, guide the DQN agent training through the objective function, and obtain the reward R1, where The total delay of task i in round t is calculated, z and b represent hyperparameters;

[0062]

[0063] Among them, R1 represents the reward function, and z1 represents the total calculation delay of the adjustment task The weight parameter of the influence on the reward function, z2 is used to adjust the weight parameter of the influence of the terminal energy consumption on the reward function, e i The energy consumption of the i-th task includes communication energy consumption and computational energy consumption. b1b2 is the bias term used to adjust the baseline value of the reward function to reflect different offloading strategies or environmental conditions.

[0064] S24: The goal of the DQN agent is to maximize the cumulative reward, and the reward is calculated through the designed reward function R1, where the energy consumption e is set to a negative value, indicating that the reward obtained is the largest when the task delay is low and the energy consumption is small;

[0065] S3: After the first stage of unloading is completed, the second stage of unloading is carried out, and the cumulative reward is calculated. Next, the identified action tasks are rated for danger and grouped accordingly. In each group, the task with the longest calculation time in the previous round is selected as the representative task. A PPO agent different from the DQN agent is used to collect the status of the representative tasks of each group and calculate the unloading behavior, that is, to determine the unloading percentage of the task;

[0066] S3 specifically includes the following steps:

[0067] S31: Use a lightweight high-risk action prediction network to group the risk levels. In each group, select the task with the longest calculation time in the previous round as the representative task, and use the PPO agent to collect the status of the representative task. After round 0, calculate the target workload based on the output action of the PPO agent, and select the layer closest to the target workload as the decomposition layer;

[0068] S32: Use a lightweight high-risk action prediction network HDN to predict the danger level D of each action, and divide all action categories into three levels according to the probability of danger. The high-risk action prediction network HDN adopts an efficient network implementation and calculates the task time of transmitting the danger level D to server j. Where l(D i ) represents the amount of data of the danger level, and the total delay of executing the high-risk action prediction network HDN is composed of the calculation delay of the high-risk action prediction network HDN and the transmission time of the danger level;

[0069]

[0070] in Represents the calculation delay of HDN, that is, the time required to perform hazard level prediction on the terminal. Total latency of executing HDN.

[0071] S33: The total computational delay and local computational load percentage of the representative task are input into the PPO agent as states, and the agent is trained to generate actions based on these states. The preprocessed actions are then assigned to each group as the next round of offloading decisions, and modeled based on Markov decision making:

[0072]

[0073] Where S2 is the computational latency of the representative task and the percentage of local computational work, is the previous round of unloading decision, G is the number of groups;

[0074] S34: Each group of tasks executes consistent offloading decision logic. During the execution of each task, when it reaches a specific decomposition level on the terminal, the intermediate results generated will be transmitted to the corresponding servers. These servers will be responsible for continuing to calculate the remaining levels of the task and feeding the results back to the system after the calculation is completed;

[0075] S35: The server receives a reward R2 based on the delay in processing the task, which includes The local computation delay of task i in round t and The total delay composition is calculated for task i in round t, providing feedback to the PPO agent to adjust future unloading decisions;

[0076]

[0077] Among them It refers to the computation delay of the i-th task completed locally on the terminal.

[0078] S4: After completing the second stage of unloading, the action recognition task is distributedly executed according to the selected server and the unloading percentage. The action recognition task is completed through the dynamic point cloud graph convolutional network DPC-GCN model to obtain the required features and action recognition results;

[0079] S4 specifically includes the following steps:

[0080] S41: In the human action recognition task, in order to perform action recognition and classification, the data that needs to be transmitted and processed, that is, the serialized data of the transmitted skeleton, is ensured to ensure the integrity and security of the data through the first layer multi-input branch of the designed dynamic point cloud convolutional network;

[0081] S42: Next, the transmitted feature pair is input with feature f in Perform SGC (simple graph convolution) feature extraction, ns q sRepresents the spatial features of node q in a certain state, (q, s) represents a pair of nodes, where q and s are two nodes in the graph. Get the enhanced features ns of each node;

[0082] S43: Use the attention mechanism to calculate the weight distribution m between nodes q , based on the features of adjacent nodes, calculate the weight α between nodes i,j , focus on important nodes, according to the weight α i,j , get the weighted node feature X q ;

[0083]

[0084] Where W * represents the weight matrix, which is used for linear transformation, b is the bias term, α q,s It is represented as the attention weight between node q and its neighbor node s, It is represented as the weight vector in the attention mechanism, σ is the activation function, and δ is the conversion function.

[0085] S44: After obtaining the node features, the node clustering is performed. Xq is input into the evaluation function ADconv, and the evaluation score φ is calculated for each node. The top C nodes with the highest evaluation scores are selected as cluster centers. i,j And the natural connection between nodes, assign the nodes to the corresponding clusters, and get the clustering result P, where W′, W″, W″′ are the attention weights;

[0086] S45: The clustering result P is fed into the adaptive spatiotemporal attention module to perform spatial attention calculation, using the clustering result P and node feature f in , the spatial attention distribution f is obtained through attention calculation p , distribute the spatial attention of all clustering results P to f p Connect them together to get the overall spatial attention representation f i ' n ;

[0087]

[0088] f i ' n =concat({f p |p=1,2,…,P})

[0089] S46: By calculating the spatial attention, the spatiotemporal joint attention extraction is performed, and the overall spatial attention f i ' n and the original node feature fin Perform spatiotemporal joint attention extraction to obtain the attention score of the entire skeleton sequence:

[0090] f inner =θ * ([pool t (f′ in )||pool v (f in )]·W2)

[0091]

[0092] S47: Output the results obtained in S45 and S46 and use them as input for subsequent layers to perform further feature extraction and classification, and finally output the final feature map or classification result.

[0093] S48: Figure 2 As shown in the figure, the received joint positions, joint velocities and skeleton information are taken as input, and after batch normalization, features are extracted and enhanced through a series of feature extraction layers (including B0, B1 and B2 blocks) and global graph convolutional network (GCN) layers and adaptive spatial-temporal attention (ASTA) modules. The model further outputs the action recognition results through feature connection, global average pooling and fully connected layers.

[0094] The present invention proposes a two-stage edge computing task offloading method suitable for human motion recognition, aiming to solve the problem that the terminal device has limited processing power and is difficult to directly process complex image and video data. The method first uses edge computing technology to intelligently allocate computing tasks to terminal devices and edge servers through a carefully designed two-stage offloading strategy. In the first stage, the deep reinforcement learning algorithm DQN (Deep Q-Network) is used to perform preliminary evaluation and selection of tasks, and low global task intelligent scheduling is realized to select the optimal server to reduce the network burden. Subsequently, in the second stage, the policy optimization algorithm PPO (Proximal Policy Optimization) is used to further determine the offloading percentage of each task and offload tasks with high computing requirements to the edge server to ensure efficient execution of the task. After the task is offloaded, the present invention introduces the DPC-GCN model for human motion recognition. The DPC-GCN model can dynamically capture changes in human posture, effectively integrate the spatial relationship and time series information between joints through a graph convolutional network, and achieve high-precision recognition of complex actions. This method not only makes full use of the distributed processing advantages of edge computing, but also improves the recognition accuracy through advanced deep learning technology.

[0095] In summary, the dual-stage offloading strategy of the present invention combined with the DPC-GCN model effectively solves the problem of limited processing power of terminal devices and realizes efficient and accurate human motion recognition. Experimental results show that this method significantly improves the accuracy and robustness of motion recognition while reducing computational latency, providing an innovative solution for human motion recognition applications in the field of edge computing.

[0096] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the above embodiments, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.

Claims

1. A two-stage edge computing task offloading method suitable for human motion recognition, characterized in that: The steps include: S1: Use the terminal to collect video images to generate human behavior recognition tasks. Instead of sending all computing tasks to the edge server or cloud for processing, the terminal will first complete a part of the computing tasks locally based on its own computing capabilities, and use the DQN agent to select the server to perform the remaining task calculations; S2: In the first stage of unloading, the DQN agent is used to collect the unloading environment information, and the environment data is standardized and input into the DQN agent to obtain the unloading behavior; Then, based on the dynamic adjustment of the server load and range, it is decided to offload to the selected server or the cloud. When the task is processed on the server, the DQN agent receives positive rewards, and when it is processed in the cloud or locally, the DQN agent receives negative rewards. S3: After the first stage of unloading is completed, the second stage of unloading is carried out, and the cumulative reward is calculated. Next, the identified action tasks are rated for danger and grouped accordingly. In each group, the task with the longest calculation time in the previous round is selected as the representative task. A PPO agent different from the DQN agent is used to collect the status of the representative tasks of each group and calculate the unloading behavior, that is, to determine the unloading percentage of the task; S4: After completing the second stage of unloading, the action recognition task is distributed according to the selected server and the unloading percentage. The action recognition task is completed through the dynamic point cloud graph convolutional network DPC-GCN model to obtain the required features and action recognition results.

2. The dual-stage edge computing task offloading method suitable for human motion recognition according to claim 1 is characterized in that: S2 specifically includes the following steps: S21: Based on the received environmental data, in the first round of iteration, the server closest to the terminal is selected as the initial offloading target for offloading. In subsequent rounds, the DQN agent makes autonomous selections, and tasks that exceed the server range or are larger than the load are offloaded to the cloud. S22: The reward obtained by the DQN agent is based on the Markov decision process, which calculates the reward value by unloading information in the environment; The system state S1 includes the total task computation delay of task i The transmission rate between the terminal and the server is r i,j , computing power is better than g i and the energy consumption of the terminal's i-th task e i , M represents the number of servers, and N represents the number of tasks; S23: Train the DQN agent, guide the DQN agent training through the objective function, and obtain the reward R1, where The total delay of task i in round t is calculated, z and b represent hyperparameters; Among them, R1 represents the reward function, and z1 represents the total calculation delay of the adjustment task The weight parameter of the influence on the reward function, z2 is used to adjust the weight parameter of the influence of the terminal energy consumption on the reward function, e i Including communication energy consumption and computing energy consumption, b1b2 is a bias term used to adjust the baseline value of the reward function to reflect different offloading strategies or environmental conditions; S24: The goal of the DQN agent is to maximize the cumulative reward, and the reward is calculated through the designed reward function R1, where the energy consumption e is set to a negative value, indicating that the reward obtained is the largest when the task delay is low and the energy consumption is small.

3. The dual-stage edge computing task offloading method suitable for human motion recognition according to claim 1 is characterized in that: S3 specifically includes the following steps: S31: Use a lightweight high-risk action prediction network to group the risk levels. In each group, select the task with the longest calculation time in the previous round as the representative task, and use the PPO agent to collect the status of the representative task. After round 0, calculate the target workload based on the output action of the PPO agent, and select the layer closest to the target workload as the decomposition layer; S32: Use a lightweight high-risk action prediction network HDN to predict the danger level D of each action, and divide all action categories into three levels according to the probability of danger. The high-risk action prediction network HDN adopts an efficient network implementation and calculates the task time of transmitting the danger level D to server j. Where l(D i ) represents the amount of data of the danger level, and the total delay of executing the HDN is composed of the calculation delay of the high-risk action prediction network HDN and the transmission time of the danger level; in represents the computational delay of the high-risk action prediction network HDN, that is, the time required to perform the risk level prediction on the terminal. Perform high-risk actions to predict the total delay of the network HDN; S33: The total computational delay and local computational load percentage of the representative task are input into the PPO agent as states, and the agent is trained to generate actions based on these states. The preprocessed actions are then assigned to each group as the next round of offloading decisions, and modeled based on Markov decision making: Where S2 is the computational latency of the representative task and the percentage of local computational work, is the previous round of unloading decision, G is the number of groups; S34: Each group of tasks executes a consistent offloading decision logic. During the execution of each task, when it reaches a specific decomposition level on the terminal, the intermediate results generated will be transmitted to the corresponding servers. These servers will be responsible for continuing to calculate the remaining levels of the task and feeding the results back to the system after the calculation is completed; S35: The server receives a reward R2 based on the delay in processing the task, which includes The local calculation delay of task i in round t and The total delay composition is calculated for task i in round t, providing feedback to the PPO agent to adjust future unloading decisions; in It refers to the computation delay of the i-th task completed locally on the terminal.

4. The dual-stage edge computing task offloading method suitable for human motion recognition according to claim 1 is characterized in that: S4 specifically includes the following steps: S41: In the human action recognition task, in order to perform action recognition and classification, the data that needs to be transmitted and processed, that is, the serialized data of the transmitted skeleton, is ensured to ensure the integrity and security of the data through the first layer multi-input branch of the designed dynamic point cloud convolutional network; S42: Next, the transmitted feature pair is input with feature f in Perform simple graph convolution SGC feature extraction, ns q s Represents the spatial features of node q in a certain state, (q, s) represents a pair of nodes, where q and s are two nodes in the graph, and the enhanced features ns of each node are obtained; S43: Use the attention mechanism to calculate the weight distribution m between nodes q , based on the features of adjacent nodes, calculate the weight α between nodes i,j , focus on important nodes, according to the weight α i,j , get the weighted node feature X q ; Where W * represents the weight matrix, which is used for linear transformation, b is the bias term, α q,s It is represented as the attention weight between node q and its neighbor node s, Represented as the weight vector in the attention mechanism, σ is the activation function, and δ is the conversion function; S44: After obtaining the node features, the node clustering is performed. Xq is input into the evaluation function ADconv, and the evaluation score φ is calculated for each node. The top C nodes with the highest evaluation scores are selected as cluster centers. i,j And the natural connection between nodes, assign the nodes to the corresponding clusters, and get the clustering result P, where W′, W″, W″′ are the attention weights; S45: The clustering result P is fed into the adaptive spatiotemporal attention module to perform spatial attention calculation, using the clustering result P and node feature f in , the spatial attention distribution f is obtained through attention calculation p , distribute the spatial attention of all clustering results P to f p Connect them together to get the overall spatial attention representation f i ' n ; f′ in =concat({f p ∣p=1,2,…,P}) S46: By calculating the spatial attention, the spatiotemporal joint attention extraction is performed, and the overall spatial attention f i ' n and the original node feature f in Perform spatiotemporal joint attention extraction to obtain the attention score of the entire skeleton sequence: f inner =θ * ([pool t (f′ in )||pool v (f in )]·W2) S47: Output the results obtained in S45 and S46 and use them as input for subsequent layers to perform further feature extraction and classification, and finally output the final feature map or classification result; S48: The received joint position, joint velocity and skeleton information are taken as input. After batch normalization, features are extracted and enhanced through a series of feature extraction layers, global graph convolutional network GCN layers and adaptive spatial-temporal attention ASTA modules. The behavior recognition results are output through feature connection, global average pooling and fully connected layers.