Edge iot device based on characteristic factor network dynamic programming method

By using the feature factor network dynamic programming method for edge IoT devices and leveraging reinforcement learning models to dynamically adjust network resources, the problem of insufficient resource allocation in traditional methods is solved, achieving efficient and secure network resource management that adapts to the ever-changing IoT environment.

CN119254641BActive Publication Date: 2026-01-02E SURFING IOT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411502127.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-25
Publication Date
2026-01-02
Estimated Expiration
2044-10-25

AI Technical Summary

Technical Problem

Traditional network planning methods cannot personalize network resource configuration according to the dynamic changes and business needs of IoT scenarios, resulting in low network resource utilization efficiency, high communication latency, and insufficient consideration of business characteristics, thus failing to meet the real-time needs of different scenarios.

Method used

A feature factor-based network dynamic planning method for edge IoT devices is adopted. By extracting general and business feature factors and combining them with a deep Q-network reinforcement learning model, the network resource allocation, including bandwidth, routing and transmission strategies, is dynamically adjusted to optimize the network configuration in real time.

Benefits of technology

It improves network resource utilization, enhances network security, reduces communication latency, adapts to complex and ever-changing IoT environments, possesses high versatility and adaptability, and reduces computational complexity and energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119254641B_ABST
    Figure CN119254641B_ABST
Patent Text Reader

Abstract

The application discloses an edge Internet-of-Things device-based network dynamic programming method based on characteristic factors, comprising the following steps: S1, extraction and weight proportioning of characteristic factors; S2, construction of a dynamic programming characteristic engineering model; S3, training of the dynamic programming characteristic engineering model; and S4, optimization of the dynamic programming characteristic engineering model. The application extracts and analyzes characteristic factors composed of business scene characteristics, adjusts the allocation strategy of network resources in real time according to different scene requirements, closely combines business requirements in the scene with network resource management, and enables the edge device to flexibly adjust the network strategy according to the scene change, so as to improve the utilization efficiency of network resources, enhance the security of data transmission, and reduce the communication delay.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of Internet of Things, in particular to a network dynamic planning method based on characteristic factors of edge Internet of Things equipment. BACKGROUND

[0002] With the rapid development of Internet of Things (IoT) technology, more and more edge devices are widely used in various scenarios, such as smart cities, smart healthcare, industrial Internet of Things, etc. The edge devices in these scenarios usually need to process a large amount of business data and transmit them efficiently through the network. However, the business characteristics in different scenarios are different, and the network demand and optimization target faced by edge devices are also different. Traditional network planning methods usually configure according to network conditions and rough existing resources, lack of fine-grained perception of business data, and cannot dynamically adjust network resource configuration according to actual business demand. In this context, the traditional static network planning method cannot effectively cope with the dynamic and variable Internet of Things environment.

[0003] The complexity of Internet of Things scenarios not only lies in the increase of device quantity and the growth of data volume, but also in the strong difference of network resource demand in different business scenarios. For example, in the smart healthcare scenario, sensitive data involving personal privacy needs to be transmitted through encrypted channels; while in the intelligent factory, high-concurrency production control data requires low delay and high bandwidth guarantee. This scenario-specific demand makes it difficult for a unified network planning strategy to meet the real-time needs of different business scenarios.

[0004] To solve this problem, existing network management methods try to introduce machine learning and other technologies to predict and optimize network status. However, these methods often rely on a large amount of historical data and pre-trained models, making it difficult to flexibly respond to rapidly changing Internet of Things environments. In addition, they also lack sufficient consideration of scenario business characteristics, and cannot allocate network resources individually according to specific business needs. SUMMARY

[0005] The network dynamic planning method based on characteristic factors of edge Internet of Things equipment provided by the present application closely combines business demand in different scenarios with network resource management to improve network resource utilization efficiency, enhance network security, and reduce communication delay, which can at least solve one of the above technical problems.

[0006] To solve the above technical problems, the present application adopts the following technical solution: a network dynamic planning method based on characteristic factors of edge Internet of Things equipment, comprising the following steps:

[0007] S1, extraction and weight distribution of characteristic factors: extract general characteristic factors and business characteristic factors, and distribute weights to the extracted characteristic factors to complete the scheduling assistance of network resources;

[0008] S2, constructing a dynamic programming feature engineering model: based on the feature factors obtained in S1, the dynamic programming feature engineering model is constructed, the dynamic programming feature engineering model adopts a deep Q network of a reinforcement learning model to process a network dynamic programming problem, adjusts and outputs an optimal network resource allocation strategy, and makes network resources dynamically adapt to business demands;

[0009] S3, training a dynamic programming feature engineering model: based on the dynamic programming feature engineering model obtained in S2, at least one or more of an experience replay mechanism, Q value updating, or a target network is used for model training;

[0010] S4, optimizing the dynamic programming feature engineering model: the feature factor data input into the dynamic programming feature engineering model is updated in real time, the dynamic programming feature engineering model re-plans network resource allocation under network state changes or business demand fluctuation states, and is gradually optimized and updated during operation.

[0011] Further, in S1, the general feature factor is used to represent general business indicators in a general business scenario, and at least includes a concurrent bandwidth adjustment factor, a response time adjustment factor, and a connection stability factor;

[0012] The concurrent bandwidth adjustment factor is composed of a business service concurrency number, a current bandwidth occupancy rate, and a service request rate. When the system detects an increase in the business service concurrency number, in combination with the current bandwidth occupancy rate and the service request rate, the system evaluates and allocates more bandwidth resources to high concurrency services in real time, ensures smooth transmission of data, and avoids congestion and delay;

[0013] The response time adjustment factor is combined with a business service response time, a current network delay, and a current load. When the business service response time exceeds a set threshold, the system extracts the features of the current network delay and the current load to determine whether the routing path needs to be adjusted or the bandwidth resources need to be increased, optimizes the network configuration, and reduces the delay of service response;

[0014] The connection stability factor is used to monitor the stability of the network and the reliability of the service connection, and is composed of a business service connection loss rate, a data packet loss rate, and a network transmission error rate. When the business service connection loss rate or the data packet loss rate is frequent, the system detects and adjusts the transmission strategy of the network to ensure the stable operation of the key business.

[0015] Further, in S1, the business-specific factor is used to represent specific business indicators in a specific business scenario, and at least includes a data sensitivity factor, a real-time demand factor, and a concurrent number factor of inter-backup communication;

[0016] The data sensitivity factor is used to determine whether the data needs to be transmitted through an encrypted channel;

[0017] The real-time requirement factor is used to control the low delay requirement of the signal to ensure the instant response of the business operation;

[0018] The concurrency factor of inter-site communication is used to determine the dynamic allocation requirement of bandwidth.

[0019] Further, the specific business scenario represented by the business feature factor at least includes an intelligent transportation scenario, an intelligent medical scenario, and an intelligent factory scenario, and the corresponding feature factors are a traffic flow adjustment factor, a medical data sensitivity factor, and a production load adjustment factor, respectively.

[0020] The traffic flow adjustment factor is composed of real-time traffic flow, vehicle density, and camera video data frame rate. In the intelligent transportation scenario, when the traffic flow increases, the system extracts this factor to dynamically adjust the bandwidth resources. The vehicle density and the camera video data frame rate jointly determine the demand for bandwidth. The system increases the bandwidth allocation to ensure the real-time performance and clarity of the video stream, and avoids delays or stalls in the monitoring picture due to insufficient bandwidth.

[0021] The medical data sensitivity factor is composed of patient data sensitivity level, real-time monitoring requirement, and data transmission index requirement. In the intelligent medical scenario, the system extracts this factor to evaluate the sensitivity of patient data. When the system detects that the patient data sensitivity level is high-sensitivity data, the medical data sensitivity factor triggers an encrypted transmission strategy, and according to the real-time monitoring requirement, high-bandwidth and low-delay network resources are preferentially allocated to ensure that the data transmission index requirement meets the standard.

[0022] The production load adjustment factor is combined with line equipment load rate, real-time production data volume, and operation task priority. In the intelligent factory scenario, when the line equipment load rate increases, the system extracts this factor to evaluate the real-time load condition of the line equipment and the real-time production data volume, to ensure that the high-level key tasks in the operation task priority meet the high-bandwidth and low-delay network resource requirements.

[0023] Further, in the S1, the weight ratio of the feature factors is optimized by a multi-objective optimization algorithm based on historical data.

[0024] Further, the S2 further includes:

[0025] S21, data cleaning: before constructing the dynamic programming feature engineering model, the extracted feature factors are preprocessed to ensure the accuracy and consistency of the data.

[0026] The operation of data cleaning at least includes missing data processing, abnormal data filtering and duplicate data removal, wherein the missing data processing is to fill in the missing data using the average value or the median, and the abnormal data filtering and the duplicate data removal are to identify and filter outlier data through z-score or clustering-based method;

[0027] S22, model construction: the core of the dynamic programming feature engineering model is network dynamic programming based on the feature factors, and a reinforcement learning model deep Q network is used to process the network dynamic programming problem, and the network dynamic programming includes state space definition, action space definition and reward mechanism design;

[0028] The dynamic programming feature engineering model outputs an optimal network resource allocation strategy based on the input feature factors, and the network resource allocation strategy includes bandwidth allocation, routing path selection and transmission strategy adjustment.

[0029] Further, in the S22, in the network dynamic programming scenario, the state space is composed of a plurality of feature factors including the general feature factors and the service feature factors, and a specific state can be represented as a vector including at least the current network condition, network resource allocation and service demand information;

[0030] The action space defines various network resource allocation strategies that the system can take, that is, for each state, the reinforcement learning model deep Q network needs to select one of the network resource allocation strategies and give an action;

[0031] The reward mechanism, as a core part of the network dynamic programming, defines the feedback brought by each action, and in the network dynamic programming scenario, the reward mechanism is designed according to the following indexes:

[0032] ① If the selected action improves the service response speed, reduces the delay or reduces the packet loss rate, a positive reward is given;

[0033] ② If the selected action optimizes the bandwidth utilization or reduces the resource waste, a positive reward is given;

[0034] ③ In sensitive data transmission, ensuring the secure transmission of data can obtain an additional reward.

[0035] Further, in the S22, the bandwidth allocation refers to dynamically adjusting the bandwidth allocation strategy between edge devices according to the bandwidth demand of each service in the network, and preferentially increasing the bandwidth allocation for critical tasks;

[0036] The routing path selection refers to selecting a low-delay or high-security routing path according to the weight of the feature factors to ensure that data is transmitted in an optimal manner;

[0037] The transmission strategy refers to deciding whether to enable encrypted transmission or selecting other suitable transmission modes to ensure data security and transmission efficiency.

[0038] Further, in the S3, the experience replay mechanism refers to storing the tuple of state, action, reward and next state in the experience replay pool after each action execution;

[0039] Q value updating refers to randomly extracting a batch of data from the experience replay pool for training each time, and updating the Q value function through the Q-learning algorithm;

[0040] The target network refers to two neural networks used by the deep Q network of the reinforcement learning model, one being the main network responsible for current Q value prediction, and the other being the target network periodically updated to stabilize the training process.

[0041] Further, the S4 further comprises:

[0042] S41, real-time monitoring: the system acquires the characteristic factor data of the current network in real time through the sensors or monitoring tools deployed on the edge devices, and inputs them into the dynamic programming feature engineering model;

[0043] S42, real-time decision-making: after sufficient training of the dynamic programming feature engineering model, the deep Q network of the reinforcement learning model receives the current state information in real time, and selects the optimal action based on the learned strategy; when the system detects changes in network state or business demand fluctuations, the deep Q network of the reinforcement learning model quickly adjusts the network resource allocation to adapt to the new environment;

[0044] S43, self-updating: when significant changes in network environment or business demand are detected, the system will automatically re-plan the network resource allocation according to the new characteristic factor weight, and gradually optimize the dynamic programming feature engineering model through the online learning algorithm adaptive gradient method during operation;

[0045] S44, feedback mechanism: the system compares the results of each network planning with the expected target, and continuously adjusts the model parameters through the feedback mechanism to ensure that the allocation of network resources is always optimal.

[0046] The beneficial effects of the present application are:

[0047] 1. Improve network resource utilization: by dynamically adjusting bandwidth, routing and transmission strategy, the present application can reasonably allocate network resources according to actual business needs, maximize network utilization efficiency and reduce congestion and delay.

[0048] 2. Enhance network security: The application can automatically select an encrypted channel for data transmission by identifying sensitive information in business data, ensuring the security of data during transmission, especially suitable for privacy protection scenarios such as medical, financial, etc.

[0049] 3. Reduce communication delay and improve service quality: For high concurrency or real-time requirement of business scenarios, the application can dynamically allocate bandwidth and optimize routing according to the characteristic factor, guarantee the low delay transmission of key data, and improve the user experience and service quality of business.

[0050] 4. Strong universality and adaptability: The method is suitable for different Internet of Things scenarios and edge devices, with high universality and adaptability. Through real-time sensing of business characteristics and flexible network planning, the application can adapt to complex and variable Internet of Things environment.

[0051] 5. Reduce energy consumption and computing overhead: By accurately extracting and selecting key characteristic factors, the application reduces the computational complexity and communication overhead of the system under the premise of ensuring network performance, suitable for resource-constrained edge devices. BRIEF DESCRIPTION OF DRAWINGS

[0052] The drawings described herein are used to provide further understanding of the present application, and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation on the present application.

[0053] Figure 1 is a network dynamic planning method flowchart of an embodiment of the application.

[0054] Figure 2 is a network dynamic planning method operation flowchart of an embodiment of the application.

[0055] Figure 3 is a structural block diagram of a computer device of an embodiment of the application. DETAILED DESCRIPTION

[0056] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only a part of the embodiments of the application, not all the embodiments. The embodiments in the application and the features in the embodiments can be combined with each other without conflict. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.

[0057] It should be noted that the meaning of "and / or" appearing in the entire text includes three parallel schemes, for example, "A and / or B" includes A scheme, or B scheme, or A and B scheme.

[0058] It should be noted that those skilled in the art can understand that all or part of the steps implemented in the embodiments of the present application can be implemented by software, hardware, firmware or any combination thereof. When implemented by hardware, it can be implemented in the form of purchased standard components or modified components, in whole or in part. When implemented by software, it can be implemented in the form of a computer program product, in whole or in part. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (for example, floppy disk, hard disk, magnetic tape), optical media (for example, DVD), or semiconductor media (for example, solid state disk (SSD)) and the like.

[0059] Referring to Figures 1-2 The embodiments of the present application provide a network dynamic programming method based on characteristic factors of edge Internet of Things equipment, comprising the following steps:

[0060] S1, extraction and weight distribution of characteristic factors: extracting general characteristic factors and business characteristic factors, and distributing weights to the extracted characteristic factors, to complete the scheduling assistance of network resources;

[0061] S2, constructing a dynamic programming feature engineering model: constructing the dynamic programming feature engineering model based on the characteristic factors obtained in S1, the dynamic programming feature engineering model uses a deep Q network of reinforcement learning model to process network dynamic programming problems, adjusts and outputs the optimal network resource allocation strategy, so that the network resources dynamically adapt to business demands;

[0062] S3, training a dynamic programming feature engineering model: based on the dynamic programming feature engineering model obtained in S2, at least one or more of the following is used for model training: experience replay mechanism, Q value update, or target network;

[0063] S4, optimizing the dynamic programming feature engineering model: real-time updating the feature factor data input into the dynamic programming feature engineering model, the dynamic programming feature engineering model re-plans network resource allocation under network state change or business demand fluctuation state, and gradually optimizes and updates in the running process.

[0064] The application provides a network dynamic programming method based on feature factors, first, according to the business characteristics in different Internet of Things scenarios, the core general feature factors and business feature factors are extracted, these factors can dynamically reflect the current business demand of the scene through the weight distribution ratio, guiding the system to perform network planning and resource allocation according to the specific business demand, secondly, through real-time analysis of the business scene feature factors, the system allocates network resources such as bandwidth allocation, routing selection and transmission channel encryption according to the dynamic adjustment mechanism, in this way, the planning method of the application can efficiently run on the edge device with limited resources, and has strong adaptive ability, can adjust the network planning in real time according to the changes of network state and business demand, and ensure that the system can maintain efficient and stable operation in the dynamic Internet of Things environment.

[0065] In this embodiment, in S1, the general feature factor is used to represent the general business indicators in the general business scene, and at least includes a concurrent bandwidth adjustment factor, a response time adjustment factor and a connection stability factor;

[0066] The concurrent bandwidth adjustment factor is composed of business service concurrency, current bandwidth occupancy rate and service request rate, when the system detects that the business service concurrency increases, in combination with the current bandwidth occupancy rate and the service request rate, the system evaluates and allocates more bandwidth resources to the high concurrency service in real time, ensures the smooth transmission of data, and avoids congestion and delay;

[0067] The response time adjustment factor is combined with business service response time, current network delay and current load, when the business service response time exceeds the set threshold, the system extracts the features of the current network delay and the current load, judges whether it is necessary to adjust the routing path or increase the bandwidth resources, optimizes the network configuration, and reduces the delay of service response;

[0068] The connection stability factor is used to monitor the stability of the network and the reliability of the service connection, which consists of service connection loss rate, packet loss rate and network transmission error rate. When the service connection loss rate or the packet loss rate is frequent, the system detects and adjusts the transmission strategy of the network, such as re-routing, increasing redundant paths or improving connection priority, to ensure the stable operation of critical services.

[0069] In this application, the general characteristic factor can be applied to the network resource demand of multiple scenes, such as network delay, bandwidth occupation, device load, packet loss rate, etc. These characteristics are usually used to measure the basic indicators of network performance, and have high universality in all scenes. By monitoring these indicators in real time, the system can master the basic operating conditions of the network.

[0070] In this embodiment, the S1, the service characteristic factor is used to represent specific business indicators in specific business scenarios, including at least data sensitivity factor, real-time demand factor and concurrent number of inter-backup communication factor;

[0071] The data sensitivity factor is used to determine whether the data needs to be transmitted through an encrypted channel;

[0072] The real-time demand factor is used to control the low delay demand of the signal to ensure the immediate response of the business operation;

[0073] The concurrent number of inter-backup communication factor is used to determine the dynamic allocation demand of bandwidth;

[0074] The specific business scenarios represented by the service characteristic factor include at least intelligent transportation scenarios, intelligent medical scenarios and intelligent factory scenarios, and the corresponding characteristic factors are traffic flow adjustment factor, medical data sensitivity factor and production load adjustment factor;

[0075] The traffic flow adjustment factor consists of real-time traffic flow, vehicle density and camera video data frame rate. In the intelligent transportation scenario, when the traffic flow increases, the system extracts this factor to dynamically adjust the bandwidth resources. Vehicle density and camera video data frame rate jointly determine the demand for bandwidth. The system increases bandwidth allocation to ensure the real-time performance and clarity of video streams, and avoids delays or stalls in the monitoring picture due to insufficient bandwidth;

[0076] The medical data sensitivity factor is composed of patient data sensitivity level, real-time monitoring requirement and data transmission index requirement. In the intelligent medical scene, the system extracts the factor to evaluate the sensitivity of patient data. When the system detects that the patient data sensitivity level is high-sensitivity data (such as personal health data or surgical information), the medical data sensitivity factor triggers the encrypted transmission strategy, and according to the real-time monitoring requirement (such as remote surgery or health monitoring), high bandwidth and low delay network resources are preferentially allocated to ensure that the data transmission index requirement meets the standard.

[0077] The production load adjustment factor is combined with line equipment load rate, real-time production data volume and operation task priority. In the intelligent factory scene, when the line equipment load rate increases, the system extracts the factor to evaluate the real-time load condition of the line equipment and the real-time production data volume, and ensures that the high-level key tasks in the operation task priority meet the network resource demand of high bandwidth and low delay. The system can automatically adjust the network resources to ensure high-speed communication between the line equipment and avoid production interruption or delay.

[0078] In this application, the business feature factor is different from the above-mentioned general feature factor. It is extracted according to the demand in a specific business scene. Common specific business scenes include the intelligent traffic scene, the intelligent medical scene and the intelligent factory scene mentioned above, and other business scenes should also be included. These business feature factors change dynamically and determine the specific planning strategy of the network according to the demand of different business scenes.

[0079] In this application, the extraction of the feature factor is the basis of network dynamic planning, which ensures that the system can make decisions according to different feature factors in the business scene.

[0080] Referring to Figure 2 In this embodiment, in the S1, the weight ratio of the feature factor is adjusted by a multi-objective optimization algorithm based on a historical data weighted linear model. The extracted general feature factor and the business feature factor need to be weighted to reflect their relative importance to network resource scheduling. The weight ratio is adjusted according to different business scenes. For example, in a scene with high real-time requirement, the weight of delay feature will be higher than that of bandwidth utilization, and in a high data sensitivity scene, the weight of encrypted transmission will be increased.

[0081] Referring to Figures 1-2 In this embodiment, the S2 further includes:

[0082] S21, data cleaning: before constructing the dynamic planning feature engineering model, the extracted feature factors are preprocessed to ensure the accuracy and consistency of the data;

[0083] The operation of data cleaning at least includes missing data processing, abnormal data filtering and duplicate data removal, wherein the missing data processing is to fill in the missing data using the average value or the median, and the abnormal data filtering and the duplicate data removal are to identify and filter outlier data by z-score or clustering-based method;

[0084] S22, model construction: the core of the dynamic programming feature engineering model is based on the network dynamic programming of the feature factor, and a reinforcement learning model deep Q network (DQN) is used to process the network dynamic programming problem. The network dynamic programming includes state space definition, action space definition and reward mechanism design;

[0085] Further, regarding the state space definition: in the scenario of network dynamic programming, the state space is composed of multiple feature factors including the general feature factors (such as concurrent bandwidth adjustment factor, response delay adjustment factor and connection stability factor) and the service feature factors (such as traffic flow adjustment factor, medical data sensitivity factor and production load adjustment factor), and a specific state can be represented as a vector containing at least the current network condition, network resource allocation and service demand information;

[0086] Regarding the action space definition: the action space defines various network resource allocation strategies that the system can take, i.e. for each state, the reinforcement learning model deep Q network (DQN) needs to select one of the network resource allocation strategies and assign an action to it;

[0087] Regarding the reward mechanism design: the reward mechanism, as a core part of network dynamic programming, defines the feedback brought by each action. In the scenario of network dynamic programming, the reward mechanism is designed according to the following indicators:

[0088] ① If the selected action improves the service response speed, reduces the delay or reduces the packet loss rate, a positive reward is given;

[0089] ② If the selected action optimizes the bandwidth utilization or reduces resource waste, a positive reward is given;

[0090] ③ In sensitive data transmission, ensuring the secure transmission of data can obtain additional rewards;

[0091] The dynamic programming feature engineering model outputs the optimal network resource allocation strategy based on the input feature factors, and the network resource allocation strategy includes bandwidth allocation, routing path selection and transmission strategy adjustment;

[0092] Further, the bandwidth allocation refers to dynamically adjusting the bandwidth allocation strategy between edge devices according to the bandwidth demand of each service in the network, and preferentially increasing bandwidth allocation for critical tasks;

[0093] The route path selection refers to selecting a route path with low latency or high security according to the weight of the characteristic factor, to ensure that data is transmitted in an optimal manner;

[0094] The transmission strategy refers to deciding whether to enable encrypted transmission, or selecting other suitable transmission methods to ensure data security and transmission efficiency.

[0095] Referring to Figure 2 In this embodiment, in the S3, the experience replay mechanism refers to storing a tuple of state, action, reward and next state in the experience replay pool after each execution of an action, which enables the dynamic programming feature engineering model to learn from past experiences, thereby breaking the correlation between data and improving learning efficiency;

[0096] The Q value update refers to randomly extracting a batch of data from the experience replay pool for training each time, and updating the Q value function through the Q-learning algorithm. The DQN uses a deep neural network to approximate the Q value function, and the updating process includes:

[0097]

[0098] Wherein, s is the current state, a is the current action, r is the immediate reward, s' is the next state, alpha is the learning rate, and gamma is the discount factor;

[0099] The target network refers to two neural networks used by the deep Q network of the reinforcement learning model, one being the main network responsible for current Q value prediction, and the other being the target network periodically updated to stabilize the training process. This mechanism helps to improve the convergence of the dynamic programming feature engineering model.

[0100] Referring to Figures 1-2 In this embodiment, the S4 further includes:

[0101] S41, real-time monitoring: the system acquires the characteristic factor data of the current network in real time through sensors or monitoring tools deployed on edge devices, and inputs them into the dynamic programming feature engineering model. The specific model parameters of the selected sensors or monitoring tools are not specially limited in this embodiment, as long as they can achieve the collection of characteristic factor data;

[0102] S42, real-time decision making: after being fully trained, the dynamic programming feature engineering model receives current state information in real time through the deep Q network (DQN) of the reinforcement learning model, and selects the optimal action based on the learned strategy. When the system detects changes in network state or fluctuations in business demand, the deep Q network of the reinforcement learning model quickly adjusts the allocation of network resources to adapt to the new environment, thereby maintaining the high quality and stability of services;

[0103] S43, self-updating: when significant changes in network environment or business requirements are monitored (such as bandwidth tension or rising security requirements), the system will automatically re-plan network resource allocation according to new characteristic factor weights, and gradually optimize the dynamic programming feature engineering model during operation through an online learning algorithm, Adagrad, adaptive gradient method;

[0104] S44, feedback mechanism: the system compares the results of each network planning with the expected target, and continuously adjusts the model parameters through the feedback mechanism to ensure that the allocation of network resources is always optimal.

[0105] In the following, the present application will be described in detail in combination with a plurality of specific business scenarios provided in the above embodiments:

[0106] Example one, intelligent transportation scenario:

[0107] In the intelligent city traffic monitoring system, the edge device needs to process a large amount of real-time video stream data from the road camera, and the transmission of these data must ensure low delay and high bandwidth to realize the real-time and accuracy of monitoring.

[0108] Through the dynamic programming method of the present application, the system can adjust the allocation of network resources in real time according to the changes of current traffic flow. Specifically, when the traffic flow is large, the real-time requirement of monitoring video data increases, the system will automatically identify and extract traffic flow and video transmission delay as key characteristic factors, and based on these key characteristic factors, the system will preferentially allocate more network bandwidth to ensure smooth transmission and processing of monitoring data; on the contrary, when the traffic flow is small, the system will automatically reduce the allocation of bandwidth resources, thereby reducing network load and saving resources. This dynamic adjustment enables the traffic monitoring system to work efficiently under different traffic conditions and optimizes resource use.

[0109] Example two, intelligent medical scenario:

[0110] In the intelligent medical scenario, the edge device is responsible for processing and transmitting a large amount of patient health data, which is sensitive and has high requirements for real-time transmission.

[0111] Through the dynamic planning method of the application, the system can dynamically allocate network resources based on extracted feature factors (such as data sensitivity, delay requirements, etc.) in different medical service scenarios. Specifically, when the system detects that the patient data to be transmitted belongs to sensitive information, it automatically enables an encrypted transmission channel to ensure data privacy and security. At the same time, in emergency medical scenarios such as remote surgery or real-time monitoring, the system will preferentially allocate high-bandwidth and low-latency network resources to ensure timely transmission of critical medical information. Conversely, in ordinary examination data or non-sensitive scenarios, the system reduces the occupation of network resources and optimizes bandwidth allocation, so that the medical system can still achieve efficient resource management and transmission protection under the premise of ensuring patient privacy and security.

[0112] Embodiment three, intelligent factory scenario:

[0113] In the production process of an intelligent factory, edge devices need to transmit a large amount of production status data in real time, and the system needs to dynamically adjust network resources as the factory operation changes.

[0114] Through the dynamic planning method of the application, the system can extract feature factors such as device load and data sensitivity according to different scenarios of production load. Specifically, when the factory is in a high-load production phase, the demand for data communication between devices increases, and the system will identify the priority of low latency and high bandwidth and allocate more network resources to ensure the stable operation of the production control system. At the same time, if the data processed by the device involves sensitive information during the production process, the system will automatically select an encrypted channel for transmission to ensure the security of the data. Conversely, in the case of non-high-load or non-sensitive data transmission, the system reduces bandwidth allocation to reduce energy consumption, thereby improving the overall communication efficiency and security of the factory.

[0115] The application also provides a computer readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the above method.

[0116] Referring to Figure 3 The application also provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the computer program, when executed by the processor, causes the processor to perform the steps of the above method.

[0117] The application also provides a computer program product comprising instructions which, when executed on a computer, cause the computer to perform the steps of the above method.

[0118] It can be understood that the system, device and storage medium provided by the embodiments of the present application correspond to the method provided by the embodiments of the present application, and the explanation, examples and beneficial effects of the related content can refer to the corresponding part in the above method.

[0119] In summary, the present application proposes a network dynamic programming method based on feature factors, which is specifically used for network operation and management of edge devices in Internet of Things scenarios. By extracting and analyzing feature factors in different business scenarios, dynamic adjustment and optimization of network resources are achieved. The present application aims to solve the problem of lack of perception and individualized scheduling of specific business scenario requirements in existing network planning methods, so that edge devices can flexibly adjust network strategies according to scenario changes to improve the utilization efficiency of network resources, enhance the security of data transmission and reduce communication delay.

[0120] Specifically, the present method uses feature factors for dynamic programming, and the system can automatically identify business requirements in different scenarios, such as enabling encrypted channels in data-sensitive scenarios, automatically adjusting bandwidth in high-concurrency scenarios, and prioritizing critical data transmission in low-latency scenarios. This method not only improves the adaptability of network planning, but also efficiently runs on resource-limited edge devices to meet the dynamic needs of complex business scenarios.

[0121] It should be understood that the examples and embodiments described herein are only for illustration and do not limit the present application, and those skilled in the art can make various modifications or changes according to it, any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A network dynamic programming method for edge IoT devices based on characteristic factors, characterized in that, The method comprises the following steps: S1, extraction and weight distribution of characteristic factors: extracting general characteristic factors and business characteristic factors, and distributing weights to the extracted characteristic factors to complete the scheduling assistance of network resources; S2, constructing a dynamic programming feature engineering model: constructing the dynamic programming feature engineering model based on the characteristic factors obtained in S1, using a deep Q network of a reinforcement learning model to process network dynamic programming problems, adjusting and outputting an optimal network resource allocation strategy to make network resources dynamically adapt to business demands; S3, training the dynamic programming feature engineering model: based on the dynamic programming feature engineering model obtained in S2, at least one or more of the experience replay mechanism, Q value update, or target network is used for model training; S4, optimizing the dynamic programming feature engineering model: real-time updating the characteristic factor data input into the dynamic programming feature engineering model, and the dynamic programming feature engineering model re-plans network resource allocation under network state changes or business demand fluctuations, and gradually optimizes and updates during operation; In S1, the general characteristic factors are used to represent general business indicators in general business scenarios, and at least include a concurrent bandwidth adjustment factor, a response time delay adjustment factor, and a connection stability factor; The concurrent bandwidth adjustment factor is composed of the number of concurrent services, the current bandwidth occupancy rate, and the service request rate; The response time delay adjustment factor is combined with the business service response time, the current network delay, and the current load; The connection stability factor is used to monitor the stability of the network and the reliability of the service connection, and is composed of the business service connection loss rate, the data packet loss rate, and the network transmission error rate; In S1, the business characteristic factors are used to represent specific business indicators in specific business scenarios, and at least include a data sensitivity factor, a real-time demand factor, and a concurrent number factor of inter-backup communication; The data sensitivity factor is used to determine whether the data needs to be transmitted through an encrypted channel; The real-time demand factor is used to control the low-delay demand of the signal to ensure the immediate response of the business operation; The concurrent number factor of inter-backup communication is used to judge the dynamic allocation demand of bandwidth; The specific business scenarios represented by the business characteristic factors at least include intelligent transportation scenarios, intelligent medical scenarios, and intelligent factory scenarios, and the corresponding characteristic factors are traffic flow adjustment factors, medical data sensitivity factors, and production load adjustment factors; The traffic flow adjustment factor is composed of real-time traffic flow, vehicle density, and camera video data frame rate; The medical data sensitivity factor is composed of patient data sensitivity level, real-time monitoring demand, and data transmission index requirement; The production load adjustment factor is combined with the line equipment load rate, real-time production data volume, and operation task priority. 2.The edge IoT device network dynamic planning method based on feature factors according to claim 1, wherein, In S1, the weight distribution of the characteristic factors is optimized by a multi-objective optimization algorithm weighting linear model based on historical data. 3.The edge IoT device network dynamic planning method based on feature factors according to claim 1, wherein, The S2 further comprises: S21, data cleaning: before constructing the dynamic planning feature engineering model, the extracted feature factors are preprocessed to ensure the accuracy and consistency of the data; The operation of data cleaning at least includes missing data processing, abnormal data filtering and repeated data removal, wherein the missing data processing is to fill in the missing data using the average value or the median value, and the abnormal data filtering and repeated data removal are to identify and filter outlier data through z-score or clustering-based method; S22, model construction: the core of the dynamic planning feature engineering model is network dynamic planning based on the feature factors, and a deep Q network of reinforcement learning model is used to process the network dynamic planning problem, which includes state space definition, action space definition and reward mechanism design; The dynamic planning feature engineering model outputs the optimal network resource allocation strategy based on the input feature factors, which includes bandwidth allocation, routing path selection and transmission strategy adjustment. 4.The edge IoT device network dynamic planning method based on feature factors according to claim 3, wherein, In the S22, in the scenario of network dynamic planning, the state space is composed of multiple feature factors including the general feature factors and the service feature factors, and a specific state can be represented as a vector containing at least the current network condition, network resource allocation and service demand information; The action space defines various network resource allocation strategies that the system can take, i.e. for each state, the deep Q network of reinforcement learning model needs to select one of the network resource allocation strategies and assign an action to it; The reward mechanism, as the core part of network dynamic planning, defines the feedback of each action, and in the scenario of network dynamic planning, the reward mechanism is designed according to the following indicators: (1) if the selected action improves the service response speed, reduces the delay or reduces the packet loss rate, a positive reward is given; (2) if the selected action optimizes the bandwidth utilization or reduces the resource waste, a positive reward is given; (3) in sensitive data transmission, ensuring the secure transmission of data can obtain additional rewards. 5.The edge IoT device network dynamic planning method based on feature factors according to claim 3, wherein, In the S22, the bandwidth allocation refers to dynamically adjusting the bandwidth allocation strategy between edge devices according to the bandwidth demand of each service in the network, and preferentially increasing the bandwidth allocation for critical tasks; The routing path selection refers to selecting a low-delay or high-security routing path according to the weight of the feature factors to ensure that the data is transmitted in the optimal way; The transmission strategy refers to deciding whether to enable encrypted transmission or selecting other appropriate transmission methods to ensure data security and transmission efficiency. 6.The edge IoT device network dynamic planning method based on feature factors according to claim 4, wherein, In the S3, the experience replay mechanism refers to storing the tuple of state, action, reward and next state in the experience replay pool after each action is performed; Q value update refers to randomly extracting a batch of data from the experience replay pool for training each time, and updating the Q value function through the Q-learning algorithm; The target network refers to the two neural networks used by the deep Q network of reinforcement learning model, one is the main network responsible for current Q value prediction, and the other is the target network used for regular updating to stabilize the training process. 7.The edge IoT device network dynamic planning method based on feature factors according to claim 1, wherein, The S4 further comprises: S41, real-time monitoring: the system collects the current network characteristic factor data in real time through sensors or monitoring tools deployed on edge devices and inputs them into the dynamic programming feature engineering model; S42, real-time decision-making: after sufficient training, the dynamic programming feature engineering model receives current state information in real time through the reinforcement learning model deep Q network and selects the optimal action based on the learned strategy. When the system detects changes in network status or fluctuations in business demand, the reinforcement learning model deep Q network quickly adjusts the allocation of network resources to adapt to the new environment; S43, self-updating: when significant changes in network environment or business demand are detected, the system will automatically re-plan the allocation of network resources according to new characteristic factor weights, and gradually optimize the dynamic programming feature engineering model through online learning algorithms and adaptive gradient methods during operation; S44, feedback mechanism: the system compares the results of each network planning with the expected target and continuously adjusts the model parameters through the feedback mechanism to ensure that the allocation of network resources is always optimal.

Citation Information

Patent Citations

  • Network resource selection method and device based on deep Q network, and storage medium

    CN113015179A

  • Service scheduling deployment method for edge computing node resources of Internet of Things

    CN114077485A