Task allocation method and device and electronic equipment
By receiving tasks in the edge computing network, obtaining and computing the operating data of edge agent nodes, and using deep reinforcement learning model to generate strategies, the tasks are allocated to the most suitable nodes, solving the problem of inaccurate resource allocation under the fixed strategy method, and achieving efficient and flexible task allocation and network efficiency improvement.
Patent Information
- Application Number
- CN202510435118.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-29
AI Technical Summary
When using fixed strategy methods to allocate node tasks in the prior art, resource allocation accuracy and flexibility are low, making it difficult to cope with dynamic and complex environmental changes in edge computing networks.
By receiving the tasks to be executed, determining the running data of the initial edge proxy node, calculating feature data, and using the policy generation model to generate target policies, sending the task to the most suitable edge proxy node for execution, the policy generation model is obtained through deep reinforcement learning training.
It improves the accuracy and flexibility of task allocation, ensures that tasks can be processed efficiently and in a timely manner, and improves the overall operation efficiency and service quality of edge computing networks.
Smart Images

Figure CN120386624A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular, to a task allocation method, apparatus, and electronic device. Background Art
[0002] In the fields of edge computing and cloud computing, resource scheduling is a key link to ensure the efficient execution of tasks and improve service quality. Traditional resource scheduling methods, such as First-Come-First-Served (FCFS) and Round Robin, queue and allocate tasks based on preset rules. The FCFS strategy is simple and intuitive, processing tasks in the order of arrival; Round Robin scheduling ensures fairness between tasks by cyclically allocating processing time to avoid a certain task occupying resources for a long time. These methods have been widely used in early computing systems due to their simplicity of implementation and low computational complexity.
[0003] However, although traditional scheduling methods perform well in simple static environments, they show obvious limitations in modern rapidly developing edge computing networks. The current edge computing network faces a dynamic and complex environment, with diverse task types and rapidly changing resource states, which poses higher requirements for the intelligence, flexibility, and real-time performance of scheduling algorithms. Fixed-strategy scheduling methods lack the ability to perceive environmental dynamic changes and self-adjustment mechanisms, making it difficult to effectively cope with fluctuations in resource requirements and task suddenness, resulting in a series of problems such as uneven resource allocation, slow task response, and degradation of service quality and user experience.
[0004] Regarding the problem of low resource allocation accuracy and flexibility in node task allocation using fixed-strategy methods in related technologies, no effective solution has been proposed yet. Summary of the Invention
[0005] The present application provides a task allocation method, apparatus, and electronic device to solve the problem of low resource allocation accuracy and flexibility in node task allocation using fixed-strategy methods in related technologies.
[0006] According to one aspect of the present application, a task allocation method is provided. The method includes: receiving a task to be executed, and determining M initial edge agent nodes, where the initial edge agent nodes are edge agent nodes with the execution ability to execute the task to be executed, and M is a positive integer; obtaining the operation data of each initial edge agent node to obtain M sets of operation data, and calculating the feature data of each initial edge agent node according to each set of operation data to obtain M sets of feature data; inputting the M sets of feature data and the task to be executed into a policy generation model to obtain the target policy output by the policy generation model; and sending the task to be executed to the target edge agent node indicated by the target policy, where the target edge agent node is used to execute the task to be executed.
[0007] Optionally, determining the M initial edge proxy nodes includes: obtaining the task type and task level of the task to be executed; screening multiple preset edge proxy nodes according to the task level and task type to obtain the M initial edge proxy nodes.
[0008] Optionally, calculating the feature data of each initial edge proxy node according to each group of operation data to obtain M groups of feature data includes: for any group of operation data, grouping each operation data in the group of operation data according to the time stamp to obtain a plurality of operation data sets; obtaining the calculation formulas of each initial feature, and substituting the data in the plurality of operation data sets into the calculation formulas to obtain the feature data of each initial feature; calculating the correlation between each initial feature and the duration of the initial edge proxy node to process the task, and determining the feature data of the initial feature with the correlation greater than the preset correlation as a group of feature data.
[0009] Optionally, the policy generation model is trained in the following manner: obtaining the processing duration of each sample edge proxy node in the sample edge proxy node set at N historical moments, and the feature data of each sample edge proxy node, where N is a positive integer; using the processing duration of each sample edge proxy node to process the sample task and the feature data of each sample edge proxy node at each historical moment as a group of sample data to obtain N groups of sample data; iteratively training the deep reinforcement learning model with each group of sample data in turn to obtain the policy generation model.
[0010] Optionally, obtaining the operation data of each initial edge proxy node to obtain M groups of operation data includes: for any initial edge proxy node, obtaining the initial operation data of the initial edge proxy node, and determining whether there are outliers in the initial operation data; in the case of the existence of outliers, deleting the outliers, and obtaining all missing values in the initial operation data; determining the data types of the missing values, and calculating the missing values according to the other operation data under the data types to obtain filling values, and using the filling values to replace the missing values in the initial operation data to obtain the updated operation data; performing a normalization operation on the updated operation data to obtain the operation data of the initial edge proxy node.
[0011] Optionally, after sending the task to be executed to the target edge proxy node indicated by the target policy, the method further includes: obtaining the execution result of the target edge proxy node to execute the task to be executed, and determining whether the execution result is successful; in the case of the execution result being a failure, obtaining the feature data of each initial edge proxy node, and regenerating the execution policy of the task to be executed through the policy generation model; reselecting the edge proxy node to execute the task to be executed according to the execution policy.
[0012] Optionally, obtaining the operation data of each initial edge proxy node includes: identifying the interface information of the initial edge proxy node; determining a data transmission decryption policy according to the interface information; and decrypting the transmission data through the data transmission decryption policy to obtain the operation data when the transmission data sent by the initial edge proxy node is received.
[0013] According to another aspect of the present application, a task allocation device is provided. The device includes: a receiving unit, configured to receive a task to be executed and determine M initial edge proxy nodes, where the initial edge proxy node is an edge proxy node having the execution ability to execute the task to be executed, and M is a positive integer; a first obtaining unit, configured to obtain the operation data of each initial edge proxy node to obtain M sets of operation data, and calculate the characteristic data of each initial edge proxy node according to each set of operation data to obtain M sets of characteristic data; a first generating unit, configured to input the M sets of characteristic data and the task to be executed into a policy generation model to obtain a target policy output by the policy generation model; and a sending unit, configured to send the task to be executed to a target edge proxy node indicated by the target policy, where the target edge proxy node is used to execute the task to be executed.
[0014] According to another aspect of the present invention, a computer program product is further provided, including a computer program, where the computer program, when executed by a processor, implements a task allocation method provided in the foregoing embodiments of the present application.
[0015] According to another aspect of the present invention, an electronic device is further provided, including one or more processors and a memory; the memory stores computer-readable instructions, and the processor is configured to run the computer-readable instructions, where the computer-readable instructions, when running, execute a task allocation method provided in the foregoing embodiments.
[0016] Through the present application, the following steps are adopted: receiving a task to be executed, and determining M initial edge agent nodes, where the initial edge agent nodes are edge agent nodes with the execution ability to execute the task to be executed, and M is a positive integer; obtaining the operation data of each initial edge agent node, obtaining M sets of operation data, and calculating the characteristic data of each initial edge agent node according to each set of operation data, obtaining M sets of characteristic data; inputting the M sets of characteristic data and the task to be executed into a policy generation model to obtain the target policy output by the policy generation model; sending the task to be executed to the target edge agent node indicated by the target policy, where the target edge agent node is used to execute the task to be executed. This solves the problem of low resource allocation accuracy and flexibility in the related art when using a fixed policy method for node task allocation. By receiving the task to be executed, determining the operation data of each initial edge agent node at the moment when the task to be executed is received, further determining the characteristic data according to the operation data, and determining the policy for executing the task to be executed according to the characteristic data and the policy generation model, and then sending the task to be executed to the corresponding target edge agent node for execution according to the policy, the technical effect of improving the accuracy and flexibility of the allocation of the task to be executed is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings that form a part of this application are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0018] Figure 1 is a flowchart of a task allocation method provided according to an embodiment of this application;
[0019] Figure 2 is a flowchart of an optional task allocation method provided according to an embodiment of this application;
[0020] Figure 3 is a schematic diagram of a task allocation system provided according to an embodiment of this application;
[0021] Figure 4 is a schematic diagram of a task allocation device provided according to an embodiment of this application;
[0022] Figure 5 is a schematic diagram of an electronic device provided according to an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the accompanying drawings and describe this application in detail with reference to the embodiments.
[0024] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to implement the embodiments of this application described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0026] It should be noted that the task allocation method, device, and electronic device determined by this disclosure can be used in the field of artificial intelligence, and can also be used in any field other than the field of artificial intelligence. The application fields of the task allocation method, device, and electronic device determined by this disclosure are not limited.
[0027] It should be noted that the information collected, user information (including but not limited to user device information, user personal information, etc.), and data (including but not limited to data for analysis, stored data, displayed data, etc.) used in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, complies with the relevant laws, regulations, and standards of the relevant regions, takes necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or refuse to use. If the user chooses to refuse, the expert decision-making process will be entered. For example, there is an interface between this system and relevant users or institutions. Before obtaining relevant information, a request for acquisition needs to be sent to the aforementioned users or institutions through the interface, and relevant information can be obtained after receiving the consent information feedback from the aforementioned users or institutions.
[0028] The embodiments or examples of the present disclosure are not exhaustive. They are only schematic representations of some embodiments or examples and do not constitute specific limitations on the protection scope of the present disclosure. Without contradiction, each step in a certain embodiment or example can be implemented as an independent embodiment, and the steps can be combined arbitrarily. For example, a solution obtained by removing some steps in a certain embodiment or example can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment or example can be exchanged arbitrarily. Additionally, the optional ways or optional examples in a certain embodiment or example can be combined arbitrarily; furthermore, the embodiments or examples can be combined arbitrarily. For example, some or all of the steps of different embodiments or examples can be combined arbitrarily, and a certain embodiment or example can be combined arbitrarily with the optional ways or optional examples of other embodiments or examples.
[0029] For ease of description, some nouns or terms related to the embodiments of the present application are explained below:
[0030] Edge Computing: It is a computing architecture concept that emphasizes "marginalizing" the computing capabilities of data processing and application services, that is, deploying these capabilities at the edge of the network, such as near intelligent devices, sensors, access points, or in a local network, rather than on a traditional centralized cloud computing server. Its core goal is to reduce the time delay of data transmission to a remote data center, improve the real-time performance and efficiency of data processing, and at the same time reduce the demand and cost for network bandwidth.
[0031] According to an embodiment of the present application, a task allocation method is provided.
[0032] Figure 1 is a flowchart of the task allocation method provided according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:
[0033] Step S101, receive a task to be executed and determine M initial edge agent nodes, where the initial edge agent nodes are edge agent nodes with the execution ability to execute the task to be executed, and M is a positive integer.
[0034] It should be noted that the execution subject of this embodiment can be a central control unit, which is used to control the edge computing network to process tasks. The edge computing network can include multiple edge intelligent agent nodes. The task to be executed refers to a task waiting to be processed in the edge computing network, and these tasks can be of various types such as data processing, analysis, and storage, and usually carry certain resource requirements and priority information. The initial edge agent nodes can be nodes in the edge computing network that have the computing, storage, and network resources required to execute tasks. These nodes can carry and process tasks and are candidate objects for resource scheduling.
[0035] Specifically, when a task to be executed is submitted to the central control unit, the system first needs to identify and classify the received task to determine its resource requirements (such as CPU (Central Processing Unit), memory, bandwidth, etc.) and priority. Next, the system filters out M edge proxy nodes from the entire edge computing network. Among them, the filtered edge proxy nodes need to have the execution ability to meet the task requirements, so as to ensure that the task to be executed can be successfully executed.
[0036] For example, assume that a video surveillance task requires high computing power and a low-latency network connection. The system filters out M edge proxy nodes with GPU acceleration capabilities and low-latency network interfaces from the entire edge computing network as the initial edge proxy nodes for executing this task.
[0037] Step S102: Obtain the operation data of each initial edge proxy node to get M groups of operation data, and calculate the feature data of each initial edge proxy node based on each group of operation data to get M groups of feature data.
[0038] It should be noted that the operation data can be the real-time operation status data of the edge proxy node, including but not limited to CPU utilization, memory occupancy, network bandwidth, latency, storage space, etc. The feature data can be the result after preprocessing and feature extraction of the operation data, which can more intuitively reflect the current state and resource availability of the node. For example, the feature information can be: average utilization rate, task completion rate, average network latency, etc.
[0039] Specifically, in the case of obtaining multiple initial edge proxy nodes, the system will collect the operation data of the M initial edge proxy nodes, including the resource occupancy of the nodes, the current task queue length, network conditions, etc. The collected data will then be preprocessed to clean outliers, missing values and perform data normalization, and then through feature extraction, the original operation data will be converted into a set of feature data that helps the model understand, such as calculating key metrics such as average CPU utilization, memory usage rate, and current network latency.
[0040] By obtaining the operation data, the system can monitor the resource status and operation efficiency of the edge proxy node in real time, and convert the original data into feature data, providing more meaningful input for the policy generation model to facilitate the model's analysis and decision-making.
[0041] For example, the system collects the real-time operation data of the M initial edge proxy nodes, including CPU utilization, memory usage rate, network latency, etc., and further calculates the current resource occupancy ratio, task queue length and average latency time of each node to form M groups of feature data.
[0042] Step S103: Input the M groups of feature data and the task to be executed into the policy generation model to obtain the target policy output by the policy generation model.
[0043] It should be noted that the policy generation model can be a model developed based on artificial intelligence technologies such as deep learning and reinforcement learning, which is used to generate a resource scheduling policy according to the input feature data and task requirements, that is, to determine which edge agent node executes a specific task. The target policy can be the optimal resource allocation and task scheduling plan output by the policy generation model after analyzing the M groups of feature data and the task to be executed.
[0044] Specifically, after obtaining the feature information of each initial edge agent node, the execution information or execution requirements of the task to be executed can be obtained, and the information such as the resource requirements and priorities of the M groups of feature data and the task to be executed can be used as model inputs and input into the policy generation model, so that the model automatically analyzes the matching degree between the resource status of the node and the task requirements through deep learning and reinforcement learning algorithms, and finally outputs a target policy to guide the task to be executed to be assigned to the most suitable edge agent node.
[0045] Through the automatic learning and analysis of the policy generation model, the intelligentization of resource scheduling can be realized, the resource utilization rate and task processing efficiency can be improved, it is ensured that tasks can be assigned to the nodes with the most suitable resources, the task waiting time can be reduced, and the service quality can be improved.
[0046] For example, the above-mentioned collected and processed M groups of feature data and the specific resource requirements of the video surveillance task are sent into the trained deep reinforcement learning model. After analysis and calculation by the model, a target policy is output, indicating that the task is assigned to the node with the lowest current CPU utilization rate and the minimum network latency.
[0047] Step S104: Send the task to be executed to the target edge agent node indicated by the target policy, where the target edge agent node is used to execute the task to be executed.
[0048] It should be noted that the target edge agent node is also the node selected according to the target policy output by the policy generation model and is used to execute the task to be executed.
[0049] Specifically, after the relevant information is input into the policy generation model, the policy generation model will input the target policy, which contains the indication information on which edge agent node the task to be executed is assigned to. The system can determine the target edge agent node, that is, the node most suitable for executing the task, according to the indication information in the target policy, and send the task to be executed to the target edge node. After receiving the task, the target edge agent node starts to execute the task according to its own resource status and scheduling policy, thereby completing the execution operation of the task to be executed, ensuring that the task can be processed efficiently and in a timely manner, and further improving the overall operation efficiency of the edge computing network.
[0050] For example, the system sends the video surveillance task to the node with the lowest current CPU utilization rate and the minimum network latency according to the policy output by the deep reinforcement learning model. After receiving the task, the node uses its high computing power and low-latency network interface to start processing video data efficiently and perform real-time surveillance analysis, thereby completing the task.
[0051] Figure 2 is a flowchart of an optional task allocation method provided according to an embodiment of the present application. As Figure 2 shown, the edge agent node sends the operation data to the central control unit. The policy generation model in the central control unit generates a target policy according to the operation data and the received task to be executed, and issues the task to the target edge agent node according to the target policy, and monitors the execution situation of the target edge agent node, and further updates the policy generation model according to the execution situation and the execution result, thereby completing the operations of task allocation and model iteration update.
[0052] The task allocation method provided by the embodiments of the present application includes receiving a task to be executed and determining M initial edge agent nodes, where the initial edge agent nodes are edge agent nodes with the ability to execute the task to be executed, and M is a positive integer; obtaining the operation data of each initial edge agent node to obtain M sets of operation data, and calculating the characteristic data of each initial edge agent node according to each set of operation data to obtain M sets of characteristic data; inputting the M sets of characteristic data and the task to be executed into a policy generation model to obtain the target policy output by the policy generation model; and sending the task to be executed to the target edge agent node indicated by the target policy, where the target edge agent node is used to execute the task to be executed. This solves the problem of low resource allocation accuracy and flexibility in the related art when using a fixed policy method for node task allocation. By receiving the task to be executed and determining the operation data of each initial edge agent node at the moment when the task to be executed is received, and then determining the characteristic data according to the operation data, and determining the policy for executing the task to be executed according to the characteristic data and the policy generation model, so as to send the task to be executed to the corresponding target edge agent node for execution according to the policy, thereby achieving the technical effect of improving the accuracy and flexibility of the allocation of the task to be executed.
[0053] Optionally, in the task allocation method provided by the embodiments of the present application, determining M initial edge agent nodes includes: obtaining the task type and task level of the task to be executed; screening multiple preset edge agent nodes according to the task level and task type to obtain M initial edge agent nodes.
[0054] It should be noted that the task type refers to the specific type or attribute of the task. For different types of tasks such as video processing, image analysis, and data storage, there are differences in resource requirements and processing methods. The task level is used to represent the priority or importance level of the task, usually determined according to factors such as the urgency of the task and the resource demand. A high task level means that the task needs to be processed first.
[0055] Specifically, when determining the initial edge agent nodes, it is first necessary to obtain the task type and task level of the task to be executed. The task type reflects the basic attributes of the task and determines the type and amount of resources required; the task level clarifies the urgency and importance of the task and affects the priority of resource allocation.
[0056] Furthermore, the system screens the preset edge agent nodes according to the task level and task type. The screening process considers whether the node has sufficient resources (such as CPU, GPU, etc.) to process this type of task, and whether it can meet the quality of service requirements (such as latency, throughput, etc.) required by the task level. The system selects M most compliant nodes from all available nodes as candidates for task execution, thereby obtaining the initial edge agent nodes.
[0057] For example, the system receives a real-time video analysis task, which is a high-computation-intensive task and is marked as high-priority, indicating that the task requires not only a large amount of computing resources but also needs to be processed as soon as possible. At this time, the system filters out a large number of pre-set edge proxy nodes (e.g., 100) in the edge computing network, and finally selects 10 nodes. These nodes have high-performance GPU resources and a network latency of less than 5 milliseconds, meeting the high-computation requirements and low-latency requirements of real-time video analysis. Through this process, the system can specifically select the most suitable nodes for processing specific task types from the edge computing network, ensuring that tasks can be executed quickly, efficiently, and in accordance with priorities, thereby improving the overall network resource utilization efficiency and service quality.
[0058] In this embodiment, through refined task type and level screening, precise matching and optimal allocation of edge proxy node resources are achieved, which not only improves the accuracy of resource scheduling but also ensures the timely processing of high-priority tasks, greatly enhancing the network response speed and overall service quality.
[0059] Optionally, in the task allocation method provided in the embodiment of the present application, characteristic data of each initial edge proxy node is calculated according to each group of operation data, and M groups of characteristic data are obtained, including: for any group of operation data, each operation data in the group of operation data is grouped according to the time stamp to obtain a plurality of operation data sets; calculation formulas of each initial characteristic are obtained, and the data in the plurality of operation data sets are substituted into the calculation formulas to obtain the characteristic data of each initial characteristic; the correlation between each initial characteristic and the duration of the initial edge proxy node for processing tasks is calculated, and the characteristic data of the initial characteristics with a correlation greater than a preset correlation is determined as a group of characteristic data.
[0060] It should be noted that the operation data set is a series of operation data grouped by time stamp, representing the operation status of the edge proxy node at different time points. The preliminary characteristic is the original characteristic data generated during the calculation process, covering all aspects of the edge node resources, such as CPU utilization, memory occupancy, network bandwidth, latency, etc. The characteristic data is the refined result of the preliminary characteristic obtained through the calculation formula and is an important input for the policy generation model. The correlation is an index to measure the degree of closeness of the relationship between the preliminary characteristic and the task processing duration of the node, helping to screen out the characteristics that have the greatest impact on task processing. The preset correlation threshold is the standard set by the system for determining whether a characteristic is important enough, and only the characteristics with a correlation exceeding this threshold will be selected for the next step of analysis.
[0061] Specifically, after the system receives the operation data from M initial edge agent nodes, it will obtain M sets of operation data, with each set corresponding to an initial edge agent node. Further, each set of operation data may include multiple operation data at multiple collection times. For example, a set of operation data may include data such as CPU occupancy rate, memory occupancy, network bandwidth, and latency at time A, and may also include data such as CPU occupancy rate, memory occupancy, network bandwidth, and latency at time B. At this time, the operation data can be grouped according to the timestamp to obtain multiple operation data sets, and each operation data set includes operation data of different feature types at a collection time.
[0062] Further, after obtaining multiple operation data sets, a preset preliminary feature calculation formula can be obtained, and the specific data in each operation data set is substituted into the corresponding formula to calculate the specific value of each preliminary feature. For example, calculate the average value of CPU utilization rate, the maximum value of network latency, etc., so as to obtain the initial feature data of each initial edge agent node.
[0063] For example, feature extraction aims to extract information useful for the model from the preprocessed data to improve the prediction performance of the model, which may include: time series features, task features, and network features. Among them, time series features may include average utilization rate and data fluctuation. The calculation method of the average utilization rate can be: calculate the average value of the resource utilization rate data according to the time window, and the determination method of the data fluctuation can be to calculate the variance or standard deviation of the utilization rate to measure the volatility. Task features may include: task arrival rate, task type ratio. Among them, the calculation method of the task arrival rate can be: count the number of newly arrived tasks per unit time, and the calculation method of the task type ratio can be: calculate the proportion of each type of task in the total tasks to form a task distribution feature vector. Network features may include: network latency, bandwidth utilization rate. Among them, the calculation method of network latency can be: obtain the average latency with the target node through Ping test or network measurement tools, and the calculation method of bandwidth utilization rate can be: calculate the usage rate of the current bandwidth, that is, the ratio of the used bandwidth to the total bandwidth.
[0064] Further, since the obtained multiple initial feature data may have a poor correlation with the efficiency of the initial edge agent node in processing tasks, it is necessary to further screen the multiple initial feature data to obtain multiple feature data with a higher degree of association with the task processing operation. At this time, the system will evaluate the correlation between each preliminary feature and the node processing task duration to determine which features have the most significant impact on the task processing efficiency. This step calculates the correlation coefficient between the feature and the task processing duration through statistical methods, such as the Pearson correlation coefficient, so as to quantify the association strength between the feature and the task processing duration.
[0065] Finally, the system can set a preset relevance threshold, and retain the preliminary features whose relevance to the task processing duration is greater than this threshold as the final feature data, enabling the system to extract the core features that truly affect the task processing efficiency from a large amount of operation data, avoiding the interference of non-critical factors, providing more accurate and valuable inputs for the decision-making of the subsequent policy generation model, thereby improving the accuracy and efficiency of resource scheduling, reducing the input dimension of the model, and improving the model efficiency.
[0066] Through the above process, this embodiment realizes the efficient conversion from the original operation data to the key feature data, not only reducing the computational burden of the policy generation model, but also significantly improving the accuracy and timeliness of resource scheduling decisions.
[0067] Optionally, in the task allocation method provided in the embodiment of the present application, the policy generation model is trained in the following manner: obtaining the processing duration of each sample edge proxy node in the sample edge proxy node set at N historical moments for processing sample tasks, and the feature data of each sample edge proxy node, where N is a positive integer; taking the processing duration of each sample edge proxy node in processing sample tasks and the feature data of each sample edge proxy node at each historical moment as a set of sample data to obtain N sets of sample data; and iteratively training the deep reinforcement learning model with each set of sample data in turn to obtain the policy generation model.
[0068] It should be noted that the historical moment can refer to a certain past time point, which usually includes the running state and task processing situation of the edge proxy node during the historical task execution period and is used for model training. The sample edge proxy node can be an edge proxy node selected from the historical dataset for training the model, and its running data and task processing results are used as learning cases. The processing duration can be the time taken by the sample edge proxy node to process a specific historical task, reflecting the task processing efficiency of the node under different running conditions. The N sets of sample data can be a set of examples composed of the processing duration and feature data of the sample edge proxy node in processing tasks at N historical moments, and are used for multiple iterative trainings of the deep reinforcement learning model.
[0069] Specifically, when training the model, it is first necessary to determine the training sample data. The system first obtains from the database the processing duration of each node in the sample edge proxy node set for processing a specific task at N historical moments, and the feature data of the node at that time. The feature data here includes the resource status of the node (such as CPU utilization, memory occupancy), network status, task level, task type, etc., and the processing duration is the actual time for the node to complete the historical task.
[0070] Further, the processing duration of the sample edge proxy node's processing tasks at each historical moment is combined with the feature data at that moment to form a set of sample data. By doing so, the system converts the data collected at N historical moments into N sets of example data that can be used for model training. For example, for edge proxy node A, its processing duration at 3:00 pm on a certain day is 4 seconds. At the same time, the CPU utilization rate at that time is 75%, and the network latency is 100 milliseconds. These data combined constitute a sample data.
[0071] In the case of obtaining the sample data, the constructed N sets of sample data can be used to iteratively train the deep reinforcement learning model. In each iteration, the model will learn how to make optimal resource allocation and task scheduling decisions based on the features and processing durations of the nodes in the sample data to reduce the latency of task processing. The training process includes state observation, action selection, reward feedback, and model parameter adjustment. Through techniques such as experience replay and target network, the model training is stabilized, gradually approaching the optimal strategy. Thus, the deep reinforcement learning model can learn the operation rules and task processing characteristics of the edge proxy node through a large amount of historical data, and can make more intelligent and dynamically adaptable resource scheduling decisions when processing new tasks, improving task processing efficiency and resource utilization rate.
[0072] For example, the following is an optional establishment and training process of a policy generation model:
[0073] Step 1: First, perform a model selection operation. A deep reinforcement learning model can be adopted, which combines a deep neural network and a reinforcement learning algorithm to process high-dimensional state and action spaces.
[0074] Step 2: Define the state, action, and reward functions:
[0075] Among them, the state space (S) includes: the resource status of the edge node (CPU, memory, bandwidth, etc.), the network status (latency, bandwidth), and task information (task type, requirements, priority).
[0076] The action space (A) includes: task allocation strategy (which node to allocate the task to), resource allocation plan (how much computing and storage resources to allocate), and adjustment of task scheduling order.
[0077] The reward function (R) includes: aiming at "maximizing resource utilization rate, minimizing task completion time, and ensuring service quality", designing the reward function. For example: if the task is completed on time, the reward is +1; if the resource utilization rate increases, the reward increases proportionally; if the task is delayed, the penalty is -1; if there is resource overload, the penalty is -1.
[0078] Step 3: Model training, which includes: Data preparation: Using the collected global data to construct a training dataset. Training process: Policy iteration: Interact with the environment according to the current policy to generate empirical data. Experience replay: Store the empirical data in the experience pool and randomly sample a small batch of data for training to break data correlation. Model update: Use optimization algorithms such as gradient descent to update the parameters of the neural network. Target network: Introduce a target network to stabilize the training process and prevent model oscillation. Training iteration: Continuously repeat policy iteration and model update until the model converges or reaches the preset performance metrics.
[0079] Step 4: Model deployment: Deploy the trained model in the system for real-time decision-making.
[0080] Through the above process, this embodiment trains a policy generation model capable of intelligent decision-making using historical data. When processing new tasks, this model can predict and optimize resource scheduling policies based on the real-time feature data of edge proxy nodes, reduce task processing latency, and improve resource utilization efficiency.
[0081] Optionally, in the task allocation method provided in the embodiment of the present application, obtaining the operation data of each initial edge proxy node to obtain M groups of operation data includes: For any initial edge proxy node, obtaining the initial operation data of the initial edge proxy node and determining whether there are outliers in the initial operation data; in the case of the existence of outliers, deleting the outliers and obtaining all missing values in the initial operation data; determining the data types of the missing values and calculating the missing values according to other operation data under the data types to obtain filled values, and using the filled values to replace the missing values in the initial operation data to obtain updated operation data; performing a normalization operation on the updated operation data to obtain the operation data of the initial edge proxy node.
[0082] It should be noted that outliers can be values that deviate too far from the normal range in the collected operation data. Missing values are data points that are not recorded or lost in the operation dataset. Data types are also different categories of operation data, such as CPU utilization, memory usage, network latency, etc. Different types of data require different filling strategies. Filled values are values calculated by analyzing other healthy data under the data types and used to replace the missing values to maintain the integrity of the dataset. The normalization operation is also a mathematical processing of the operation data to make its values distributed within a fixed range, aiming to eliminate the dimensionality impact between different data types and facilitate model processing.
[0083] Specifically, when obtaining operation data, since the operation data may be missing or incorrect, after the system collects the operation status data of M initial edge agent nodes in real time, it is necessary to detect whether there are outliers, that is, data points that significantly deviate from the normal range. For example, if negative numbers or values exceeding 100% appear in the CPU utilization data, these data points will be marked as outliers.
[0084] In the case of identifying outliers, the system will remove these values from the data set, and at the same time search for and identify all missing values in the data, and then decide which method to use to calculate the filling values according to the data types where the missing values are located.
[0085] Furthermore, for each determined missing value, the system can calculate reasonable filling values based on other normal data points of the same data type. This may include methods such as using the average value, median, nearest neighbor interpolation, etc. Replace the original missing values with the calculated filling values to keep the data set complete, and perform standardization operations on the updated operation data to convert different types of data to the same numerical scale. For example, use min-max normalization or standard score (Z-score) to scale the data to eliminate the influence of dimensions and ensure that the data is in a comparable and analyzable state before inputting into the model. Through outlier detection and missing value filling, ensure the accuracy and integrity of the operation data, and avoid the negative impact of abnormal data on model training and resource scheduling decisions. The standardization operation further improves the data processing efficiency, eliminates the dimension difference, and enables the model to analyze and utilize these data more efficiently.
[0086] For example, the following is an optional operation process for obtaining and processing operation data:
[0087] Step 1: Collect the resource status information of the edge nodes in real time, including but not limited to: CPU, GPU utilization; memory and storage space occupancy; network bandwidth and latency; current task queue length and task type; physical parameters such as energy consumption and temperature.
[0088] Step 2: Process the collected data:
[0089] 1. Data cleaning: Remove outliers and missing values. Including:
[0090] 1) Remove missing values
[0091] Missing value detection: Scan the data set to detect whether there are missing values (NaN, Null, etc.) in each feature. Algorithm: Traverse each column in the data set and count the number and proportion of missing values.
[0092] Processing Strategy: Deletion Strategy: Condition: If the missing values in a certain record exceed the set threshold (e.g., 50%), then delete that record. Algorithm: Traverse the dataset, calculate the proportion of missing values for each record, and delete the records that exceed the threshold. Filling Strategy: Mean Filling: Applicable Scenario: Numerical features with relatively uniform data distribution and no obvious skewness. Algorithm: Calculate the mean of the feature column and replace the missing values with this mean. Median Filling: Applicable Scenario: Numerical features with outliers or skewed distributions. Algorithm: Calculate the median of the feature column and replace the missing values with this median. Mode Filling: Applicable Scenario: Categorical features. Algorithm: Statistically find the value with the highest frequency in the feature column and replace the missing values with it. Interpolation Method: Applicable Scenario: Time series data. Algorithm: Use the average value or linear interpolation of adjacent valid data before and after to fill the missing values.
[0093] 2) Handling Outliers
[0094] Outlier Detection: Box Plot Method: Algorithm: Calculate the first quartile (Q1) and the third quartile (Q3), calculate the interquartile range, determine the outlier range (upper and lower bounds), and mark the data outside the above range as outliers.
[0095] Standard Deviation Method (Z-score Method): Algorithm: Calculate the mean and standard deviation of the feature column, calculate the Z-score for each data point, set a threshold, and mark the data that exceeds the threshold as outliers.
[0096] Processing Strategy: Deleting Outliers: Algorithm: Directly delete the detected outlier records from the dataset.
[0097] Modified Outlier Truncation Method: Algorithm: Replace the outliers with the threshold boundaries (upper or lower bounds).
[0098] Mean or Median Substitution: Algorithm: Replace the outliers with the mean or median of the non-outlier values.
[0099] 2. Normalization: Map the data to a unified numerical range for convenient model processing. The purpose of data normalization is to eliminate the differences between different feature dimensions and facilitate the training and convergence of the model. The following methods can be used for normalization operations: Min-Max Normalization, Standardization (Z-score Standardization), and Decimal Scaling Normalization.
[0100] In this embodiment, by performing preprocessing steps on the data, the quality of the running data is improved, laying a solid foundation for subsequent feature calculation and resource scheduling decisions.
[0101] Optionally, in the task allocation method provided in the embodiments of the present application, after sending the task to be executed to the target edge proxy node indicated by the target policy, the method further includes: obtaining the execution result of the target edge proxy node executing the task to be executed, and determining whether the execution result is successful; in the case where the execution result is a failure, obtaining the characteristic data of each initial edge proxy node, and regenerating the execution policy of the task to be executed through a policy generation model; and reselecting the edge proxy node for executing the task to be executed according to the execution policy.
[0102] It should be noted that the execution result refers to the actual effect after the edge proxy node executes the task to be executed, including indicators such as whether the task is successfully completed, the execution duration, and resource consumption.
[0103] Specifically, after sending the task to be executed to the target edge proxy node, the system continuously monitors the execution status of the task, including but not limited to the execution duration of the task, the node resource consumption, and whether the task output meets the expectations. In the case of obtaining the execution status, it is necessary to determine whether the final result of the edge proxy node executing the task to be executed is successful, that is, whether the task is executed completely and correctly according to the expected goal.
[0104] In the case where the task execution fails, the system needs to handle the task rescheduling. First, re-obtain the characteristic data of the M initial edge proxy nodes at the current moment, which updates the real-time resource status and network conditions of the nodes. Then, input the updated characteristic data into the policy generation model again to regenerate the execution policy of the task to be executed. The model will dynamically adjust the resource allocation and node selection of the task according to the latest node status and resource availability.
[0105] In the case of obtaining the execution policy output by the policy generation model according to the re-obtained characteristic data, the system will determine a new target edge proxy node according to the newly generated execution policy, and re-send the task to be executed to this node for execution, thereby ensuring that the task can be processed in a timely and effective manner, avoiding resource waste and user service interruption.
[0106] This embodiment significantly improves the reliability and dynamic adaptability of edge computing network resource scheduling by monitoring the execution status of the task and re-determining the target edge proxy node in the case of abnormal execution status.
[0107] Optionally, in the task allocation method provided in the embodiments of the present application, obtaining the operation data of each initial edge proxy node includes: identifying the interface information of the initial edge proxy node; determining the data transmission decryption policy according to the interface information; and in the case of receiving the transmission data sent by the initial edge proxy node, decrypting the transmission data through the data transmission decryption policy to obtain the operation data.
[0108] It should be noted that interface information refers to the protocol, port, encryption method, and other information used by the initial edge agent node to communicate with the central control system or other network components. Data transmission decryption strategies refer to the methods and processes pre-set by the central control unit to decrypt operational data to restore the original data, based on different encryption methods and security requirements. Transmission data refers to operational status information sent by the initial edge agent node to the central control unit in the form of encrypted data packets, including resource utilization, network status, and task queue data.
[0109] Specifically, when receiving the operating data sent by the edge agent node, the central control unit first identifies the interface information for communicating with each initial edge agent node through registration or configuration files, which includes the communication protocol, port number, encryption method, etc., and determines the corresponding data transmission decryption strategy based on the obtained interface information.
[0110] Furthermore, when the central control unit receives encrypted transmission data from the initial edge proxy node, it applies the previously determined decryption strategy to decrypt the data packet and restore the original operational data. By identifying the interface information and applying the data transmission decryption strategy, the central control unit ensures the secure and efficient reception and processing of operational data from each initial edge proxy node, preventing data leakage and tampering during transmission and ensuring data security and the accuracy of subsequent processing.
[0111] Through the above process, this embodiment strengthens the secure data transmission mechanism in the edge computing network, ensures the accurate reception of operation data, and provides a solid data foundation for subsequent feature data calculation and resource scheduling decisions.
[0112] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0113] In this embodiment, Figure 3 is a schematic diagram of a task allocation system provided according to an embodiment of the present application, such as Figure 3 As shown, the aforementioned task allocation method is executed with an optional task allocation system as the execution body, and the task allocation system at least includes: an edge agent node 301, a central control unit 302, and a network communication module 303.
[0114] The edge proxy node 301 includes: a data collection module, a data pre-processing module, a local policy execution module, and local cache and storage.
[0115] Among them, the data collection module: is used to collect the real-time operation data of edge nodes, such as CPU utilization, memory occupancy, network bandwidth, latency, task queue length, etc.
[0116] The data preprocessing module: is used to perform preprocessing such as cleaning, normalization, and feature extraction on the collected data to form a standardized data format.
[0117] The local policy execution module: is used to execute task allocation and resource scheduling according to the received scheduling policy.
[0118] The local cache and storage: is used to store local operation data and temporary model parameters.
[0119] The functions of the edge proxy node 301 are: real-time monitoring of the resource status and task conditions of edge nodes, data preprocessing, providing high-quality input data for the model, executing the scheduling policy, completing task processing and resource allocation, and communicating with the central control unit 302, uploading local data, and receiving global policies.
[0120] The central control unit 302 includes: a data aggregation module, an artificial intelligence model training module, a policy generation and distribution module, and a model update and management module.
[0121] Among them,
[0122] The data aggregation module: is used to collect and aggregate data from each edge proxy node 301 to form a global network status view.
[0123] The artificial intelligence model training module: is used to train a resource scheduling model using deep reinforcement learning algorithms.
[0124] The policy generation and distribution module: is used to generate a scheduling policy according to the trained model and distribute it to each edge proxy node 301.
[0125] The model update and management module: is used to be responsible for the version control, update, and optimization of the model.
[0126] The functions of the central control unit 302 are: global data analysis to grasp the resource status and task requirements of the entire edge computing network. Model training and optimization to generate the optimal resource scheduling policy. Policy distribution to send the scheduling policy to each edge proxy node 301 to guide resource scheduling. Model update to continuously optimize the model performance according to real-time data and feedback.
[0127] The network communication module 303 includes: a data transmission interface, a security encryption module, and an exception handling module.
[0128] Among them,
[0129] Data transmission interface: used to define the data transmission protocol and interface between the edge proxy node 301 and the central control unit 302.
[0130] Security encryption module: used to encrypt the transmitted data to ensure the security and reliability of communication.
[0131] Exception handling module: used to monitor the communication status and handle network exceptions and faults.
[0132] The functions of the network communication module 303 are: to realize the two-way transmission of data, including status data, scheduling policies, model parameters, etc. To ensure the security and stability of communication and prevent data leakage and communication interruption.
[0133] The embodiment of the present application also provides a task allocation device. It should be noted that the task allocation device of the embodiment of the present application can be used to execute the task allocation method provided by the embodiment of the present application. The following introduces the task allocation device provided by the embodiment of the present application.
[0134] Figure 4 is a schematic diagram of the task allocation device provided by the embodiment of the present application. As Figure 4 shown, the device includes: a receiving unit 41, a first obtaining unit 42, a first generating unit 43, and a sending unit 44.
[0135] The receiving unit 41 is used to receive the task to be executed and determine M initial edge proxy nodes, where the initial edge proxy node is an edge proxy node with the execution ability to execute the task to be executed, and M is a positive integer.
[0136] The first obtaining unit 42 is used to obtain the operation data of each initial edge proxy node, obtain M groups of operation data, and calculate the characteristic data of each initial edge proxy node according to each group of operation data to obtain M groups of characteristic data.
[0137] The first generating unit 43 is used to input the M groups of characteristic data and the task to be executed into the policy generation model to obtain the target policy output by the policy generation model.
[0138] The sending unit 44 is used to send the task to be executed to the target edge proxy node indicated by the target policy, where the target edge proxy node is used to execute the task to be executed.
[0139] The task allocation device provided by the embodiment of the present application receives a task to be executed through a receiving unit 41, and determines M initial edge proxy nodes, where the initial edge proxy nodes are edge proxy nodes with the execution ability to execute the task to be executed, and M is a positive integer; a first obtaining unit 42 obtains the operation data of each initial edge proxy node, obtains M groups of operation data, and calculates the characteristic data of each initial edge proxy node according to each group of operation data to obtain M groups of characteristic data; a first generating unit 43 inputs the M groups of characteristic data and the task to be executed into a policy generation model to obtain a target policy output by the policy generation model; a sending unit 44 sends the task to be executed to the target edge proxy node indicated by the target policy, where the target edge proxy node is used to execute the task to be executed, solving the problem of low resource allocation accuracy and flexibility in node task allocation using a fixed policy method in the related art. By receiving the task to be executed and determining the operation data of each initial edge proxy node at the moment when the task to be executed is received, and then determining the characteristic data according to the operation data, and determining the policy for executing the task to be executed according to the characteristic data and the policy generation model, and thus sending the task to be executed to the corresponding target edge proxy node for execution according to the policy, thereby achieving the technical effect of improving the accuracy and flexibility of the allocation of the task to be executed.
[0140] Optionally, in the task allocation device provided by the embodiment of the present application, the receiving unit 41 includes: a first obtaining module, configured to obtain the task type and task level of the task to be executed; a screening module, configured to screen multiple preset edge proxy nodes according to the task level and task type to obtain M initial edge proxy nodes.
[0141] Optionally, in the task allocation device provided by the embodiment of the present application, the first obtaining unit 42 includes: a grouping module, configured to group each operation data in the group of operation data according to the time stamp for any group of operation data to obtain a plurality of operation data sets; a second obtaining module, configured to obtain the calculation formula of each initial feature, and substitute the data in the plurality of operation data sets into the calculation formula to obtain the characteristic data of each initial feature; a calculation module, configured to calculate the correlation between each initial feature and the duration of the initial edge proxy node for processing the task, and determine the characteristic data of the initial feature with a correlation greater than a preset correlation as a group of characteristic data.
[0142] Optionally, in the task allocation device provided in the embodiments of the present application, the policy generation model is trained in the following manner: a second acquisition unit, configured to acquire the processing duration of each sample edge proxy node in the sample edge proxy node set at N historical moments, and the feature data of each sample edge proxy node, where N is a positive integer; a second generation unit, configured to use the processing duration of each sample edge proxy node in each historical moment and the feature data of each sample edge proxy node as a set of sample data to obtain N sets of sample data; a training unit, configured to iteratively train the deep reinforcement learning model with each set of sample data in turn to obtain the policy generation model.
[0143] Optionally, in the task allocation device provided in the embodiments of the present application, the first acquisition unit 42 includes: a third acquisition module, configured to, for any initial edge proxy node, acquire the initial operation data of the initial edge proxy node and determine whether there is an outlier in the initial operation data; a deletion module, configured to, in the case where there is an outlier, delete the outlier and acquire all missing values in the initial operation data; a first determination module, configured to determine the data type of the missing values, calculate the missing values according to other operation data under the data type to obtain filled values, and use the filled values to replace the missing values in the initial operation data to obtain updated operation data; a normalization module, configured to perform a normalization operation on the updated operation data to obtain the operation data of the initial edge proxy node.
[0144] Optionally, in the task allocation device provided in the embodiments of the present application, after sending the task to be executed to the target edge proxy node indicated by the target policy, the device further includes: a third acquisition unit, configured to acquire the execution result of the target edge proxy node executing the task to be executed and determine whether the execution result is a successful execution; a third generation unit, configured to, in the case where the execution result is an execution failure, acquire the feature data of each initial edge proxy node and regenerate the execution policy of the task to be executed through the policy generation model; an execution unit, configured to reselect the edge proxy node for executing the task to be executed according to the execution policy.
[0145] Optionally, in the task allocation device provided in the embodiments of the present application, the first acquisition unit 42 includes: an identification module, configured to identify the interface information of the initial edge proxy node; a second determination module, configured to determine the data transmission decryption policy according to the interface information; a decryption module, configured to, in the case of receiving the transmission data sent by the initial edge proxy node, decrypt the transmission data through the data transmission decryption policy to obtain the operation data.
[0146] The above task allocation device includes a processor and a memory. The above receiving unit 41, first acquisition unit 42, first generation unit 43, sending unit 44, etc. are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement corresponding functions.
[0147] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problems of low resource allocation accuracy and flexibility in node task allocation using a fixed strategy method in the related art are solved.
[0148] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one storage chip.
[0149] An embodiment of the present invention provides a computer-readable storage medium with a program stored thereon, and when the program is executed by a processor, it implements the task allocation method.
[0150] An embodiment of the present invention provides a processor, and the processor is used to run a program. When the program runs, it executes the task allocation method.
[0151] Figure 5 is a schematic diagram of an electronic device provided according to an embodiment of the present application. As Figure 5 shown, an embodiment of the present invention provides an electronic device. The electronic device 50 includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the above task allocation method. The devices herein can be servers, PCs, PADs, mobile phones, etc.
[0152] The present application also provides a computer program product, which is suitable for executing a program initialized with the steps of the above task allocation method when executed on a data processing device.
[0153] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0154] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the specified functions in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0155] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the specified functions in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0156] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0157] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0158] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0159] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0160] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0161] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A task allocation method, characterized in that Including: Receiving a task to be executed and determining M initial edge proxy nodes, where the initial edge proxy nodes are edge proxy nodes with the execution ability to execute the task to be executed, and M is a positive integer; Obtaining the operation data of each initial edge proxy node to obtain M sets of operation data, and calculating the feature data of each initial edge proxy node according to each set of operation data to obtain M sets of feature data; Inputting the M sets of feature data and the task to be executed into a policy generation model to obtain a target policy output by the policy generation model; Sending the task to be executed to a target edge proxy node indicated by the target policy, where the target edge proxy node is used to execute the task to be executed.
2. The method according to claim 1, wherein Determining the M initial edge proxy nodes includes: Obtaining the task type and task level of the task to be executed; Filtering multiple preset edge proxy nodes according to the task level and the task type to obtain the M initial edge proxy nodes.
3. The method according to claim 1, wherein Calculating the feature data of each initial edge proxy node according to each set of operation data to obtain M sets of feature data includes: For any set of operation data, grouping each operation data in the set of operation data according to the time stamp to obtain a plurality of operation data sets; Obtaining the calculation formula of each initial feature, and substituting the data in the plurality of operation data sets into the calculation formula to obtain the feature data of each initial feature; Calculating the correlation degree between each initial feature and the duration of the initial edge proxy node to process the task, and determining the feature data of the initial feature with the correlation degree greater than a preset correlation degree as a set of feature data.
4. The method according to claim 1, wherein The policy generation model is trained in the following manner: Obtaining the processing duration of each sample edge proxy node in the sample edge proxy node set at N historical moments, and the feature data of each sample edge proxy node, where N is a positive integer; Taking the processing duration of each sample edge proxy node to process the sample task and the feature data of each sample edge proxy node at each historical moment as a set of sample data to obtain N sets of sample data; Iteratively training the deep reinforcement learning model with each set of sample data in turn to obtain the policy generation model.
5. The method according to claim 1, characterized in that, Obtaining the operation data of each initial edge proxy node to obtain M sets of operation data includes: For any initial edge proxy node, obtaining the initial operation data of the initial edge proxy node and determining whether there are outliers in the initial operation data; In the case of the existence of the outliers, deleting the outliers and obtaining all missing values in the initial operation data; Determining the data type of the missing values, calculating the missing values according to other operation data under the data type to obtain filling values, and replacing the missing values in the initial operation data with the filling values to obtain updated operation data; Performing a standardization operation on the updated operation data to obtain the operation data of the initial edge proxy node.
6. The method according to claim 1, characterized in that, After sending the task to be executed to the target edge proxy node indicated by the target policy, the method further includes: Obtain the execution result of the target edge proxy node for executing the task to be executed, and determine whether the execution result is successful execution; In the case where the execution result is execution failure, obtain the feature data of each initial edge proxy node, and regenerate the execution policy of the task to be executed through the policy generation model; Re-select the edge proxy node for executing the task to be executed according to the execution policy.
7. The method according to claim 1, wherein Obtaining the operation data of each initial edge proxy node includes: Identify the interface information of the initial edge proxy node; Determine the data transmission decryption policy according to the interface information; In the case of receiving the transmission data sent by the initial edge proxy node, perform a decryption operation on the transmission data through the data transmission decryption policy to obtain the operation data.
8. A task allocation device, characterized in that, Includes: A receiving unit, configured to receive a task to be executed, and determine M initial edge proxy nodes, where the initial edge proxy node is an edge proxy node having the execution ability to execute the task to be executed, and M is a positive integer; A first obtaining unit, configured to obtain the operation data of each initial edge proxy node, obtain M sets of operation data, and calculate the feature data of each initial edge proxy node according to each set of operation data to obtain M sets of feature data; A first generating unit, configured to input the M sets of feature data and the task to be executed into a policy generation model, and obtain a target policy output by the policy generation model; A sending unit, configured to send the task to be executed to the target edge proxy node indicated by the target policy, where the target edge proxy node is used to execute the task to be executed.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, where when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the task allocation method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, Includes: A memory, storing an executable program; A processor, configured to run the program, where when the program runs, it executes the task allocation method according to any one of claims 1 to 7.