A multi-domain Internet of Things task allocation optimization model training method, device, equipment and medium

By generating a multi-dimensional state tensor and deep reinforcement learning training task allocation optimization model, the problem of low task processing efficiency in multi-domain Internet of Things is solved, and efficient, low-latency and low-energy task allocation optimization is achieved under diverse needs.

CN120499018BActive Publication Date: 2025-09-26CRSC COMM & INFORMATION GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510976135.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-09-26
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

Existing technologies cannot effectively meet the diverse needs of IoT devices with limited computing resources and energy consumption in multi-domain IoT under different task processing requirements, resulting in low task processing efficiency.

Method used

By generating a multi-dimensional state tensor and using deep reinforcement learning strategies to train the task allocation optimization model, the distribution of tasks between local devices and edge servers is optimized. Combining the task allocation identification matrix, node feature matrix and three-dimensional network feature tensor, the task processing cost objective function and reward function are constructed to achieve adaptive optimization of task allocation.

Benefits of technology

Taking into account different task processing requirements and reasonable allocation of computing resources, the task allocation of multi-domain IoT is optimized, the task processing efficiency is improved, and low-latency and high-energy-efficiency task processing is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499018B_ABST
    Figure CN120499018B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of Internet of Things, and discloses a multi-domain Internet of Things task allocation optimization model training method, device, equipment and medium. According to the task allocation information, operation performance indicators, task attribute information and task communication quality indicators of multiple single-domain Internet of Things and edge server groups, a first multi-dimensional state tensor is generated and input into the task allocation optimization model to be trained, so that the task allocation optimization model to be trained performs task allocation optimization based on the first multi-dimensional state tensor, obtains an allocation optimization strategy, and updates the task allocation optimization model to be trained based on the allocation optimization strategy to obtain a trained task allocation optimization model. The task allocation optimization model trained by the present invention can optimize the task allocation of the multi-domain Internet of Things while taking into account the diverse needs such as the processing requirements of different tasks and the reasonable allocation of computing resources, and effectively achieve diverse goals such as the different task processing requirements and the reasonable allocation of computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet of Things, and in particular to a multi-domain Internet of Things task allocation optimization model training method, device, equipment and medium. Background Art

[0002] With the development of science and technology, Internet of Things technology continues to improve.

[0003] The booming development of the Internet of Things (IoT) and its complex applications are driving numerous computationally intensive and latency-sensitive tasks. These tasks may span multiple independent yet interconnected domains. Each domain may deploy a large number of heterogeneous devices, such as sensors, cameras, and actuators, generating diverse data types, such as structured data, images, and video streams. IoT devices have limited computing resources and energy consumption, resulting in low task processing efficiency. Related technologies are improving task processing efficiency by deploying edge servers with high processing capabilities.

[0004] Different tasks may have different processing requirements. When the processing requirements are low latency and small computational load, the IoT can directly perform local computations. When the processing requirements are high reliability and low latency, the processing can be offloaded to edge servers with idle resources.

[0005] Therefore, in order to meet diverse demands such as the processing requirements of different tasks and the reasonable allocation of computing resources, the multi-domain Internet of Things urgently needs an effective task allocation optimization method. Summary of the Invention

[0006] The present invention provides a multi-domain Internet of Things task allocation optimization model training method, device, equipment and medium, which are used to solve the defect that multi-domain Internet of Things task allocation in related technologies cannot meet diversified needs, optimize multi-domain Internet of Things task allocation, and meet diversified needs.

[0007] In a first aspect, the present invention provides a multi-domain Internet of Things task allocation optimization model training method, comprising:

[0008] Generate a first multidimensional state tensor based on task assignment information, operation performance indicators, task attribute information, and task communication quality indicators of multiple single-domain Internet of Things and edge server groups; wherein the single-domain Internet of Things is used to assign corresponding pending tasks to local devices or the edge server groups for processing;

[0009] Inputting the first multidimensional state tensor into a task allocation optimization model to be trained, so that the task allocation optimization model to be trained performs task allocation optimization based on the first multidimensional state tensor to obtain an allocation optimization strategy;

[0010] Based on the allocation optimization strategy, optimizing the allocation of the to-be-processed tasks by at least one of the single-domain Internet of Things, and generating a second multidimensional state tensor based on the optimized task allocation information, the operation performance index, the task attribute information, and the task communication quality index;

[0011] Determining an optimized reward value based on the task assignment information, the optimized task assignment information, the constructed task processing cost objective function and the reward function;

[0012] Based on the optimization reward value, the second multidimensional state tensor, the first multidimensional state tensor and the deep reinforcement learning strategy, the task allocation optimization model to be trained is updated to obtain a trained task allocation optimization model.

[0013] Optionally, generating a first multidimensional state tensor according to task allocation information, operation performance indicators, task attribute information, and task communication quality indicators of multiple single-domain IoTs and edge server groups includes:

[0014] Generate a task assignment identification matrix, a node feature matrix, a task feature matrix, and a three-dimensional network feature tensor based on the task assignment information, operation performance indicators, task attribute information, and task communication quality indicators of the multiple single-domain Internet of Things and the edge server group;

[0015] The task allocation identification matrix, the node feature matrix, the task feature matrix and the three-dimensional network feature tensor are taken as the first multi-dimensional state tensor as a whole.

[0016] Optionally, the edge server group includes multiple edge servers; the task attribute information includes key attribute information of each of the pending tasks; and the task communication quality indicators include estimated communication quality indicators of any of the pending tasks in the corresponding single-domain Internet of Things and each of the edge servers.

[0017] Optionally, generating a task allocation identification matrix, a node feature matrix, a task feature matrix, and a three-dimensional network feature tensor based on the task allocation information, operation performance indicators, task attribute information, and task communication quality indicators of the multiple single-domain Internet of Things and the edge server group includes:

[0018] Based on the task allocation information of the multiple single-domain Internet of Things and the multiple edge servers, construct the task allocation identification matrix, where each row of data in the task allocation identification matrix is ​​used to identify whether the single-domain Internet of Things allocates the to-be-processed task to the local device or the edge server;

[0019] Using the operating performance indicators of each of the single-domain Internet of Things and each of the edge servers, constructing the node feature matrix, wherein each row of data in the node feature matrix includes the operating performance indicator of the single-domain Internet of Things or the edge server;

[0020] Based on the key attribute information of each task to be processed, a task feature matrix is ​​constructed, wherein each row of the task feature matrix includes the key attribute information of the task to be processed;

[0021] A three-dimensional network feature tensor is constructed using the estimated communication quality indicators of each of the tasks to be processed in the corresponding single-domain Internet of Things and each of the edge servers. Each row of data in the three-dimensional network feature tensor includes the estimated communication quality indicators of the tasks to be processed in the corresponding single-domain Internet of Things and each of the edge servers.

[0022] Optionally, determining the optimized reward value according to the task allocation information, the optimized task allocation information, the constructed task processing cost objective function and the reward function includes:

[0023] determining a first task processing cost based on the task allocation information and the task processing cost objective function; and determining a second task processing cost based on the optimized task allocation information and the task processing cost objective function;

[0024] subtracting the first task processing cost from the second task processing cost to obtain a cost difference;

[0025] The cost difference is input into the reward function to perform reward calculation to determine the optimized reward value.

[0026] Optionally, inputting the cost difference into the reward function to perform reward calculation to determine the optimized reward value includes:

[0027] Inputting the cost difference into the reward function so that the reward function: when the cost difference is greater than 0, determines a first set value greater than 0 as the optimized reward value; when the cost difference is equal to 0, determines a second set value less than 0 as the optimized reward value; and when the cost difference is less than 0, determines a third set value less than 0 as the optimized reward value;

[0028] The absolute value of the third set value is equal to the first set value, and the absolute value of the third set value is greater than the absolute value of the second set value.

[0029] Optionally, when the task allocation optimization model to be trained is a deep Q network DQN model, the task allocation optimization model to be trained includes an estimation network, a target network and an error function;

[0030] The updating of the task allocation optimization model to be trained based on the optimization reward value, the second multidimensional state tensor, the first multidimensional state tensor, and the deep reinforcement learning strategy to obtain a trained task allocation optimization model includes:

[0031] Inputting the first multidimensional state tensor into the estimation network to perform action benefit prediction to obtain a first benefit value; inputting the second multidimensional state tensor into the target network to perform expected benefit prediction to obtain a second benefit value;

[0032] Inputting the first benefit value, the second benefit value, and the optimization reward value into the error function, so that the error function: performs a weighted summation of the optimization reward value and the second benefit value based on a set weight to obtain a target benefit value, and determines a loss function value based on a difference between the target benefit value and the first benefit value;

[0033] The task allocation optimization model to be trained is updated based on the loss function value to obtain a trained task allocation optimization model.

[0034] In a second aspect, the present invention provides a multi-domain Internet of Things task allocation optimization model training device, comprising:

[0035] A first generating unit is configured to generate a first multidimensional state tensor based on task allocation information, operation performance indicators, task attribute information, and task communication quality indicators of multiple single-domain Internet of Things and edge server groups; wherein the single-domain Internet of Things is configured to allocate corresponding pending tasks to a local device or the edge server group for processing;

[0036] an input unit, configured to input the first multidimensional state tensor into a task allocation optimization model to be trained, so that the task allocation optimization model to be trained performs task allocation optimization based on the first multidimensional state tensor to obtain an allocation optimization strategy;

[0037] an optimization unit, configured to optimize the allocation of the to-be-processed tasks by at least one of the single-domain Internet of Things based on the allocation optimization strategy;

[0038] A second generating unit is used to generate a second multidimensional state tensor based on the optimized task allocation information, the operation performance index, the task attribute information and the task communication quality index;

[0039] a determining unit, configured to determine an optimized reward value based on the task allocation information, the optimized task allocation information, the constructed task processing cost objective function, and the reward function;

[0040] An updating unit is used to update the task allocation optimization model to be trained based on the optimization reward value, the second multidimensional state tensor, the first multidimensional state tensor and the deep reinforcement learning strategy to obtain a trained task allocation optimization model.

[0041] In a third aspect, the present invention provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to thereby execute the multi-domain Internet of Things task allocation optimization model training method of the first aspect or any corresponding embodiment thereof.

[0042] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the multi-domain Internet of Things task allocation optimization model training method of the first aspect or any corresponding embodiment thereof.

[0043] The multi-domain Internet of Things task allocation optimization model training method, device, equipment and medium provided by the present invention can generate a first multidimensional state tensor based on the task allocation information, operation performance indicators, task attribute information and task communication quality indicators of multiple single-domain Internet of Things and edge server groups, and input the first multidimensional state tensor into the task allocation optimization model to be trained, so that the task allocation optimization model to be trained performs task allocation optimization based on the first multidimensional state tensor to obtain an allocation optimization strategy. Based on the allocation optimization strategy, the allocation of at least one single-domain Internet of Things to be processed tasks is optimized, and based on the optimized task allocation information, operation performance indicators, task attribute information and task communication quality indicators, a second multidimensional state tensor is generated. According to the task allocation information, the optimized task allocation information, the constructed task processing cost objective function and reward function, the optimized reward value is determined. Based on the optimized reward value, the second multidimensional state tensor, the first multidimensional state tensor and the deep reinforcement learning strategy, the task allocation optimization model to be trained is updated to obtain a trained task allocation optimization model. The task allocation optimization model trained by the present invention can optimize the task allocation of multi-domain Internet of Things while taking into account diverse requirements such as the processing requirements of different tasks and the reasonable allocation of computing resources, and effectively achieve diverse goals such as different task processing requirements and the reasonable allocation of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following is a brief introduction to the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0045] Figure 1 A flowchart of a multi-domain IoT task allocation optimization model training method provided by an embodiment of the present invention;

[0046] Figure 2 A flowchart of another multi-domain IoT task allocation optimization model training method provided by an embodiment of the present invention;

[0047] Figure 3 A schematic diagram of the structure of a multi-domain Internet of Things task allocation optimization model training device provided by an embodiment of the present invention;

[0048] Figure 4 A schematic structural diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0049] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0050] The following combination Figure 1-Figure 2 The present invention describes a multi-domain Internet of Things task allocation optimization model training method.

[0051] like Figure 1 As shown, this embodiment proposes a first multi-domain IoT task allocation optimization model training method, which may include the following steps:

[0052] S101. Generate a first multidimensional state tensor based on task allocation information, operational performance indicators, task attribute information, and task communication quality indicators of multiple single-domain Internet of Things (IoTs) and edge server groups; wherein the single-domain IoT is used to allocate corresponding pending tasks to local devices or edge server groups for processing.

[0053] Each single-domain IoT can host multiple IoT devices. Different single-domain IoTs are independent yet interconnected. It should be noted that multiple independent yet interconnected single-domain IoTs constitute a multi-domain IoT.

[0054] Specifically, an edge server group can be used to perform edge computing on multi-domain IoT tasks. The edge server group can include multiple edge servers, each of which can be used to handle tasks offloaded from a single domain IoT.

[0055] The task allocation information may include information about the assignment of tasks generated by each single-domain IoT to a local device or a specific edge server for processing. It should be understood that a local device is the IoT device within the corresponding single-domain IoT that processes the task. Specifically, when a single-domain IoT assigns a task to a local device for processing, it means that the single-domain IoT assigns the task to the IoT device within that single-domain IoT that processes the task.

[0056] Among them, the operating performance indicators may include indicator data related to the operating performance of each single-domain IoT and each edge server, such as CPU utilization, memory, number of queued tasks, signal-to-noise ratio, bandwidth and transmission rate.

[0057] Optionally, in other multi-domain Internet of Things task allocation optimization model training methods proposed in this embodiment, the edge server group includes multiple edge servers; the task attribute information includes key attribute information of each pending task; and the task communication quality indicator includes the estimated communication quality indicator of any pending task in the corresponding single-domain Internet of Things and each edge server.

[0058] Specifically, the task attribute information may include relevant attribute information of each task generated by each single-domain IoT, such as the computing requirements, data volume, and deadline of the task.

[0059] Specifically, the task communication quality indicators may include the estimated communication quality indicators of any task in each single-domain IoT and each edge server, such as signal-to-noise ratio and round-trip delay.

[0060] Specifically, this embodiment can construct a first multidimensional state tensor based on the task allocation information, operation performance indicators, task attribute information, and task communication quality indicators of each single-domain Internet of Things and each edge server.

[0061] Optionally, step S101 may include:

[0062] Generate a task assignment identification matrix, a node feature matrix, a task feature matrix, and a three-dimensional network feature tensor based on the task assignment information, operation performance indicators, task attribute information, and task communication quality indicators of multiple single-domain IoT and edge server groups;

[0063] The task assignment identification matrix, the node feature matrix, the task feature matrix and the three-dimensional network feature tensor are taken as a whole as the first multidimensional state tensor.

[0064] Among them, the task allocation identification matrix is ​​a matrix used to describe task allocation information, the node feature matrix is ​​a matrix including the operating performance indicators of each single-domain Internet of Things and each edge server, the task feature matrix is ​​a matrix including the task attribute information of each task, and the three-dimensional network feature tensor includes the estimated communication quality indicators of any task in each single-domain Internet of Things and each edge server.

[0065] Optionally, the above generates a task allocation identification matrix, a node feature matrix, a task feature matrix, and a three-dimensional network feature tensor based on the task allocation information, operation performance indicators, task attribute information, and task communication quality indicators of multiple single-domain IoTs and edge server groups, including:

[0066] Based on the task allocation information of multiple single-domain IoTs and multiple edge servers, a task allocation identification matrix is ​​constructed. Each row of data in the task allocation identification matrix is ​​used to identify whether the single-domain IoT allocates the task to be processed to a local device or an edge server.

[0067] Using the operating performance indicators of each single-domain IoT and each edge server, a node feature matrix is ​​constructed. Each row of the node feature matrix includes the operating performance indicator of the single-domain IoT or edge server.

[0068] Based on the key attribute information of each task to be processed, a task feature matrix is ​​constructed, and each row of the task feature matrix includes the key attribute information of the task to be processed;

[0069] A three-dimensional network feature tensor is constructed using the estimated communication quality indicators of each pending task in the corresponding single-domain Internet of Things and each edge server. Each row of data in the three-dimensional network feature tensor includes the estimated communication quality indicators of the pending task in the corresponding single-domain Internet of Things and each edge server.

[0070] Specifically, the task assignment identification matrix can be expressed as in, M The total number of tasks representing multi-domain IoT, N Represents the total number of nodes available for computing, including each single-domain IoT and each edge server. When the current task is at the node n Execute, otherwise This matrix representation can not only intuitively reflect the distribution of tasks, but also facilitate efficient state updates through matrix operations, thereby reducing computational complexity.

[0071] Specifically, the task feature matrix can be expressed as Each row of data corresponds to a task m The attribute vector , including key parameters such as computing requirements, data volume, and deadline.

[0072] Specifically, the node feature matrix can be expressed as Each row of data represents a node n The real-time status vector can include performance indicators such as CPU utilization, memory, and available bandwidth.

[0073] Among them, the three-dimensional network feature tensor can be expressed as . Three-dimensional matrix elements Describe the task m With node n The estimated communication quality between them, such as signal-to-noise ratio and round-trip delay.

[0074] Specifically, in this embodiment, the constructed task allocation identification matrix, node feature matrix, task feature matrix and three-dimensional network feature tensor can be used as a whole as the first multi-dimensional state tensor.

[0075] Among them, the first multidimensional state tensor can be expressed as .

[0076] S102: Input the first multidimensional state tensor into the task allocation optimization model to be trained, so that the task allocation optimization model to be trained performs task allocation optimization based on the first multidimensional state tensor to obtain an allocation optimization strategy.

[0077] Among them, the task allocation optimization model to be trained can be a model with a deep reinforcement network architecture, which is to undergo deep reinforcement learning to achieve task allocation optimization capabilities.

[0078] Optionally, the task allocation optimization model to be trained may be a Deep Q-Network (DQN) model or another type of deep reinforcement network architecture model.

[0079] Specifically, this embodiment can input the constructed first multidimensional state tensor into the task allocation optimization model to be trained, so that the task allocation optimization model performs task allocation optimization based on the first multidimensional state tensor, and obtains and outputs the allocation optimization strategy.

[0080] S103: Optimize the allocation of tasks to be processed by at least one single-domain Internet of Things based on the allocation optimization strategy.

[0081] Specifically, this embodiment can optimize the task allocation of each single-domain IoT according to the allocation optimization strategy, that is, adjust some tasks from local devices to edge servers for processing, and adjust some tasks from edge servers to local devices for processing.

[0082] S104: Generate a second multidimensional state tensor based on the optimized task allocation information, operation performance index, task attribute information, and task communication quality index.

[0083] Specifically, this embodiment can obtain the task allocation information, operation performance indicators, task attribute information and task communication quality indicators of the multi-domain Internet of Things and the edge server group after the allocation optimization after optimizing the task allocation of the multi-domain Internet of Things based on the allocation optimization strategy, and refer to the specific process of generating the first multidimensional state tensor in step S101, and generate the corresponding second multidimensional state tensor based on the task allocation information, operation performance indicators, task attribute information and task communication quality indicators of the multi-domain Internet of Things and the edge server group after the allocation optimization.

[0084] It can be understood that the second multi-dimensional state tensor is also composed of the task allocation identification matrix, the node feature matrix, the task feature matrix and the three-dimensional network feature tensor.

[0085] S105 : Determine an optimized reward value based on the task allocation information, the optimized task allocation information, the constructed task processing cost objective function, and the reward function.

[0086] Specifically, this embodiment can construct a task processing cost objective function and a reward function according to actual conditions.

[0087] Specifically, this embodiment can determine the corresponding optimization reward value based on the task allocation information of the multi-domain Internet of Things, the optimized task allocation information, the task processing cost objective function and the reward function.

[0088] Optionally, step S105 may include:

[0089] determining a first task processing cost based on the task allocation information and the task processing cost objective function; and determining a second task processing cost based on the optimized task allocation information and the task processing cost objective function;

[0090] Subtract the first task processing cost from the second task processing cost to obtain a cost difference;

[0091] The cost difference is input into the reward function for reward calculation to determine the optimized reward value.

[0092] Specifically, in this embodiment, the task processing cost objective function may be used to calculate based on the task allocation information to obtain the first task processing cost, and the task processing cost objective function may be used to calculate based on the optimized task allocation information to obtain the second task processing cost.

[0093] Optionally, the cost difference is input into the reward function to perform reward calculation to determine the optimized reward value, including:

[0094] The cost difference is input into the reward function so that the reward function: when the cost difference is greater than 0, a first set value greater than 0 is determined as the optimized reward value; when the cost difference is equal to 0, a second set value less than 0 is determined as the optimized reward value; when the cost difference is less than 0, a third set value less than 0 is determined as the optimized reward value;

[0095] The absolute value of the third setting value is equal to the first setting value, and the absolute value of the third setting value is greater than the absolute value of the second setting value.

[0096] S106. Based on the optimized reward value, the second multidimensional state tensor, the first multidimensional state tensor and the deep reinforcement learning strategy, the task allocation optimization model to be trained is updated to obtain a trained task allocation optimization model.

[0097] Optionally, in other multi-domain IoT task allocation optimization model training methods proposed in this embodiment, when the task allocation optimization model to be trained is a deep Q-network (DQN) model, the task allocation optimization model to be trained includes an estimation network, a target network, and an error function. In this case, step S106 may include:

[0098] Input the first multidimensional state tensor into the estimation network to predict the action benefit and obtain a first benefit value; input the second multidimensional state tensor into the target network to predict the expected benefit and obtain a second benefit value;

[0099] Inputting the first benefit value, the second benefit value, and the optimized reward value into an error function, so that the error function: performs a weighted summation of the optimized reward value and the second benefit value based on a set weight to obtain a target benefit value, and determines a loss function value based on the difference between the target benefit value and the first benefit value;

[0100] The task allocation optimization model to be trained is updated based on the loss function value to obtain a trained task allocation optimization model.

[0101] Specifically, the task allocation optimization model to be trained is a DQN model, which includes an estimation network, a target network, and a DQN error function.

[0102] Specifically, this embodiment considers the multi-domain IoT and edge server group as an environment. The network state of the environment at time t, i.e., a first multidimensional state tensor, is input into the estimation network to predict the action benefit, obtaining a first benefit value output by the estimation network. The network state of the environment at time t+1, i.e., a second multidimensional state tensor, is input into the target network to predict the expected benefit, obtaining a second benefit value.

[0103] Specifically, this embodiment can input the first benefit value, the second benefit value, and the optimization reward value into the DQN error function, so that the DQN error function can perform a weighted sum of the optimization reward value and the second benefit value based on a set weight to obtain a target benefit value. Based on the difference between the target benefit value and the first benefit value and the set loss function, a corresponding loss function value is calculated. Subsequently, when the loss function value is not less than a set threshold, this embodiment can update the parameters of the task allocation optimization model to be trained to obtain an updated task allocation optimization model.

[0104] Afterwards, this embodiment can use the updated task allocation optimization model as the task allocation optimization model to be trained, return to step S101, and continue to update the task allocation optimization model until the calculated loss function value is less than the set threshold, or the number of iterations meets the requirements, and obtain the latest task allocation optimization model. Afterwards, this embodiment can test and verify the latest task allocation optimization model. When the test and verification pass, the latest task allocation optimization model can be determined as the trained task allocation optimization model. If the test and verification fail, the latest task allocation optimization model can continue to be trained and updated until a trained task allocation optimization model is obtained.

[0105] The multi-domain Internet of Things task allocation optimization model training method proposed in this embodiment can generate a first multidimensional state tensor based on the task allocation information, operation performance indicators, task attribute information and task communication quality indicators of multiple single-domain Internet of Things and edge server groups, and input the first multidimensional state tensor into the task allocation optimization model to be trained, so that the task allocation optimization model to be trained performs task allocation optimization based on the first multidimensional state tensor to obtain an allocation optimization strategy. Based on the allocation optimization strategy, the allocation of at least one single-domain Internet of Things to be processed tasks is optimized, and based on the optimized task allocation information, operation performance indicators, task attribute information and task communication quality indicators, a second multidimensional state tensor is generated. According to the task allocation information, the optimized task allocation information, the constructed task processing cost objective function and the reward function, the optimized reward value is determined. Based on the optimized reward value, the second multidimensional state tensor, the first multidimensional state tensor and the deep reinforcement learning strategy, the task allocation optimization model to be trained is updated to obtain a trained task allocation optimization model. The task allocation optimization model trained in this embodiment can optimize the task allocation of the multi-domain Internet of Things while taking into account diverse requirements such as the processing requirements of different tasks and the reasonable allocation of computing resources, thereby effectively achieving diverse goals such as different task processing requirements and the reasonable allocation of computing resources.

[0106] based on Figure 1This embodiment proposes another multi-domain Internet of Things task allocation optimization model training method. This embodiment can construct a task processing cost objective function and a reward function according to the diverse needs in the multi-domain Internet of Things task allocation scenario.

[0107] Specifically, in the process of establishing the task processing cost objective function, the network model can be set to be D A single domain IoT and E edge servers, where different single-domain IoT can be represented as , the edge server is represented as . It is assumed that each single-domain IoT will generate separate tasks, which can be offloaded to the edge server.

[0108] equipment Generate Tasks ,in Represents the task input value, Represents the resources required to process the task, such as different numbers of CPU cycles.

[0109] The total energy consumption of executing the task can be expressed as:

[0110] ;

[0111] in, is the fixed frequency of the CPU, that is, the CPU cycle frequency required to process 1 bit of data, The energy consumption ratio corresponding to each CPU. is the task size.

[0112] The local computing latency is expressed as:

[0113] ;

[0114] in, Represents single-domain IoT The amount of data that a CPU can process per cycle.

[0115] The cost function is expressed as:

[0116] ;

[0117] in, A weight representing the cost of delay.

[0118] In the edge offloading task, the communication delay can be expressed as:

[0119] ;

[0120] ;

[0121] in, For single domain IoT and edge servers The communication delay between For single domain IoT and edge servers bandwidth channels between them. Representative tasks are assigned to edge servers . W is the shared bandwidth between the hypothetical edge server and the connected single-domain IoT. represent The channel gain of the single domain IoT With edge servers The communication gap between them. Represents single-domain IoT To the edge server The received signal power. It represents the use of complex Gaussian model to deal with the noise in the data. Represents the signal-to-noise ratio.

[0122] The transmission delay between the device and the server can be expressed as:

[0123] ;

[0124] in, Represents a single domain IoT To the edge server The amount of data.

[0125] Regarding the task computation model, when an edge server receives two or more offloaded tasks from different single-domain IoTs, it is assumed that the computation resources are evenly shared between the tasks. In this case, the amount of computation resources allocated to each single-domain IoT can be expressed as:

[0126] ;

[0127] .

[0128] in, Represents the edge server The task processing capability of all servers is assumed to be consistent. Represents the CPU cycles required by the edge server to process a task, Represents the energy consumption per CPU cycle. Represents the frequency of the CPU, To calculate the cost.

[0129] The latency of edge server executing tasks can be expressed as:

[0130] ;

[0131] in, Represents the CPU frequency.

[0132] The transmission delay of a task transmitted from a single domain IoT to an edge server can be expressed as:

[0133] .

[0134] Among them, when When the IoT device Successfully offloaded the task to Calculation is performed on the server to satisfy the constraints:

[0135] ;

[0136] Finally, the cost function of the edge computing model is expressed as

[0137] ;

[0138] in, P Represents the transmission energy consumption of the device, specifically the fixed transmission power of the device during communication.

[0139] The cost function can be expressed as:

[0140] .

[0141] Penalty function Applicable to task execution failures caused by violation of computational constraints.

[0142] Optimize the objective function and construct the task processing cost objective function:

[0143] .

[0144] Specifically, this embodiment can calculate the first task processing cost of the multi-domain Internet of Things and edge server group based on the task processing cost objective function. And after the multi-domain Internet of Things and edge server group adjust the task allocation according to the allocation optimization strategy, the second task processing cost of the multi-domain Internet of Things and edge server after task allocation optimization can be calculated based on the task processing cost objective function. The cost difference is obtained by subtracting the first task processing cost from the second task processing cost. .

[0145] Specifically, the reward function constructed in this embodiment is:

[0146] .

[0147] in, and is the set optimization reward value, all greater than 0, and Greater than .

[0148] Specifically, this embodiment can be based on the cost difference And reward function, determine the corresponding optimized reward value.

[0149] In related technologies, cloud computing suffers from high latency when processing remote tasks. Edge computing addresses this issue by placing computing tasks closer to the data source. The dynamic and heterogeneous nature of multi-domain IoT (such as smart agriculture and smart cities) makes task offloading an NP-hard problem. The key challenge is how to balance local computing with cloud offloading in real time in resource-constrained, dynamic, and ever-changing environments to simultaneously meet the requirements of low latency, high energy efficiency, and low cost. The inventors of this embodiment have discovered that due to the inherent complexity and dynamic nature of task offloading in multi-domain IoT, related technologies struggle to handle the high-dimensional state space (e.g., device load, network latency, energy consumption, etc.) of multi-domain IoT. DQN, on the other hand, uses a deep neural network to approximate the Q-value function, automatically extracting key features and achieving end-to-end adaptive decision-making. Its experience replay mechanism breaks data correlation and improves training stability, making it particularly suitable for cross-domain experience sharing. The target network mitigates the problem of Q-value overestimation, ensuring stable optimization of policies across multiple objectives such as latency and energy consumption. Furthermore, DQN has low computational overhead and can be deployed on edge devices through a lightweight design. It can learn long-term optimal policies without pre-defined rules. Therefore, this embodiment adopts the DQN model.

[0150] Furthermore, considering the complexity and continuous nature of edge computing, traditional Q-learning is difficult to use. As the search space increases, the speed of model convergence and optimal strategy acquisition decreases significantly. Therefore, a deep convolutional neural network is designed to estimate Q values, which can both accelerate convergence and reduce dimensionality.

[0151] It is understandable that this embodiment can be combined with a distributed deep reinforcement learning method to train the DQN model.

[0152] like Figure 2 As shown, when the DQN model to be trained includes an estimation network, a target network, and a DQN error function. Both the estimation network and the target network can be deep convolutional neural networks. In this embodiment, the multi-domain Internet of Things and the edge server group can be regarded as the environment as a whole, and the first multi-dimensional state tensor constructed can be regarded as the network state of the environment at time t. In this embodiment, the first multi-dimensional state tensor is input into the DQN model to perform task allocation optimization, obtain the allocation optimization strategy and return it to the environment. The allocation optimization strategy is regarded as an action, and the environment responds to the action to generate the corresponding second multi-dimensional state tensor, that is, the network state of the environment at time t+1. The first multi-dimensional state tensor and the second multi-dimensional state tensor are both stored in the experience pool. In this embodiment, the calculated optimization reward value r t Also stored in the experience pool.

[0153] In this embodiment, the first multi-dimensional state tensor and the second multi-dimensional state tensor in the experience pool are input into the estimation network and the target network respectively. The estimation network predicts the action benefit based on the first multi-dimensional state tensor to obtain the first benefit value, i.e., the estimated Q value. The target network predicts the expected benefit based on the second multi-dimensional state tensor to obtain the second benefit value, i.e., the target Q value. The first benefit value, the second benefit value, and the determined optimization reward value r are combined into a target reward value. t This is input into the DQN error function for calculation to obtain the corresponding loss function value. The parameters in the estimation network are updated based on this loss function value, and the parameters in the estimation network are periodically synchronized and updated to the target network. This embodiment can then continue training the updated DQN model until a fully trained DQN model is obtained.

[0154] This embodiment uses a distributed deep reinforcement learning framework to model the Internet of Things as an intelligent agent and optimize task offloading strategies through a Markov decision process. The DQN model and convolutional neural network are used as function approximators to accelerate convergence and process high-dimensional state spaces. Dynamic resource allocation strategies are designed, taking into account device computing power, bandwidth, energy status, and other factors. This embodiment designs an energy-efficient distributed task offloading system and constructs an offloading knowledge model that enables efficient, real-time load decision-making, thereby minimizing idle resources and response time.

[0155] This embodiment enables multi-domain IoT network architecture design and supports dynamic task offloading. A distributed deep reinforcement learning algorithm allows devices to make independent decisions without requiring global information. Deep convolutional neural networks compress the state space, improving learning efficiency and convergence speed.

[0156] This embodiment can achieve efficient and adaptive task offloading through distributed deep reinforcement learning in the dynamic and resource-constrained environment of the multi-domain Internet of Things. It needs to overcome difficulties such as high-dimensional decision-making, multi-objective optimization, and cross-domain collaboration, and ultimately achieve a low-latency, low-energy, and highly reliable edge intelligent system.

[0157] This embodiment proposes a multi-domain IoT task allocation optimization model training method and a distributed deep reinforcement learning task offloading optimization approach. By constructing a knowledge model that dynamically balances local computing with edge offloading, it achieves lower latency, higher energy efficiency, and faster convergence. It also provides a scalable and adaptive solution for implementation in complex IoT scenarios.

[0158] like Figure 3 As shown, this embodiment proposes a multi-domain Internet of Things task allocation optimization model training device, which may include:

[0159] A first generating unit 301 is configured to generate a first multidimensional state tensor based on task allocation information, operation performance indicators, task attribute information, and task communication quality indicators of multiple single-domain Internet of Things and edge server groups; wherein the single-domain Internet of Things is configured to allocate corresponding pending tasks to local devices or edge server groups for processing;

[0160] An input unit 302 is configured to input the first multidimensional state tensor into the task allocation optimization model to be trained, so that the task allocation optimization model to be trained performs task allocation optimization based on the first multidimensional state tensor to obtain an allocation optimization strategy;

[0161] An optimization unit 303 is configured to optimize the allocation of tasks to be processed by at least one single-domain Internet of Things based on the allocation optimization strategy;

[0162] A second generating unit 304 is configured to generate a second multidimensional state tensor based on the optimized task allocation information, the operation performance index, the task attribute information, and the task communication quality index;

[0163] A determination unit 305 is configured to determine an optimized reward value based on the task allocation information, the optimized task allocation information, the constructed task processing cost objective function, and the reward function;

[0164] The updating unit 306 is used to update the task allocation optimization model to be trained based on the optimization reward value, the second multidimensional state tensor, the first multidimensional state tensor and the deep reinforcement learning strategy to obtain a trained task allocation optimization model.

[0165] It should be noted that the processing of the first generating unit 301, the input unit 302, the optimizing unit 303, the second generating unit 304, the determining unit 305 and the updating unit 306 and the beneficial effects thereof can be referred to in the respective Figure 1 Steps S101 to S105 in the above are not described in detail.

[0166] Optionally, the first generating unit 301 is further configured to:

[0167] Generate a task assignment identification matrix, a node feature matrix, a task feature matrix, and a three-dimensional network feature tensor based on the task assignment information, operation performance indicators, task attribute information, and task communication quality indicators of multiple single-domain IoT and edge server groups;

[0168] The task assignment identification matrix, the node feature matrix, the task feature matrix and the three-dimensional network feature tensor are taken as a whole as the first multidimensional state tensor.

[0169] Optionally, the edge server group includes multiple edge servers; the task attribute information includes key attribute information of each pending task; and the task communication quality indicator includes the estimated communication quality indicator of any pending task in the corresponding single-domain Internet of Things and each edge server.

[0170] Optionally, the first generating unit 301 is further configured to:

[0171] Based on the task allocation information of multiple single-domain IoTs and multiple edge servers, a task allocation identification matrix is ​​constructed. Each row of data in the task allocation identification matrix is ​​used to identify whether the single-domain IoT allocates the task to be processed to a local device or an edge server.

[0172] Using the operating performance indicators of each single-domain IoT and each edge server, a node feature matrix is ​​constructed. Each row of the node feature matrix includes the operating performance indicator of the single-domain IoT or edge server.

[0173] Based on the key attribute information of each task to be processed, a task feature matrix is ​​constructed, and each row of the task feature matrix includes the key attribute information of the task to be processed;

[0174] A three-dimensional network feature tensor is constructed using the estimated communication quality indicators of each pending task in the corresponding single-domain Internet of Things and each edge server. Each row of data in the three-dimensional network feature tensor includes the estimated communication quality indicators of the pending task in the corresponding single-domain Internet of Things and each edge server.

[0175] Optionally, the determining unit 305 is further configured to:

[0176] determining a first task processing cost based on the task allocation information and the task processing cost objective function; and determining a second task processing cost based on the optimized task allocation information and the task processing cost objective function;

[0177] Subtract the first task processing cost from the second task processing cost to obtain a cost difference;

[0178] The cost difference is input into the reward function for reward calculation to determine the optimized reward value.

[0179] Optionally, the determining unit 305 is further configured to:

[0180] The cost difference is input into the reward function so that the reward function: when the cost difference is greater than 0, a first set value greater than 0 is determined as the optimized reward value; when the cost difference is equal to 0, a second set value less than 0 is determined as the optimized reward value; when the cost difference is less than 0, a third set value less than 0 is determined as the optimized reward value;

[0181] The absolute value of the third setting value is equal to the first setting value, and the absolute value of the third setting value is greater than the absolute value of the second setting value.

[0182] Optionally, when the task allocation optimization model to be trained is a deep Q network DQN model, the task allocation optimization model to be trained includes an estimation network, a target network, and an error function;

[0183] The updating unit 306 is further configured to:

[0184] Input the first multidimensional state tensor into the estimation network to predict the action benefit and obtain a first benefit value; input the second multidimensional state tensor into the target network to predict the expected benefit and obtain a second benefit value;

[0185] Inputting the first benefit value, the second benefit value, and the optimized reward value into an error function, so that the error function: performs a weighted summation of the optimized reward value and the second benefit value based on a set weight to obtain a target benefit value, and determines a loss function value based on the difference between the target benefit value and the first benefit value;

[0186] The task allocation optimization model to be trained is updated based on the loss function value to obtain a trained task allocation optimization model.

[0187] The multi-domain IoT task allocation optimization model training device proposed in this embodiment can generate a first multi-dimensional state tensor based on the task allocation information, operating performance indicators, task attribute information, and task communication quality indicators of multiple single-domain IoTs and edge server groups, and input it into the task allocation optimization model to be trained, so that the task allocation optimization model to be trained performs task allocation optimization based on the first multi-dimensional state tensor, obtains an allocation optimization strategy, and updates the task allocation optimization model to be trained based on the allocation optimization strategy to obtain a trained task allocation optimization model. The trained task allocation optimization model of this embodiment can optimize the task allocation of the multi-domain IoT while taking into account diverse requirements such as the processing requirements of different tasks and the reasonable allocation of computing resources, effectively achieving diverse goals such as the processing requirements of different tasks and the reasonable allocation of computing resources.

[0188] The multi-domain IoT task allocation optimization model training device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0189] The embodiment of the present invention also provides a computer device having the above Figure 3 The multi-domain IoT task allocation optimization model training device shown.

[0190] See also Figure 4 , a structural diagram of a computer device provided by an optional embodiment of the present invention, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses for communication, and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed in the computer device, including instructions stored in or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 4 A processor 10 is taken as an example.

[0191] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.

[0192] The memory 20 stores instructions that can be executed by at least one processor 10, so as to enable at least one processor 10 to execute the method shown in the above embodiment.

[0193] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0194] The memory 20 may include volatile memory, such as random access memory. The memory may also include non-volatile memory, such as flash memory, a hard disk, or a solid-state drive. The memory 20 may also include a combination of the above types of memory.

[0195] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0196] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0197] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A multi-domain Internet of Things task allocation optimization model training method, characterized in that: include: Generate a first multidimensional state tensor based on task assignment information, operation performance indicators, task attribute information, and task communication quality indicators of multiple single-domain Internet of Things and edge server groups; wherein the single-domain Internet of Things is used to assign corresponding pending tasks to local devices or the edge server groups for processing; Inputting the first multidimensional state tensor into a task allocation optimization model to be trained, so that the task allocation optimization model to be trained performs task allocation optimization based on the first multidimensional state tensor to obtain an allocation optimization strategy; Based on the allocation optimization strategy, optimizing the allocation of the to-be-processed tasks by at least one of the single-domain Internet of Things, and generating a second multidimensional state tensor based on the optimized task allocation information, the operation performance index, the task attribute information, and the task communication quality index; Determining an optimized reward value based on the task assignment information, the optimized task assignment information, the constructed task processing cost objective function and the reward function; Based on the optimization reward value, the second multidimensional state tensor, the first multidimensional state tensor and the deep reinforcement learning strategy, the task allocation optimization model to be trained is updated to obtain a trained task allocation optimization model.

2. The method according to claim 1, characterized in that The generating of a first multidimensional state tensor according to task allocation information, operation performance indicators, task attribute information, and task communication quality indicators of multiple single-domain IoTs and edge server groups includes: Generate a task assignment identification matrix, a node feature matrix, a task feature matrix, and a three-dimensional network feature tensor based on the task assignment information, operation performance indicators, task attribute information, and task communication quality indicators of the multiple single-domain Internet of Things and the edge server group; The task allocation identification matrix, the node feature matrix, the task feature matrix and the three-dimensional network feature tensor are taken as the first multi-dimensional state tensor as a whole.

3. The method according to claim 2, characterized in that The edge server group includes multiple edge servers; the task attribute information includes key attribute information of each of the pending tasks; and the task communication quality index includes an estimated communication quality index of any of the pending tasks in the corresponding single-domain Internet of Things and each of the edge servers.

4. The method according to claim 3, characterized in that The generating of a task assignment identification matrix, a node feature matrix, a task feature matrix, and a three-dimensional network feature tensor based on the task assignment information, operation performance indicators, task attribute information, and task communication quality indicators of the multiple single-domain Internet of Things and the edge server group includes: Based on the task allocation information of the multiple single-domain Internet of Things and the multiple edge servers, construct the task allocation identification matrix, where each row of data in the task allocation identification matrix is ​​used to identify whether the single-domain Internet of Things allocates the to-be-processed task to the local device or the edge server; Using the operating performance indicators of each of the single-domain Internet of Things and each of the edge servers, constructing the node feature matrix, wherein each row of data in the node feature matrix includes the operating performance indicator of the single-domain Internet of Things or the edge server; Based on the key attribute information of each task to be processed, construct the task feature matrix, wherein each row of data in the task feature matrix includes the key attribute information of the task to be processed; The three-dimensional network feature tensor is constructed using the estimated communication quality indicators of each of the tasks to be processed in the corresponding single-domain Internet of Things and each of the edge servers, wherein each row of data of the three-dimensional network feature tensor includes the estimated communication quality indicators of the tasks to be processed in the corresponding single-domain Internet of Things and each of the edge servers.

5. The method according to claim 1, wherein The step of determining an optimized reward value based on the task allocation information, the optimized task allocation information, the constructed task processing cost objective function, and the reward function includes: determining a first task processing cost based on the task allocation information and the task processing cost objective function; and determining a second task processing cost based on the optimized task allocation information and the task processing cost objective function; subtracting the first task processing cost from the second task processing cost to obtain a cost difference; The cost difference is input into the reward function to perform reward calculation to determine the optimized reward value.

6. The method according to claim 5, characterized in that Inputting the cost difference into the reward function to perform reward calculation to determine the optimized reward value includes: Inputting the cost difference into the reward function so that the reward function: when the cost difference is greater than 0, determines a first set value greater than 0 as the optimized reward value; when the cost difference is equal to 0, determines a second set value less than 0 as the optimized reward value; and when the cost difference is less than 0, determines a third set value less than 0 as the optimized reward value; The absolute value of the third set value is equal to the first set value, and the absolute value of the third set value is greater than the absolute value of the second set value.

7. The method according to any one of claims 1 to 6, characterized in that When the task allocation optimization model to be trained is a deep Q network DQN model, the task allocation optimization model to be trained includes an estimation network, a target network and an error function; The updating of the task allocation optimization model to be trained based on the optimization reward value, the second multidimensional state tensor, the first multidimensional state tensor, and the deep reinforcement learning strategy to obtain a trained task allocation optimization model includes: Inputting the first multidimensional state tensor into the estimation network to perform action benefit prediction to obtain a first benefit value; inputting the second multidimensional state tensor into the target network to perform expected benefit prediction to obtain a second benefit value; Inputting the first benefit value, the second benefit value, and the optimization reward value into the error function, so that the error function: performs a weighted summation of the optimization reward value and the second benefit value based on a set weight to obtain a target benefit value, and determines a loss function value based on a difference between the target benefit value and the first benefit value; The task allocation optimization model to be trained is updated based on the loss function value to obtain a trained task allocation optimization model.

8. A multi-domain Internet of Things task allocation optimization model training device, characterized in that: include: A first generating unit is configured to generate a first multidimensional state tensor based on task allocation information, operation performance indicators, task attribute information, and task communication quality indicators of multiple single-domain Internet of Things and edge server groups; wherein the single-domain Internet of Things is configured to allocate corresponding pending tasks to a local device or the edge server group for processing; an input unit, configured to input the first multidimensional state tensor into a task allocation optimization model to be trained, so that the task allocation optimization model to be trained performs task allocation optimization based on the first multidimensional state tensor to obtain an allocation optimization strategy; an optimization unit, configured to optimize the allocation of the to-be-processed tasks by at least one of the single-domain Internet of Things based on the allocation optimization strategy; A second generating unit is used to generate a second multidimensional state tensor based on the optimized task allocation information, the operation performance index, the task attribute information and the task communication quality index; a determining unit, configured to determine an optimized reward value based on the task allocation information, the optimized task allocation information, the constructed task processing cost objective function, and the reward function; An updating unit is used to update the task allocation optimization model to be trained based on the optimization reward value, the second multidimensional state tensor, the first multidimensional state tensor and the deep reinforcement learning strategy to obtain a trained task allocation optimization model.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the multi-domain Internet of Things task allocation optimization model training method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the multi-domain Internet of Things task allocation optimization model training method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Adaptive multi-type task scheduling and multi-domain resource configuration joint design method

    CN117032928A

  • Scalability of reinforcement learning by separation of concerns

    US20180165602A1