Task sorting method, system and equipment and storage medium

By combining data fusion and reinforcement learning methods for enterprise tasks and dynamically adjusting task sorting, the problem of insufficient task type similarity and static data adaptability in the existing technology is solved, and the efficiency and accuracy of task sorting are improved.

CN119940780APending Publication Date: 2025-05-06粟健明
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411873046.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When sorting tasks in enterprises, the prior art ignores the similarity of task types, and relying on static data cannot adapt to the dynamic operation environment, resulting in low sorting efficiency and poor accuracy.

Method used

By obtaining the historical data of the tasks to be sorted and the corresponding task types, data fusion is carried out to generate a first order table of the importance of the task type, and sorting the tasks to be sorted under the same task type to generate a second order table. In combination with reinforcement learning method, dynamically adjust the initial order to obtain the final sort.

Benefits of technology

Through stratified optimization, avoid overall calculation consumption, obtain a reasonable initial sort, and flexibly regulate according to the current state and resource conditions through reinforcement learning, improving the accuracy of task sorting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940780A_ABST
    Figure CN119940780A_ABST
Patent Text Reader

Abstract

The invention provides a task sorting method, system and device and a storage medium, and the method comprises the steps: obtaining to-be-sorted tasks and historical data of corresponding task types; performing data combination on the basis of historical data, generating a piece of corresponding data for each task type to serve as fusion data, and inputting the fusion data into a pre-established regression model to obtain a first sequence table for sorting importance degrees of the task types; filling the to-be-sorted tasks under the corresponding task types in the first sequence table, and sorting the to-be-sorted tasks belonging to the same task type by using a regression model to obtain a corresponding second sequence table under each task type; and taking the first sequence table as a first priority and the second sequence table as a second priority to obtain an initial sequence of the to-be-sorted tasks, and adjusting the initial sequence by using a reinforcement learning method to obtain a final sequence. According to the method, the initial sequence of the specific tasks is dynamically adjusted through interaction with the environment by using reinforcement learning, and the accuracy of task sorting is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of reinforcement learning technology, and in particular to a task sequencing method, system, device and storage medium. Background Art

[0002] Prioritizing tasks in an enterprise is crucial. It not only helps improve resource utilization efficiency, enhance productivity and performance, but also helps enterprises better respond to emergencies and enhance decision-making quality.

[0003] In related technologies, with the development of big data technology, existing technologies often use machine learning, correlation analysis and other technical means to score each task based on the existing static historical data of each task, and sort them based on the importance score, arranging tasks with higher importance scores first, and then arranging tasks with lower importance scores; however, the above technical means have the following defects: on the one hand, if a large number of tasks are directly sorted for enterprise managers, it may cause a lot of computing loss, and there are similarities between tasks belonging to the same task type. For example, accounting work is an important part of the enterprise, and operations management is less urgent and important than accounting work. Therefore, accounting, accounting supervision and other tasks belonging to accounting work may be more important than the specific tasks under operations management. The existing sorting method ignores the similarity of task types; on the other hand, it only relies on static importance score sorting without using other dynamic optimization methods, and cannot adapt to the dynamic operating environment of the enterprise.

[0004] Based on the above analysis of the development status of this technology field, the existing technology lacks the importance of prioritizing different task types and adopts reinforcement learning algorithms to dynamically adjust the solution based on the initial order of static data. Summary of the Invention

[0005] The purpose of the present invention is to provide a task scheduling method, system, device and storage medium, aiming to solve the above-mentioned problems in the prior art.

[0006] According to a first aspect of an embodiment of the present invention, there is provided a task sequencing method, comprising:

[0007] Get the historical data of tasks to be sorted and corresponding task types;

[0008] Perform data fusion based on historical data to generate a corresponding fused data for each task type. Input the fused data into the pre-established regression model to obtain a first-order table of task type importance.

[0009] Fill the tasks to be sorted into the corresponding task type in the first sequence table, use the regression model to sort the tasks to be sorted belonging to the same task type, and obtain the second sequence table corresponding to each task type;

[0010] The first sequence table is used as the first priority, and the second sequence table is used as the second priority to obtain the initial order of the tasks to be sorted. The reinforcement learning method is used to adjust the initial order to obtain the final order.

[0011] According to a second aspect of an embodiment of the present invention, there is provided a task sequencing device, comprising:

[0012] Data acquisition module, used to obtain historical data of tasks to be sorted and corresponding task types;

[0013] The type sorting module is used to perform data fusion based on historical data, generate a corresponding fused data for each task type, input the fused data into the pre-established regression model, and obtain the first order table of task type importance;

[0014] A task sorting module is used to fill the tasks to be sorted into the corresponding task type in the first sequence table, and use the regression model to sort the tasks to be sorted belonging to the same task type to obtain the second sequence table corresponding to each task type;

[0015] The adjustment module is used to obtain the initial order of the tasks to be sorted with the first sequence table as the first priority and the second sequence table as the second priority, and use the reinforcement learning method to adjust the initial order to obtain the final order.

[0016] According to a third aspect of an embodiment of the present invention, an electronic device is provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the task sorting method provided in the first aspect of the present disclosure are implemented.

[0017] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which a program for implementing information transmission is stored. When the program is executed by a processor, the steps of the task sorting method provided in the first aspect of the present disclosure are implemented.

[0018] The technical solution provided by the embodiment of the present invention includes the following beneficial effects: on the basis of sorting the importance of task types, the specific tasks to be sorted are sorted, so as to avoid the huge consumption brought by the overall calculation by using a hierarchical optimization method, and obtain a reasonable initial order in a short period of time; reinforcement learning is used to dynamically adjust the initial order of specific tasks through interaction with the environment, which can be flexibly adjusted according to the current status and resource conditions, thereby improving the accuracy of task sorting.

[0019] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate one or more embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 is a flowchart of a task sequencing method according to an embodiment of the present invention;

[0022] Figure 2 is a schematic diagram of a sorting architecture according to an embodiment of the present invention;

[0023] Figure 3 is a schematic diagram of a task sequencing system according to an embodiment of the present invention;

[0024] Figure 4 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this document.

[0026] Method Example

[0027] According to an embodiment of the present invention, a task sorting method is provided. Figure 1 is a flowchart of a task sequencing method according to an embodiment of the present invention. Figure 1 As shown, the task sequencing method according to an embodiment of the present invention specifically includes:

[0028] In step S110, historical data of tasks to be sorted and corresponding task types are obtained, specifically including:

[0029] In this embodiment of the present invention, the task types include accounting tasks and operations tasks. The tasks to be sorted include: accounting tasks, accounting supervision tasks, accounting basic tasks 1, accounting basic tasks 2, content operation tasks, and activity operation tasks, a total of six. There are two accounting basic tasks among the tasks to be sorted, so they are distinguished by different labels.

[0030] Obtain task data for tasks to be sorted, count all task types included in the task data as query tags, resample equal amounts of data based on the query tags to obtain historical data; wherein both task data and historical data indicators include task type, duration, dependency, number of dependencies, and number of resources; preferably, additional indicator types can be added based on actual conditions;

[0031] A dependency relationship means that the start or completion of one task depends on the start of another task. In daily work, the rule is to carry out accounting tasks first, and then accounting supervision tasks. Therefore, the dependency relationship of accounting supervision tasks is accounting tasks, and the dependency number is 1. For the sake of description, it is assumed that there is no dependency relationship between other tasks to be sorted.

[0032] Traditionally, the number of resources refers to CPU or memory, etc. In the embodiment of the present invention, the number of resources refers to the number of manpower used. Each person can only process one task at each moment, so it can be similarly represented as a processing block in an operating system.

[0033] In step S120, data fusion is performed based on historical data to generate a corresponding fused data for each task type. The fused data is input into a pre-established regression model to obtain a first order table of task type importance, specifically including:

[0034] Extract data corresponding to the task type from historical data as statistical samples. If the indicator is numerical, calculate the average value based on the statistical samples. If the indicator is non-numerical, randomly extract a data based on the statistical samples to obtain fused data.

[0035] Since this step only requires ranking the task types by importance, the six specific tasks are not considered at this time. Instead, we only need to extract statistical samples of equal amounts of accounting work and operations work. We do not care whether each task type includes tasks that are not among the tasks to be ranked.

[0036] The regression model specifically includes:

[0037] The quantitative value of each indicator is used as the independent variable. Some indicators, such as qualitative descriptions such as dependency, cannot be expressed as numerical values. Therefore, they can be quantified as time lag or other methods. Time lag indicates the time interval between two dependent tasks. A linear regression model is used with the importance score as the dependent variable.

[0038] The regression model is a pre-established model, which describes the relationship between each indicator and the importance score. In the embodiment of the present invention, the accounting work score S1 is higher than the operation work score S2, so the first order table is: accounting work, operation work.

[0039] In step S130, the tasks to be sorted are added to the corresponding task types in the first sequence table, and the tasks to be sorted belonging to the same task type are sorted using the regression model to obtain a second sequence table corresponding to each task type, specifically including:

[0040] The regression model used in this step is the same as that used in step S120. The six tasks to be sorted need to be input into the regression model to obtain the corresponding importance scores.

[0041] In the embodiment of the present invention, the second order list under the accounting work is: accounting basic task 1, accounting basic task 2, accounting calculation task, accounting supervision task; that is, the importance scores of these four specific tasks are ranked from high to low, and the importance scores are a, b, c, d in descending order;

[0042] Similarly, the second priority list under operational work is: content operation tasks, event operation tasks, and the importance scores are e and f respectively.

[0043] Figure 2 Schematic diagram of the sorting architecture of an embodiment of the present invention, such as Figure 2 As shown, the relationship between the first sequence table and the second sequence table is demonstrated, that is, a hierarchical structure is formed. The importance score order of the task types is maintained under the first sequence table. The importance score of task type 1 is higher than that of task type 2. n represents the number of task types. Under each task type, the second sequence table corresponding to the specific task is maintained.

[0044] In step S140, the first order table is used as the first priority, and the second order table is used as the second priority to obtain the initial order of the tasks to be sorted. The reinforcement learning method is used to adjust the initial order to obtain the final order, which specifically includes:

[0045] Traverse each task type in the first sequence table. When a task type is traversed, obtain the second sequence table under the task type and add the order in the second sequence table to the end of the existing initial order until the traversal is completed and the initial order including the order of all tasks to be sorted is obtained;

[0046] Before the traversal begins, the initial sequence is empty. When the first traversal reaches accounting tasks, the second sequence list under accounting tasks: accounting basic tasks 1, accounting basic tasks 2, accounting calculation tasks, accounting supervision tasks, is added to the empty initial sequence. When the second traversal reaches operations tasks, the second sequence list under operations tasks: content operations tasks, event operations tasks, is added to the end of the existing initial sequence.

[0047] The initial order of all tasks to be sorted is: basic accounting task 1, basic accounting task 2, accounting calculation task, accounting supervision task, content operation task, and activity operation task.

[0048] Q-Learning is a reinforcement learning method that uses trial and error to gradually learn which sequences of tasks can be completed faster or more efficiently. Specifically, it attempts to determine the next task to perform and assigns rewards or penalties based on the results. Over time, Q-Learning selects the optimal action to achieve optimal scheduling. The Q-Learning process requires defining three key elements: state, action, and reward. The specific implementation is as follows:

[0049] Use Q learning method to iteratively adjust the initial order to get the final order:

[0050] In traditional Q-learning methods, the initial value is often 0 or a random value. In the embodiment of the present invention, the initial value is set based on the importance score. On the one hand, it is hoped that the learning will be based on the results of the initial order, and on the other hand, the convergence speed of reinforcement learning will be effectively accelerated.

[0051] Calculate the sum of the importance scores of the specific task and the task type to which it belongs as the corresponding initial Q value. For the initial order, the initial Q values ​​are: S1+a, S1+b, S1+c, S1+d, S2+e, S2+f. Preferably, normalization can be selected. The purpose of calculating the sum of the importance scores is to make the initial Q value consistent with the order of the initial order as much as possible, so that the basis of Q learning also takes into account the task type itself.

[0052] After initializing the Q value, a table is maintained to record the updated Q value;

[0053] In the embodiment of the present invention, the status is similar to: which tasks have been completed at the current moment, which tasks have not yet started, how much time has passed, the current available resources, etc.; assuming that the current status is that the accounting task has been completed;

[0054] During the iteration process, the next task is selected based on the Q value. The task with the largest Q value among the remaining tasks is selected as the next task with probability 1-ε. That is, the accounting supervision task with the largest remaining Q value is selected with probability 1-ε. Other tasks are randomly selected as the next task with probability ε. That is, one task is randomly selected from the content operation task and the activity operation task as the next task, which is the ε-greedy exploration strategy in Q learning.

[0055] If the selected task is the accounting supervision task, simulate the execution of the next task and update the Q-value score of the next task based on the new status and immediate reward updated by the simulation. The criteria for immediate reward include whether there is delay, whether there is dependency violation, and whether there is resource margin. At this time, the next task is updated to the completed task and the Q-value is updated using Formula 1, that is, the Q-value of the accounting supervision task is updated:

[0056]

[0057] Among them, Q(s,a) represents the Q-value score of task a in state s; s represents the current state, a represents the task selected for execution; r represents the immediate reward obtained after simulating task a in state s; s' represents the new state, a' represents the unexecuted task; α represents the learning rate, which controls the weight of new information; γ represents the discount factor, which controls the degree of discount of future rewards; max represents the maximum Q-value among all possible tasks in the new state, that is, the maximum Q-value among content operation tasks and event operation tasks, which is used to estimate the maximum future return;

[0058] In the embodiments of the present invention, the Q-value update logic in Q-Learning reinforcement learning is not changed.

[0059] The Q value updated in each round of iteration is used as the initial value of the next round of iteration. The Q value corresponding to each task is updated in turn in each round of iteration. After the iteration is completed, the final order is obtained by sorting by Q value.

[0060] To sum up, in response to the problems existing in the current situation, the task sorting method of this invention sorts the specific tasks to be sorted on the basis of sorting the importance of task types, so as to avoid the huge consumption brought by the overall calculation by using a hierarchical optimization method, and obtain a reasonable initial order in a short period of time, and takes into account the differences between different task types, that is, the importance of tasks of the same type has commonalities; Q-Learning reinforcement learning is used to dynamically adjust the initial order of specific tasks through interaction with the environment, which can be flexibly adjusted according to the current state and resource conditions, thereby improving the accuracy of task sorting; in the process of using reinforcement learning to optimize the initial order, the sum of the importance scores of the specific task and the task type to which it belongs is calculated as the corresponding initialization Q value, which effectively accelerates the convergence speed of reinforcement learning.

[0061] System Example

[0062] According to an embodiment of the present invention, a task sorting system is provided. Figure 3 Schematic diagram of a task sequencing system according to an embodiment of the present invention. Figure 3 As shown, the task sequencing system according to an embodiment of the present invention specifically includes:

[0063] The data acquisition module 30 is used to obtain historical data of tasks to be sorted and corresponding task types, specifically for:

[0064] Obtain the task data of the tasks to be sorted, count all task types included in the task data as query tags, resample the same amount of data based on the query tags to obtain historical data; among them, both task data and historical data indicators include task type, duration, dependency, number of dependencies, and number of resources.

[0065] The type sorting module 32 is used to perform data fusion based on historical data, generate a corresponding fused data for each task type, input the fused data into a pre-established regression model, and obtain a first order table of task type importance, specifically for:

[0066] The data corresponding to the task type is extracted from the historical data as a statistical sample. If the indicator is numerical, the average value is calculated based on the statistical sample. If the indicator is non-numerical, a data is randomly extracted based on the statistical sample to obtain the fused data.

[0067] The regression model specifically includes:

[0068] A linear regression model was established with the quantitative value of each indicator as the independent variable and the importance score as the dependent variable.

[0069] The task sorting module 34 is used to fill the tasks to be sorted into the corresponding task types in the first sequence table, and use the regression model to sort the tasks to be sorted belonging to the same task type to obtain a second sequence table corresponding to each task type;

[0070] The adjustment module 36 is used to obtain an initial order of the tasks to be sorted using the first order table as the first priority and the second order table as the second priority, and to adjust the initial order using the reinforcement learning method to obtain a final order, specifically for:

[0071] Traverse each task type in the first sequence table. When traversing to a certain task type, obtain the second sequence table under the task type, and add the order in the second sequence table to the end of the existing initial order until the traversal is completed to obtain the initial order including the order of all tasks to be sorted.

[0072] Use Q learning method to iteratively adjust the initial order to get the final order:

[0073] Calculate the sum of the importance scores of the specific task and the task type to which it belongs as the corresponding initial Q value;

[0074] During the iteration process, the next task is selected based on the Q value. The task with the largest Q value among the remaining tasks is selected as the next task with probability 1-S, and other tasks are randomly selected as the next task with probability ε.

[0075] Simulate the execution of the next task and update the Q-value score of the next task based on the new state updated by the simulation and the immediate reward. The criteria for the immediate reward include whether there is a delay, whether there is a dependency violation, and whether there is a resource margin.

[0076] The Q value updated in each round of iteration is used as the initial value of the next round of iteration, and the final order is obtained after the iteration is completed.

[0077] To sum up, in response to the problems existing in the status quo, the task sorting system of this invention sorts the specific tasks to be sorted on the basis of sorting the importance of task types, and uses a hierarchical optimization method to avoid the huge consumption brought by the overall calculation, and obtains a reasonable initial order in a short period of time, and takes into account the differences between different task types, that is, the importance of tasks of the same type has commonalities; Q-Learning reinforcement learning is used to dynamically adjust the initial order of specific tasks through interaction with the environment, which can be flexibly adjusted according to the current state and resource conditions, thereby improving the accuracy of task sorting; in the process of using reinforcement learning to optimize the initial order, the sum of the importance scores of the specific task and the task type to which it belongs is calculated as the corresponding initialization Q value, which effectively accelerates the convergence speed of reinforcement learning.

[0078] Electronic device embodiment

[0079] Figure 4 is a schematic diagram of an electronic device according to an embodiment of the present invention. Electronic device 400 may include at least one processor 410 and memory 420. Processor 410 can execute instructions stored in memory 420. Processor 410 is communicatively coupled to memory 420 via a data bus. In addition to memory 420, processor 410 may also be communicatively coupled to input device 430, output device 440, and communication device 450 via the data bus.

[0080] The processor 410 may be any conventional processor, such as a commercially available CPU. The processor may also include a graphics processing unit (GPU), a field programmable gate array (FPGA), a system on chip (SOC), an application specific integrated circuit (ASIC), or a combination thereof.

[0081] The memory 420 may be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0082] In an embodiment of the present disclosure, executable instructions are stored in the memory 420 , and the processor 410 can read the executable instructions from the memory 420 and execute the instructions to implement all or part of the steps of any task sequencing method in the above exemplary embodiments.

[0083] Computer readable storage medium embodiments

[0084] In addition to the above-mentioned methods and devices, the exemplary embodiments of the present disclosure may also be a computer program product or a computer-readable storage medium storing the computer program product, wherein the computer product includes computer program instructions that can be executed by a processor to implement all or part of the steps described in any of the task sorting methods in the above-mentioned exemplary embodiments.

[0085] The computer program product may be written in any combination of one or more programming languages ​​to write program code for performing the operations of the embodiments of the present application, including object-oriented programming languages ​​such as Java, C++, etc., as well as conventional procedural programming languages ​​such as "C" or similar programming languages ​​and scripting languages ​​(e.g., Python). The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0086] Computer-readable storage media can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or components, or any combination thereof. More specific examples of readable storage media include: static random access memory (SRAM) electrically connected with one or more wires, electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk, or any suitable combination thereof.

[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A task sequencing method, characterized in that: include: Get the historical data of the tasks to be sorted and the corresponding task types; Perform data fusion based on the historical data, generate a corresponding fused data for each task type, input the fused data into a pre-established regression model, and obtain a first order table of task type importance ranking; Fill the tasks to be sorted into the corresponding task types in the first sequence table, and use the regression model to sort the tasks to be sorted belonging to the same task type to obtain a second sequence table corresponding to each task type; The first sequence table is used as the first priority, and the second sequence table is used as the second priority to obtain an initial sequence of tasks to be sorted, and the reinforcement learning method is used to adjust the initial sequence to obtain a final sequence.

2. The method according to claim 1, characterized in that The acquisition of the historical data of the tasks to be sorted and the corresponding task types specifically includes: Obtain task data of the tasks to be sorted, count all task types included in the task data as query tags, resample equal amounts of data based on the query tags, and obtain the historical data; wherein the task data and the historical data indicators both include task type, time consumption, dependency, dependency quantity, and resource quantity.

3. The method according to claim 1, characterized in that The data fusion based on the historical data to generate a corresponding fusion data for each task type specifically includes: Data corresponding to the task type is extracted from the historical data as a statistical sample. If the indicator is numerical, an average value is calculated based on the statistical sample. If the indicator is non-numerical, a data is randomly extracted based on the statistical sample to obtain the fused data.

4. The method according to claim 1, characterized in that: The regression model specifically includes: A linear regression model with the quantitative value of each indicator as the independent variable and the importance score as the dependent variable.

5. The method according to claim 1, characterized in that The step of taking the first sequence table as the first priority and the second sequence table as the second priority to obtain the initial sequence of the tasks to be sorted specifically includes: Traverse each task type in the first sequence table. When traversing to a certain task type, obtain the second sequence table under the task type, and add the sequence in the second sequence table to the end of the existing initial sequence until the traversal is completed to obtain the initial sequence including the sequences of all tasks to be sorted.

6. The method according to claim 1, characterized in that The using of reinforcement learning method to adjust the initial sequence to obtain the final sequence specifically includes: using Q learning method to iteratively adjust the initial sequence to obtain the final sequence.

7. The method according to claim 6, characterized in that The iterative adjustment of the initial sequence using the Q learning method to obtain the final sequence specifically includes: Calculate the sum of the importance scores of the specific task and the task type to which it belongs as the corresponding initialization Q value; In the iteration process, the next task is selected according to the Q value, the task with the largest Q value among the remaining tasks is selected as the next task with probability 1-ε, and other tasks are randomly selected as the next task with probability ε; Simulate the execution of the next task, and update the Q-value score of the next task according to the new state updated by the simulation and the immediate reward, wherein the criteria of the immediate reward include whether there is a delay, whether there is a violation of dependency, and whether there is a resource margin; The Q value updated in each round of iteration is used as the initial value of the next round of iteration, and the final order is obtained after the iteration is completed.

8. A task sequencing system, characterized in that: include: The data acquisition module is used to obtain the historical data of the tasks to be sorted and the corresponding task types; A type sorting module is used to perform data fusion based on the historical data, generate a corresponding fused data for each task type, input the fused data into a pre-established regression model, and obtain a first order table of task type importance; A task sorting module, used to fill the tasks to be sorted into the corresponding task types in the first sequence table, and use the regression model to sort the tasks to be sorted belonging to the same task type to obtain a second sequence table corresponding to each task type; The adjustment module is used to obtain the initial sequence of the tasks to be sorted by taking the first sequence table as the first priority and the second sequence table as the second priority, and use the reinforcement learning method to adjust the initial sequence to obtain the final sequence.

9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the task sequencing method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores an implementation program for information transmission, and when the program is executed by a processor, the steps of the task sorting method according to any one of claims 1 to 7 are implemented.