Project scheduling method and device based on reinforcement learning

Through the project scheduling method based on reinforcement learning, the project plan and task arrangement are dynamically adjusted, and the real-time changes in the project management system are solved, improving project success rate and resource utilization efficiency.

CN120494373APending Publication Date: 2025-08-15HONGTA TOBACCO (GROUP) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510573881.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing project management system is difficult to adapt to real-time changes in the project implementation process, resulting in project progress delays, cost overruns and even failures, and lacks dynamic adaptability.

Method used

The project scheduling method based on reinforcement learning is adopted, and the current status vector of the target project is obtained and the preset project scheduling model is used for dynamic adjustment, including the time, resources, task priority and member allocation of engineering and service projects, and the model parameters are optimized based on the project reward function.

Benefits of technology

It has achieved dynamic adjustment of project plans based on real-time progress of the project and changes in the external environment, reducing the risks of delays and failures, improving project success rate, improving resource utilization efficiency and team member satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494373A_ABST
    Figure CN120494373A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a project scheduling method and device based on reinforcement learning. The method comprises the following steps: acquiring a current state vector of a target project; wherein the target project comprises engineering and service projects of a tobacco manufacturing enterprise; inputting the current state vector into a preset project scheduling model to output an execution action corresponding to the current state vector; wherein the preset project scheduling model is a convergent preset project scheduling model obtained by training by adopting a reinforcement learning algorithm and adjusting model parameters by adopting a pre-constructed project reward function; and performing dynamic project scheduling on the target project according to the execution action. By adopting the technical scheme of the embodiment of the invention, the plan and task arrangement of the target project can be dynamically adjusted according to the real-time progress of the target project, the risk factor and the external environment change, and various changes can be timely coped with, so that the project delay and failure risks are reduced, and the success rate of the project is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of project management, and in particular to a project scheduling method and device based on reinforcement learning. Background Art

[0002] The project management system is a popular management discipline. It is based on modern management science and efficiently integrates financial management, personnel management, risk assessment, quality management, etc. in corporate management, thereby achieving the effect of completing corporate project work with high efficiency, high quality and low cost.

[0003] In traditional project management, project scheduling is typically based on pre-defined plans, making it difficult to adapt to real-time changes during project implementation, such as the emergence of risk factors and changes in the external environment. These changes can lead to project delays, cost overruns, and even project failure. Most existing project management systems lack the ability to adapt to dynamic changes and are unable to adjust project plans and task schedules in a timely manner based on real-time circumstances.

[0004] Therefore, how to construct a project scheduling solution that can respond to changes in real time is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0005] The embodiments of the present invention provide a project scheduling method and device based on reinforcement learning to dynamically adjust the plan and task arrangement of the target project, respond to various changes in a timely manner, reduce the risk of project delays and failures, and improve the success rate of the project.

[0006] In a first aspect, an embodiment of the present invention provides a project scheduling method based on reinforcement learning, comprising:

[0007] Obtaining a current state vector of a target project; wherein the target project includes engineering and service projects of a tobacco manufacturing enterprise;

[0008] Inputting the current state vector into a preset project scheduling model to output an execution action corresponding to the current state vector; wherein the preset project scheduling model is trained using a reinforcement learning algorithm and the model parameters are adjusted using a pre-built project reward function to obtain a converged preset project scheduling model;

[0009] According to the execution action, dynamic project scheduling is performed on the target project.

[0010] In a second aspect, an embodiment of the present invention further provides a project scheduling device based on reinforcement learning, comprising:

[0011] A current state acquisition module is used to obtain the current state vector of a target project; wherein the target project includes engineering and service projects of a tobacco manufacturing enterprise;

[0012] an execution action determination module, configured to input the current state vector into a preset project scheduling model to output an execution action corresponding to the current state vector; wherein the preset project scheduling model is trained using a reinforcement learning algorithm and adjusts model parameters using a pre-built project reward function to obtain a converged preset project scheduling model;

[0013] The project dynamic scheduling module is used to dynamically schedule the target project according to the execution action.

[0014] In a third aspect, an embodiment of the present invention further provides an electronic device, the electronic device comprising:

[0015] one or more processors;

[0016] a storage device for storing one or more programs;

[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the reinforcement learning-based project scheduling method described in any embodiment of the present invention.

[0018] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the reinforcement learning-based project scheduling method described in any embodiment of the present invention.

[0019] In a fifth aspect, an embodiment of the present invention further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the reinforcement learning-based project scheduling method as described in any embodiment of the present invention.

[0020] Embodiments of the present invention provide a reinforcement learning-based project scheduling method, device, electronic device, and storage medium. These methods obtain the current state vectors of engineering and service projects at a tobacco manufacturing enterprise; input these current state vectors into a pre-set project scheduling model to output the corresponding execution actions; and dynamically schedule target projects based on these execution actions. The technical solutions of these embodiments enable dynamic adjustment of project plans and task schedules based on real-time project progress, risk factors, and changes in the external environment, enabling timely response to various changes, reducing the risk of project delays and failures, and improving project success rates. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Other features, objects, and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings. The drawings are for the purpose of illustrating preferred embodiments only and are not to be considered as limiting the present invention. Like reference characters are used throughout the drawings to denote like parts. In the drawings:

[0022] Figure 1 is a flowchart of a project scheduling method based on reinforcement learning provided in an embodiment of the present invention;

[0023] Figure 2 is a flowchart of another project scheduling method based on reinforcement learning provided in an embodiment of the present invention;

[0024] Figure 3 This is a flowchart of a project scheduling model training method based on reinforcement learning provided in an embodiment of the present invention;

[0025] Figure 4 is a flowchart of another method for training a project scheduling model based on reinforcement learning provided in an embodiment of the present invention;

[0026] Figure 5 is a flow chart of a value table updating method provided in an embodiment of the present invention;

[0027] Figure 6 is a structural diagram of a project scheduling device based on reinforcement learning provided in an embodiment of the present invention;

[0028] Figure 7 It is a structural diagram of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0029] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.

[0030] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the various operations (or steps) as sequential processes, many of the operations (or steps) therein can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the various operations can be rearranged. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0031] The acquisition, storage, use, and processing of data in the technical solution of this application are in compliance with the relevant provisions of national laws and regulations. It should be noted that in the embodiments of this application, certain software, components, or models, etc., which are already available in the industry, may be mentioned. These should be considered as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. It does not mean that the applicant has already or necessarily used such a solution.

[0032] It should be noted that the terms "first," "second," and the like in the description and claims of the present invention and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, such that the embodiments of the present invention described herein can be practiced in an order other than that illustrated or described herein.

[0033] Example 1

[0034] Figure 1 This is a flow chart of a reinforcement learning-based project scheduling method provided in an embodiment of the present invention. This embodiment is applicable to situations where adaptive project scheduling based on reinforcement learning is performed in a project management system. The method of this embodiment can be executed by a reinforcement learning-based project scheduling device, which can be implemented in hardware and / or software. The device can be configured in a server of the reinforcement learning-based project scheduling method. The method specifically includes the following steps:

[0035] S110. Obtain the current state vector of the target project.

[0036] Among them, the embodiment of the present invention is applied to the tobacco manufacturing industry, and the target project may refer to engineering and service projects included in the tobacco manufacturing enterprise, and the engineering and service projects include but are not limited to tobacco manufacturing investment, tobacco production and manufacturing, purchase of tobacco manufacturing equipment and maintenance of tobacco manufacturing equipment.

[0037] The current state vector may refer to the state vector of the target project at the current moment, and is obtained by quantifying and characterizing at least two pieces of current state information. Optionally, the at least two pieces of state information for the target project include, but are not limited to, real-time project progress information, basic information such as project category and procurement method, external environmental factors, and project team member skill level information.

[0038] S120: Input the current state vector into a preset project scheduling model to output an execution action corresponding to the current state vector.

[0039] The preset project scheduling model is trained using a reinforcement learning algorithm, and the model parameters are adjusted using a pre-built project reward function to obtain a converged preset project scheduling model. The reinforcement learning algorithm may refer to an intelligent agent that continuously tries different actions through interaction with the environment and adjusts its own behavior strategy based on the rewards fed back by the environment, in order to achieve the goal of maximizing the cumulative reward. In the embodiment of the present invention, the reinforcement learning algorithm is used to continuously execute different actions corresponding to the current state vector of the target project, and continuously adjust the corresponding execution actions through the project reward function to achieve the maximum project reward value.

[0040] A project reward function quantifies the immediate reward for an agent taking a specific action in a specific state. By accumulating these immediate rewards, the agent can optimize its strategy and maximize long-term benefits. In this embodiment of the present invention, rewards are assigned based on the impact of the current action on the target project in its current state. These rewards are accumulated to obtain the reward value for the current action taken in the current state.

[0041] The execution action corresponding to the current state vector may be the action that, under the current state vector, yields the maximum cumulative reward value output by the preset project scheduling model. In this embodiment of the present invention, the current state vector is input into the preset project scheduling model, and the corresponding cumulative reward value for the project is obtained by continuously trying new actions. The execution action is updated based on the cumulative reward value for the project. When the maximum cumulative reward value for the project is determined, the execution action corresponding to the current state vector is determined.

[0042] S130: Dynamically schedule the target project according to the execution action.

[0043] The execution actions include, but are not limited to, project time adjustment, resource reallocation, project task priority change, and project member reallocation. In an embodiment of the present invention, the target project is dynamically scheduled based on the execution actions. For example, if the execution action is project time adjustment, the target project's current time and the time of subsequent stage projects are dynamically adjusted. If the execution action is project member composition allocation, project members are reallocated to the target project based on project requirements, project member skill levels, and workload.

[0044] An embodiment of the present invention provides a project scheduling method based on reinforcement learning, which obtains the current state vector of a target project; wherein the target project includes engineering and service projects of a tobacco manufacturing enterprise; inputs the current state vector into a preset project scheduling model to output an execution action corresponding to the current state vector; wherein the preset project scheduling model is trained using a reinforcement learning algorithm and uses a pre-built project reward function to adjust the model parameters to obtain a converged preset project scheduling model; based on the execution action, the target project is dynamically scheduled. Using the technical solution of the embodiment of the present invention, it is possible to dynamically adjust project plans and task arrangements based on the real-time progress of the project, risk factors, and changes in the external environment, and respond to various changes in a timely manner to reduce the risk of project delays and failures and improve the success rate of the project.

[0045] Example 2

[0046] Figure 2 This is a flowchart of a project scheduling method based on reinforcement learning provided in an embodiment of the present invention. The embodiment of the present invention further optimizes the above embodiment on the basis of the above embodiment, and the embodiment of the present invention can be combined with various optional solutions in one or more of the above embodiments. Figure 2 As shown, the project scheduling method based on reinforcement learning provided in an embodiment of the present invention may include the following steps:

[0047] S210: Obtain at least two pieces of current state information of a target project, quantify and represent the at least two pieces of current state information, and combine them into a current state vector.

[0048] The current state information includes real-time project progress information, external environment change information, basic project attribute information, and project member skills and workload information. In this embodiment of the present invention, the at least two pieces of current state information are quantitatively represented and combined into a current state vector. The quantitative representation can be a numerical description of the at least two pieces of current state information for the target project.

[0049] The real-time project progress information may refer to the completion progress of each stage of the target project and the remaining time of the project. The completion progress of the stage project can be obtained by calculating the completion status of the weekly progress; for example, if the first stage of the target project lasts for 10 weeks, 3 weeks have been completed, and the weekly progress is 10%, 12%, and 15% respectively, then the completion progress of this stage is (10% + 12% + 15%) = 37%. The remaining time of the project can be calculated based on the start time and end time of the project; let the start time of the project be t start , the end time is t end , the current time is t now , then the remaining time tremain =t end -t now .

[0050] Information on external environmental changes includes, but is not limited to, industry technological advancements and adjustments to policies and regulations. For industry technological advancements, data such as the frequency of industry technology updates and the application of new technologies can be collected and quantified into a numerical indicator. For adjustments to policies and regulations, a score can be assigned based on the degree of impact of the policy on the project and converted into a numerical value.

[0051] The basic attribute information of the project may refer to basic attributes such as the category of the target project and the procurement method. The project category and procurement method may be one-hot encoded and converted into a numerical vector for inclusion in the state vector.

[0052] Project member skill and workload information may refer to information such as project member skill levels and workload, and can be obtained by querying personnel information and task assignments in the project management system. Quantifying project member skill and workload information may include scoring project members based on their past project experience and professional skill certifications. Workload can be calculated based on the number of currently assigned tasks and the estimated duration of each task.

[0053] The current state vector is obtained by combining the target project's real-time progress information, external environment change information, basic project attributes, and quantitative representations of project member skills and workload information. The quantified current state vector obtained in this embodiment facilitates the subsequent calculation of the target project's project reward in its current state to determine the corresponding execution action.

[0054] In an alternative embodiment of the present invention, the reinforcement learning-based project scheduling method constructed in the embodiment of the present invention is applied to a project management system; the project management system includes, but is not limited to, the Django project management system, an open-source web application framework written in Python. Using the Django project management system as an example, the embodiment of the present invention obtains at least two pieces of current status information for a target project and stores them in a system database, facilitating modification of the status information during subsequent project scheduling.

[0055] S220: Input the current state vector into a preset project scheduling model to output an execution action corresponding to the current state vector.

[0056] S230: Dynamically schedule the target project according to the execution action.

[0057] The execution actions include, but are not limited to, adjusting the start time of a phased project, reallocating resources, changing project task priorities, and reallocating project members. After determining the execution action corresponding to the current state vector of the target project, dynamic project scheduling is performed on the target project based on the execution action.

[0058] As an optional but non-limiting implementation, the dynamic project scheduling of the target project according to the execution action includes but is not limited to steps A1-A4:

[0059] Step A1: If the execution action is to adjust the start time of a stage project, the start time of the stage project and the start time of subsequent stage projects of the stage project are dynamically adjusted.

[0060] Step A2: If the execution action is resource reallocation, then the resource allocation amounts of the projects at different stages are readjusted.

[0061] Step A3: If the execution action is to change the priority of the project task, dynamic adjustment is performed according to the priority of the phase project.

[0062] Step A4: If the execution action is project member reallocation, dynamic adjustment is performed based on the project members' skill levels, workloads, and work requirements.

[0063] If the execution action involves adjusting the start time of a phase project, the start time of the corresponding phase project and the schedule of subsequent phase projects are modified. Specifically, through object-relational mapping operations in the project management system, the start time field of the phase project in the database is updated, and the start and end times of subsequent phase projects are automatically adjusted based on the dependencies between the phase projects.

[0064] If the execution action is resource reallocation, the resource allocation ratio among different projects or phase projects is adjusted, and the resource allocation amount of each project or phase project is updated by modifying the relevant records of the resource allocation table in the database.

[0065] If the execution action is a change in the priority of a project task, the priority order of the project task is updated. In the database, the priority field of the task table is updated to ensure that the project tasks are sorted and executed according to the new priority order.

[0066] If the execution action is to reassign project members, then reassign project members to different tasks or projects, update the association table between project members and tasks or projects in the database, and assign members to new tasks or projects. For example, if a project member has a high level of skills in a certain field, they should be assigned to tasks that require this skill in action decision-making. For example, in a software development project, a project member who is good at front-end development should be assigned to related tasks such as front-end page design and interactive function implementation. This can be achieved by establishing a skill matching constraint. Let the skills required for task j be S j , member i's skill level is Skill i , you can define the skill matching When selecting the "Project Member Reassignment" action, select Match first. ij Members with higher values are assigned to task j.

[0067] Workload affects resource allocation for action decisions. When a project member's workload is too high, in order to ensure project quality and progress, we should avoid assigning more tasks, or even reallocate some tasks to other project members with lower workloads. Assume that the total estimated time required for the currently assigned tasks of project member i is Workload i , the team's overall workload is AvgWorkload, and the workload threshold can be set. i >Threshold×AvgWorkload, consider the "project member reallocation" action and allocate part of the tasks of project member i to Workload i <(1-Threshold)×AvgWorkload of project members k, ensuring workload balance and improving project execution efficiency.

[0068] An embodiment of the present invention provides a project scheduling method based on reinforcement learning. The method obtains at least two pieces of current state information of a target project, quantifies and represents the at least two pieces of current state information, and combines them into a current state vector. The current state vector is input into a preset project scheduling model to output an execution action corresponding to the current state vector. Based on the execution action, the target project is dynamically scheduled. The technical solution of the embodiment of the present invention can reduce the risk of project delays and failures and improve the success rate of projects by adjusting and optimizing project scheduling in real time. By rationally allocating resources to different projects and stages, resource utilization efficiency is improved and project costs are reduced. By optimizing the allocation of team members, the work satisfaction and work efficiency of team members are improved.

[0069] Example 3

[0070] Figure 3This is a flowchart of a project scheduling model training method based on reinforcement learning provided in an embodiment of the present invention. The embodiment of the present invention further optimizes the above embodiment on the basis of the above embodiment, and the embodiment of the present invention can be combined with various optional solutions in one or more of the above embodiments. Figure 3 As shown, the project scheduling model training method based on reinforcement learning provided in an embodiment of the present invention may include the following steps:

[0071] S310 , executing the project status definition, project action definition, and project reward function definition tasks for the project scheduling model to be trained.

[0072] When training the project scheduling model, it is first necessary to obtain a training data set; the training data set includes but is not limited to historical data of engineering and service projects of tobacco manufacturing enterprises, and the historical data is preprocessed to construct the training data set.

[0073] As an optional but non-limiting implementation, before training the preset project scheduling model, the method further includes but is not limited to steps B1-B2:

[0074] Step B1: Acquire project data of a target project from a local storage file, and store the project data in a database.

[0075] Step B2: Preprocess the project data and quantify the text data included in the project data.

[0076] The target project data is collected from local storage or other data sources, including but not limited to basic project attribute information (such as project name, category, procurement method, and amount), project phase information (such as phase name, start time, end time, number of weeks, etc.), and real-time project progress information. The project scheduling model is trained in a project management system. For example, using the Django project management system as an example, project data is obtained from a JSON file in the Django project management system and stored in the Django project management system's database. JSON is a data exchange format, and the JSON file is used to store the obtained project data.

[0077] Optional, see Figure 4 After obtaining the project data of the target project, the project data is preprocessed and the text data is quantified. The preprocessing includes outlier processing and missing value processing, the text data includes project progress data, and the quantification includes converting the text description of the project progress into a project progress percentage.

[0078] Optionally, missing value handling for project data can involve using the mean filling method for numerical data, such as project amounts and the duration of phased projects. For example, the average amount or average duration of projects in that category can be calculated and used to fill missing values. For categorical data, such as project categories and procurement methods, the mode filling method can be used, that is, the most frequently occurring category can be used to fill missing values.

[0079] Outlier processing can refer to identifying outliers using a standard deviation-based approach. For a numerical feature, its mean μ and standard deviation δ are calculated. If a data point x satisfies |x-μ|>3δ, the data point is considered an outlier. For identified outliers, if their deviation from a reasonable range is small, they can be corrected to a boundary value within the 3δ range; if the deviation is excessive, they can be deleted. Optionally, the specific value of the standard deviation is not specifically limited in this embodiment of the present invention.

[0080] Quantifying text data can involve converting the weekly progress text descriptions into progress percentages using text classification methods from natural language processing. Specifically, a Naive Bayes classifier is trained using data derived from weekly progress text descriptions of historical projects and their corresponding actual progress percentages. After preprocessing the text data through word segmentation and stop word removal, term frequency-inverse document frequency (TF-IDF) features are extracted and fed into the Naive Bayes classifier for training. Once training is complete, the new weekly progress text descriptions can be classified to obtain the corresponding progress percentages.

[0081] Among them, in the embodiment of the present invention, the project data is preprocessed and the text data is quantified, missing values are supplemented and outliers are processed, thereby improving the integrity and accuracy of the project data; the quantification of the text data facilitates the subsequent use of the text data and improves the speed of data processing.

[0082] After preprocessing and quantifying the project data, a training dataset is constructed and used to train the project scheduling model. During the project scheduling model training phase, the project status definition, project action definition, and project reward function definition tasks must first be performed.

[0083] Optionally, the project status definition task involves combining at least two pieces of project status information into a project status vector. For example, at least two pieces of project status information, such as the target project's real-time progress, risk factors, and external environmental changes, are obtained, and the at least two pieces of project status information are quantified and combined into a project status vector.

[0084] Optionally, the project action definition defines a set of selectable actions within the model. Each action has corresponding execution logic. The action set includes adjusting the start time of a phase project, reallocating resources, changing the priority of project tasks, and reallocating project members. Adjusting the start time of a phase project can refer to adjusting the start time of a phase project to optimize the project schedule. A target project can include at least one phase project. For example, if a phase project is behind schedule due to insufficient preparatory work, the start time of that phase project can be appropriately postponed to allow for more preparation time. Resource reallocation can refer to reallocating resources to different projects or phases to ensure their proper utilization. Resources include human, material, and financial resources. Based on project priorities and current progress, resources can be redeployed from projects or phases with faster progress or excess resources to projects or phases with slower progress or limited resources. Changing the priority of project tasks can refer to changing the priority of project tasks to address emergencies or optimize project processes. For example, when an urgent and important project task arises, the priority of that task can be increased to ensure its completion first. Project member reallocation can refer to adjusting the allocation of project members, optimizing them according to task requirements and project member skill levels, and allocating project members with high skill matching to corresponding project tasks to improve work efficiency.

[0085] As an optional but non-limiting implementation, the project scheduling model to be trained performs the project reward function definition task, including but not limited to steps C1-C5:

[0086] Step C1: Define a first reward function and a second reward function based on the degree of impact of the actions taken by the model on the project progress; wherein the first reward function is a positive reward function obtained when the actions taken shorten the project progress; the second reward function is a negative reward function obtained when the actions taken delay the project progress.

[0087] Step C2: Define a third reward function and a fourth reward function based on the degree of impact of the actions taken by the model on the project risk; wherein the third reward function is a positive reward function obtained by taking actions that reduce project risk; and the fourth reward function is a negative reward function obtained by taking actions that increase project risk.

[0088] Step C3: Define the fifth reward function based on the utilization of project resources by the actions taken by the model.

[0089] Step C4: Define the sixth reward function based on the satisfaction of project members with the actions taken by the model.

[0090] Step C5: superimpose the first reward function, the second reward function, the third reward function, the fourth reward function, the fifth reward function, and the sixth reward function to define an item reward function.

[0091] The embodiment of the present invention defines a project reward function that rewards the model based on the impact of its actions on the project. The definition principle of the project reward function is as follows:

[0092] If the actions taken by the model enable the project to be completed ahead of schedule or reduce the risk, a positive reward will be given. Let the project completion time be T plan , the actual completion time is T actual , if T plan >T actual , then the first reward function is represented as R1=k1×(T plan -T actual ), where k1 is the reward coefficient for completing the task ahead of time. For risk reduction, rewards can be given based on the changes in risk assessment indicators. Suppose the risk assessment indicator changes from R before Reduce to R after , then the third reward function is represented as R2=k2×(R before -R after ), where k2 is the reward coefficient for risk reduction.

[0093] If the model takes an action that causes project delays or increases risk, a negative reward is given. plan <T actual , then a penalty is given, and the second reward function is represented by P1=-k3×(T actual -T plan ), where k3 is the penalty coefficient for the delay time. If the risk assessment index increases, a penalty is imposed. The fourth reward function is represented by P2 = -k4 × (R after -R before ), where k4 is the penalty coefficient for increased risk.

[0094] Considering factors such as resource utilization and project members’ job satisfaction, the project reward function is refined. Resource utilization can be measured by calculating the difference between the actual resource utilization and the ideal utilization. Let the actual resource utilization be U actual , the ideal utilization rate is U ideal , then the reward function is represented as R3=k5×(U actual -U ideal ), where k5 is the bonus coefficient for resource utilization efficiency. Let the actual input of resource r in project stage s be Actual rs , the actual usage is The ideal input amount is Ideal rs, then the utilization rate of resource r in project stage s is The overall project resource utilization index can be expressed as OverallUtilization = ∑ r ∑ s w rs ×Utilization rs , where w rs is the weight of resource r in project stage s; wherein, the weight is determined according to the importance of the resource and the requirements of the project stage. When the action increases OverallUtilization, the new resource utilization fifth reward function can be expressed as R 31 =k 51 ×(OverallUtilization new -OverallUtilization old ), where k 51 is the reward coefficient.

[0095] Define the sixth reward function of project member satisfaction, and let the difficulty of task j be Difficulty j , the estimated time required for member i to complete task j is EstimatedTime ij The average difficulty of the team's tasks is AvgDifficulty, the average estimated time is AvgEstimatedTime, and the member's career development opportunity assessment is Development i The satisfaction index of member i can be expressed as Among them, β1, β2, and β3 are weight coefficients. The average satisfaction of project members can be expressed as n is the number of project members. When improved, the sixth reward function can be expressed as Among them, k 52 is the reward coefficient.

[0096] By superimposing each reward function, the project reward function can be defined, which can be expressed as R = R1 + R2 + R 31 +R 32 +P1+P2.

[0097] The embodiment of the present invention constructs a project reward function from multiple dimensions based on the actual situation of the project and the execution results of the actions; based on the project reward function, the execution results of the actions are evaluated to provide a basis for the state update at the next moment.

[0098] S320: Based on the project status definition, the project action definition, and the project reward function definition, a reinforcement learning algorithm is used to perform a model training task on the project scheduling model to be trained to obtain a project scheduling model.

[0099] Among them, the embodiment of the present invention takes the reinforcement learning algorithm as the Q-learning algorithm as an example, and performs a model training task on the project scheduling model to be trained. The Q-learning algorithm refers to updating the value corresponding to each state-action by continuously interacting with the environment, so as to obtain the optimal strategy. The Q-learning algorithm includes a value table, namely the Q table. The Q table is a two-bit table. The rows represent states, the columns represent actions, and each cell Q(s, a) represents the value estimate of taking action a under state s. The Q-learning algorithm executes the selected action based on the current state, observes the changes in the project state, obtains the state at the next moment and the reward obtained by the current action; updates the Q table based on the state at the next moment and the reward obtained by the current action, so as to perform the model training task on the project scheduling model to be trained and obtain a project scheduling model.

[0100] As an optional but non-limiting implementation, based on the project status definition, project action definition, and project reward function definition, a reinforcement learning algorithm is used to perform a model training task on the project scheduling model to be trained to obtain a project scheduling model, including but not limited to steps D1-D4:

[0101] Step D1: Get the current project state at the current time step and select the first project action based on the greedy strategy.

[0102] Step D2: Execute the first project action, and obtain the second project action and the project reward obtained by executing the first project action based on the change in the project status after the first project action.

[0103] Step D3: Update the value table in the reinforcement learning algorithm based on the second project action and the project reward obtained by executing the first project action; the value table is composed of project states and corresponding project actions, and the value table is used to represent the reward that can be obtained by executing the project action corresponding to the project state.

[0104] Step D4: Repeat the task of updating the value table to perform the model training task on the project scheduling model to be trained to obtain the project scheduling model.

[0105] Among them, see Figure 5 First, initialize the parameters of the project scheduling model to be trained. Taking the Q-learning algorithm as an example, initialize the Q table and set all values in the Q table to 0. At each time step, obtain the current project state st , select an action a according to the ε-greedy strategy t In the ε-greedy strategy, ε represents the exploration probability, and its initial value is set to 0.9. As the training progresses, ε gradually decreases and finally stabilizes at 0.1. Specifically, with probability ε, an action is randomly selected for exploration, and with probability 1, an action is randomly selected from the Q table (s t , a) The action with the largest value is used. In an optional solution of the embodiment of the present invention, the method of selecting the first action includes but is not limited to the greedy strategy, which is not specifically limited in the embodiment of the present invention. Execute the selected first action a t , observe the changes in the current project status and get the next state s t+1 , and calculate the project reward R obtained by executing the first project action according to the project reward function t . According to the project reward R t and the next state s t+1 Update the Q table. Repeat the task of updating the value table to perform the model training task on the project scheduling model to be trained to obtain the project scheduling model.

[0106] Optionally, the update formula of the Q table is expressed as:

[0107]

[0108] Among them, α is the learning rate, the initial value is set to 0.1, which is used to control the step size of each update; γ is the discount factor, the initial value is set to 0.9, which is used to measure the importance of future rewards.

[0109] The embodiment of the present invention executes a model training task of a project scheduling model to be trained through a reinforcement learning algorithm to obtain a project scheduling model; the project scheduling model is used to schedule engineering and service projects of a tobacco manufacturing enterprise, thereby improving the flexibility and efficiency of scheduling engineering and service projects of the tobacco manufacturing enterprise and reducing the risks of project delays and failures.

[0110] S330 : Adjust the model parameters of the project scheduling model according to the project reward function definition task and the model training task to obtain a converged preset project scheduling model.

[0111] Among them, after executing the project reward function definition task and the model training task, the model parameters of the constructed project scheduling model are continuously updated according to the project reward function to obtain a converged preset project scheduling model to improve the performance of the model.

[0112] As an optional but non-limiting implementation, after obtaining a converged preset project scheduling model, the method further includes but is not limited to steps E1-E2:

[0113] Step E1: Using preset evaluation indicators to perform performance evaluation on the converged preset project scheduling model to obtain evaluation results; the preset evaluation indicators include project completion time, input cost, project risk level, resource utilization and project members' salary satisfaction.

[0114] Step E2: According to the evaluation result, the model parameters of the converged preset project scheduling model are adjusted, and the project reward function is optimized.

[0115] After obtaining a converged preset project scheduling model, the performance of the converged preset project scheduling model needs to be regularly evaluated using preset evaluation indicators. For example, performance evaluation can be performed using at least one of the preset evaluation indicators: project completion time, input cost, project risk level, resource utilization, and project member salary satisfaction.

[0116] Based on the evaluation results, adjust the reinforcement learning model parameters and optimize the project reward function. These parameters include the learning rate and discount factor. For example, if the model converges slowly, increase the learning rate; if the model is too short-sighted, increase the discount factor. Additionally, adjust the various coefficients in the project reward function based on any issues identified during the evaluation process to improve the model's performance and adaptability.

[0117] The embodiment of the present invention improves the performance and accuracy of the model by regularly evaluating the model.

[0118] An embodiment of the present invention provides a method for training a project scheduling model based on reinforcement learning. The method comprises the following steps: executing project status definition, project action definition, and project reward function definition tasks on a project scheduling model to be trained; performing a model training task on the project scheduling model to be trained based on the project status definition, project action definition, and project reward function definition, thereby obtaining a project scheduling model; and adjusting model parameters of the project scheduling model based on the project reward function definition task and the model training task to obtain a converged preset project scheduling model. By adopting the technical solution of the embodiment of the present invention, a project scheduling model based on reinforcement learning is trained to obtain a converged preset project scheduling model, thereby scheduling engineering and service projects of tobacco manufacturing enterprises, reducing the risk of project delays and failures, and improving the success rate of projects.

[0119] Example 4

[0120] Figure 6This is a schematic diagram of the structure of a project scheduling device based on reinforcement learning provided in an embodiment of the present invention. The technical solution of this embodiment can be applied to the case of adaptive project scheduling based on reinforcement learning in a project management system. The device can be implemented by software and / or hardware and is generally integrated into any electronic device with network communication function, including but not limited to: servers, computers, personal digital assistants and other devices. Figure 6 As shown, the project scheduling device based on reinforcement learning provided in this embodiment may include: a current state acquisition module 610, an execution action determination module 620 and a project dynamic scheduling module 630; wherein,

[0121] The current state acquisition module 610 is used to obtain the current state vector of the target project; wherein the target project includes engineering and service projects of the tobacco manufacturing enterprise;

[0122] An execution action determination module 620 is configured to input the current state vector into a preset project scheduling model to output an execution action corresponding to the current state vector; wherein the preset project scheduling model is trained using a reinforcement learning algorithm and adjusts model parameters using a pre-established project reward function to obtain a converged preset project scheduling model;

[0123] The project dynamic scheduling module 630 is used to dynamically schedule the target project according to the execution action.

[0124] Based on the above embodiment, optionally, the current state acquisition module is specifically configured to:

[0125] Acquire at least two pieces of current state information of a target project, quantify and represent the at least two pieces of current state information, and combine them into a current state vector;

[0126] The current status information includes real-time project progress information, external environment change information, basic project attribute information, project member skills and load information.

[0127] Based on the above embodiment, optionally, the execution action determination module includes a model training unit, and the model training unit is specifically configured to:

[0128] For the project scheduling model to be trained, perform the tasks of defining project status, project action, and project reward function.

[0129] According to the project status definition, the project action definition, and the project reward function definition, a reinforcement learning algorithm is used to perform a model training task on the project scheduling model to be trained to obtain a project scheduling model;

[0130] According to the project reward function definition task and the model training task, the model parameters of the project scheduling model are adjusted to obtain a converged preset project scheduling model.

[0131] Based on the above embodiment, optionally, the model training unit is further configured to:

[0132] The project status definition refers to combining at least two project status information into a project status vector; the project action definition refers to defining a set of selectable actions for the model, each action has corresponding execution logic, and the action set includes adjustment of the stage project start time, resource reallocation, change of project task priority, and reallocation of project members.

[0133] Based on the above embodiment, optionally, the model training unit is further configured to:

[0134] Based on the degree of impact of the actions taken by the model on the project progress, a first reward function and a second reward function are defined; the first reward function is a positive reward function obtained by taking an action that shortens the project progress; the second reward function is a negative reward function obtained by taking an action that delays the project progress;

[0135] Based on the degree of impact of the actions taken by the model on the project risk, a third reward function and a fourth reward function are defined; wherein the third reward function is a positive reward function obtained when the actions taken reduce the project risk; and the fourth reward function is a negative reward function obtained when the actions taken increase the project risk;

[0136] Define the fifth reward function based on the utilization of project resources by the actions taken by the model;

[0137] Define the sixth reward function based on the satisfaction of project members with the actions taken by the model;

[0138] The first reward function, the second reward function, the third reward function, the fourth reward function, the fifth reward function, and the sixth reward function are superimposed to define an item reward function.

[0139] Based on the above embodiment, optionally, the model training unit is further configured to:

[0140] Get the current project state at the current time step and select the first project action based on the greedy strategy;

[0141] Execute the first project action, and obtain the second project action and the project reward obtained by executing the first project action based on the change in the project status after the first project action;

[0142] Updating a value table in the reinforcement learning algorithm based on the second project action and the project reward obtained by performing the first project action; the value table is composed of project states and corresponding project actions, and the value table is used to represent the reward that can be obtained by performing the project action corresponding to the project state;

[0143] The task of updating the value table is repeatedly executed to perform the model training task on the project scheduling model to be trained, thereby obtaining the project scheduling model.

[0144] Based on the above embodiment, optionally, the model training unit is further configured to:

[0145] Obtain project data of the target project from the local storage file, and store the project data in the database;

[0146] Preprocessing the project data and quantifying the text data included in the project data;

[0147] The preprocessing includes outlier processing and missing value processing, the text data includes project progress data, and the quantification processing includes converting the text description data of the project progress into a percentage of the project progress.

[0148] Based on the above embodiment, optionally, the model training unit is further configured to:

[0149] Using preset evaluation indicators to perform performance evaluation on the converged preset project scheduling model to obtain evaluation results; the preset evaluation indicators include project completion time, input cost, project risk level, resource utilization rate, and project member salary satisfaction;

[0150] Based on the evaluation results, the model parameters of the converged preset project scheduling model are adjusted, and the project reward function is optimized.

[0151] Based on the above embodiment, optionally, the project dynamic scheduling module is specifically used to:

[0152] If the execution action is to adjust the start time of a stage project, the start time of the stage project and the start time of subsequent stage projects of the stage project are dynamically adjusted;

[0153] If the execution action is resource reallocation, then the resource allocation amounts of projects at different stages are readjusted;

[0154] If the execution action is a change in the priority of a project task, dynamic adjustment is performed based on the priority of the stage project;

[0155] If the execution action is project member reallocation, dynamic adjustments are made based on the project members' skill levels, workloads, and work requirements.

[0156] The reinforcement learning-based project scheduling device provided in the embodiment of the present invention can execute the reinforcement learning-based project scheduling method provided in any of the above-mentioned embodiments of the present invention, and has the corresponding functions and beneficial effects of executing the reinforcement learning-based project scheduling method. For detailed processes, please refer to the relevant operations of the reinforcement learning-based project scheduling method in the above-mentioned embodiments.

[0157] Example 5

[0158] Figure 7 1 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0159] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0160] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0161] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the project scheduling method based on reinforcement learning.

[0162] In some embodiments, the reinforcement learning-based project scheduling method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the reinforcement learning-based project scheduling method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the reinforcement learning-based project scheduling method in any other appropriate manner (e.g., by means of firmware).

[0163] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0164] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0165] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0166] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0167] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0168] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0169] Example 6

[0170] An embodiment of the present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the reinforcement learning-based project scheduling method provided in any embodiment of the present application.

[0171] The computer program product may be implemented by writing computer program code for performing the operations of the present invention in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0172] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0173] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A project scheduling method based on reinforcement learning, characterized in that: The method comprises: Obtaining a current state vector of a target project; wherein the target project includes engineering and service projects of a tobacco manufacturing enterprise; Inputting the current state vector into a preset project scheduling model to output an execution action corresponding to the current state vector; wherein the preset project scheduling model is trained using a reinforcement learning algorithm and the model parameters are adjusted using a pre-built project reward function to obtain a converged preset project scheduling model; According to the execution action, dynamic project scheduling is performed on the target project.

2. The method according to claim 1, characterized in that The obtaining of the current state vector of the target project includes: Acquire at least two pieces of current state information of a target project, quantify and represent the at least two pieces of current state information, and combine them into a current state vector; The current status information includes real-time project progress information, external environment change information, basic project attribute information, project member skills and load information.

3. The method according to claim 1, characterized in that The preset project scheduling model is trained in the following way: For the project scheduling model to be trained, perform the tasks of defining project status, project action, and project reward function. According to the project status definition, the project action definition, and the project reward function definition, a reinforcement learning algorithm is used to perform a model training task on the project scheduling model to be trained to obtain a project scheduling model; According to the project reward function definition task and the model training task, the model parameters of the project scheduling model are adjusted to obtain a converged preset project scheduling model.

4. The method according to claim 3, characterized in that The project status definition refers to combining at least two project status information into a project status vector; the project action definition refers to defining a set of selectable actions for the model, each action has corresponding execution logic, and the action set includes adjustment of the stage project start time, resource reallocation, change of project task priority, and reallocation of project members.

5. The method according to claim 3, characterized in that The project scheduling model to be trained performs the project reward function definition task, including: Based on the degree of impact of the actions taken by the model on the project progress, a first reward function and a second reward function are defined; the first reward function is a positive reward function obtained by taking an action that shortens the project progress; the second reward function is a negative reward function obtained by taking an action that delays the project progress; Based on the degree of impact of the actions taken by the model on the project risk, a third reward function and a fourth reward function are defined; wherein the third reward function is a positive reward function obtained when the actions taken reduce the project risk; and the fourth reward function is a negative reward function obtained when the actions taken increase the project risk; Define the fifth reward function based on the utilization of project resources by the actions taken by the model; Define the sixth reward function based on the satisfaction of project members with the actions taken by the model; The first reward function, the second reward function, the third reward function, the fourth reward function, the fifth reward function, and the sixth reward function are superimposed to define an item reward function.

6. The method according to claim 3, characterized in that The project scheduling model to be trained is obtained by using a reinforcement learning algorithm to perform a model training task on the project scheduling model to be trained based on the project status definition, the project action definition, and the project reward function definition, including: Get the current project state at the current time step and select the first project action based on the greedy strategy; Execute the first project action, and obtain the second project action and the project reward obtained by executing the first project action based on the change in the project status after the first project action; Updating a value table in the reinforcement learning algorithm based on the second project action and the project reward obtained by performing the first project action; the value table is composed of project states and corresponding project actions, and the value table is used to represent the reward that can be obtained by performing the project action corresponding to the project state; The task of updating the value table is repeatedly executed to perform the model training task on the project scheduling model to be trained, thereby obtaining the project scheduling model.

7. The method according to claim 3, characterized in that Before training the preset project scheduling model, the method further includes: Obtain project data of the target project from the local storage file, and store the project data in the database; Preprocessing the project data and quantifying the text data included in the project data; The preprocessing includes outlier processing and missing value processing, the text data includes project progress data, and the quantification processing includes converting the text description data of the project progress into a percentage of the project progress.

8. The method according to claim 3, characterized in that After obtaining a converged preset project scheduling model, the method further includes: Using preset evaluation indicators to perform performance evaluation on the converged preset project scheduling model to obtain evaluation results; the preset evaluation indicators include project completion time, input cost, project risk level, resource utilization rate, and project member salary satisfaction; Based on the evaluation results, the model parameters of the converged preset project scheduling model are adjusted, and the project reward function is optimized.

9. The method according to claim 1, characterized in that The dynamically scheduling the target project according to the execution action includes: If the execution action is to adjust the start time of a stage project, the start time of the stage project and the start time of subsequent stage projects of the stage project are dynamically adjusted; If the execution action is resource reallocation, then the resource allocation amounts of projects at different stages are readjusted; If the execution action is a change in the priority of a project task, dynamic adjustment is performed based on the priority of the stage project; If the execution action is project member reallocation, dynamic adjustments are made based on the project members' skill levels, workloads, and work requirements.

10. A project scheduling device based on reinforcement learning, characterized in that: The device comprises: A current state acquisition module is used to obtain the current state vector of a target project; wherein the target project includes engineering and service projects of a tobacco manufacturing enterprise; an execution action determination module, configured to input the current state vector into a preset project scheduling model to output an execution action corresponding to the current state vector; wherein the preset project scheduling model is trained using a reinforcement learning algorithm and adjusts model parameters using a pre-built project reward function to obtain a converged preset project scheduling model; The project dynamic scheduling module is used to dynamically schedule the target project according to the execution action.

Citation Information

Cited By

  • Financial data risk analysis method and device based on reinforcement learning and medium

    CN120952994A