Operation case distribution method based on deep reinforcement learning

Through the in-depth reinforcement learning training agent, combined with distribution rewards and overall rewards, the scheduling lag problem of surgical case allocation method in dynamic environments is solved, intelligent and real-time surgical scheduling is realized, and the utilization rate and operational efficiency of medical resources are improved.

CN120376074APending Publication Date: 2025-07-25HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510393219.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing surgical case allocation methods lack dynamic adaptability in dynamic environments, and are difficult to meet real-time requirements, and cannot implement feedback dynamic optimization scheduling strategies, resulting in scheduling lag and waste of resources.

Method used

The surgical case allocation method based on deep reinforcement learning is adopted, and the agent is trained through the deep Q network, and the reward function is formed by combining distributed rewards and overall rewards to realize intelligent and real-time scheduling decisions, and the exploration ability and convergence speed are improved by using greed strategies and epsilon-decreasing mechanisms.

Benefits of technology

In a dynamic environment, efficient and autonomous scheduling decisions have been achieved, artificial intervention has been reduced, computing efficiency and generalization capabilities have been improved, rapid changes in surgical plans have been adapted to the utilization rate of medical resources and operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_12
    Figure SMS_12
  • Figure SMS_13
    Figure SMS_13
  • Figure SMS_16
    Figure SMS_16
Patent Text Reader

Abstract

The invention discloses a surgical case allocation method based on deep reinforcement learning. The method comprises the following steps: step 1, obtaining a historical surgical case allocation data set and dividing the historical surgical case allocation data set into a training set and a test set; a deep reinforcement learning agent is established, a state space is formed by case information features and operating room information features in the agent, an action space of two layers of rules is defined, and the agent forms a reward function for action execution through combination of distributed rewards and overall rewards; step 2, training the intelligent agent by using a deep Q network to obtain a distribution model; and step 3, deploying the distribution model obtained in the step 2 in a hospital for operation case distribution. The method has the remarkable advantages in the aspects of dynamic adaptability, calculation efficiency, autonomous learning ability and the like, and effective support is provided for intelligent medical treatment and intelligent operation scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of surgical case allocation methods, and specifically, to a surgical case allocation method based on deep reinforcement learning. Background Art

[0002] In the surgical scheduling management of modern hospitals, reasonable and efficient surgical case allocation is crucial, directly affecting the utilization rate of medical resources and the working efficiency of medical staff. However, during the execution of surgical plans, they are often affected by many dynamic factors, such as the temporary cancellation of surgeries, the addition of emergency surgeries, the change of surgical time, and the scheduling problems of medical staff. These uncertain factors make the allocation and scheduling of surgical cases extremely challenging, and traditional surgical scheduling methods are difficult to handle efficiently.

[0003] Currently, hospitals mainly use rule-driven methods for surgical scheduling, such as scheduling based on heuristic algorithms (such as genetic algorithms, ant colony algorithms) or optimization models (such as integer programming, constraint programming). These methods can provide certain optimization effects in static scheduling situations, but in the face of frequent changes in surgical plans, they often require a large amount of manual intervention or recalculation, resulting in scheduling lags, resource waste, and even surgical conflicts, affecting the overall medical operation efficiency.

[0004] Current surgical case allocation mainly relies on scheduling methods based on rules, heuristic algorithms, or optimization models. Although these methods have improved the utilization rate of surgical resources to a certain extent, there are still many deficiencies in the dynamic surgical scheduling environment. First, most existing methods are based on static scheduling assumptions and lack dynamic adaptability, making it difficult to handle emergencies. Second, traditional optimization methods have high computational complexity and are difficult to meet real-time requirements. Although heuristic algorithms can improve computational efficiency to a certain extent, they usually rely on manually set rules and lack the ability of autonomous optimization, making it difficult to handle complex and changeable scheduling requirements. In addition, existing methods cannot dynamically optimize scheduling strategies based on historical data and real-time feedback.

[0005] Therefore, how to achieve an intelligent and real-time adjustable surgical case allocation method in a dynamic environment has become a research hotspot. Summary of the Invention

[0006] The present invention provides a surgical case allocation method based on deep reinforcement learning to solve the problems of the existing surgical case allocation method, such as the lack of dynamic adaptability, the difficulty in meeting real-time requirements, and the inability to implement feedback for dynamic optimization of scheduling strategies.

[0007] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0008] A surgical case allocation method based on deep reinforcement learning, comprising the following steps:

[0009] Step 1: Obtain the historical surgical case allocation dataset and divide the historical surgical case allocation dataset into a training set and a test set;

[0010] And establish an agent for deep reinforcement learning. In the agent, the state space is composed of case information features and operating room information features; A two-layer rule action space is defined in the agent, where the first layer of rules is used to select appropriate cases, and the second layer of rules is used to select the allocated operating room and surgical time; In the agent, a reward function for executing actions is formed by combining distributed rewards and overall rewards;

[0011] Step 2: Train the agent using a deep Q-network. The training process is as follows:

[0012] Process the data in the training set and use it as the given state S0 to input into the main network of the deep Q-network. After passing through the ReLU activation function of the hidden layer, the Q value of each action is obtained at the output layer;

[0013] Based on the Q value output by the agent for the current state, select an action a based on the greedy policy;

[0014] After executing the selected action a, update the environment, transfer to the next state S1, and return an immediate reward r(t);

[0015] Use the target Q-network to calculate the target Q value, calculate the loss function based on the Q value of the main network and the target Q value of the target Q-network, and update the parameters θ of the backbone network through the backpropagation algorithm based on the calculation result of the loss function;

[0016] Use the experience replay mechanism to put the experience quadruple of the agent, namely state S0, action a, reward r(t), and next state S1, into the temporary experience replay buffer pool, and randomly sample a batch of experiences during the training process to update the network parameters;

[0017] Loop the above process until an episode ends; After that, calculate the overall reward R2. By combining the overall reward and the step-by-step reward, form the final reward R, and replace r(t) in the experience quadruple in the temporary experience replay buffer pool with R and store it in the experience replay buffer pool;

[0018] Repeat the above episode process until the loss function of the deep Q-network converges, thereby completing the training of the agent and using the trained agent as the allocation model;

[0019] Step 3: Deploy the allocation model obtained in Step 2 to the hospital, and the allocation model performs surgical case allocation according to the current case information and operating room information in the hospital.

[0020] In the further step 1, when the agent selects cases based on the first - layer rules of the action space, it follows the following rules: (1) Forced cases have priority; (2) Cases that may be rejected have priority; (3) Newly inserted cases have priority; (4) Cases that match the remaining time have priority.

[0021] In the further step 1, when the agent selects the operating room and operation time based on the second - layer rules of the action space, it follows the following rules: (1) For the earliest available operating room, select the operating room that can schedule the current case and has the earliest time; (2) For the operating room that matches the operation duration, select the operating room whose remaining normal operation time is closest to the operation duration of the current case; (3) Select the operation with the longest remaining available time; (4) Select the operating room with the lowest utilization rate.

[0022] In the further step 1, the agent uses the comparison result between the end time of the operating room after each step of decision - making and the fixed normal working hours of the operating room to give step - by - step rewards to evaluate the immediate impact of each step of decision - making on the final scheduling result;

[0023] The agent calculates the overall reward through the difference between the total cost of the operating room center caused after the scheduling is completed and the total cost of the operating room center under the ideal situation.

[0024] In the further step 1, the agent calculates the weighted sum of the distributed reward and the overall reward as the reward function for performing actions.

[0025] In the further step 2, the action selection strategy is a greedy strategy.

[0026] Furthermore, during the training in step 2, an epsilon - decreasing mechanism is adopted to gradually reduce the exploration probability.

[0027] Furthermore, during the training in step 2, a Bayesian optimization method is adopted to adjust and optimize the parameters of the deep Q - network.

[0028] The present invention realizes the intelligence, real - time and high - efficiency of surgical scheduling through the adaptive optimization ability of reinforcement learning.

[0029] The present invention adopts the deep Q - network (DQN) algorithm, combines the advantages of deep learning and reinforcement learning, improves the generalization ability of decision - making, and can quickly adapt to the changes of surgical plans in a dynamic environment. In addition, in order to enhance the exploration ability of the agent and improve the convergence speed of the scheduling strategy, the present invention adopts a greedy strategy. Through the epsilon - decreasing mechanism, the agent fully explores in the initial stage of training to avoid falling into local optimal solutions, and gradually reduces the random exploration ratio in the later stage of training, so that the scheduling decision tends to converge to the global optimum.

[0030] The present invention has significant advantages in aspects such as dynamic adaptability, computational efficiency, and autonomous learning ability, providing effective support for intelligent healthcare and intelligent surgical scheduling. Compared with the prior art, it has the following advantages:

[0031] 1) Adapt to dynamic changes and improve generalization ability. When faced with dynamic adjustments, the prior art may require re-design or adjustment. By continuously learning environmental changes, the present invention can adapt to dynamic adjustments, does not require artificial setting of fixed rules, and has stronger generalization ability.

[0032] 2) Stronger autonomous decision-making ability and reduced human intervention. The present invention makes decisions directly based on historical data and environmental feedback, reducing human intervention and improving decision-making efficiency.

[0033] 3) Higher computational efficiency and faster convergence speed. The present invention accelerates training through parallel computing and can quickly converge through reinforcement learning strategies, with higher computational efficiency in complex tasks. Detailed implementation manners

[0034] The present invention will be further described below in conjunction with embodiments.

[0035] This embodiment discloses a surgical case allocation method based on deep reinforcement learning, including the following steps:

[0036] Step 1, obtain a historical surgical case allocation data set, and divide the historical surgical case allocation data set into a training set and a test set.

[0037] Establish an agent for deep reinforcement learning, including defining the state space, action space, and reward function of the agent, which are specifically described as follows:

[0038] In this embodiment, it is defined that the state space is composed of case information features and operating room information features. The initial case information and operating room information are integrated and processed to extract features to describe the overall state of the case and the operating room. In this embodiment, the following state matrix S is constructed RP :

[0039] S RP =[S P ,S R

[0040] In the formula: S P =[S p1 ,S p2 ,S p3 ,S p4 ,S p5 ,S p6 is the extracted case feature.

[0041] S p1 ​For the progress of scheduling, it is the proportion of the number of cases in the scheduled case set SW to the total number of all cases n.

[0042] S p2 and S p3 are the mean and standard deviation of the operation durations of the waiting case set SW respectively, and there is P i is the operation duration of patient i, SW is the waiting case set, and |SW| is the number of patients in the waiting case set.

[0043] S p4 is the average due date of the waiting case set SW, and there is C i is the due date of patient i.

[0044] S p5 and S p6 are the shortest and longest operation durations in the waiting case set SW respectively, and there is S p5 = min i∈ sw P i , S p6 = max i∈SW P i .

[0045] S R = [S r1 , S r2 , S r3 , S r4 , S r5 is the operating room characteristic.

[0046] S r1 and S r2 are the mean and standard deviation of the utilization rate of the operating room respectively. The utilization rate U of the operating room kt is defined as The mean of the utilization rate U of the operating room kt The standard deviation of the utilization rate U of the operating room The utilization rate U of the operating room kt The standard deviation of the utilization rate U of the operating room SD kt is the actual operation time of operating room k on the t-th day; T kt is the set normal available operation time of operating room k on the t-th day; m is the number of operating rooms; h is the planning period time.

[0047] S r3 is the average of the remaining maximum available operation time of the operating room, and there is O kt is the set maximum allowable overtime of operating room k on the t-th day.

[0048] S r4 is the average of the remaining normal operation time in the operating room, and there is

[0049] S r5 is the total execution cost of the operating room, that is, the total operation cost of the operations caused by the scheduled cases, and there is WC kt is the waste cost of operating room k on the t-th day; OC kt is the overtime cost of operating room k on the t-th day.

[0050] In this embodiment, an action space of two-layer rules is defined to ensure the rationality of the case allocation process. Among them, the first-layer rule is used to select appropriate cases, and the second-layer rule is used to select the allocated operating room and operation time. Through the combination of the two-layer rules, the action space can cover a variety of flexible scheduling schemes.

[0051] Among them, when the agent performs case selection actions based on the first-layer rule of the action space, it needs to comprehensively consider the arrival date, deadline, operation duration, and case type of the case to ensure that all mandatory cases can be completed on time, specifically following the following rules:

[0052] (1) Mandatory cases first, select the mandatory case with the smallest deadline from the waiting case set SW, as shown in the following formula:

[0053] Case = min i∈SW C i . Among them, C i is the deadline of patient i

[0054] (2) Cases that may be rejected first, select the cases whose operation duration exceeds the average remaining available time of the operating room from the waiting case set SW, as shown in the following formula:

[0055] Case = min i∈SW P i > S r4 ,

[0056] Among them, case refers to the selected case, P i is the operation duration of patient i, S r4 is the average of the remaining normal operation time in the operating room

[0057] (3) Newly inserted cases first, when there are newly inserted cases, give priority to selecting this case, as shown in the following formula:

[0058] Case = i ∈ SI,

[0059] Among them, SI is the set of newly inserted cases;

[0060] (4) Give priority to the cases that match the remaining time. Select the case in the waiting case set SW that is closest to the average remaining normal operating time of the operating room, as shown in the following formula:

[0061] Case = min i∈SW |S r4 -P i |.

[0062] When the agent selects the operating room and operating time based on the second layer of rules of the action space, it is assumed that each operating room has a maximum available overtime, and the end time of the operation in the operating room cannot exceed this time. Therefore, when selecting the operating room, it is necessary to ensure that the end time of the operation of the case is within the maximum overtime. After the case is determined, starting from the optimization goal and constraints, the selection of the operating room follows the following rules:

[0063] (1) The earliest available operating room. Select the operating room that can schedule the current case and has the earliest time, as shown in the following formula:

[0064]

[0065] Among them, room refers to the operating room selected by the rule, case represents the selected surgical case, t is the number of days of the selected operating room, Case(C i ) represents the deadline of the selected case, Case(P i ) represents the surgical duration of the selected case, SD kt is the actual operating time of operating room k on the t-th day; T kt is the set normal available operating time of operating room k on the t-th day, O kt is the set maximum allowable overtime of operating room k on the t-th day.

[0066] (2) The operating room that matches the surgical duration. Select the operating room with the remaining normal operating time closest to the surgical duration of the current case, as shown in the following formula:

[0067] Room = min|T kt -Case(P i )|

[0068] (3) Select the operation with the longest remaining available time, as shown in the following formula:

[0069] Room = max(T kt +O kt -SD kt )

[0070] (4) Select the operating room with the lowest utilization rate, as shown in the following formula:

[0071] Room = minU kt Among them, U kt represents the utilization rate of the operating room

[0072] In this embodiment, the agent forms a reward function for executing actions by combining distributed rewards and overall rewards.

[0073] Specifically, in this embodiment, the comparison result between the end time of the operating room after each step of decision-making and the fixed normal working hours of the operating room is used for step-by-step rewards to evaluate the immediate impact of each step of decision-making on the final scheduling result. Through the current state S j and the next state S j+1 The state characteristics change, and the step-by-step rewards are dynamically calculated. The specific calculation rules of the step-by-step rewards r(t) are shown in Table 1, where RTK(j + 1) is the end time of the operating room after scheduling, and T kt is the fixed normal working hours of the operating room. Table 1 is as follows:

[0074] Table 1 Specific calculation rules of step-by-step rewards r(t)

[0075]

[0076] In this embodiment, the overall reward is used to measure the completion effect of the entire scheduling task, which is closely related to the optimization goal of this embodiment, minimizing the total cost of the operating room center. The overall reward is calculated by the difference between the total cost of the operating room center after the scheduling is completed and the total cost of the operating room center under ideal conditions. First, the optimal cost C b needs to be calculated, and its calculation formula is as follows:

[0077]

[0078] Among them, n represents the number of patients, m represents the number of operating rooms, and t represents the planned number of days. represents the sum of the normal available times of all operating rooms during the planned period, represents the sum of the surgical durations of all patients, and α and β are the cost coefficients in two different situations.

[0079] Then, the overall reward is calculated by the normalized difference between the optimal cost and the actual cost cost, as shown in the following formula:

[0080]

[0081] Among them, cost is the actual cost, and C b is the optimal cost under ideal conditions.

[0082] Since the dimensions of the step-by-step reward and the overall reward are different, directly adding them will affect the convergence effect of the model. Therefore, in this embodiment, the step-by-step reward is normalized to have the same dimension as the overall reward. The specific normalization formula is as follows:

[0083]

[0084] where r(t) is the step-by-step reward, is the number of arranged cases.

[0085] After that, the cumulative value of the step-by-step reward within the entire round is calculated through the following formula:

[0086]

[0087] Finally, the reward R for the execution action obtained by the agent after the end of the entire round is the weighted sum of the distributed reward and the overall reward, and the calculation formula is as follows:

[0088] R = r1×R1 + r2×R2

[0089] where: r1 and r2 are adjustable parameters used to control the weight ratio of the step-by-step reward to the overall reward. By adjusting the values of r1 and r2, the reward function designed in this paper can effectively prevent the model from falling into a local optimal solution, thereby improving the global optimization ability.

[0090] Step 2: Train the agent using a deep Q-network.

[0091] The deep Q-network adopted in this embodiment includes a main network and a target Q-network. The main network is used to select action targets; the target Q-network is used to calculate the target Q value and update the main network. The deep Q-network adopts a seven-layer fully connected neural network, including an input layer, five hidden layers, and an output layer, and its structure is shown in Table 2:

[0092] Table 2 Neural network structure table

[0093]

[0094]

[0095] The training process is as follows:

[0096] The data in the training set is processed (the state space setting below) and then input as the given state S0 into the main network of the deep Q-network. After passing through the ReLU activation function of the hidden layer, the Q value of each action is obtained at the output layer;

[0097] The agent selects an action a based on the greedy strategy according to the Q value output by the current state;

[0098] After performing the selected action a, update the environment, transition to the next state S1, and return an immediate reward r(t);

[0099] Use the target Q-network to calculate the target Q-value, and calculate the loss function (mean squared error) based on the Q-value of the main network and the target Q-value of the target Q-network. Update the parameters θ of the backbone network through the backpropagation algorithm based on the calculation result of the loss function;

[0100] Use the experience replay mechanism to put the agent's experience quadruple (i.e., state S0, action a, reward r(t), next state S1) into the temporary experience replay buffer pool, and randomly sample a batch of experiences during training to update the network parameters.

[0101] Loop the above process until an episode ends. After that, calculate the overall reward R2, form the final reward R by combining the overall reward and the step-by-step reward, and replace r(t) in the experience quadruple in the temporary experience replay buffer pool with R and store it in the experience replay buffer pool;

[0102] Repeat the above episode process until the loss function of the deep Q-network converges, thus completing the training of the agent, and using the trained agent as the allocation model.

[0103] During training in this embodiment, the action selection strategy adopted by the agent is the greedy strategy (epsilon-greedy).

[0104] During training in this embodiment, the epsilon-decreasing mechanism is used to gradually reduce the exploration probability. Specifically, first set the initial exploration probability ε0 and the minimum exploration probability ε min . At each time step t, update the exploration probability ε according to the following decreasing strategy.

[0105] ε t =max(ε min , ε0 - slope×t)

[0106] where slope is the amount of decrease per step.

[0107] According to the current exploration probability ε, the rules for the agent to select actions are as follows:

[0108]

[0109] The specific algorithm flow of the deep reinforcement learning model is shown in Table 3:

[0110] Table 3 Algorithm Flow of the Model

[0111]

[0112] During the training in this embodiment, Bayesian optimization is used to adjust and optimize the hyperparameters of the deep Q-network, including the learning rate, discount factor, batch size, experience pool capacity, exploration rate decay parameter, etc. The hyperparameter settings for the deep reinforcement learning model training are shown in Table 4:

[0113] Table 4 Hyperparameter Settings for Model Training

[0114]

[0115]

[0116] After the training is completed, the effectiveness of the model is verified using the test set, and the scheduling total cost index is used to measure the model's performance. By comparing the scheduling total cost with other methods, it shows that this model can find the known best solution and exhibits significant effectiveness advantages, as shown in Table 5:

[0117] Table 5 Performance Comparison between DDQN and Six Other Solving Methods

[0118]

[0119] Step 3: Deploy the allocation model obtained in Step 2 to the intelligent execution system of the hospital. The allocation model performs surgical case allocation based on the current case information and operating room information in the hospital to achieve online surgical scheduling optimization. When situations such as surgery cancellation or insertion occur, the system automatically adjusts the scheduling plan.

[0120] The preferred embodiments of the present invention have been described in detail above. The embodiments described in the present invention are merely descriptions of the preferred embodiments of the present invention, and do not limit the concept and scope of the present invention. Among the various specific technical features described in the above specific embodiments, they can be combined in any suitable way without contradiction. As long as such a combination does not violate the idea of the present invention, it should also be regarded as the content disclosed in this disclosure. To avoid unnecessary repetition, the present invention does not further explain various possible combination methods.

[0121] The present invention is not limited to the specific details in the above embodiments. Within the technical concept scope of the present invention and without departing from the design idea of the present invention, various variations and improvements made by those skilled in the art to the technical solutions of the present invention should all fall within the protection scope of the present invention. The technical content claimed by the present invention has been fully recorded in the claims.

Claims

1. A surgical case allocation method based on deep reinforcement learning, characterized in that It includes the following steps: Step 1: Obtain the historical surgical case allocation dataset and divide the historical surgical case allocation dataset into a training set and a test set; And establish an agent for deep reinforcement learning. In the agent, the state space is composed of case information features and operating room information features; A two-layer rule action space is defined in the agent, where the first-layer rule is used to select appropriate cases, and the second-layer rule is used to select the allocated operating room and surgical time; In the agent, a reward function for executing actions is formed by combining distributed rewards and overall rewards; Step 2: Use a deep Q-network to train the agent. The training process is as follows: Process the data in the training set and input it as the given state S0 into the main network of the deep Q-network. After passing through the ReLU activation function of the hidden layer, obtain the Q value of each action at the output layer; The agent selects an action a based on the greedy policy according to the Q value output by the current state; After executing the selected action a, update the environment, transfer to the next state S1, and return an immediate reward r(t); Use the target Q-network to calculate the target Q value, calculate the loss function according to the Q value of the main network and the target Q value of the target Q-network, and update the parameters θ of the backbone network through the backpropagation algorithm based on the calculation result of the loss function; Use the experience replay mechanism to put the experience quadruple of the agent, namely state S0, action a, reward r(t), and next state S1, into the temporary experience replay buffer pool, and randomly sample a batch of experiences during the training process to update the network parameters; Loop the above process until an episode ends; After that, calculate the overall reward R2. By combining the overall reward and the step-by-step reward, form the final reward R, and replace r(t) in the experience quadruple in the temporary experience replay buffer pool with R and store it in the experience replay buffer pool; Repeat the above episode process until the loss function of the deep Q-network converges, thus completing the training of the agent, and use the trained agent as the allocation model; Step 3: Deploy the allocation model obtained in Step 2 to the hospital, and the allocation model allocates surgical cases according to the current case information and operating room information in the hospital.

2. The surgical case allocation method based on deep reinforcement learning according to claim 1, wherein In Step 1, when the agent makes a case selection action based on the first-layer rule of the action space, it follows the following rules: (1) Force cases first; (2) Cases that may be rejected first; (3) Newly inserted cases first; (4) Cases that match the remaining time first.

3. The surgical case allocation method based on deep reinforcement learning according to claim 1, characterized in that, In Step 1, when the agent makes a selection of the operating room and surgical time based on the second-layer rule of the action space, it follows the following rules: (1) The earliest available operating room, select the operating room that can arrange the current case and has the earliest time; (2) The operating room that matches the surgical duration, select the operating room whose remaining normal surgical time is closest to the surgical duration of the current case; (3) Select the surgery with the longest remaining available time; (4) Select the operating room with the lowest utilization rate.

4. The surgical case allocation method based on deep reinforcement learning according to claim 1, wherein In Step 1, the agent uses the comparison result between the end time of the operating room after each step of decision-making and the fixed normal working hours of the operating room to perform step-by-step rewards to evaluate the immediate impact of each step of decision-making on the final scheduling result; The agent calculates the overall reward based on the difference between the total cost of the operating room center after the scheduling is completed and the total cost of the operating room center under the ideal situation.

5. The surgical case allocation method based on deep reinforcement learning according to claim 1 or 4, characterized in that, In step 1, the agent calculates the weighted sum of the distributed reward and the overall reward as the reward function for performing the action.

6. The surgical case allocation method based on deep reinforcement learning according to claim 1, wherein, In step 2, the action selection strategy is a greedy strategy.

7. The surgical case allocation method based on deep reinforcement learning according to claim 1, wherein During the training in step 2, an epsilon-decreasing mechanism is adopted to gradually reduce the exploration probability.

8. The surgical case allocation method based on deep reinforcement learning according to claim 1, wherein During the training in step 2, the Bayesian optimization method is used to adjust and optimize the parameters of the deep Q-network.