Acetabular periarticular osteotomy planning method and device based on reinforcement learning
Through deep Q network training based on reinforcement learning, the osteotomy position around the acetabular is optimized, and the problem of inaccurate setting of the osteotomy surface is solved, and the precise adjustment and rapid recovery of the acetabular position are achieved.
Patent Information
- Application Number
- CN202411042156.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2044-07-31
AI Technical Summary
The existing setting of osteotomy surface position during periacetabular osteotomy depends on doctor's experience, resulting in excessive fluctuations in the effect and lack of accuracy and consistency.
Using a reinforcement learning-based method, the osteotomy position around the acetabular is iteratively trained through deep Q network learning, and the osteotomy position with the highest cumulative reward is selected as the planning, and the position of the osteotomy surface is determined and optimized using hip medical images.
It improves the accuracy of osteotomy position setting, reduces postoperative recovery time, ensures that the acetabular has the smallest gap when it transfers to the adjusted position, and improves the surgical effect.
Smart Images

Figure CN118806432B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of medical image processing, and in particular to a method and device for planning periacetabular osteotomy based on reinforcement learning. Background Art
[0002] Periacetabular osteotomy (PAO) is a surgical procedure used to treat developmental dysplasia of the hip. PAO improves hip joint coverage by repositioning the acetabulum, thereby alleviating pain and delaying or preventing the progression of hip degenerative disease.
[0003] Among them, the position setting of the osteotomy surface during periacetabular osteotomy has a great impact on the success of the operation and postoperative recovery. Currently, it is only manually set based on the doctor's personal experience, which has great uncertainty. Summary of the Invention
[0004] The problem to be solved by this application is that the current manual setting of the osteotomy surface position results in excessive fluctuations in the setting effect.
[0005] To solve the above problems, the first aspect of the present application provides a periacetabular osteotomy planning method based on reinforcement learning, comprising:
[0006] Obtain medical images of the hip joint;
[0007] determining adjustment information of the acetabulum based on the hip joint medical image;
[0008] Based on reinforcement learning and acetabular adjustment information, the osteotomy position around the acetabulum is trained;
[0009] After the training is completed, the osteotomy position with the highest cumulative reward is selected as the periacetabular osteotomy plan;
[0010] The reinforcement learning is deep Q network learning.
[0011] A second aspect of the present application provides a periacetabular osteotomy planning device based on reinforcement learning, comprising:
[0012] An image acquisition module, which is used to acquire medical images of the hip joint;
[0013] an adjustment determination module, configured to determine adjustment information of the acetabulum based on the hip joint medical image;
[0014] A reinforcement learning module, which is used to train periacetabular osteotomy positions based on reinforcement learning and acetabular adjustment information;
[0015] The osteotomy planning module is used to select the osteotomy position with the highest cumulative reward as the periacetabular osteotomy planning after training is completed; the reinforcement learning is deep Q network learning.
[0016] A third aspect of the present application provides an electronic device, comprising: a memory and a processor;
[0017] The memory is used to store programs;
[0018] The processor, coupled to the memory, is configured to execute the program to:
[0019] Obtain medical images of the hip joint;
[0020] determining adjustment information of the acetabulum based on the hip joint medical image;
[0021] Based on reinforcement learning and acetabular adjustment information, the osteotomy position around the acetabulum is trained;
[0022] After the training is completed, the osteotomy position with the highest cumulative reward is selected as the periacetabular osteotomy plan;
[0023] The reinforcement learning is deep Q network learning.
[0024] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the program is executed by a processor to implement the above-mentioned reinforcement learning-based periacetabular osteotomy planning method.
[0025] In this application, the periacetabular osteotomy position is iterated and trained by reinforcement learning, so that the obtained periacetabular osteotomy position ensures the accuracy of the osteotomy position setting.
[0026] In the present application, by iterating and training the osteotomy position around the acetabulum, the acetabulum can be transferred to the adjusted position with the smallest gap after osteotomy at the osteotomy position, thereby accelerating postoperative recovery and surgical effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a flow chart of a periacetabular osteotomy planning method according to an embodiment of the present application;
[0028] Figure 2 A schematic diagram of the osteotomy surface arrangement of the periacetabular osteotomy planning method according to an embodiment of the present application;
[0029] Figure 3 4 is an architecture diagram of a deep Q-network model for a periacetabular osteotomy planning method according to an embodiment of the present application;
[0030] Figure 4is a structural block diagram of a periacetabular osteotomy planning device according to an embodiment of the present application;
[0031] Figure 5 2 is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0032] To make the above-mentioned objects, features, and advantages of the present application more clearly understood, specific embodiments of the present application are described in detail below with reference to the accompanying drawings. Although the accompanying drawings show exemplary embodiments of the present application, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0033] It should be noted that, unless otherwise specified, the technical or scientific terms used in this application should have the common meanings understood by those skilled in the art to which this application belongs.
[0034] To address the above issues, the present application provides a new acetabular osteotomy planning scheme based on reinforcement learning, which can iterate and optimize the osteotomy position through reinforcement learning, solving the problem of excessive fluctuation in the current manual setting of the osteotomy surface position.
[0035] The embodiment of the present application provides a method for planning periacetabular osteotomy based on reinforcement learning, the specific scheme of the method is as follows: Figure 1-Figure 3 As shown, the method can be performed by a periacetabular osteotomy planning device based on reinforcement learning, and the periacetabular osteotomy planning device based on reinforcement learning can be integrated into electronic devices such as computers, servers, computers, server clusters, and data centers. Figure 1 , which is a flow chart of a periacetabular osteotomy planning method based on reinforcement learning according to one embodiment of the present application; wherein the periacetabular osteotomy planning method based on reinforcement learning includes:
[0036] S100, acquiring medical images of the hip joint;
[0037] The hip joint medical image can be an X-ray medical image, a CT medical image, an MRI medical image, or the like, as long as the hip joint can be segmented therefrom and a three-dimensional point cloud of the bones can be generated. It should be noted that the X-ray medical image can be obtained by performing bone segmentation and projection on the CT medical image or the MRI medical image.
[0038] S200, determining adjustment information of the acetabulum based on the hip joint medical image;
[0039] S300, training periacetabular osteotomy locations based on reinforcement learning and acetabular adjustment information;
[0040] Among them, the osteotomy position around the acetabulum is the osteotomy surface distributed around the acetabulum. Since it is planar, only four planes need to be determined. The intersection of these four planes and the three-dimensional model of the acetabulum is the outer edge of the corresponding osteotomy surface.
[0041] S400, after the training is completed, the osteotomy position with the highest cumulative reward is selected as the periacetabular osteotomy planning; the reinforcement learning is deep Q network learning.
[0042] Among them, training is the training of deep Q network learning. After the training is completed, each action of the new hip joint medical image can be predicted and executed based on the trained deep Q network learning. Finally, the action sequence with the highest cumulative reward is selected. The osteotomy position corresponding to this action sequence is the osteotomy plane of the periacetabular osteotomy planning.
[0043] In this application, the periacetabular osteotomy position is iterated and trained by reinforcement learning, so that the obtained periacetabular osteotomy position ensures the accuracy of the osteotomy position setting.
[0044] In the present application, by iterating and training the osteotomy position around the acetabulum, the acetabulum can be transferred to the adjusted position with the smallest gap after osteotomy at the osteotomy position, thereby accelerating postoperative recovery and surgical effect.
[0045] In one embodiment, the step S200 of determining adjustment information of the acetabulum based on the hip joint medical image includes:
[0046] Performing image segmentation and three-dimensional reconstruction on the medical image of the hip joint to obtain a three-dimensional model of the hip joint;
[0047] Identifying key points of the three-dimensional model of the hip joint;
[0048] Analyze the three-dimensional model of the hip joint to evaluate the direction of the acetabulum opening, the location and extent of coverage of the defect;
[0049] Based on the evaluation results and the standard acetabular inclination angle, adjusted acetabular angle information and acetabular position information are determined.
[0050] In this application, the key points of the three-dimensional model of the hip joint are identified in order to accurately position the three-dimensional model so that the computer can understand the precise positions of the three-dimensional model; the acetabulum, femoral head and related anatomical landmarks are marked on the three-dimensional model. These marks help to accurately align when adjusting the position of the acetabulum.
[0051] The specific key points can be adjusted according to the indicators to be evaluated. In this application, the key point identification method and which key points to identify can refer to the conventional identification method and will not be described in detail.
[0052] In this application, the specific adjustment of the acetabulum is to perform a resection operation on the iliac bone component and the osteotome component to obtain a pelvic and acetabulum bone block model after osteotomy, and then rotate the free affected acetabulum outward with the center of the affected femoral head as the rotation point to meet the correction of the lateral center edge angle (LCEA) of the acetabulum, the acetabulum top inclination angle (AIA) and the head and acetabulum coverage (HAI).
[0053] Among them, the direction of the acetabular opening is evaluated by the lateral center edge angle (LCEA) and the acetabular inclination angle (AIA) (in fact, the direction of the acetabular opening is also used to evaluate the coverage rate, but it is an indirect evaluation), and the location and degree of coverage defect are evaluated by the head and acetabulum coverage ratio (HAI).
[0054] The lateral center-edge angle (LCEA) is an angle measured on hip radiographs to assess the degree of coverage of the femoral head by the acetabulum. The LCEA is the angle between a vertical line drawn from the lateral rim of the acetabulum (i.e., the top of the acetabulum) and a line drawn from the center of the femoral head to the lateral rim of the acetabulum.
[0055] The normal value of LCEA is between 25° and 40°.
[0056] The acetabular index angle (AIA) is another angle measured on hip radiographs. It is used to assess the orientation of the acetabulum roof in the coronal plane and its coverage of the superior and lateral aspect of the femoral head. The weight-bearing surface of the acetabulum appears on radiographs as a sclerotic band resembling an "eyebrow arch." The AIA is the angle between a line drawn from the apex of the acetabulum to the ischial tuberosity (the acetabular index line) and a horizontal line drawn from the ischial tuberosity (the pelvic horizontal line).
[0057] Among them, the normal value of AIA is generally between 0° and 10°.
[0058] The Head Acetabulum Index (HAI) refers to the proportion of the femoral head covered by the acetabulum, indicating the degree of coverage of the femoral head by the acetabulum. Hip X-rays are used to measure the percentage of the femoral head covered by the acetabulum relative to the total area of the femoral head.
[0059] Among them, the normal value of HAI is generally above 70%.
[0060] Among them, based on the evaluation results and the standard acetabular inclination angle, the adjusted acetabular angle information and acetabular position information are determined; that is, the acetabular direction and rotation position of the acetabulum are adjusted so that the lateral center edge angle (LCEA), acetabular roof inclination angle (AIA) and head and acetabulum coverage (HAI) of the adjusted acetabulum are all within the normal range.
[0061] Among them, the standard acetabular inclination angle can be the standard angle of the lateral center edge angle (LCEA) of the acetabulum, the standard angle of the acetabular roof inclination angle (AIA), or the standard angle of other acetabular evaluation angles, such as: acetabular incidence angle (PI), acetabular anteversion angle (Acetabular Anteversion Angle) and acetabular inclination angle (Acetabular Inclination Angle).
[0062] Among them, during the specific adjustment process, on the basis of meeting the lateral center edge angle (LCEA), acetabular inclination angle (AIA) and head acetabular coverage (HAI), the smaller the rotation angle of the acetabulum is, the better, so as to avoid the generation of excessive gaps during acetabular osteotomy, which is not conducive to postoperative recovery.
[0063] The step S300, based on reinforcement learning and acetabulum adjustment information, trains the osteotomy position around the acetabulum, including:
[0064] Construct simulation environment and state space based on hip joint medical images;
[0065] Setting an action space, wherein the action space includes a plurality of osteotomy surfaces, each osteotomy surface having an ascending and descending action along a normal direction and a front-back and left-right rotation of the osteotomy surface;
[0066] Set up a reward function, build a deep Q network model and an experience replay pool, which includes the agent's actions, states, rewards, and next states;
[0067] Sampling from the experience replay pool and calculating the Q value based on the deep Q network model;
[0068] By minimizing the loss function, the weights of the deep Q network model and the weights of the target network are updated until the training is completed, and the deep Q network model and the target network remain consistent.
[0069] In this application, the specific training process includes:
[0070] Define the environment and state space
[0071] Environment: Defines a simulation environment that can provide the current state, receive actions, and return the new state, reward, and information about whether it ends.
[0072] State space: The state can be a CT or MRI image of the knee joint, which is preprocessed as input, and the four osteotomy planes of the image.
[0073] Defining the action space
[0074] Action: Action is defined as different manipulations to adjust the osteotomy position, such as moving the osteotomy position up, down, left, right, or rotating it a certain angle.
[0075] In this application, the action is the adjustment direction of the four osteotomy surfaces, wherein the adjustment of one osteotomy surface is used as an example for explanation:
[0076] For an osteotomy surface, a coordinate system is set with the plane as a plane, and the z-axis is the normal to the osteotomy surface; thus, the movement of the osteotomy surface is decomposed into: rising and falling movements along the z-axis (rising and falling movements along the normal direction), clockwise and counterclockwise rotation movements around the x-axis, and clockwise and counterclockwise rotation movements around the y-axis (rotation of the osteotomy surface forward and backward and left and right), a total of six movements, where the moving distance of each movement can be a unit distance, for example, rising and falling 1 mm, and rotating 1°.
[0077] Based on this, the action space in this application includes twenty-four actions, which are corresponding actions of the four osteotomy surfaces.
[0078] Designing the reward function
[0079] Reward function: The reward function should reflect the size of the resection area. For example, if the goal is to minimize the resection area, it can be designed to be a negative resection area. For example: R = -A, where A is the resection area after resection.
[0080] Building and training a deep Q-network
[0081] Model Architecture: A convolutional neural network (CNN) is used to process the input CT or MRI image, extract features, and output the Q value of each action. In this application, the input data is extracted from the state space and is a high-dimensional representation of the state. In this application, each state in the state space needs to be converted into the input data format of the deep convolutional model through preprocessing such as normalization, cropping, and resizing.
[0082] Experience Replay
[0083] Experience replay pool: stores the agent's experience, including state, action, reward, next state, and whether it ends.
[0084] Random sampling: Randomly sample a small batch from the experience replay pool for training to break data correlation and improve training efficiency and stability.
[0085] Target Network
[0086] Target network: Use an independent target network to calculate the target Q value, slowing down the change of the target and making the training more stable. Synchronous update: Every certain number of steps, the weights of the current Q network are copied to the target network.
[0087] Model training
[0088] Loss function: Use the loss function to calculate the difference between the actual Q value and the target Q value. You can also use the Adam optimizer to update the network weights.
[0089] During the specific training process,
[0090] Initialize the Q network and target network, setting their initial states. Execute actions in the environment, obtaining state transitions and rewards. Store experience in a replay pool. Sample mini-batches of experience from the replay pool and calculate the target Q value. Update the Q network's weights by minimizing the loss function. Periodically update the target network's weights.
[0091] In one embodiment, the loss function is:
[0092]
[0093] Among them, θ is the weight of the deep Q network model, L(θ) is the loss, N is the number of samples, and y i is the target Q value of the i-th sample, Q(s i ,a i ; θ) is the Q value of the i-th sample.
[0094] In one embodiment, the reward function is that the gap between the acetabulum side osteotomy surface and the retention side osteotomy surface after rotation is minimized, and the distance between the recess between the anterior superior iliac spine and the anterior inferior iliac spine and the corresponding osteotomy surface is less than 1 cm, and the distance between the recess in front of the greater sciatic notch and the corresponding osteotomy surface is greater than 1 cm, and the distance between the edge of the preset nerve sensitive area at the junction of the osteotomy surface and the outer surface of the hip joint is greater than the preset distance.
[0095] In this application, the above-mentioned reward function includes four parts, each of which has a different distribution weight, and the sum of the products based on the weights is used as the overall reward function.
[0096] In one embodiment, an initial state of the state space is set. In the initial state, the four osteotomy planes are: osteotomy plane 1, which is a plane 1 cm medial to the iliopectineal eminence and oblique to the midline of the body, for osteotomy of the pubic bone; osteotomy plane 2, which is defined as a plane from between the anterior superior and inferior iliac spines toward the greater sciatic notch, for osteotomy of the ilium at the upper edge of the acetabulum; osteotomy plane 3, which is defined as a plane from the obturator foramen in front of the ischial body at the lower edge of the acetabulum toward the ischial spine, for osteotomy of the ischium at the lower edge of the acetabulum; osteotomy plane 4, which is defined as about 1 cm away from the edge of the greater sciatic notch, connects osteotomy planes 2 and 3, and forms an angle of about 25° with the standard coronal plane, for osteotomy of the posterior pelvic column at the posterior edge of the acetabulum.
[0097] In the present application, the osteotomy surface is set to constrain the initial position of the osteotomy, thereby avoiding large deviations in the osteotomy position.
[0098] In one embodiment, the convergence standard, Q-value convergence, or a predetermined number of training rounds may be used as the training completion standard.
[0099] Among them, in actual application after the training, it is preferred to use the above four osteotomy surfaces as the initial state and a predetermined number of iterative rounds as the completion standard, so as to make adjustments based on the four osteotomy surfaces. On the one hand, better osteotomy position adjustment can be achieved, and on the other hand, large deviations from the four osteotomy surfaces can be avoided during the adjustment process.
[0100] Among them, the average reward is stable: when the average reward of the agent in several training rounds (episodes) tends to be stable, the training can be considered complete.
[0101] Specific method: Plot and monitor the total reward of the agent in each round, calculate the average reward of the last N rounds, and observe whether it fluctuates within a certain range.
[0102] Among them, Q value convergence: when the Q value (predicted state-action value) no longer changes significantly during the training process, it means that the agent's strategy tends to be stable.
[0103] Specific method: Monitor the rate of change of the Q value, calculate its mean and variance, and observe whether it tends to converge.
[0104] In one embodiment, after the training is completed, S400, selecting the osteotomy position with the highest cumulative reward as the periacetabular osteotomy plan includes:
[0105] Obtain new medical images of the hip joint and information on adjustments to the acetabulum;
[0106] Load the trained deep Q network model in the simulation environment and gradually execute corresponding actions based on the deep Q network model until the task is completed;
[0107] The osteotomy position at the end of the task was selected as the periacetabular osteotomy plan.
[0108] After the training is completed, the trained convolutional model is used in subsequent tasks. The specific process is as follows:
[0109] Model loading and initialization: You need to load the trained model and make sure it is in inference mode.
[0110] Environment initialization: Initialize the same environment as during training to ensure that the agent can interact with the environment.
[0111] State preprocessing: Before using the model for inference, ensure that the input state undergoes the same preprocessing steps as during training. The initial state remains consistent with the initial state during training.
[0112] Action selection strategy: During inference, the agent selects an action based on its current state. In practical applications, a greedy policy is often used, which selects the action with the highest Q value. The Q value is obtained using a trained convolutional model.
[0113] Execute tasks: The agent performs tasks in the environment, uses the trained model to select actions, executes and updates the state until the task is completed or a termination condition is reached.
[0114] In the present application, the termination condition is preferably a preset number of executions. The task is completed after the number of executions reaches the preset number, thereby achieving a local optimum and avoiding excessive deviation of the osteotomy position from the initial state.
[0115] After the task is completed, the action sequence during the task is obtained. According to the initial state and the action sequence, the position of the osteotomy plane when the task is completed can be determined, which is the periacetabular osteotomy planning.
[0116] In one embodiment, combined Figure 3 As shown, the deep Q network model includes a low-level feature extraction structure and a high-level feature extraction structure set in parallel; after the feature map processed by the low-level feature extraction structure is added to the feature map processed by the high-level feature extraction structure, convolution processing and multiple fully connected layer processing are performed in sequence to obtain the corresponding Q value.
[0117] In this application, the low-level feature extraction structure is composed of multiple consecutive convolutional layers; through continuous convolution, the underlying features are extracted, that is, the basic structural information of the input data; and it is also possible to extract more abstract and complex features based on the gradual combination of the underlying features, including high-level semantic understanding of objects, shapes and patterns.
[0118] In this application, the high-level feature extraction structure includes a downward-sized convolution layer and an upward-sized convolution layer; through the downward-sized convolution, the spatial resolution of the feature map is reduced, and the receptive field is increased, thereby extracting high-level features.
[0119] In this application, the feature maps extracted in parallel by different structures are added and convolved, and then processed through multiple consecutive fully connected layers to output the Q value corresponding to each action in the action space.
[0120] In this application, continuous convolution: the convolution layer close to the input extracts low-level features such as edges and textures; the convolution layer close to the output extracts high-level features such as objects and patterns.
[0121] In this application, down-convolution is performed first and then up-convolution: the down-convolution stage (encoder) extracts high-level features because it integrates more global information; the up-convolution stage (decoder) is mainly used to restore spatial resolution and maintain the previously extracted high-level semantic information.
[0122] Preferably, in the present application, after processing by multiple fully connected layers, the Q value corresponding to each action is output through an activation function similar to a SoftMax operation.
[0123] In one embodiment, the activation function of the SoftMax-like operation is:
[0124]
[0125] Among them, SoftMax1(x) is the activation function, x i 、x j is the element in the input vector, i and j are the element numbers.
[0126] In this application, SoftMax is a mathematical function that is often used to convert a set of arbitrary real numbers into real numbers that represent a probability distribution. It is essentially a normalization function that can convert a set of arbitrary real values into probability values between [0,1]. Because SoftMax converts them to values between 0 and 1, they can be interpreted as probabilities. If one of the inputs is small or negative, SoftMax turns it into a small probability, and if the input is large, it turns it into a large probability, but it will always remain between 0 and 1.
[0127] However, for the standard SoftMax function, since the input is mapped between 0 and 1, and the sum of all output values is 1, this means that even if some input values are very small, they will have a non-zero output value after being processed by the SoftMax function. This will cause the noise to be amplified, resulting in the final output being affected by greater noise.
[0128] In this application, a 1 is added to the denominator of the SoftMax-like function; this change means that when the input value is very small, their output value can be closer to zero. This allows the corresponding output to tend to zero when there is no valuable information to add, thus greatly reducing unnecessary noise.
[0129] An embodiment of the present application provides a periacetabular osteotomy planning device based on reinforcement learning, which is used to execute the periacetabular osteotomy planning method based on reinforcement learning described above in the present application. The periacetabular osteotomy planning device based on reinforcement learning is described in detail below.
[0130] like Figure 4 As shown, the acetabular periarthritis osteotomy planning device based on reinforcement learning includes:
[0131] An image acquisition module 101 is used to acquire medical images of the hip joint;
[0132] an adjustment determination module 102, configured to determine adjustment information of the acetabulum based on the hip joint medical image;
[0133] A reinforcement learning module 103 is used to train the osteotomy position around the acetabulum based on reinforcement learning and acetabulum adjustment information;
[0134] The osteotomy planning module 104 is used to select the osteotomy position with the highest cumulative reward as the periacetabular osteotomy planning after training is completed; the reinforcement learning is deep Q network learning.
[0135] In one embodiment, the adjustment determination module 102 is further configured to:
[0136] Perform image segmentation and three-dimensional reconstruction on the medical image of the hip joint to obtain a three-dimensional model of the hip joint; identify key points of the three-dimensional model of the hip joint; analyze the three-dimensional model of the hip joint to evaluate the direction of the acetabulum opening, the location and degree of coverage of the defect; and determine the adjusted acetabulum angle information and acetabulum position information based on the evaluation results and the standard acetabulum inclination angle.
[0137] In one embodiment, the reinforcement learning module 103 is further configured to:
[0138] A simulation environment and state space are constructed based on medical images of the hip joint; an action space is set, wherein the action space contains multiple osteotomy surfaces, each of which has an ascending and descending action along the normal direction and a forward, backward, left, and right rotation of the osteotomy surface; a reward function is set, and a deep Q network model and an experience replay pool are constructed, wherein the experience replay pool includes the action, state, reward, and next state of the intelligent agent; sampling is performed from the experience replay pool, and the Q value is calculated based on the deep Q network model; by minimizing the loss function, the weights of the deep Q network model and the weights of the target network are updated until the training is completed and the deep Q network model and the target network remain consistent.
[0139] In one embodiment, the deep Q network model includes a low-level feature extraction structure and a high-level feature extraction structure arranged in parallel; after the feature map obtained by processing the low-level feature extraction structure is added to the feature map obtained by processing the high-level feature extraction structure, convolution processing and multiple fully connected layer processing are performed in sequence to obtain the corresponding Q value.
[0140] In one embodiment, the loss function is:
[0141]
[0142] Among them, θ is the weight of the deep Q network model, L(θ) is the loss, N is the number of samples, and y i is the target Q value of the i-th sample, Q(s i ,a i ; θ) is the Q value of the i-th sample.
[0143] In one embodiment, the reward function is that the gap between the acetabulum side osteotomy surface and the retention side osteotomy surface after rotation is minimized, and the distance between the recess between the anterior superior iliac spine and the anterior inferior iliac spine and the corresponding osteotomy surface is less than 1 cm, and the distance between the recess in front of the greater sciatic notch and the corresponding osteotomy surface is greater than 1 cm, and the distance between the edge of the preset nerve sensitive area at the junction of the osteotomy surface and the outer surface of the hip joint is greater than the preset distance.
[0144] In one embodiment, the osteotomy planning module 104 is further configured to:
[0145] Acquire new medical images of the hip joint and acetabulum adjustment information; load the trained deep Q-network model in a simulation environment, and gradually execute corresponding actions based on the deep Q-network model until the task is completed; select the osteotomy position at the end of the task as the periacetabular osteotomy plan.
[0146] The reinforcement learning-based periacetabular osteotomy planning device provided in the above-mentioned embodiment of the present application has a corresponding relationship with the reinforcement learning-based periacetabular osteotomy planning method provided in the embodiment of the present application. Therefore, the specific content in the device has a corresponding relationship with the periacetabular osteotomy planning method. The specific content can refer to the records in the periacetabular osteotomy planning method, and will not be repeated in this application.
[0147] The reinforcement learning-based periacetabular osteotomy planning device provided in the above-mentioned embodiment of the present application and the reinforcement learning-based periacetabular osteotomy planning method provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.
[0148] The above describes the internal functions and structure of the acetabular osteotomy planning device based on reinforcement learning, such as Figure 5 As shown, in practice, the acetabular peri-osteotomy planning device based on reinforcement learning can be implemented as an electronic device, including: a memory 301 and a processor 303 .
[0149] The memory 301 may be configured to store programs.
[0150] In addition, the memory 301 may also be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application or method operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.
[0151] The memory 301 can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0152] The processor 303 is coupled to the memory 301 and is configured to execute a program in the memory 301 to:
[0153] Obtain medical images of the hip joint;
[0154] determining adjustment information of the acetabulum based on the hip joint medical image;
[0155] Based on reinforcement learning and acetabular adjustment information, the osteotomy position around the acetabulum is trained;
[0156] After the training is completed, the osteotomy position with the highest cumulative reward is selected as the periacetabular osteotomy plan;
[0157] The reinforcement learning is deep Q network learning.
[0158] In one embodiment, the processor 303 is further configured to:
[0159] Perform image segmentation and three-dimensional reconstruction on the medical image of the hip joint to obtain a three-dimensional model of the hip joint; identify key points of the three-dimensional model of the hip joint; analyze the three-dimensional model of the hip joint to evaluate the direction of the acetabulum opening, the location and degree of coverage of the defect; and determine the adjusted acetabulum angle information and acetabulum position information based on the evaluation results and the standard acetabulum inclination angle.
[0160] In one embodiment, the processor 303 is further configured to:
[0161] A simulation environment and state space are constructed based on medical images of the hip joint; an action space is set, wherein the action space contains multiple osteotomy surfaces, each of which has an ascending and descending action along the normal direction and a forward, backward, left, and right rotation of the osteotomy surface; a reward function is set, and a deep Q network model and an experience replay pool are constructed, wherein the experience replay pool includes the action, state, reward, and next state of the intelligent agent; sampling is performed from the experience replay pool, and the Q value is calculated based on the deep Q network model; by minimizing the loss function, the weights of the deep Q network model and the weights of the target network are updated until the training is completed and the deep Q network model and the target network remain consistent.
[0162] In one embodiment, the deep Q network model includes a low-level feature extraction structure and a high-level feature extraction structure arranged in parallel; after the feature map obtained by processing the low-level feature extraction structure is added to the feature map obtained by processing the high-level feature extraction structure, convolution processing and multiple fully connected layer processing are performed in sequence to obtain the corresponding Q value.
[0163] In one embodiment, the loss function is:
[0164]
[0165] Among them, θ is the weight of the deep Q network model, L(θ) is the loss, N is the number of samples, and y i is the target Q value of the i-th sample, Q(s i ,a i ; θ) is the Q value of the i-th sample.
[0166] In one embodiment, the reward function is that the gap between the acetabulum side osteotomy surface and the retention side osteotomy surface after rotation is minimized, and the distance between the recess between the anterior superior iliac spine and the anterior inferior iliac spine and the corresponding osteotomy surface is less than 1 cm, and the distance between the recess in front of the greater sciatic notch and the corresponding osteotomy surface is greater than 1 cm, and the distance between the edge of the preset nerve sensitive area at the junction of the osteotomy surface and the outer surface of the hip joint is greater than the preset distance.
[0167] In one embodiment, the processor 303 is further configured to:
[0168] Acquire new medical images of the hip joint and acetabulum adjustment information; load the trained deep Q-network model in a simulation environment, and gradually execute corresponding actions based on the deep Q-network model until the task is completed; select the osteotomy position at the end of the task as the periacetabular osteotomy plan.
[0169] In this application, the processor is also specifically used to execute all processes and steps of the above-mentioned reinforcement learning-based periacetabular osteotomy planning method. The specific content can be referred to the records in the periacetabular osteotomy planning method, and will not be repeated in this application.
[0170] In this application, Figure 5 Only some components are shown schematically, which does not mean that the electronic device only includes Figure 5 Components shown.
[0171] The electronic device provided in this embodiment is based on the same inventive concept as the acetabular peri-osteotomy planning method based on reinforcement learning provided in the embodiment of the present application, and has the same beneficial effects as the method adopted, run or implemented by the application stored therein.
[0172] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0173] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0174] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0175] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0176] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0177] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0178] The present application also provides a computer-readable storage medium corresponding to the reinforcement learning-based periacetabular osteotomy planning method provided in the aforementioned embodiment, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the reinforcement learning-based periacetabular osteotomy planning method provided in any of the aforementioned embodiments.
[0179] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0180] The computer-readable storage medium provided in the above-mentioned embodiment of the present application and the acetabular peri-osteotomy planning method based on reinforcement learning provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by the application program stored therein.
[0181] It should be noted that, in the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known structures and technologies are not shown in detail so as not to obscure the understanding of this description.
[0182] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0183] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A method for planning periacetabular osteotomy based on reinforcement learning, characterized in that: include: Obtain medical images of the hip joint; determining adjustment information of the acetabulum based on the hip joint medical image; Based on reinforcement learning and acetabular adjustment information, the osteotomy position around the acetabulum is trained; After the training is completed, the osteotomy position with the highest cumulative reward is selected as the periacetabular osteotomy plan; The reinforcement learning is deep Q network learning; The training of the periacetabular osteotomy position based on reinforcement learning and acetabulum adjustment information includes: Construct simulation environment and state space based on hip joint medical images; Setting an action space, wherein the action space includes a plurality of osteotomy surfaces, each osteotomy surface having an ascending and descending action along a normal direction and a front-back and left-right rotation of the osteotomy surface; Set up a reward function, build a deep Q network model and an experience replay pool, which includes the agent's actions, states, rewards, and next states; Sampling from the experience replay pool and calculating the Q value based on the deep Q network model; By minimizing the loss function, the weights of the deep Q network model and the weights of the target network are updated until the training is completed, and the deep Q network model and the target network remain consistent.
2. The acetabular osteotomy planning method based on reinforcement learning according to claim 1, characterized in that: Determining adjustment information of the acetabulum based on the hip joint medical image includes: Performing image segmentation and three-dimensional reconstruction on the medical image of the hip joint to obtain a three-dimensional model of the hip joint; Identifying key points of the three-dimensional model of the hip joint; Analyze the three-dimensional model of the hip joint to evaluate the direction of the acetabulum opening, the location and extent of coverage of the defect; Based on the evaluation results and the standard acetabular inclination angle, adjusted acetabular angle information and acetabular position information are determined.
3. The acetabular osteotomy planning method based on reinforcement learning according to claim 1, characterized in that: The deep Q network model includes a low-level feature extraction structure and a high-level feature extraction structure set in parallel; after the feature map processed by the low-level feature extraction structure is added to the feature map processed by the high-level feature extraction structure, convolution processing and multiple fully connected layer processing are performed in sequence to obtain the corresponding Q value.
4. The acetabular osteotomy planning method based on reinforcement learning according to claim 1, characterized in that: The loss function is: Among them, θ is the weight of the deep Q network model, L(θ) is the loss, N is the number of samples, and y i is the target Q value of the i-th sample, Q(s i ,a i ; θ) is the Q value of the i-th sample.
5. The acetabular osteotomy planning method based on reinforcement learning according to claim 1, characterized in that: The reward function is that the gap between the acetabulum side osteotomy surface and the retention side osteotomy surface after rotation is minimized, and the distance between the recess between the anterior superior iliac spine and the anterior inferior iliac spine and the corresponding osteotomy surface is less than 1 cm, and the distance between the recess in front of the greater sciatic notch and the corresponding osteotomy surface is greater than 1 cm, and the distance between the edge of the preset nerve sensitive area at the junction of the osteotomy surface and the outer surface of the hip joint is greater than the preset distance.
6. The acetabular periarthritis osteotomy planning method based on reinforcement learning according to any one of claims 1-2, characterized in that: After the training is completed, the osteotomy position with the highest cumulative reward is selected as the periacetabular osteotomy plan, including: Obtain new medical images of the hip joint and information on adjustments to the acetabulum; Load the trained deep Q network model in the simulation environment and gradually execute corresponding actions based on the deep Q network model until the task is completed; The osteotomy position at the end of the task was selected as the periacetabular osteotomy plan.
7. A periacetabular osteotomy planning device based on reinforcement learning, characterized in that: include: An image acquisition module, which is used to acquire medical images of the hip joint; an adjustment determination module, configured to determine adjustment information of the acetabulum based on the hip joint medical image; A reinforcement learning module, which is used to train periacetabular osteotomy positions based on reinforcement learning and acetabular adjustment information; An osteotomy planning module, which is used to select the osteotomy position with the highest cumulative reward after training is completed as the periacetabular osteotomy plan; The reinforcement learning is deep Q network learning; The training of the periacetabular osteotomy position based on reinforcement learning and acetabulum adjustment information includes: Construct simulation environment and state space based on hip joint medical images; Setting an action space, wherein the action space includes a plurality of osteotomy surfaces, each osteotomy surface having an ascending and descending action along a normal direction and a front-back and left-right rotation of the osteotomy surface; Set up a reward function, build a deep Q network model and an experience replay pool, which includes the agent's actions, states, rewards, and next states; Sampling from the experience replay pool and calculating the Q value based on the deep Q network model; By minimizing the loss function, the weights of the deep Q network model and the weights of the target network are updated until the training is completed, and the deep Q network model and the target network remain consistent.
8. An electronic device, characterized in that: include: memory and processor; The memory is used to store programs; The processor, coupled to the memory, is configured to execute the program to: Obtain medical images of the hip joint; determining adjustment information of the acetabulum based on the hip joint medical image; Based on reinforcement learning and acetabular adjustment information, the osteotomy position around the acetabulum is trained; After the training is completed, the osteotomy position with the highest cumulative reward is selected as the periacetabular osteotomy plan; The reinforcement learning is deep Q network learning; The training of the periacetabular osteotomy position based on reinforcement learning and acetabulum adjustment information includes: Construct simulation environment and state space based on hip joint medical images; Setting an action space, wherein the action space includes a plurality of osteotomy surfaces, each osteotomy surface having an ascending and descending action along a normal direction and a front-back and left-right rotation of the osteotomy surface; Set up a reward function, build a deep Q network model and an experience replay pool, which includes the agent's actions, states, rewards, and next states; Sampling from the experience replay pool and calculating the Q value based on the deep Q network model; By minimizing the loss function, the weights of the deep Q network model and the weights of the target network are updated until the training is completed, and the deep Q network model and the target network remain consistent.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the acetabular peri-osteotomy planning method based on reinforcement learning according to any one of claims 1 to 6.
Citation Information
Patent Citations
Total hip replacement preoperative planning method and device based on deep learning and X rays
CN111888059A
Osteotomy adjustment planning method based on visual image, electronic equipment and medium
CN114587585A
Improved n-step TD error priority experience playback training method
CN118114745A