A virtual reality-based arthroscopic hip surgery assisted teaching system

Through the neural network-based image segmentation model and 3D rendering technology, the registration position of hip arthroscopy surgery can be determined quickly and accurately, solving the problems of inaccurate position and low efficiency in traditional teaching systems and improving teaching effects.

CN116959307BActive Publication Date: 2025-10-17CHINA JAPAN FRIENDSHIP HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311112508.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-31
Publication Date
2025-10-17
Estimated Expiration
2043-08-31

AI Technical Summary

Technical Problem

In traditional virtual reality-based hip arthroscopy surgery assisted teaching systems, teaching doctors rely on their own experience to determine the registration position, which is inaccurate and inefficient, resulting in poor teaching results.

Method used

A neural network-based image segmentation model is used to extract features through a cascaded residual block structure. The proportional branch, integral branch, and differential branch are combined to quickly and accurately determine the registration position of the hip joint and perform three-dimensional model rendering and display.

Benefits of technology

It can quickly and accurately determine the registration position, improve the teaching effect, and provide intuitive three-dimensional space virtual display to assist teaching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116959307B_ABST
    Figure CN116959307B_ABST
Patent Text Reader

Abstract

The application provides a virtual reality-based hip arthroscopy surgery auxiliary teaching system. The virtual reality-based hip arthroscopy surgery auxiliary teaching system comprises: a first image segmentation module configured to input a hip CT image into a preset image segmentation model and output a hip bone image; a second image segmentation module configured to input a hip MRI image into the preset image segmentation model and output a hip part image; a registration fusion reconstruction module configured to determine a registration position based on first pelvis position information, first femur position information, second pelvis position information, second femur position information and soft tissue position information, perform registration, and fuse and reconstruct to obtain a hip part multi-modal information image; and a rendering display module configured to perform three-dimensional model rendering based on the hip part multi-modal information image, and perform three-dimensional space virtual display auxiliary teaching. According to the embodiment of the application, the registration position can be determined quickly and accurately, and the teaching effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of deep learning intelligent identification, and particularly relates to a hip arthroscopy surgery auxiliary teaching system and method based on virtual reality, an electronic device and a computer readable storage medium. BACKGROUND

[0002] Hip arthroscopy is considered a new minimally invasive surgical technique with great potential clinical value. Traditional virtual reality-based hip arthroscopy surgery auxiliary teaching systems still rely on teaching doctors to determine the registration position using medical instruments based on their own experience, and perform three-dimensional reconstruction virtual display to assist teaching. However, the registration position determined by the teaching doctors using medical instruments based on their own experience is inaccurate and inefficient, which leads to poor teaching effect.

[0003] Therefore, how to quickly and accurately determine the registration position and improve the teaching effect is a technical problem that needs to be solved by those skilled in the art. SUMMARY

[0004] The application provides a virtual reality-based hip arthroscopy surgery auxiliary teaching system and method, an electronic device and a computer readable storage medium, which can quickly and accurately determine the registration position and improve the teaching effect.

[0005] In a first aspect, the application provides a virtual reality-based hip arthroscopy surgery auxiliary teaching system, which comprises:

[0006] A first image segmentation module is configured to input a hip joint CT image into a preset image segmentation model and output a hip joint bone image, wherein the hip joint bone image comprises first pelvic position information and first femur position information;

[0007] A second image segmentation module is configured to input a hip joint MRI image into a preset image segmentation model and output a hip joint part image, wherein the hip joint part image comprises second pelvic position information, second femur position information and soft tissue position information;

[0008] A registration fusion reconstruction module is configured to determine a registration position based on the first pelvic position information, the first femur position information, the second pelvic position information, the second femur position information and the soft tissue position information, and perform registration and fusion reconstruction to obtain a hip joint part multi-modal information image;

[0009] A rendering display module is configured to perform three-dimensional model rendering based on the hip joint part multi-modal information image, and perform three-dimensional space virtual display to assist teaching;

[0010] The preset image segmentation model is obtained by model training based on a neural network, a structure of the neural network uses a cascaded residual block as a backbone network, and different depth and width networks are used to extract features; the structure of the neural network comprises: three convolutional layers and pooling layers, which are used for downsampling operation, reducing image size, reducing calculation amount, and accelerating model inference speed; three network branches connected respectively: a scale branch, which is used for being responsible for analyzing and retaining detailed information in a high-resolution feature map; an integral branch, which is used for being responsible for aggregating local and global context information to capture long-distance dependence; and a differential branch, which is used for being responsible for extracting high-frequency features to predict a boundary region.

[0011] Optionally, the feature map obtained after the scale branch passes through the convolutional layer and the feature map obtained after the differential branch passes through the convolutional layer are input into a pixel attention guiding module and then output after passing through a convolutional layer, and a total of three convolutional layers are passed through;

[0012] The integral branch outputs a feature map after passing through three convolutional layers and pooling layers; the feature map output by the integral branch is reduced in size because it is output after passing through three pooling layers, and is then restored to the original size by passing through a pyramid pooling module;

[0013] The feature map obtained after the differential branch passes through the convolutional layer and the feature map obtained after the scale branch passes through the convolutional layer are combined and output to a next convolutional layer, and a total of three convolutional layers are passed through to output a feature map;

[0014] The output results of the three branches are collectively input into a boundary attention guiding module, and a feature map output after passing through a convolutional layer, a BN layer and a RELU activation function is output as a predicted mask.

[0015] Optionally, the method further comprises:

[0016] The loss function calculation module is configured to: perform boundary detection on a real mask to obtain a boundary mask; calculate a loss function Loss1 based on the boundary mask and the output result of the differential branch; calculate a loss function Loss2 based on the predicted mask and the real mask; calculate a loss function Loss3 based on the predicted mask, the real mask and the output result of the differential branch; calculate a loss function Loss4 based on the output result of the scale branch and the real mask; and comprehensively calculate a final loss function based on the loss functions Loss1, Loss2, Loss3 and Loss4.

[0017] Optionally, the method further comprises:

[0018] The pixel attention guiding module is configured to interactively enhance the feature maps of the scale branch and the differential branch by using an attention mechanism;

[0019] Two inputs of the pixel attention guiding module are an output of the proportional branch and an output of the differential branch; the output of the proportional branch is subjected to a convolution layer (3x3), a BN layer and a RELU activation function to obtain feature maps T1 and T2;

[0020] The output of the differential branch is subjected to a convolution layer (3x3), a BN layer and a RELU activation function to obtain feature maps T4 and T5;

[0021] The feature map obtained after the convolution layer is applied to the feature map T1, and the feature map obtained after the convolution layer is applied to the feature map T4 to obtain a feature map T3;

[0022] The feature map obtained after the convolution layer is applied to the feature map T2, and the feature map T3 are fused to obtain a feature map T6;

[0023] The feature map obtained after the convolution layer is applied to the feature map T5, and the feature map T3 are fused to obtain a feature map T7;

[0024] The feature map T7 and the feature map T6 are merged to obtain a feature map T8, and the output 1 is obtained after a RELU activation function;

[0025] The output of the differential branch is subjected to a convolution layer (3x3), a BN layer and a RELU activation function to obtain an output 2.

[0026] Optionally, the pixel attention guiding module further comprises:

[0027] The pyramid pooling module is used to aggregate the context information of different regions to improve the ability of the network to obtain global information;

[0028] The pyramid pooling module is used to: use different scales of pooling on the original feature map to obtain a plurality of feature maps of different sizes, then splice the feature maps in the channel dimension, and then splice the original feature map, and finally output a composite feature map that integrates multiple scales, so as to achieve the purpose of balancing global semantic information and local detail information.

[0029] Optionally, the pixel attention guiding module further comprises:

[0030] The pyramid pooling module is used to: perform different scale pooling operations on the original feature map to obtain a plurality of feature maps of different sizes (using 5 branches); perform an upsampling operation on the obtained feature maps to restore to the size of the original feature map (6x6), and finally splice in the channel dimension to obtain the final composite feature map;

[0031] The first branch: using (6x6) pooling, the output size is (1x1), and then up-sampling to (6x6) through bilinear interpolation;

[0032] The second branch: using (3x3) pooling, the output size is (2x2), and then up-sampling to (6x6) by bilinear interpolation;

[0033] The third branch: using (2x2) pooling, the output size is (3x3), and then up-sampling to (6x6) by bilinear interpolation;

[0034] The fourth branch: using (1x1) pooling, the output size is (6x6);

[0035] The fifth branch: representing the input original feature map to play a role of residual connection;

[0036] The output results of the first four branches are spliced and then spliced with the fifth branch, and then the output result is output.

[0037] Optionally, the method further comprises:

[0038] The boundary attention guiding module is configured to: obtain two branch outputs after the differential branch input is subjected to a Sigmoid activation function, wherein one branch output feature is fused with the feature of the differential branch input, and then merged with the feature of the proportion branch input, and the other branch output feature is fused with the feature of the proportion branch input, and then merged with the feature of the differential branch; after the two merged features are respectively subjected to a convolution layer (3x3 convolution kernel) and a BN layer, an output is obtained, and finally the features are merged.

[0039] In a second aspect, an embodiment of the present application provides a hip arthroscopy surgery auxiliary teaching method based on virtual reality, comprising:

[0040] Inputting the hip joint CT image into a preset image segmentation model to output a hip joint bone image; wherein the hip joint bone image comprises first pelvic position information and first femur position information;

[0041] Inputting the hip joint MRI image into the preset image segmentation model to output a hip joint part image; wherein the hip joint part image comprises second pelvic position information, second femur position information, and soft tissue position information;

[0042] Determine a registration position based on the first pelvic position information, the first femur position information, the second pelvic position information, the second femur position information, and the soft tissue position information, and perform registration to fuse and reconstruct a hip joint part multi-modal information image;

[0043] Based on the hip joint part multi-modal information image, perform three-dimensional model rendering to perform three-dimensional space virtual display auxiliary teaching;

[0044] The preset image segmentation model is obtained by model training based on a neural network, a structure of the neural network uses a cascaded residual block as a backbone network, and different depth and width networks are used to extract features; the structure of the neural network comprises three convolutional layers and pooling layers, which are used for downsampling operation, reducing image size, reducing calculation amount, and accelerating model inference speed; three network branches connected respectively: a scale branch, which is responsible for analyzing and retaining detailed information in a high-resolution feature map; an integral branch, which is responsible for aggregating local and global context information to capture long-distance dependencies; and a differential branch, which is responsible for extracting high-frequency features to predict boundary regions.

[0045] In a third aspect, an electronic device is provided, and the electronic device comprises a processor and a memory storing computer program instructions.

[0046] The processor, when executing the computer program instructions, implements the virtual reality-based hip arthroscopy surgery auxiliary teaching method of the second aspect.

[0047] In a fourth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer program instructions, and the computer program instructions, when executed by a processor, implement the virtual reality-based hip arthroscopy surgery auxiliary teaching method of the second aspect.

[0048] The virtual reality-based hip arthroscopy surgery auxiliary teaching system and method, the electronic device, and the computer-readable storage medium can quickly and accurately determine the registration position, thereby improving the teaching effect.

[0049] The virtual reality-based hip arthroscopy surgery auxiliary teaching system comprises: a first image segmentation module configured to input a hip joint CT image into a preset image segmentation model and output a hip joint bone image; wherein the hip joint bone image comprises first pelvic position information and first femur position information; a second image segmentation module configured to input a hip joint MRI image into the preset image segmentation model and output a hip joint part image; wherein the hip joint part image comprises second pelvic position information, second femur position information and soft tissue position information; a registration fusion reconstruction module configured to determine a registration position based on the first pelvic position information, the first femur position information, the second pelvic position information, the second femur position information and the soft tissue position information, perform registration, and fuse and reconstruct to obtain a hip joint part multi-modal information image; and a rendering display module configured to perform three-dimensional model rendering based on the hip joint part multi-modal information image, and perform three-dimensional space virtual display auxiliary teaching; wherein the preset image segmentation model is obtained by model training based on a neural network, a structure of the neural network uses a cascaded residual block as a backbone network, and different depth and width networks are used to extract features; the structure of the neural network comprises: three convolutional layers and pooling layers configured to perform downsampling operation, reduce image size, reduce calculation amount, and accelerate model inference speed; and three network branches connected respectively: a scale branch configured to be responsible for analyzing and retaining detailed information in a high-resolution feature map; an integral branch configured to be responsible for aggregating local and global context information to capture long-distance dependencies; and a differential branch configured to be responsible for extracting high-frequency features to predict boundary regions.

[0050] The virtual reality-based hip arthroscopy surgery auxiliary teaching system, method, electronic device and computer readable storage medium provided by the embodiments of the present application can quickly and accurately determine a registration position, thereby improving teaching effect. BRIEF DESCRIPTION OF DRAWINGS

[0051] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the drawings needed in the specific embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0052] Figure 1 is a structural schematic diagram of a virtual reality-based hip arthroscopy surgery auxiliary teaching system provided by an embodiment of the present application;

[0053] Figure 2 is a framework schematic diagram of a virtual reality-based hip arthroscopy surgery auxiliary teaching method provided by an embodiment of the present application;

[0054] Figure 3is a structural schematic diagram of an image segmentation model provided by an embodiment of the present application;

[0055] Figure 4 is a structural schematic diagram of a pixel attention guiding module provided by an embodiment of the present application;

[0056] Figure 5 is a structural schematic diagram of a pyramid pooling module provided by an embodiment of the present application;

[0057] Figure 6 is a structural schematic diagram of a boundary attention guiding module provided by an embodiment of the present application;

[0058] Figure 7 is a flow schematic diagram of a virtual reality-based hip arthroscopy surgery auxiliary teaching method provided by an embodiment of the present application;

[0059] Figure 8 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0060] The features and exemplary embodiments of various aspects of the present application will be described below in detail, in order to make the purposes, technical solutions and advantages of the present application more clear and apparent, the present application will be further described in detail below in combination with the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, but not to limit the present application. The present application can be implemented without some of these specific details by those skilled in the art. The following description of the embodiments is only to provide a better understanding of the present application by showing examples of the present application.

[0061] It should be noted that in this paper, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the elements defined by the statement "include" do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0062] Hip arthroscopy is considered a new minimally invasive surgical technique with great potential clinical value. Traditional virtual reality-based hip arthroscopy teaching systems rely on the teaching physician's experience to determine registration positions using medical instruments and then perform 3D reconstruction and virtual display to assist in teaching. However, this experience-based registration is inaccurate and inefficient, leading to poor teaching outcomes.

[0063] To solve the problems of the prior art, the present invention provides a virtual reality-based hip arthroscopy surgery assisted teaching system, method, electronic device, and computer-readable storage medium. The virtual reality-based hip arthroscopy surgery assisted teaching system provided by the present invention is first introduced below.

[0064] Figure 1 FIG1 shows a schematic diagram of the structure of a hip arthroscopy surgery auxiliary teaching system based on virtual reality provided by an embodiment of the present application. Figure 1 As shown, the virtual reality-based hip arthroscopy surgery auxiliary teaching system includes:

[0065] The first image segmentation module 101 is configured to input the hip joint CT image into a preset image segmentation model and output a hip joint bone image; wherein the hip joint bone image includes first pelvic position information and first femoral position information;

[0066] The second image segmentation module 102 is configured to input the hip joint MRI image into a preset image segmentation model and output a hip joint part image; wherein the hip joint part image includes second pelvic position information, second femoral position information, and soft tissue position information;

[0067] A registration fusion reconstruction module 103 is configured to determine a registration position and perform registration based on the first pelvis position information, the first femur position information, the second pelvis position information, the second femur position information, and the soft tissue position information, and to fuse and reconstruct a multimodal information image of the hip joint;

[0068] The rendering and display module 104 is used to render a three-dimensional model based on the multimodal information image of the hip joint to perform three-dimensional space virtual display-assisted teaching;

[0069] The preset image segmentation model is obtained by model training based on a neural network, a structure of the neural network uses a cascaded residual block as a backbone network, and different depth and width networks are used to extract features; the structure of the neural network comprises: three convolutional layers and pooling layers, which are used for downsampling operation, reducing image size, reducing calculation amount, and accelerating model inference speed; three network branches connected respectively: a scale branch, which is used for being responsible for analyzing and retaining detailed information in a high-resolution feature map; an integral branch, which is used for being responsible for aggregating local and global context information to capture long-distance dependence; and a differential branch, which is used for being responsible for extracting high-frequency features to predict a boundary region.

[0070] Specifically, the image segmentation model is used for segmenting the CT and MRI images of the patient's hip joint, and the segmented CT image is used to obtain a hip joint bone image (including a pelvis and a femur), and the segmented MRI image is used to obtain a hip joint part image (including the pelvis, the femur, and soft tissue). Figure 1 A corresponding framework schematic diagram of the virtual reality-based hip arthroscopy surgery auxiliary teaching method is shown in FIG. 1. Figure 2 As shown in FIG. 1, the CT and MRI data of the patient's hip joint are input; the CT and MRI images of the patient are segmented by a neural network, the CT image is segmented to obtain a hip joint bone image (including a pelvis and a femur), and the MRI image is segmented to obtain a hip joint part image (including the pelvis, the femur, and soft tissue). The hip joint bone image obtained by the CT image segmentation and the hip joint bone image obtained by the MRI image segmentation are registered and fused to reconstruct a hip joint part multi-modal information image. With the aid of a VR head-mounted device, a three-dimensional model of the hip joint is imported and rendered on the VR head-mounted display device, providing an intuitive, three-dimensional, and realistic three-dimensional space virtual display for simulating surgery operation planning and reference.

[0071] In one embodiment, the feature map obtained after the scale branch passes through the convolutional layer and the feature map obtained after the differential branch passes through the convolutional layer are input into a pixel attention guiding module and then output after passing through a convolutional layer, and a total of three convolutional layers are passed through;

[0072] The feature map output by the integral branch is reduced in size after passing through three pooling layers, and is restored to the original size after passing through a pyramid pooling module;

[0073] The feature map obtained after the differential branch passes through the convolutional layer and the feature map obtained after the scale branch passes through the convolutional layer are combined and output to the next convolutional layer, and a total of three convolutional layers are passed through to output the feature map;

[0074] The output results of the three branches are collectively input into a boundary attention guiding module, and the output feature map is output after passing through a convolutional layer, a BN layer, and a RELU activation function.

[0075] Specifically, a structure schematic diagram of the image segmentation model is shown in FIG. 2. Figure 3As shown, in order to improve the segmentation accuracy, we adopt a cascade network structure, first adopt three convolutional layers and pooling layers for down-sampling operation, the purpose is to reduce the image size and reduce the amount of calculation, speed up the model inference speed.

[0076] After that, three branches are used in the network: the proportion branch (P): responsible for analyzing and preserving detailed information in high-resolution feature maps; the integral branch (I): responsible for aggregating local and global context information to capture long-range dependencies; the differential branch (D): responsible for extracting high-frequency features to predict boundary regions. The entire model uses a cascade residual block as the backbone network, and uses networks with different depths and widths to extract features.

[0077] The network structure inputs three-dimensional CT data (MRI data), and the features obtained after three convolutional layers and pooling layers are input into three branches. The features obtained after the convolutional layer of the proportion branch and the features obtained after the convolutional layer of the differential branch are input into the pixel attention guiding module, and then output after the convolutional layer, a total of three convolutional layers; the integral branch is output after three convolutional layers and pooling layers; the features obtained after the convolutional layer of the differential branch and the features obtained after the convolutional layer of the proportion branch are combined and output to the next convolutional layer, a total of three convolutional layers are output. The features output by the integral branch are reduced in size after being output by three pooling layers, and are restored to the original size by the pyramid pooling module. The output results of the three branches are jointly input into the boundary attention guiding module, and the output features are output after the convolutional layer, BN layer and RELU activation function. The prediction result mask is output.

[0078] In one embodiment, further comprising:

[0079] The loss function calculation module is configured to: perform boundary detection on the real mask to obtain a boundary mask; calculate a loss function Loss1 based on the boundary mask and the output result of the differential branch; calculate a loss function Loss2 based on the prediction mask and the real mask; calculate a loss function Loss3 based on the prediction mask, the real mask and the output result of the differential branch; calculate a loss function Loss4 based on the output result of the proportion branch and the real mask; and calculate a final loss function based on the loss functions Loss1, Loss2, Loss3 and Loss4.

[0080] Specifically, in order to enhance the segmentation accuracy, a multi-loss fusion form is adopted, in order to enhance the prediction ability of the boundary, a boundary mask is obtained by detecting the boundary of the real mask, and the loss function Loss1 is calculated with the result output by the differential branch, the loss function Loss2 is calculated with the predicted mask and the real mask, the loss function Loss3 is calculated with the predicted mask, the real mask and the result output by the differential branch, the loss function Loss4 is calculated with the proportion branch output result and the real mask.

[0081] Finally, the loss functions Loss1, LossL2, Loss3 and Loss4 are comprehensively calculated to obtain the final loss function.

[0082] The output of the differential branch is:

[0083]

[0084] Wherein, k mn The mth value of the convolution kernel in the mth layer. In the integral branch, I[i-1], I[i] and I[i+1] are set to more than 70% of the total number of items, the purpose is to pay more attention to local information. In the proportion branch and the differential branch, I[i-1], I[i] and I[i+1] are set to less than 30% of the total number of items, the purpose is to make the two branches pay more attention to the surrounding information.

[0085] In one embodiment, it further comprises:

[0086] The pixel attention guiding module is used to interact and enhance the feature maps of the proportion branch and the differential branch by using the attention mechanism;

[0087] The two inputs of the pixel attention guiding module are the output of the proportion branch and the output of the differential branch; the output of the proportion branch is obtained after the convolution layer (3x3), the BN layer and the RELU activation function, and the feature map T1 and the feature map T2 are obtained;

[0088] The output of the differential branch is obtained after the convolution layer (3x3), the BN layer and the RELU activation function, and the feature map T4 and the feature map T5 are obtained;

[0089] The feature map obtained after the convolution layer of the feature map T1 and the feature map obtained after the convolution layer of the feature map T4 are merged to obtain the feature map T3;

[0090] The feature map obtained after the convolution layer of the feature map T2 and the feature map T3 are fused to obtain the feature map T6;

[0091] The feature map obtained after the convolution layer of the feature map T5 and the feature map T3 are fused to obtain the feature map T7;

[0092] The feature map T7 and the feature map T6 are merged to obtain a feature map T8, and an output 1 is obtained after a RELU activation function.

[0093] The output of the differential branch is subjected to a convolution layer (3x3), a BN layer and a RELU activation function to obtain an output 2.

[0094] Specifically, a structure diagram of the pixel attention guiding module is as shown in Figure 4 The vector corresponding to the pixels of the proportional branch and the differential branch feature map is defined as: Then the output of the Sigmoid, that is, the feature map T6 / T7 can be expressed as:

[0095]

[0096] Wherein, sigma represents the possibility that the two pixels belong to the same object. If sigma is high, we are more confident Because the integral branch is more accurate in semantics, the output 1 of the pixel attention guiding module can be expressed as:

[0097]

[0098] In an embodiment, further comprising:

[0099] The pyramid pooling module is used to aggregate the context information of different regions to improve the ability of the network to obtain global information.

[0100] The pyramid pooling module is used to: use different scales of pooling on the original feature map to obtain a plurality of feature maps of different sizes, then splice the feature maps in the channel dimension, and then splice the original feature map, and finally output a composite feature map that combines multiple scales, so as to achieve the purpose of balancing global semantic information and local detail information.

[0101] In an embodiment, further comprising:

[0102] The pyramid pooling module is used to: perform different scale pooling operations on the original feature map to obtain a plurality of feature maps of different sizes (using 5 branches); perform an upsampling operation on the obtained feature maps to restore to the original feature map size (6x6), and finally splice in the channel dimension to obtain the final composite feature map.

[0103] The first branch: using (6x6) pooling, the output size is (1x1), and then up-sampling to (6x6) through bilinear interpolation;

[0104] The second branch: using (3x3) pooling, the output size is (2x2), and then up-sampling to (6x6) through bilinear interpolation;

[0105] The third branch: using (2x2) pooling, the output size is (3x3), and then up-sampling to (6x6) by bilinear interpolation;

[0106] The fourth branch: using (1x1) pooling, the output size is (6x6);

[0107] The fifth branch: representing the input original feature map, which plays a role of residual connection;

[0108] The output results of the first four branches are spliced with the fifth branch, and then the output result is output.

[0109] Specifically, the pyramid pooling module is used to aggregate the context information of different regions to improve the ability of the network to obtain global information, and the network structure is as shown in Figure 5 The specific method is: using different scales of pooling on the original feature map to obtain a plurality of feature maps of different sizes, then splicing these feature maps in the channel dimension, and then splicing with the original feature map, and finally outputting a composite feature map that integrates multiple scales, so as to achieve the purpose of considering global semantic information and local detail information.

[0110] The original feature map is subjected to different scale pooling operations to obtain a plurality of feature maps of different sizes (5 branches are used in this paper). The obtained feature maps are up-sampled to restore to the size of the original feature map (6x6), and finally spliced in the channel dimension to obtain the final composite feature map;

[0111] The first branch: using (6x6) pooling, the output size is (1x1), and then up-sampling to (6x6 by bilinear interpolation;

[0112] The second branch: using (3x3) pooling, the output size is (2x2), and then up-sampling to (6x6 by bilinear interpolation;

[0113] The third branch: using (2x2) pooling, the output size is (3x3), and then up-sampling to (6x6 by bilinear interpolation;

[0114] The fourth branch: using (1x1) pooling, the output size is (6x6).

[0115] The fifth branch: representing the input feature map, which plays a role of residual connection.

[0116] The output results of the first four branches are spliced with the fifth branch, and then the output result is output.

[0117] In one embodiment, it further comprises:

[0118] The boundary attention guiding module is used for: after the differential branch input is subjected to a Sigmoid activation function, two branch outputs are obtained, one of which is fused with the feature of the differential branch input, and then merged with the feature of the proportional branch input, and the other is fused with the feature of the proportional branch input, and then merged with the feature of the differential branch; after the two merged features are respectively subjected to a convolution layer (3x3 convolution kernel) and a BN layer, the final feature is output.

[0119] Specifically, the structural diagram of the boundary attention guiding module is as shown in Figure 6 The module has three inputs, which are the outputs of the first three branches of the network. After the input of the differential branch is subjected to a Sigmoid activation function, two branch outputs are obtained, one of which is fused with the feature of the differential branch input, and then merged with the feature of the proportional branch input, and the other is fused with the feature of the proportional branch input, and then merged with the feature of the differential branch; after the two merged features are respectively subjected to a convolution layer (3x3 convolution kernel) and a BN layer, the final feature is output.

[0120] The module is used to guide the fusion of context information by using boundary features to achieve better semantic segmentation effect. Although the context information has semantic accuracy, it loses too many geometric details in the boundary area and small objects, so the model is forced to trust the differential branch more in the boundary area and to strengthen the attention to the boundary.

[0121] The vectors corresponding to the pixels of the feature maps of the proportional branch, the integral branch and the differential branch are defined as: Then the output of the Sigmoid, the output Out of the boundary attention guiding module is:

[0122]

[0123]

[0124] Where fout represents the combination of convolution, batch normalization and ReLU. When σ>0.5, the model trusts the detailed features more, otherwise it is more inclined to detect the context information.

[0125] The loss function is described in detail as follows:

[0126] The loss function in the network designed by us is a composite function, which is composed of four parts:

[0127] First, the output position of the first pixel attention guiding module is generated by two convolution operations, and an additional semantic loss Loss1 is generated to better optimize the entire network.

[0128]

[0129] where y is the label value and y' is the predicted value.

[0130] Second, in order to deal with the imbalance problem in boundary detection, a weighted binary cross-entropy loss Loss2 is used instead of Dice Loss, because this can make the network more inclined to use rough boundaries to highlight boundary areas and enhance the features of small objects.

[0131]

[0132]

[0133] where the loss of a single sample is calculated as Loss(i), is the true label corresponding to the category, and is 1 for the kth category and 0 otherwise, many items will be shielded and not involved in the calculation. is the predicted probability after the softmax function.

[0134] In order to make some pixel points more important, w(x) is introduced. We pre-compute a weight map for each labeled image to compensate for the different frequencies of each class of pixels in the training set, so that the network pays more attention to learning small segmentation boundaries that are in contact with each other. The weight map calculation formula is as follows:

[0135]

[0136] where d1 represents the distance to the nearest boundary, and d2 represents the distance to the second nearest boundary. Based on experience, we set w0=10 and σ≈5 pixel values.

[0137] Third, Loss3 and Loss4 represent cross-entropy loss, respectively. Here, the output boundary head is used to coordinate the semantic segmentation and boundary detection tasks, and to enhance the function of the boundary attention guide module. Therefore, the loss can be defined as:

[0138]

[0139] where t represents a predefined threshold, b i , s i,c , are the output of the boundary, the true value of the segmentation, and the prediction result of the ith pixel for the category, respectively.

[0140] Therefore, the final network loss function is represented as:

[0141] Loss = λ1L1 + λ2L2 + λ3L3 + λ4L4

[0142] The training loss parameter is set as λ1=0.4, λ2=0.2, λ3=0.2, λ4=0.2, and t=0.8. A larger λ1 is set to enhance the boundary learning ability.

[0143] VR auxiliary display: the CT and MRI segmentation results of the patient are generated into a three-dimensional model, input into the VR device, and the three-dimensional model is rendered on the VR head-mounted display device to provide intuitive, stereoscopic and realistic three-dimensional space virtual display for simulation operation planning and reference.

[0144] Figure 7 It is an embodiment of the present application to provide a flowchart of a virtual reality-based hip arthroscopy surgery auxiliary teaching method, the virtual reality-based hip arthroscopy surgery auxiliary teaching method comprising:

[0145] S701, inputting a hip joint CT image into a preset image segmentation model to output a hip joint bone image; wherein the hip joint bone image comprises first pelvic position information and first femur position information;

[0146] S702, inputting a hip joint MRI image into a preset image segmentation model to output a hip joint part image; wherein the hip joint part image comprises second pelvic position information, second femur position information and soft tissue position information;

[0147] S703, determining a registration position based on the first pelvic position information, the first femur position information, the second pelvic position information, the second femur position information and the soft tissue position information, and performing registration to fuse and reconstruct a hip joint part multi-modal information image;

[0148] S704, three-dimensional model rendering based on the hip joint part multi-modal information image to perform three-dimensional space virtual display auxiliary teaching;

[0149] The preset image segmentation model is obtained by model training based on a neural network, the structure of the neural network uses a cascaded residual block as a backbone network, and different depth and width networks are used to extract features; the structure of the neural network comprises: three convolutional layers and pooling layers, which are used for downsampling operation, reducing image size to reduce calculation amount and accelerating model inference speed; three network branches connected respectively: a scaling branch, which is responsible for analyzing and retaining detailed information in a high-resolution feature map; an integration branch, which is responsible for aggregating local and global context information to capture long-distance dependencies; and a differentiation branch, which is responsible for extracting high-frequency features to predict boundary regions.

[0150] Figure 8 The structure of the electronic device provided by the embodiment of the present application is shown.

[0151] The electronic device can include a processor 801 and a memory 802 having stored computer program instructions.

[0152] In particular, the processor 801 described above can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or can be configured to implement one or more integrated circuits that embody the embodiments of the present application.

[0153] The memory 802 can include a mass storage for data or instructions. By way of example and not limitation, the memory 802 can include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive or a combination of two or more of these. The memory 802 can be removable and / or non-removable (or fixed) as appropriate. The memory 802 can be internal or external to the electronic device as appropriate. In a particular embodiment, the memory 802 can be a non-volatile solid-state memory.

[0154] In one embodiment, the memory 802 can be a read only memory (ROM). In one embodiment, the ROM can be a mask programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these.

[0155] The processor 801 implements any of the above-mentioned virtual reality-based hip arthroscopy surgery auxiliary teaching methods by reading and executing the computer program instructions stored in the memory 802.

[0156] In one example, the electronic device can further include a communication interface 803 and a bus 810. As shown, the processor 801, the memory 802, and the communication interface 803 are connected through the bus 810 and complete communication with each other. Figure 8

[0157] The communication interface 803 is mainly used to realize the communication between the modules, devices, units and / or equipment in the embodiments of the present application.

[0158] ​Bus 810 includes a hardware, software, or both that couples components of electronic device to each other. As an example and not by way of limitation, bus can include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand (IB) interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or another suitable bus or a combination of two or more of these. Where appropriate, bus 810 can include one or more buses. Although this application describes and shows a particular bus, this application contemplates any suitable bus or interconnect.

[0159] In addition, in combination with the above-mentioned hip arthroscopy surgery auxiliary teaching method based on virtual reality, the embodiments of the present application can provide a computer readable storage medium for implementation. The computer readable storage medium has computer program instructions stored thereon; the computer program instructions are executed by a processor to implement any of the above-mentioned hip arthroscopy surgery auxiliary teaching methods based on virtual reality.

[0160] It needs to be clear that the present application is not limited to the specific configurations and processes described above and shown in the drawings. For the sake of brevity, detailed descriptions of well-known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between steps, after understanding the spirit of the present application.

[0161] The functional modules shown in the structure block diagram described above can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present application are program or code segments used to perform the required tasks. The program or code segments can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave over a transmission medium or communication link. The "machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segments can be downloaded via a computer network such as the Internet, an intranet, etc.

[0162] It should also be noted that the example embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiments, or in an order different from the embodiments, or several steps can be performed simultaneously.

[0163] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other processing device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other processing device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0164] The above only describes specific implementation manners of the present application. For the convenience and brevity of description, the specific working process of the system, module and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described herein. It should be understood that the protection scope of the present application is not limited in this way. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed in the present application, and these modifications or replacements should be covered by the protection scope of the present application.

Claims

1. A hip arthroscopy surgery auxiliary teaching system based on virtual reality, characterized in that: include: A first image segmentation module is configured to input the hip joint CT image into a preset image segmentation model and output a hip joint bone image; wherein the hip joint bone image includes first pelvic position information and first femoral position information; a second image segmentation module, configured to input the hip joint MRI image into a preset image segmentation model and output a hip joint part image; wherein the hip joint part image includes second pelvis position information, second femur position information, and soft tissue position information; A registration fusion reconstruction module is used to determine the registration position and perform registration based on the first pelvis position information, the first femur position information, the second pelvis position information, the second femur position information and the soft tissue position information, and fuse and reconstruct to obtain a multimodal information image of the hip joint; A rendering and display module is used to render a three-dimensional model based on multimodal information images of the hip joint to perform three-dimensional space virtual display-assisted teaching; Among them, the preset image segmentation model is obtained through model training based on a neural network. The structure of the neural network uses cascaded residual blocks as the backbone network, and uses networks of different depths and widths to extract features; the structure of the neural network includes: three convolutional layers and pooling layers, which are used to perform downsampling operations, reduce the image size, reduce the amount of calculation, and speed up the model inference speed; three network branches are connected respectively: the proportional branch, which is responsible for parsing and retaining detailed information in the high-resolution feature map; the integral branch, which is responsible for aggregating local and global contextual information to capture long-distance dependencies; the differential branch, which is responsible for extracting high-frequency features to predict boundary areas.

2. The virtual reality-based hip arthroscopy surgery auxiliary teaching system according to claim 1 is characterized in that: The feature map obtained by the proportional branch after passing through the convolution layer and the feature map obtained by the differential branch after passing through the convolution layer are input into the pixel attention guidance module and then output through the convolution layer, a total of three convolution layers; The integral branch outputs a feature map after three convolutional layers and a pooling layer; The feature map output by the integral branch is reduced in size after being output through three pooling layers, and then restored to its original size through the pyramid pooling module; The feature map obtained by the differential branch after passing through the convolution layer and the feature map obtained by the proportional branch after passing through the convolution layer are merged and output to the next convolution layer. After a total of three convolution layers, the feature map is output; The output results of the three branches are jointly input into the boundary attention guidance module, and the output feature map is output as a prediction mask after passing through the convolution layer, BN layer and RELU activation function.

3. The virtual reality-based hip arthroscopy surgery auxiliary teaching system according to claim 2 is characterized in that: Also includes: The loss function calculation module is used to: perform boundary detection on the real mask to obtain the boundary mask; calculate the loss function Loss1 based on the boundary mask and the output of the differential branch; The predicted mask and the true mask calculate the loss function Loss2; The predicted mask, the true mask and the differential branch output results are used to calculate the loss function Loss3; The output result of the proportional branch and the real mask are used to calculate the loss function Loss4; the final loss function is calculated based on the loss functions Loss1, LossL2, Loss3 and Loss4.

4. The virtual reality-based hip arthroscopy surgery auxiliary teaching system according to claim 3 is characterized in that: Also includes: The pixel attention guidance module is used to interactively enhance the feature maps of the scale branch and the differential branch using the attention mechanism; The two inputs of the pixel attention guidance module are the output of the ratio branch and the output of the differential branch; The output of the proportional branch passes through the convolution layer (3x3), the BN layer and the RELU activation function to obtain the feature map T1 and the feature map T2; The output of the differential branch passes through the convolution layer (3x3), the BN layer and the RELU activation function to obtain feature maps T4 and T5; The feature map T1 obtained after the convolution layer and the feature map T4 obtained after the convolution layer are merged to obtain the feature map T3; The feature map T2 obtained after the convolution layer is fused with the feature map T3 to obtain the feature map T6; The feature map T5 obtained after the convolution layer is fused with the feature map T3 to obtain the feature map T7; Feature map T7 and feature map T6 are merged to obtain feature map T8, which is activated by RELU function to get output 1. The output of the differential branch passes through the convolution layer (3x3), BN layer and RELU activation function to obtain output 2.

5. The virtual reality-based hip arthroscopy surgery auxiliary teaching system according to claim 4 is characterized in that: Also includes: The pyramid pooling module is used to aggregate contextual information from different regions to improve the network's ability to obtain global information; The pyramid pooling module is used to apply pooling of different scales to the original feature map to obtain multiple feature maps of different sizes. These feature maps are then concatenated in the channel dimension and then concatenated with the original feature map. Finally, a composite feature map that combines multiple scales is output, thereby achieving the goal of taking into account both global semantic information and local detail information.

6. The virtual reality-based hip arthroscopy surgery auxiliary teaching system according to claim 5, characterized in that: Also includes: The pyramid pooling module is used to perform pooling operations of different scales on the original feature map to obtain multiple feature maps of different sizes (using 5 branches); The obtained feature map is upsampled to restore the original feature map size (6×6), and finally spliced ​​in the channel dimension to obtain the final composite feature map; First branch: Use (6×6) pooling, the output size is (1×1), and then upsample to (6×6) through bilinear interpolation; Second branch: using (3×3) pooling, the output size is (2×2), and then upsampled to (6×6) through bilinear interpolation; The third branch uses (2×2) pooling, with an output size of (3×3), which is then upsampled to (6×6) through bilinear interpolation. The fourth branch uses (1×1) pooling and the output size is (6×6); The fifth branch: represents the input original feature map to play the role of residual connection; The output results of the first four branches are subjected to feature splicing and then spliced ​​with the fifth branch, and then the results are output.

7. The virtual reality-based hip arthroscopy surgery auxiliary teaching system according to claim 6, characterized in that: Also includes: The boundary attention guidance module is used to obtain two branch outputs after the differential branch input passes through the Sigmoid activation function. The output features of one branch are fused with the features of the differential branch input and then merged with the features of the proportional branch input. The output features of the other branch are fused with the features of the proportional branch input and then merged with the features of the differential branch. The two merged features are output through the convolution layer (3x3 convolution kernel) and the BN layer respectively, and then merged to output the final features.

8. A virtual reality-based hip arthroscopy surgery assisted teaching method, characterized in that: include: Inputting the hip joint CT image into a preset image segmentation model to output a hip joint bone image; wherein the hip joint bone image includes first pelvic position information and first femoral position information; Inputting the hip joint MRI image into a preset image segmentation model to output a hip joint image; wherein the hip joint image includes the second pelvis position information, the second femur position information and the soft tissue position information; Determining the registration position and performing registration based on the first pelvis position information, the first femur position information, the second pelvis position information, the second femur position information, and the soft tissue position information, and fusing and reconstructing to obtain a multimodal information image of the hip joint; Rendering of a three-dimensional model based on multimodal information images of the hip joint allows for three-dimensional virtual display-assisted teaching; Among them, the preset image segmentation model is obtained through model training based on a neural network. The structure of the neural network uses cascaded residual blocks as the backbone network, and uses networks of different depths and widths to extract features; the structure of the neural network includes: three convolutional layers and pooling layers, which are used to perform downsampling operations, reduce the image size, reduce the amount of calculation, and speed up the model inference speed; three network branches are connected respectively: the proportional branch, which is responsible for parsing and retaining detailed information in the high-resolution feature map; the integral branch, which is responsible for aggregating local and global contextual information to capture long-distance dependencies; the differential branch, which is responsible for extracting high-frequency features to predict boundary areas.

9. An electronic device, characterized in that: The electronic device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the virtual reality-based hip arthroscopic surgery assisted teaching method according to claim 8 is implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the virtual reality-based hip arthroscopic surgery assisted teaching method according to claim 8.

Citation Information

Patent Citations

  • Hip arthroscopic surgery auxiliary teaching system based on virtual reality

    CN112509410A

  • Muscle segmentation model method based on multi-modal MRI image

    CN114511540A