A method, device, equipment and medium for positioning key points of spinal segment space

Through the two-stage network model, the spatial area of ​​the spinal segment is first positioned and the CT image is cropped, and then the key points of the spinal segment anatomy are accurately positioned, which solves the problem of insufficient positioning accuracy and robustness in the existing technology, and achieves efficient key points of the spinal segment anatomy.

CN119559257BActive Publication Date: 2025-06-06SUZHOU ZOEZEN ROBOT CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510115927.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-06
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

The prior art has problems of low accuracy, low robustness and high computational complexity in spinal CT images, making it difficult to accurately locate anatomical key points of each spinal segment in any field of view CT images.

Method used

Using a two-stage network model, firstly, the spatial area position information of each spinal segment is obtained through the spinal segment spatial area positioning network, and the original CT image is cropped to obtain local CT images; then, the local CT image is processed through the spinal segment anatomical key point positioning network to accurately locate the anatomical key points of the spinal segment.

Benefits of technology

It realizes the precise positioning of the anatomical key points of each spinal segment in any field of view CT images, improves the accuracy and robustness of the algorithm, and reduces the computational complexity, and is suitable for clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559257B_ABST
    Figure CN119559257B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, equipment and medium for positioning key points of spinal segment space, and relates to the field of image data processing technology. The method comprises: inputting the original CT image of the spine into the spinal segment space region positioning network to obtain the position information of each spinal segment space region, and based on the position information of each spinal segment space region, the original CT image of the spine is cropped to obtain the local CT image of each spinal segment; the local CT image corresponding to the diseased spinal segment is input into the spinal segment anatomical key point positioning network to obtain the position information of the anatomical key points of the diseased spinal segment. The present invention has high precision, strong robustness and low computational complexity, and can realize the precise positioning of the spinal segment anatomical key points in any field of view CT image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a method, device, equipment and medium for positioning key points of a spinal segment space. Background Art

[0002] The localization of spinal anatomical key points in CT images is an important basis for realizing autonomous path planning of spinal surgical robots. Manual selection of spinal anatomical key points is often subjective, and inexperienced doctors often find it difficult to ensure the accuracy of key point selection. Existing spinal anatomical key point localization algorithms often have problems with low accuracy and weak robustness, and are limited by the uncertainty of the CT scanning area. Existing methods cannot accurately locate the anatomical key points of each spinal segment in CT images of any field of view. In addition, existing CT image processing algorithms often have the problem of high computational complexity, which places high demands on the hardware system equipped with the algorithm. Therefore, it is necessary to study a method with high accuracy, strong robustness, and low computational complexity that can accurately locate the anatomical key points of spinal segments in CT images of any field of view. Summary of the invention

[0003] In view of the above-mentioned defects or deficiencies in the prior art, the present invention provides a method, device, equipment and medium for locating key points in spinal segment space, which can fully or partially solve the above-mentioned technical problems.

[0004] One aspect of the present invention provides a method for locating spatial key points of spinal segments, comprising: inputting an original spinal CT image into a spinal segment spatial region positioning network to obtain position information of each spinal segment spatial region, and based on the position information of each spinal segment spatial region, cropping the original spinal CT image to obtain a local CT image of each spinal segment; wherein the spinal segment spatial region positioning network comprises a backbone network and a first decoder network, wherein the backbone network is used to extract feature information of the original spinal CT image, and the first decoder network is used to decode and obtain coordinate information of each spinal segment spatial region in the original spinal CT image according to the feature information; inputting the local CT image corresponding to the diseased spinal segment into the spinal segment anatomical key point positioning network to obtain the position information of the anatomical key points of the diseased spinal segment; wherein the spinal segment anatomical key point positioning network comprises an encoder network and a second decoder network.

[0005] Another aspect of the present invention provides a spinal segment spatial key point positioning device, including: a spinal segment spatial area positioning module, configured to input the original spinal CT image into the spinal segment spatial area positioning network to obtain the position information of each spinal segment spatial area, and based on the position information of each spinal segment spatial area, crop the original spinal CT image to obtain the local CT image of each spinal segment; wherein the spinal segment spatial area positioning network includes a backbone network and a first decoder network, the backbone network is used to extract the feature information of the original spinal CT image, and the first decoder network is used to decode according to the feature information to obtain the coordinate information of each spinal segment spatial area in the original spinal CT image; and a spinal segment anatomical key point positioning module, configured to input the local CT image corresponding to the diseased spinal segment into the spinal segment anatomical key point positioning network to obtain the position information of the anatomical key points of the diseased spinal segment; wherein the spinal segment anatomical key point positioning network includes an encoder network and a second decoder network.

[0006] Another aspect of the present invention provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement any one of the methods described above.

[0007] Another aspect of the present invention provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, any of the above-mentioned methods is implemented.

[0008] The method, device, equipment and medium for positioning key points of spinal segment space provided by the present invention have the following beneficial effects:

[0009] (1) The spinal segment spatial region positioning network model algorithm and the spinal segment anatomical key point positioning network model algorithm proposed in the present invention have the ability to extract and integrate multi-receptive field features, and the algorithm accuracy and robustness are very high;

[0010] (2) The spinal segment spatial region positioning network model algorithm and the spinal segment anatomical key point positioning network model algorithm proposed in the present invention are developed based on the lightweight structure design concept, have high algorithm operation speed, low hardware requirements, and have strong clinical practicality;

[0011] (3) The coordinated use of the spinal segment spatial region positioning network model algorithm and the spinal segment anatomical key point positioning network model algorithm proposed in the present invention can achieve accurate positioning of each spinal segment anatomical key point in any field of view. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0013] Figure 1 This is a schematic diagram of an image processing path of a method for locating key points of a spinal segment space provided by an embodiment of the present application;

[0014] Figure 2 It is a flowchart of a method for locating key points of a spinal segment space provided by an embodiment of the present application;

[0015] Figure 3 It is a schematic diagram of the structure of the spinal segment spatial region positioning network in the spinal segment spatial key point positioning method provided by an embodiment of the present application;

[0016] Figure 4 It is a schematic diagram of a grid output of a spinal segment spatial region positioning network in a spinal segment spatial key point positioning method provided by an embodiment of the present application;

[0017] Figure 5 It is a schematic diagram of the structure of the spinal segment anatomical key point positioning network in the spinal segment spatial key point positioning method provided by an embodiment of the present application;

[0018] Figure 6 It is a structural schematic diagram of a spinal segment spatial key point positioning device provided by an embodiment of the present application;

[0019] Figure 7 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0021] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.

[0022] It should be understood that although the terms first, second, third, etc. may be used to describe the acquisition modules in the embodiments of the present invention, the acquisition modules should not be limited to these terms. These terms are only used to distinguish the acquisition modules from each other.

[0023] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.

[0024] It should be noted that the directional words such as "upper", "lower", "left", and "right" described in the embodiments of the present invention are described at the angles shown in the drawings and should not be understood as limiting the embodiments of the present invention. In addition, in the context, it should also be understood that when it is mentioned that an element is formed "on" or "under" another element, it can not only be formed directly "on" or "under" another element, but also be formed "on" or "under" another element indirectly through an intermediate element.

[0025] The present application embodiment proposes a two-stage method for locating anatomical key points of spinal segments in arbitrary field of view CT images, and its image processing path is as follows: Figure 1 First, based on the spinal segment spatial region positioning network, the spatial region of each spinal segment in the CT image is accurately positioned and the corresponding coordinate information is output. On this basis, the original CT image is cropped to obtain a series of local CT images that completely contain each spinal segment. Then, the local CT image corresponding to the diseased spinal segment is input into the spinal segment anatomical key point positioning network to achieve accurate positioning of the anatomical key points of the diseased spinal segment, and finally the corresponding surgical path is autonomously generated.

[0026] See also Figure 2 The method for locating key points of spinal segment space in this embodiment includes the following steps:

[0027] Step S101, input the original spinal CT image into the spinal segment spatial area positioning network to obtain the position information of each spinal segment spatial area, and based on the position information of each spinal segment spatial area, crop the original spinal CT image to obtain the local CT image of each spinal segment; wherein the spinal segment spatial area positioning network includes a backbone network and a first decoder network, the backbone network is used to extract the feature information of the original spinal CT image, and the first decoder network is used to decode according to the feature information to obtain the coordinate information of each spinal segment spatial area in the original spinal CT image.

[0028] See also Figure 3 In order to realize the spatial area positioning of the spinal segment, the spinal segment spatial area positioning network 101 includes two parts: a backbone network 102 and a decoder network 103. Among them, the backbone network 102 is mainly responsible for extracting the feature information of the CT image, and the decoder network 103 decodes and obtains the coordinate information of each spinal segment spatial area in the CT image on this basis.

[0029] In the backbone network 102, a three-dimensional convolution module with a convolution kernel size of 3, namely: conv(3), is first used to achieve the preliminary extraction of CT image feature information, and then four basic operation units (Base Lay) connected in series are used to perform deeper feature extraction, and finally a multi-scale convolution module (MSCK) is used to integrate feature information. Preferably, each basic operation unit (Base Lay) in this embodiment includes a multi-scale convolution module (MSCK) and a conv(3), and is provided with a jump connection structure. Among them, the multi-scale convolution module (MSCK) designed in this embodiment is a multi-scale feature extraction and integration module composed of convolution modules with convolution kernels of various sizes, including three convolution modules with convolution kernel sizes of 3, 5, and 7, so as to have a variety of different feature receptive fields, so that the spinal segment spatial area positioning network 101 has a strong multi-scale feature extraction capability. The jump connection structure used in the multi-scale convolution module (MSCK) designed in this embodiment can effectively improve the training speed of the algorithm and avoid the problem of gradient disappearance. The multi-scale convolution module (MSCK) designed in this embodiment can also stack feature information of different scales in the feature channel dimension, and integrate the features by subsequent convolution, BN (Batch Norm) layer and Relu activation function module. In order to reduce the computational complexity of the spinal segment spatial region positioning network 101 of this embodiment, the multi-scale convolution module (MSCK) designed in this embodiment preferably uses dilated convolution to expand the size of the convolution kernel.

[0030] like Figure 3 As shown, in the decoder network 103, the spinal segment spatial region positioning network 101 mainly uses a multi-scale convolution module (MSCK) for feature extraction and integration to further extract and integrate multi-scale feature information. The basic principle of the spinal segment spatial region positioning network 101 is to divide the image into a series of grids and recognize the target in each grid. In order to effectively deal with the uncertainty of the spinal segment spatial size caused by various reasons, three grid outputs with different densities are set in the decoder network 103. Figure 4 As shown, a dense grid can realize the recognition of a small-sized spine such as the cervical vertebra, a medium-density grid can realize the recognition of a medium-sized spine such as the thoracic vertebra, and a sparse grid can realize the recognition of a large-sized spine such as the lumbar vertebra.

[0031] In this embodiment, a head structure (Head Block) is also designed in the decoder network 103 for predicting the position, size, category and confidence of the target area. Through the head structure (Head Block), the spinal segment spatial region positioning network 101 will output the confidence of the existence of the target object for each grid. , the location information of the center point of the space area occupied by the target , size information , category information of the target For example, you can set the confidence level of the corresponding grid .5, there is a target object in the grid. The true center point of the target and size The calculation is as follows:

[0032]

[0033]

[0034] in, is the coordinate of the upper left corner of the grid. Category information of the target object It is an array of probability distribution information that represents the predicted target object belonging to various categories. The maximum probability among them is used to determine the category to which the target object belongs.

[0035] The loss function of the further spinal segment spatial region localization network 101 is as follows:

[0036]

[0037] in is the balance coefficient, Characterize the spatial relationship between the predicted 3D box and the real 3D box, Characterizes the accuracy of the target category prediction, Characterizes the accuracy of the prediction of the existence of the target. Preferably, is the 3DIoU (Intersection over Union) loss function, and is the cross entropy loss function.

[0038] For example, see Figure 3 , the spinal segment spatial region positioning network 101 includes:

[0039] (1) Backbone network

[0040] The backbone network 102 includes a first convolution module conv(3), a first basic operation unit Baselayer, a second basic operation unit Base layer, a third basic operation unit Base layer, a fourth basic operation unit Base layer and a first multi-scale convolution module MSCK which are connected in sequence;

[0041] The first convolution module conv(3) includes a 3D convolution layer with a convolution kernel size of 3, a BN layer, and a ReLU layer connected in sequence;

[0042] The first multi-scale convolution module MSCK includes a parallel convolution layer, a first concatenation layer concat, and a second convolution module conv(3) connected in sequence, and the parallel convolution layer includes a third convolution module conv(3), a fourth convolution module conv(5), and a fifth convolution module conv(7) connected in parallel;

[0043] The first convolution module conv(3), the second convolution module conv(3) and the third convolution module conv(3) have the same network structure; the fourth convolution module conv(5) includes a 3D convolution layer with a convolution kernel size of 5, a BN layer and a Relu layer connected in sequence; the fifth convolution module conv(7) includes a 3D convolution layer with a convolution kernel size of 7, a BN layer and a Relu layer connected in sequence;

[0044] The first basic operation unit Base layer, the second basic operation unit Base layer, the third basic operation unit Base layer, and the fourth basic operation unit Base layer have the same network structure. Each basic operation unit Base layer includes a second multi-scale convolution module MSCK and a sixth convolution module conv(3) connected in sequence. The input and output of the second multi-scale convolution module MSCK are input into the sixth convolution module conv(3) after tensor addition. The sixth convolution module conv(3) has the same network structure as the first convolution module conv(3), and the second multi-scale convolution module MSCK has the same network structure as the first multi-scale convolution module MSCK.

[0045] (2) Decoder Network

[0046] The decoder network 103 includes a third multi-scale convolution module MSCK, a first upsampling module UB, a second concatenation layer concat, a fourth multi-scale convolution module MSCK, a second upsampling module UB, a third concatenation layer concat, a fifth multi-scale convolution module MSCK, a first spatial target recognition decoupling head module head, a second spatial target recognition decoupling head module head and a third spatial target recognition decoupling head module head.

[0047] Among them, the network structure of the third multi-scale convolution module MSCK, the fourth multi-scale convolution module MSCK, and the fifth multi-scale convolution module MSCK is the same as that of the first multi-scale convolution module MSCK.

[0048] Among them, the input end of the third multi-scale convolution module MSCK is connected to the output end of the first multi-scale convolution module MSCK of the backbone network 102, and the output end of the third multi-scale convolution module MSCK is connected to the input end of the third spatial target recognition decoupling head module head.

[0049] Among them, the third multi-scale convolution module MSCK, the first upsampling module UB, the second splicing layer concat, the fourth multi-scale convolution module MSCK, the second upsampling module UB, the third splicing layer concat, the fifth multi-scale convolution module MSCK and the first spatial target recognition decoupling head module head are connected in sequence.

[0050] Among them, the output end of the second basic operation unit base layer is connected to the input end of the third concatenation layer concat; the output end of the third basic operation unit base layer is connected to the input end of the second concatenation layer concat; the output end of the fourth multi-scale convolution module MSCK is connected to the input end of the second spatial target recognition decoupling head module head.

[0051] Among them, the first space target recognition decoupling head module head, the second space target recognition decoupling head module head and the third space target recognition decoupling head module head are used to predict and output target area position information, size information, category information and confidence.

[0052] Among them, the network structure of the first upsampling module UB and the second upsampling module UB is the same, each upsampling module UB includes a transposed convolution module Trans conv(3) and a seventh convolution module conv(3) connected in sequence, and the network structure of the seventh convolution module conv(3) is the same as that of the first convolution module conv(3).

[0053] Among them, the network structures of the first space target recognition decoupling head module head, the second space target recognition decoupling head module head and the third space target recognition decoupling head module head are the same. Each space target recognition decoupling head module head includes an eighth convolution module conv(3), a ninth convolution module conv(3), a parallel 3D convolution module and a fourth concatenation layer concat connected in sequence. The parallel 3D convolution module includes three parallel 3D convolution layers with a convolution kernel size of 1. The network structures of the eighth convolution module conv(3), the ninth convolution module conv(3) and the first convolution module conv(3) are the same.

[0054] Step S102, inputting the local CT image corresponding to the diseased spinal segment into the spinal segment anatomical key point positioning network to obtain the position information of the anatomical key points of the diseased spinal segment; wherein the spinal segment anatomical key point positioning network includes an encoder network and a second decoder network.

[0055] Specifically, in order to realize the positioning of the key points of spinal anatomy, this embodiment designs a lightweight three-dimensional space key point positioning network, namely: the spinal segment anatomical key point positioning network 104, the structure of which is as follows: Figure 5 shown.

[0056] The spinal segment anatomical key point localization network 104 processes the feature information extracted and sorted by the encoder network 105 and the decoder network 106 through the heat map generation module HB, and outputs a heat map of the same size as the input CT image. Each pixel in the heat map represents the probability of the corresponding pixel in the input CT image being a key point. The heat map output by the spinal segment anatomical key point localization network 104 has the same number of anatomical key points as those to be detected. The same number of feature channels. The corresponding key point can be located by searching for the pixel with the largest pixel value in each different channel. The processing method is as follows:

[0057]

[0058] in, is the formula for finding the variable value when the function value is maximum. Indicates Key points, The heat map represents feature channels, i=1,2,3... .

[0059] Furthermore, during the training process of the spinal segment anatomical key point positioning network 104, when the radius of the light spot in the target thermal map When it is larger, the algorithm is more robust but the key point positioning accuracy is lower. When the value is smaller, the robustness of the algorithm is poor and the key point positioning accuracy is high. Therefore, in order to balance the robustness of the algorithm and the positioning accuracy, in the process of training the spinal segment anatomical key point positioning network 104, the present embodiment sets Every 50 rounds, the image is scaled down by 2 pixels.

[0060] Furthermore, the spinal segment anatomical key point localization network 104 uses Mse Loss as a loss function to gradually optimize and reduce the gap between the predicted thermal map and the target thermal map. The calculation method is as follows:

[0061]

[0062] in and The predicted channels and the target heatmap channels.

[0063] Furthermore, in order to enhance the contrast of the spine relative to other soft tissues, effectively solve the impact of different imaging parameters and noise characteristics of different CT devices on the algorithm, and enhance the robustness of the algorithm, the present invention designs an automatic CT image window width and window position adjustment method. This method divides the pixel values ​​in the CT image into 256 levels and counts the cumulative distribution of pixels belonging to each level:

[0064]

[0065] in, Indicates that the pixel value in the image is less than The ratio of pixels to total pixels.

[0066] On this basis, set The minimum level greater than 0.1 in is: and The maximum level less than 0.999 is ,in, Those that satisfy greater than 0.1 or less than 0.999 .

[0067] And the lower limit of the window is obtained by mapping calculation and the window upper limit , these two values ​​are two parameters for adjusting the window width and window position of CT images. The purpose of adjusting the window width and window position is to adjust the contrast between the spine and other backgrounds. The calculation method is as follows:

[0068]

[0069] in and are the minimum and maximum pixel values ​​in the CT image, respectively. Based on this, this embodiment remaps all pixels in the CT image, and the calculation method is as follows:

[0070]

[0071] in, Indicates the pixel value of the corresponding pixel after the window width and window position are adjusted. is the original pixel value of the pixel in the image.

[0072] For example, see Figure 5 , the spinal segment anatomical key point positioning network 104 includes:

[0073] (1) Encoder Network

[0074] The encoder network 105 includes a sixth multi-scale convolution module MSCK, a seventh multi-scale convolution module MSCK, a tenth convolution module conv(3), an eighth multi-scale convolution module MSCK, a ninth multi-scale convolution module MSCK, an eleventh convolution module conv(3), a tenth multi-scale convolution module MSCK, an eleventh multi-scale convolution module MSCK, a twelfth convolution module conv(3), a twelfth multi-scale convolution module MSCK and a thirteenth multi-scale convolution module MSCK connected in sequence; wherein the network structure of the sixth multi-scale convolution module MSCK, the seventh multi-scale convolution module MSCK, the eighth multi-scale convolution module MSCK, the ninth multi-scale convolution module MSCK, the tenth multi-scale convolution module MSCK, the eleventh multi-scale convolution module MSCK, the twelfth multi-scale convolution module MSCK, the thirteenth multi-scale convolution module MSCK and the first multi-scale convolution module MSCK is the same; wherein the network structure of the tenth convolution module conv(3), the eleventh convolution module conv(3) and the twelfth convolution module conv(3) is the same as the first convolution module conv(3).

[0075] (2) Decoder Network

[0076] The decoder network 106 includes a fourteenth multi-scale convolution module MSCK, a third upsampling module UB, a fifteenth multi-scale convolution module MSCK, a fourth upsampling module UB, a sixteenth multi-scale convolution module MSCK, a fifth upsampling module UB, a seventeenth multi-scale convolution module MSCK, an eighteenth multi-scale convolution module MSCK and a heat map generation module HB connected in sequence; the fourteenth multi-scale convolution module MSCK, the fifteenth multi-scale convolution module MSCK, the sixteenth multi-scale convolution module MSCK, the seventeenth multi-scale convolution module MSCK, and the eighteenth multi-scale convolution module MSCK have the same network structure as the first multi-scale convolution module MSCK; the third upsampling module UB, the fourth upsampling module UB, and the fifth upsampling module UB have the same network structure as the first upsampling module UB.

[0077] Among them, the output end of the seventh multi-scale convolution module MSCK is connected to the input end of the seventeenth multi-scale convolution module MSCK; the output end of the ninth multi-scale convolution module MSCK is connected to the input end of the sixteenth multi-scale convolution module MSCK; the output end of the eleventh multi-scale convolution module MSCK is connected to the input end of the fifteenth multi-scale convolution module MSCK.

[0078] The heat map generation module HB includes a thirteenth convolution module conv(3), a BN module and a Softmax activation function module connected in sequence, wherein the thirteenth convolution module conv(3) has the same network structure as the first convolution module conv(3); the heat map generation module HB is used to output a heat map of the same size as the local CT image corresponding to the input diseased spinal segment, and each pixel in the heat map represents the probability of the corresponding pixel in the input image being a key point.

[0079] The corresponding key point is located by searching for the pixel with the largest pixel value in each different feature channel in the heat map, and the heat map has the same number of feature channels as the number of anatomical key points to be detected.

[0080] The spinal segment spatial region positioning network model algorithm and the spinal segment anatomical key point positioning network model algorithm proposed in this embodiment have the ability to extract and integrate multi-receptive field features, and the algorithm accuracy and robustness are very strong. They are developed based on the lightweight structure design concept, the algorithm has high computing speed, low hardware requirements, and strong clinical practicality. In addition, the coordinated use of the spinal segment spatial region positioning network model algorithm and the spinal segment anatomical key point positioning network model algorithm can achieve accurate positioning of each spinal segment anatomical key point under any field of view.

[0081] See also Figure 6 Another embodiment of the present invention further provides a spinal segment spatial key point positioning device 200, including a spinal segment spatial area positioning module 201 and a spinal segment anatomical key point positioning module 202. The spinal segment spatial key point positioning device 200 can execute the spinal segment spatial key point positioning method in the method embodiment.

[0082] Specifically, the spinal segment spatial key point positioning device 200 includes:

[0083] The spinal segment spatial region positioning module 201 is configured to input the original spinal CT image into the spinal segment spatial region positioning network to obtain the position information of each spinal segment spatial region, and based on the position information of each spinal segment spatial region, crop the original spinal CT image to obtain the local CT image of each spinal segment; wherein the spinal segment spatial region positioning network includes a backbone network and a first decoder network, the backbone network is used to extract the feature information of the original spinal CT image, and the first decoder network is used to decode and obtain the coordinate information of each spinal segment spatial region in the original spinal CT image according to the feature information; and,

[0084] The spinal segment anatomical key point positioning module 202 is configured to input the local CT image corresponding to the diseased spinal segment into the spinal segment anatomical key point positioning network to obtain the position information of the anatomical key points of the diseased spinal segment; wherein the spinal segment anatomical key point positioning network includes an encoder network and a second decoder network.

[0085] It should be noted that the spinal segment spatial key point positioning device 200 provided in this embodiment corresponds to a technical solution that can be used to execute each method embodiment, and its implementation principle and technical effect are similar to the method, which will not be repeated here.

[0086] See also Figure 7 Another embodiment of the present invention provides a schematic diagram of the structure of an electronic device 300, which is used to implement the method for locating key points of spinal segment space in the method embodiment. The electronic device 300 in the embodiment of the present invention may include but is not limited to a smart phone, a PAD, a laptop, an industrial computer, a PC, and a server. Figure 7 The electronic device 300 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0087] like Figure 7 As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 to a random access memory (RAM) 303 to implement the method of the embodiment of the present invention. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0088] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Figure 7 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0089] Another embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which can implement the above-mentioned method for locating key points of spinal segment space when executed by a processor.

[0090] The above description is only a preferred embodiment of the present invention. Those skilled in the art should understand that the scope of disclosure involved in the present invention is not limited to the technical solution formed by a specific combination of the above technical features, but also should cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present invention (but not limited to) to form a technical solution.

Claims

1. A method for locating key points of spinal segment space, characterized in that include: The original spinal CT image is input into the spinal segment spatial region positioning network to obtain the position information of each spinal segment spatial region, and based on the position information of each spinal segment spatial region, the original spinal CT image is cropped to obtain the local CT image of each spinal segment; wherein the spinal segment spatial region positioning network includes a backbone network and a first decoder network, the backbone network is used to extract the feature information of the original spinal CT image, and the first decoder network is used to decode and obtain the coordinate information of each spinal segment spatial region in the original spinal CT image according to the feature information; Inputting the local CT image corresponding to the diseased spinal segment into the spinal segment anatomical key point positioning network to obtain the position information of the anatomical key points of the diseased spinal segment; wherein the spinal segment anatomical key point positioning network includes an encoder network and a second decoder network; The backbone network includes a first convolution module, a first basic operation unit, a second basic operation unit, a third basic operation unit, a fourth basic operation unit and a first multi-scale convolution module which are connected in sequence; The first convolution module includes a 3D convolution layer with a convolution kernel size of 3, a BN layer and a ReLU layer connected in sequence; The first multi-scale convolution module includes a parallel convolution layer, a first concatenation layer, and a second convolution module connected in sequence, and the parallel convolution layer includes a third convolution module, a fourth convolution module, and a fifth convolution module connected in parallel; The first convolution module, the second convolution module and the third convolution module have the same network structure; the fourth convolution module includes a 3D convolution layer with a convolution kernel size of 5, a BN layer and a Relu layer connected in sequence; the fifth convolution module includes a 3D convolution layer with a convolution kernel size of 7, a BN layer and a Relu layer connected in sequence; The first basic operation unit, the second basic operation unit, the third basic operation unit, and the fourth basic operation unit have the same network structure, each basic operation unit includes a second multi-scale convolution module and a sixth convolution module connected in sequence, and the input and output of the second multi-scale convolution module are input to the sixth convolution module after tensor addition; wherein the sixth convolution module has the same network structure as the first convolution module, and the second multi-scale convolution module has the same network structure as the first multi-scale convolution module; The first decoder network includes a third multi-scale convolution module, a first upsampling module, a second splicing layer, a fourth multi-scale convolution module, a second upsampling module, a third splicing layer, a fifth multi-scale convolution module, a first spatial target recognition decoupling head module, a second spatial target recognition decoupling head module, and a third spatial target recognition decoupling head module; The third multi-scale convolution module, the fourth multi-scale convolution module, and the fifth multi-scale convolution module have the same network structure as the first multi-scale convolution module; The input end of the third multi-scale convolution module is connected to the output end of the first multi-scale convolution module of the backbone network, and the output end of the third multi-scale convolution module is connected to the input end of the third spatial target recognition decoupling head module; The third multi-scale convolution module, the first up-sampling module, the second splicing layer, the fourth multi-scale convolution module, the second up-sampling module, the third splicing layer, the fifth multi-scale convolution module and the first spatial target recognition decoupling head module are connected in sequence; The output end of the second basic operation unit is connected to the input end of the third splicing layer; The output end of the third basic operation unit is connected to the input end of the second splicing layer; The output end of the fourth multi-scale convolution module is connected to the input end of the second space target recognition decoupling head module; The first space target recognition decoupling head module, the second space target recognition decoupling head module and the third space target recognition decoupling head module are used to predict and output target area position information, size information, category information and confidence.

2. A method for locating key points of spinal segments in space according to claim 1, characterized in that: The first upsampling module and the second upsampling module have the same network structure, each upsampling module includes a transposed convolution module and a seventh convolution module connected in sequence, and the seventh convolution module has the same network structure as the first convolution module.

3. A method for locating key points of spinal segment space according to claim 2, characterized in that: The network structures of the first space target recognition decoupling head module, the second space target recognition decoupling head module and the third space target recognition decoupling head module are the same. Each space target recognition decoupling head module includes an eighth convolution module, a ninth convolution module, a parallel 3D convolution module and a fourth splicing layer connected in sequence. The parallel 3D convolution module includes three parallel 3D convolution layers with a convolution kernel size of 1. The network structures of the eighth convolution module, the ninth convolution module and the first convolution module are the same.

4. A method for locating key points of spinal segment space according to claim 3, characterized in that: The encoder network includes a sixth multi-scale convolution module, a seventh multi-scale convolution module, a tenth convolution module, an eighth multi-scale convolution module, a ninth multi-scale convolution module, an eleventh convolution module, a tenth multi-scale convolution module, an eleventh multi-scale convolution module, a twelfth convolution module, a twelfth multi-scale convolution module and a thirteenth multi-scale convolution module connected in sequence; wherein the sixth multi-scale convolution module, the seventh multi-scale convolution module, the eighth multi-scale convolution module, the ninth multi-scale convolution module, the tenth multi-scale convolution module, the eleventh multi-scale convolution module, the twelfth multi-scale convolution module, the thirteenth multi-scale convolution module and the first multi-scale convolution module have the same network structure; wherein the tenth convolution module, the eleventh convolution module and the twelfth convolution module have the same network structure as the first convolution module; The second decoder network includes a fourteenth multi-scale convolution module, a third upsampling module, a fifteenth multi-scale convolution module, a fourth upsampling module, a sixteenth multi-scale convolution module, a fifth upsampling module, a seventeenth multi-scale convolution module, an eighteenth multi-scale convolution module and a heat map generation module connected in sequence; the fourteenth multi-scale convolution module, the fifteenth multi-scale convolution module, the sixteenth multi-scale convolution module, the seventeenth multi-scale convolution module, and the eighteenth multi-scale convolution module have the same network structure as the first multi-scale convolution module; the third upsampling module, the fourth upsampling module, and the fifth upsampling module have the same network structure as the first upsampling module; Among them, the output end of the seventh multi-scale convolution module is connected to the input end of the seventeenth multi-scale convolution module; the output end of the ninth multi-scale convolution module is connected to the input end of the sixteenth multi-scale convolution module; the output end of the eleventh multi-scale convolution module is connected to the input end of the fifteenth multi-scale convolution module; The heat map generation module includes a thirteenth convolution module, a BN module and a Softmax activation function module connected in sequence, wherein the thirteenth convolution module has the same network structure as the first convolution module; the heat map generation module is used to output a heat map of the same size as the local CT image corresponding to the input diseased spinal segment, and each pixel in the heat map represents the probability of the corresponding pixel in the input image being a key point.

5. A method for locating key points of spinal segment space according to claim 4, characterized in that: The step of inputting the local CT image corresponding to the diseased spinal segment into the spinal segment anatomical key point positioning network to obtain the position information of the anatomical key points of the diseased spinal segment comprises: The corresponding key point is located by searching for the pixel with the maximum pixel value in each different feature channel in the heat map, and the heat map has the same number of feature channels as the number of anatomical key points to be detected.

6. A device for locating key points of spinal segments, characterized in that include: The spinal segment spatial region positioning module is configured to input the original spinal CT image into the spinal segment spatial region positioning network to obtain the position information of each spinal segment spatial region, and based on the position information of each spinal segment spatial region, crop the original spinal CT image to obtain the local CT image of each spinal segment; wherein the spinal segment spatial region positioning network includes a backbone network and a first decoder network, the backbone network is used to extract the feature information of the original spinal CT image, and the first decoder network is used to decode and obtain the coordinate information of each spinal segment spatial region in the original spinal CT image according to the feature information; and, The spinal segment anatomical key point positioning module is configured to input the local CT image corresponding to the diseased spinal segment into the spinal segment anatomical key point positioning network to obtain the position information of the anatomical key points of the diseased spinal segment; wherein the spinal segment anatomical key point positioning network includes an encoder network and a second decoder network; The backbone network includes a first convolution module, a first basic operation unit, a second basic operation unit, a third basic operation unit, a fourth basic operation unit and a first multi-scale convolution module which are connected in sequence; The first convolution module includes a 3D convolution layer with a convolution kernel size of 3, a BN layer and a ReLU layer connected in sequence; The first multi-scale convolution module includes a parallel convolution layer, a first concatenation layer, and a second convolution module connected in sequence, and the parallel convolution layer includes a third convolution module, a fourth convolution module, and a fifth convolution module connected in parallel; The first convolution module, the second convolution module and the third convolution module have the same network structure; the fourth convolution module includes a 3D convolution layer with a convolution kernel size of 5, a BN layer and a Relu layer connected in sequence; the fifth convolution module includes a 3D convolution layer with a convolution kernel size of 7, a BN layer and a Relu layer connected in sequence; The first basic operation unit, the second basic operation unit, the third basic operation unit, and the fourth basic operation unit have the same network structure, each basic operation unit includes a second multi-scale convolution module and a sixth convolution module connected in sequence, and the input and output of the second multi-scale convolution module are input to the sixth convolution module after tensor addition; wherein the sixth convolution module has the same network structure as the first convolution module, and the second multi-scale convolution module has the same network structure as the first multi-scale convolution module; The first decoder network includes a third multi-scale convolution module, a first upsampling module, a second splicing layer, a fourth multi-scale convolution module, a second upsampling module, a third splicing layer, a fifth multi-scale convolution module, a first spatial target recognition decoupling head module, a second spatial target recognition decoupling head module, and a third spatial target recognition decoupling head module; The third multi-scale convolution module, the fourth multi-scale convolution module, and the fifth multi-scale convolution module have the same network structure as the first multi-scale convolution module; The input end of the third multi-scale convolution module is connected to the output end of the first multi-scale convolution module of the backbone network, and the output end of the third multi-scale convolution module is connected to the input end of the third spatial target recognition decoupling head module; The third multi-scale convolution module, the first up-sampling module, the second splicing layer, the fourth multi-scale convolution module, the second up-sampling module, the third splicing layer, the fifth multi-scale convolution module and the first spatial target recognition decoupling head module are connected in sequence; The output end of the second basic operation unit is connected to the input end of the third splicing layer; The output end of the third basic operation unit is connected to the input end of the second splicing layer; The output end of the fourth multi-scale convolution module is connected to the input end of the second space target recognition decoupling head module; The first space target recognition decoupling head module, the second space target recognition decoupling head module and the third space target recognition decoupling head module are used to predict and output target area position information, size information, category information and confidence.

7. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Image processing method and related equipment

    CN110415291A

  • Image detection method and device, computer equipment and storage medium

    CN112967235A

  • Pedicle screw implantation operation path planning method, device and equipment

    CN114358388A

  • Lumbar vertebra CT image space positioning method based on artificial neural network

    CN114913160A

  • Intelligent planning method, device and equipment for puncture path in PVP (Polyvinyl Pyrrolidone) operation and medium

    CN116616876A