Artificial Intelligence-based Automatic Vertebral Fracture Recognition and Analysis System
Through image preprocessing and deep learning network annotation by artificial intelligence technology, combined with bimodal network and text semantic model, the accuracy and computing efficiency of vertebral fracture diagnosis are solved, and a simple and efficient fracture severity assessment and treatment recommendations are achieved.
Patent Information
- Application Number
- CN202411691245.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-11-25
AI Technical Summary
The accuracy of the prior art in the diagnosis of vertebral compression fractures depends on doctor's experience. It is affected by the accumulation of abdominal air and feces in elderly patients, and the calculation of 3D image processing is expensive and the accuracy is difficult to take into account. The existing solutions have problems of interference and low computing performance.
Using an automatic vertebral fracture recognition and analysis system based on artificial intelligence, including image rotation enhancement, affine scaling and deep learning network annotation, a bimodal network and text semantic model is built, and fracture information recognition and severity analysis are performed through the combination of images and text.
It improves the accuracy and efficiency of vertebral fracture diagnosis, reduces calculation consumption, provides concise fracture severity assessment and treatment recommendations, and is suitable for mobile applications.
Smart Images

Figure CN119205741B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and more specifically, to an automatic vertebral fracture recognition and analysis system based on artificial intelligence. Background Art
[0002] With the popularization of informatization, especially with the gradual development of artificial intelligence technology, people's use of technology products is becoming more and more extensive. Especially in the medical industry, how to use artificial intelligence technology to improve the level of medical services has become particularly important, and among them, artificial intelligence technology based on deep learning has received more and more attention and extensive use.
[0003] Vertebral compression fracture is one of the most common fracture types of osteoporotic fractures. The conventional preliminary diagnosis method relies on taking anteroposterior and lateral X-rays of the spine, and radiologists visually diagnose. The accuracy rate is closely related to the experience of radiologists. Doctors mainly make diagnoses based on the Genant visual semi-quantitative determination method. However, due to the accumulation of gas and feces in the abdomen of elderly patients, it seriously affects the accurate determination of the degree of fracture compression by doctors, thus affecting the formulation of subsequent treatment plans. In the prior art, most often, after normalizing or binarizing the input data, subsequent analysis is carried out, which brings additional interference to subsequent analysis, and the running time of most schemes is more than several minutes.
[0004] For example, Chinese Patent CN114937502A discloses an osteoporosis vertebral compression fracture evaluation method and system based on deep learning, which uses a detection model and a segmentation model for analysis, and performs a binarization operation on the input image before analysis. Since the CT images taken are not fixed, using a single threshold for binarization brings interference to subsequent analysis, and the two models are processed serially, reducing the computing performance.
[0005] At the same time, the CT multi-view dataset is open-sourced in the paper "Automatic L3 slice detection in 3D CT images using fully-convolutional networks", which proposes a deep learning scheme based on Unet for the recognition and analysis of 3D images. Although the fully convolutional network can effectively process medical images, such as CT images.
[0006] However, 3D image processing itself increases the computational consumption. At the same time, converting 3D images to 2D images will cause loss of image information, and with the influence of the environment such as angles and light, it will bring interference between image context information. Therefore, although the fully convolutional network can adapt to different input image sizes, the classification accuracy and positioning accuracy cannot be both achieved. When the receptive field is selected to be relatively large, the dimensionality reduction multiple of the corresponding pooling layer at the back.
[0007] In view of the problems in the related art, no effective solution has been proposed yet. Summary of the Invention
[0008] In view of the problems in the related art, the present invention proposes an automatic vertebral fracture recognition and analysis system based on artificial intelligence to overcome the above-mentioned technical problems existing in the existing related art.
[0009] Therefore, the specific technical solution adopted by the present invention is as follows:
[0010] An automatic vertebral fracture recognition and analysis system based on artificial intelligence, the vertebral fracture automatic recognition and analysis system includes:
[0011] A scanned image processing and annotation unit, configured to perform regulation processing operations on the tomographic scan image, and perform answer annotation based on the processed tomographic scan image;
[0012] A vertebral fracture automatic recognition unit, configured to construct an automatic recognition network structure, train the automatic recognition network structure using the tomographic scan image and the answer annotation, and recognize fracture information using the trained automatic recognition network structure;
[0013] A fracture information recognition and judgment unit, configured to analyze the severity of the fracture according to the fracture information, evaluate the impact degree of the fracture using the severity analysis result, and transmit the severity analysis result and the fracture impact degree to the mobile terminal.
[0014] Preferably, the scanned image processing and annotation unit includes:
[0015] An image rotation and enhancement module, configured to perform rotation processing on the tomographic scan image using image rotation technology, and perform image enhancement operation after rotation is completed;
[0016] An image affine scaling module, configured to perform affine transformation processing on the tomographic scan image, and perform scaling processing on the tomographic scan image using bilinear interpolation technology after affine transformation;
[0017] A scanned image annotation module, configured to perform answer annotation processing on the processed tomographic scan image according to the input form of the deep learning network.
[0018] Preferably, performing rotation processing on the tomographic scan image using image rotation technology and performing image enhancement operation after rotation is completed includes:
[0019] Performing counterclockwise rotation on the tomographic scan image using image rotation technology according to a predefined rotation point, and stopping rotation after rotating to a preset angle;
[0020] Combining the initial coordinate position of the tomographic scan image with the predefined rotation point to obtain the coordinate position of the rotated tomographic scan image;
[0021] Evaluate the image quality of the rotated tomographic scan image, and fuse the evaluation result with the noise generation technology to add noise to the tomographic scan image;
[0022] Perform an inverse transformation on the tomographic scan image with noise added to complete the noise removal operation, and obtain the enhanced tomographic scan image.
[0023] Preferably, perform an affine transformation on the tomographic scan image, and use bilinear interpolation technology to scale the tomographic scan image after the affine transformation, including:
[0024] Construct a transformation matrix after determining the type and parameters of the affine transformation based on the tomographic scan image, and define the mapping position of the tomographic scan image using the transformation matrix;
[0025] Perform an affine transformation on the tomographic scan image using the transformation matrix relationship to obtain the mapping coordinate matrix of the transformed tomographic scan image;
[0026] Enlarge or scale the tomographic scan image based on bilinear interpolation technology and the tomographic scan image in the longitudinal and transverse directions.
[0027] Preferably, perform answer annotation processing on the processed tomographic scan image according to the input form of the deep learning network, including:
[0028] Based on the rectangle annotation technology, annotate the main area frame of the vertebral body in the tomographic scan image to obtain the complete vertebral body area, and perform secondary annotation within the complete vertebral body area using pixel-level annotation to obtain the vertebral body sub-area;
[0029] Number the vertebral body sub-areas in the annotation order, and perform classification information calibration on the vertebral body sub-areas after the numbering is completed to judge the fracture condition of the vertebral body;
[0030] Obtain the personal information, injury mechanism and fracture history of the person corresponding to the tomographic scan image according to the image acquisition information, and combine with the expert annotation technology to perform fracture status evaluation and annotation on the corresponding tomographic scan image;
[0031] Integrate the personal information, injury mechanism, fracture history and fracture status evaluation annotation results of the corresponding person into an answer annotation text.
[0032] Preferably, the vertebral fracture automatic recognition unit includes:
[0033] A network structure construction module, used to construct a bimodal network and a text semantic large model structure, and obtain an automatic recognition network structure based on the construction result;
[0034] A structure parameter configuration module, used to configure the hyperparameters of the automatic recognition network structure and the training file;
[0035] An identification network structure training module for inputting tomographic images and answer annotations into an automatic identification network structure to perform network training and verify the evaluation accuracy rate of the output vertebral region fracture information;
[0036] A network structure adjustment module for adjusting the automatic identification network structure based on the evaluation accuracy rate result.
[0037] Preferably, constructing a dual-modal network and a text semantic large model structure, and obtaining the automatic identification network structure based on the construction result includes:
[0038] Selecting a deep convolutional neural network model as the basic network structure for feature extraction, setting scale feature layers in combination with a pyramid structure, and performing convolution and deconvolution operations on each scale feature layer;
[0039] Performing feature fusion operations after convolution, and defining the branch functions of each scale feature layer to obtain a dual-modal network structure including several groups of branch functions;
[0040] Based on the output result of the dual-modal network structure and the text requirements, selecting the basic structure of the text network, and splicing a long short-term memory network after performing channel pruning operations on the network basic structure to obtain a text semantic network structure;
[0041] Fusing the dual-modal network structure and the text semantic network structure to obtain the automatic identification network structure.
[0042] Preferably, inputting tomographic images and answer annotations into the automatic identification network structure to perform network training, and verifying the evaluation accuracy rate of the output vertebral region fracture information includes:
[0043] After performing word vector conversion on the answer annotation text, combining the tomographic images and dividing the training set and the validation set according to a preset ratio, and inputting the training set into the automatic identification network structure;
[0044] Using a learning rate adjustment strategy to set an initial learning rate to start the preliminary training of the automatic identification network structure, and increasing the learning rate for normal training after reaching the number of traversals;
[0045] Judging the stability of the loss function during the training process of the automatic identification network structure, and selecting the structure that meets the number of traversal requirements as the automatically identified network structure that has completed training after the loss function is stable;
[0046] Inputting the validation set into the automatically identified network structure that has completed training, and using the branch function to output the detection frame of the vertebral body main region, the number of segmented pixels in the vertebral sub-region, and the fracture classification information;
[0047] Comparing the output information with the answer annotation text, and verifying the evaluation accuracy rate of the output vertebral region fracture information based on the comparison result.
[0048] Preferably, comparing the output information with the answer annotation text, and verifying the evaluation accuracy of the output vertebral body region fracture information based on the comparison result includes:
[0049] Judging the coincidence degree between the detection frame and the corresponding main body region frame in the answer annotation text, and at the same time calculating the percentage between the number of segmented pixels in the vertebral body sub-region and the corresponding pixel-level annotation result in the answer annotation text;
[0050] Comparing the fracture classification information with the classification information calibration structure in the answer annotation text, and at the same time calculating the distance between each vertebral body sub-region based on the pixel statistics of the body sub-region;
[0051] Fusing the comparison result, the coincidence degree judgment result and the percentage calculation result to analyze the output accuracy of the trained automatic recognition network structure.
[0052] Preferably, the fracture information recognition and judgment unit includes:
[0053] A fracture severity analysis module, which is used to analyze the size and position of the fracture according to the fracture information, and judge the severity of the vertebral fracture by analyzing the distance between each vertebral joint and the adjacent sub-joints;
[0054] A fracture impact degree analysis module, which is used to input the severity of the vertebral fracture and the answer annotation text into the automatic recognition network structure to obtain the fracture impact degree;
[0055] A mobile terminal display module, which is used to convert the automatic recognition network structure into a mobile terminal computing framework by using the quantization compression processing technology, and then upload the severity analysis result and the fracture impact degree to the mobile terminal.
[0056] The beneficial effects of the present invention are:
[0057] The present invention proposes an automatic recognition and analysis system for vertebral fractures based on artificial intelligence through artificial intelligence technology. Through artificial intelligence algorithms, it analyzes the size, position and type of fractures to judge the severity, evaluates the possible impact on patients, and gives subsequent treatment suggestions. The actual operation is convenient and simple, and has important promotion and application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.
[0059] Figure 1 It is a principle block diagram of an automatic recognition and analysis system for vertebral fractures based on artificial intelligence according to an embodiment of the present invention.
[0060] Figure 2 It is a schematic flowchart in the artificial intelligence-based automatic vertebral fracture recognition and analysis system according to an embodiment of the present invention.
[0061] Figure 3 It is a flowchart of the dual-modal large model network structure in the artificial intelligence-based automatic vertebral fracture recognition and analysis system according to an embodiment of the present invention.
[0062] Figure 4 It is a flowchart of the text semantic large model network structure in the artificial intelligence-based automatic vertebral fracture recognition and analysis system according to an embodiment of the present invention.
[0063] In the figure:
[0064] 1. Scanning image processing and annotation unit; 2. Automatic vertebral fracture recognition unit; 3. Fracture information recognition and judgment unit. Detailed implementation manners
[0065] To make the above objects, features and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0066] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0067] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure or characteristic that can be included in at least one implementation manner of the present invention. The "in one embodiment" that appears in different places in this specification does not all refer to the same embodiment, nor is it an embodiment that is separately or selectively mutually exclusive with other embodiments.
[0068] The present invention is described in detail in conjunction with the schematic diagrams. When describing the embodiments of the present invention in detail, for the convenience of explanation, the cross-sectional views showing the device structure will be enlarged locally out of the general scale, and the schematic diagrams are only examples and should not limit the protection scope of the present invention here. In addition, in actual production, three-dimensional spatial dimensions including length, width and depth should be included.
[0069] Meanwhile, in the description of the present invention, it should be noted that the orientation or positional relationship indicated by terms such as "upper, lower, inner, and outer" is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation to the present invention. In addition, the terms "first, second, or third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0070] Unless otherwise clearly defined and limited in the present invention, the terms "installation, connection, and coupling" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can also be a mechanical connection, an electrical connection, or a direct connection, and can also be indirectly connected through an intermediate medium, or can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0071] Example 1, as Figures 1-4 shown, according to the artificial intelligence-based automatic vertebral fracture recognition and analysis system of the embodiment of the present invention, the vertebral fracture automatic recognition and analysis system includes:
[0072] A scan image processing and annotation unit 1, configured to perform regulation processing operations on tomographic scan images, and perform answer annotation based on the processed tomographic scan images.
[0073] It should be noted that the regulation processing of the patient's CT images (tomographic scans) collected by the mobile terminal is to make the images sent to the deep learning network as rich as possible. In actual use, due to the interference of environmental factors such as different backgrounds, different illuminations, and different angles when the mobile app captures CT images, the subsequent analysis effect is not ideal. Therefore, in order to better robust the images captured in different environments and have a better analysis effect, it is necessary to perform preprocessing operations on the CT images collected by the mobile app, including but not limited to: image rotation, image noise, image affine transformation, image scaling, etc.
[0074] At the same time, the processed images are annotated with answers, including the vertebral body main region frame (rectangular frame) in the CT image, the mask annotation (pixel-level annotation) of each sub-region of the vertebra, information such as whether there is a fracture, the patient's weight, age, injury mechanism, and fracture history, and according to the fracture severity level, the corresponding healing time, the stability of the fracture site, the risk of secondary fracture, etc. are calibrated.
[0075] In this embodiment, the scan image processing and annotation unit 1 includes:
[0076] An image rotation enhancement module, which is used to perform rotation processing on tomographic images using image rotation technology and implement image enhancement operations after the rotation is completed;
[0077] An image affine scaling module, which is used to perform affine transformation processing on tomographic images and implement scaling processing on tomographic images using bilinear interpolation technology after the affine transformation;
[0078] A scanned image annotation module, which is used to perform answer annotation processing on the processed tomographic images according to the input form of the deep learning network.
[0079] Specifically, when performing rotation processing on tomographic images using image rotation technology and implementing image enhancement operations after the rotation is completed, the image rotation technology can be used to rotate the tomographic images counterclockwise according to predefined rotation points and stop rotating after rotating to a preset angle; combine the initial coordinate position of the tomographic images with the predefined rotation points to obtain the coordinate position of the rotated tomographic images; evaluate the image quality of the rotated tomographic images, and fuse the evaluation results with noise generation technology to add noise to the tomographic images; perform inverse transformation on the tomographic images with noise added to complete the noise removal operation, and obtain the enhanced tomographic images.
[0080] It should be explained that image rotation refers to rotating the defined tomographic images around a certain point in the counterclockwise or clockwise direction by a certain angle, usually rotating counterclockwise around the center of the tomographic images.
[0081] Assume that the upper left corner of the tomographic image is (left, top) and the lower right corner is (right, bottom), then any point (x0, y0) on the tomographic image rotates counterclockwise around its center (x center , y center ) by an angle of angle, and the calculation formula for the new coordinate position (x′, y′) is:
[0082] ;
[0083] ;
[0084] Among them,
[0085] ; ;
[0086] In the formula, θ represents the angle of counterclockwise rotation of the tomographic images, and in order to better robust the scenarios in actual use, the rotation angle of the tomographic images is controlled within plus or minus 15 degrees.
[0087] Moreover, during the actual shooting process, various factors often lead to various noise points in the captured tomographic images. Therefore, in order to be robust to different noise environments, it is necessary to add and remove image noise from the captured tomographic images.
[0088] Among them, during the process of generating image noise, multiple types of noise are randomly combined (including but not limited to: Gaussian noise, salt-and-pepper noise, Gaussian filtering, mean filtering, gamma noise, etc.). Image noise removal means performing an inverse transformation on the image with added noise to obtain the image after noise removal.
[0089] Specifically, when performing affine transformation processing on the tomographic image and using bilinear interpolation technology to scale the tomographic image after affine transformation, the type and parameters of the affine transformation can be determined based on the tomographic image to construct a transformation matrix, and the mapping position of the tomographic image is defined using the transformation matrix; the mapping coordinate matrix of the transformed tomographic image is obtained by performing an affine transformation on the tomographic image using the transformation matrix relationship; the tomographic image is enlarged or scaled based on the bilinear interpolation technology and the tomographic image in the longitudinal and transverse directions.
[0090] It should be explained that an affine transformation, also known as an affine mapping, refers to a linear transformation in a vector space followed by a translation in geometry, transforming it into another vector space. An affine transformation can maintain the "flatness" of the image, including rotation, scaling, translation, and shearing operations. A common affine transformation matrix is a 2*3 matrix, that is, two rows and three columns. The elements in the third column play a role in translation, and the numbers on the diagonal of the first two columns are for scaling, and the rest are for rotation or shearing. The transformation matrix relationship is as follows:
[0091] ;
[0092] Among them, A represents the transformed coordinate matrix, B represents the affine transformation matrix, and C represents the original coordinate matrix. There are 6 unknowns. Assuming that the tomographic image rotates clockwise by σ radians around the axis (m, n) to the target image, the variables corresponding to the transformation matrix are:
[0093] ;
[0094] Among them, the 4 unknown variables a, b, d, and e in the first two columns play a role in rotation, and the 2 unknown variables c and f in the third column play a role in translation.
[0095] Meanwhile, in actual use, due to various differences in the distance of the captured images, the actual size of the vertebral body in the tomographic scan images is not fixed. The bilinear interpolation method is used to simulate the sizes of different vertebral body regions. That is, through the bilinear interpolation method, each tomographic scan image is randomly transformed, enlarged or reduced, so as to make the training data more abundant. The specific implementation of the bilinear method is as follows:
[0096] Assume the value of function g at point P = (i, j), and the values of function g at points Q 11 = (i1, j1), Q 12 = (i1, j2), Q 21 = (i2, j1), Q 22 = (i2, j2) are known. Then R1 and R2 are the interpolations in the i (horizontal) direction respectively.
[0097] First, perform linear interpolation in the i direction to obtain:
[0098] ;
[0099] ;
[0100] ;
[0101] ;
[0102] Then, perform linear interpolation in the j direction (vertical) to obtain:
[0103]
[0104] Specifically, when performing answer annotation processing on the processed tomographic scan images according to the input form of the deep learning network, the main region box of the vertebral body in the tomographic scan images can be marked based on the rectangle marking technology to obtain the complete vertebral body region, and secondary marking can be performed within the complete vertebral body region using pixel-level marking to obtain the vertebral body sub-region; the vertebral body sub-regions are numbered digitally in the marking order, and after the numbering is completed, a classification information calibration operation is performed on the vertebral body sub-regions to judge the fracture condition of the vertebral body; personal information, injury mechanism and fracture history of the person corresponding to the tomographic scan image are obtained according to the image acquisition information, and combined with the expert marking technology, the fracture state evaluation marking is performed on the corresponding tomographic scan image; the personal information, injury mechanism, fracture history and fracture state evaluation marking results of the corresponding person are integrated into the answer annotation text.
[0105] It should be noted that according to the data format required by the deep learning network, the tomographic images need to be annotated. Specifically, the main body area box of the vertebral body in the CT image needs to be annotated, that is, the circumscribed rectangle of the main body part of the vertebral body in the image. The circumscribed rectangle needs to completely include the vertebral body area. At the same time, in order to better analyze the fracture information, the tomographic images also need to be annotated with masks for each sub-region of the vertebral body (pixel-level annotation), that is, each vertebral body part needs to be annotated. For example, the area where vertebral body 1 is located is annotated as 1, the area where vertebral body 2 is located is annotated as 2, and so on for other sub-regions.
[0106] At the same time, classification information calibration needs to be carried out for all tomographic images, that is, to judge whether there is a fracture in the current tomographic image. If there is, it is calibrated as mild fracture, moderate fracture, severe fracture. If not, it is calibrated as none, and other situations are all marked as unknown. Based on the calibrated weight, age, injury mechanism, and fracture history of the patient corresponding to the current CT image, the specific injury mechanisms mainly include falling, fighting, impact, etc., and the fracture history is mainly whether there is a fracture within 3 months, whether there is a fracture within 3 to 6 months, whether there is a fracture within 6 to 9 months, whether there is a fracture within 9 to 12 months, etc.
[0107] For the fracture severity level, the corresponding healing time, the stability of the fracture site, and the risk of secondary fracture are calibrated. Specifically, a chief physician with rich clinical experience can calibrate each tomographic image, and based on personal experience, give a preliminary evaluation answer for the current CT image, including the healing time (1 to 3 months, 3 to 6 months, 6 to 9 months, 9 to 12 months), the stability of the fracture site (extremely unstable, stable, mildly stable), and the risk of secondary fracture (high, medium, low).
[0108] The vertebral fracture automatic recognition unit 2 is used to construct an automatic recognition network structure, and use the tomographic images and answer annotations to train the automatic recognition network structure, and use the trained automatic recognition network structure to recognize fracture information.
[0109] In this embodiment, the vertebral fracture automatic recognition unit 2 includes:
[0110] The network structure building module is used to construct a dual-modal network and a text semantic large model structure, and obtain an automatic recognition network structure based on the construction results;
[0111] The structure parameter configuration module is used to configure the hyperparameters of the automatic recognition network structure and the training file;
[0112] The dual-modal network training module is used to input the tomographic images and answer annotations into the automatic recognition network structure to perform network training and verify the evaluation accuracy of the fracture information in the vertebral body area output;
[0113] The network structure adjustment module is used to adjust the automatic recognition network structure based on the evaluation accuracy result.
[0114] Specifically, when constructing the dual-modal network and the text semantic large model structure and obtaining the automatic recognition network structure based on the construction result, a deep convolutional neural network model can be selected as the basic network structure for feature extraction, a pyramid structure is combined to set scale feature layers, and convolution and deconvolution operations are performed on each scale feature layer; after convolution, a feature fusion operation is performed, and the branch functions of each scale feature layer are defined to obtain a dual-modal network structure including several groups of branch functions; based on the output result of the dual-modal network structure and the text requirements, a text network basic structure is selected, and after performing channel pruning operations on the network basic structure, a long short-term memory network is spliced to obtain a text semantic network structure; the dual-modal network structure and the text semantic network structure are fused to obtain the automatic recognition network structure.
[0115] It should be explained that the dual-modal network structure adopts a multi-modal structure with text and image input at the same time. The output result of the network includes detection boxes, segmentation mask maps, classification results, etc. at the same time, which is a multi-task network structure. Specifically, this network structure includes three branches: a detection branch, a segmentation branch, and a classification branch.
[0116] In order to be suitable for mobile operation, an improved ResNet18 is used as the backbone (basic network structure), and at the same time, a multi-layer pyramid structure is adopted, which respectively includes multiple scales such as 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32, etc. Among them, the feature layers of the 1 / 8, 1 / 16, and 1 / 32 scales perform multiple convolution and deconvolution operations and then perform feature fusion as the detection branch; the feature layers of the 1 / 2, 1 / 4, and 1 / 8 scales perform multiple convolution and deconvolution operations and then perform feature fusion as the segmentation branch; after the 1 / 32 scale, multiple convolutional layers, pooling layers, regularization layers, etc. are spliced as the classification branch.
[0117] Input the text information into the network structure, and output the fracture severity, healing time, stability of the fracture site, and risk of secondary fracture in each tomographic scan image through big data modeling.
[0118] First, modify the network structure based on the original inception-v3, including deleting some channels and some convolutional layers to improve the network operation speed. Secondly, splice a bidirectional LSTM structure. During the modeling process, convert the summarized text information into word vectors and send them into the constructed network structure for training. Specifically, the text information includes information such as fracture location, fracture degree, treatment suggestions, patient age, weight, fracture history, fracture mechanism, healing time, stability of the fracture site, and risk of secondary fracture.
[0119] Specifically, when inputting the tomographic scan image and the answer annotation into the automatic recognition network structure for network training and verifying the evaluation accuracy rate of the vertebral body region fracture information, the answer annotation text can be subjected to word vector conversion and then combined with the tomographic scan image to divide the training set and the verification set according to a preset ratio, and the training set is input into the automatic recognition network structure; the learning rate adjustment strategy is used to set the initial learning rate to start the preliminary training of the automatic recognition network structure, and the learning rate is increased after reaching the number of traversals for normal training; during the training process of the automatic recognition network structure, the stability of the loss function is judged, and after the loss function is stable, the structure that meets the requirements of the number of traversals is selected as the automatically recognized network structure that has completed training; the verification set is input into the automatically recognized network structure that has completed training, and the detection box of the vertebral body main region, the number of segmented pixels of the vertebral body sub-region, and the fracture classification information are output by using the branch function; the output information is compared with the answer annotation text, and the evaluation accuracy rate of the output vertebral body region fracture information is verified based on the comparison result.
[0120] It should be explained that the hyperparameters in the network structure are set according to the initial framework of the network structure, including the size of the input tomographic scan image, the number of channels of each convolutional layer in the backbone, the size of the convolutional kernel, and the convolutional layer settings between multiple scales such as 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32.
[0121] Among them, the size of the input tomographic scan image is set to 960*960, the number of channels of the first four convolutional layers in the backbone is 64, and the size of the feature layer is 960*960; the number of channels of the fifth to eighth convolutional layers is 128, and the size of the feature layer is 480*480; the number of channels of the ninth to twelfth convolutional layers is 256, and the size of the feature layer is 240*240; the number of channels of the thirteenth to sixteenth convolutional layers is 256, and the size of the feature layer is 120*120; the number of channels of the seventeenth to twentieth convolutional layers is 512, and the size of the feature layer is 60*60; the size of the convolutional kernel of all convolutional layers is 3*3. At each scale transformation layer, the stride is 2 to achieve downsampling, and a BN layer, a Scale layer, and a Relu layer are connected after each convolutional layer. At the place where multiple scales are fused, a Concat layer is used to merge channels.
[0122] To enable the network model to converge better and obtain a stable model, it is necessary to configure hyperparameters for the answer annotation text during the training process. Among them, the optimizer uses SGD (Stochastic Gradient Descent), the learning rate uses CosineAnnealingLR (Cosine Annealing to adjust the learning rate), Tmax (maximum number of iterations) uses 10, the base learning rate is 0.0001, the number of epochs (number of traversals) is 300, 8 T4 graphics cards are used for parallel training, the batch size for each graphics card is 64, and the loss functions are BCE loss, CIOU loss, CrossEntropy loss, etc.
[0123] The calibrated text information is converted through word vectors and sent into the built network structure together with the image for training. The text information here includes information such as fracture conditions, fracture regions, each sub-module region of the spine, the front and back heights of each region, and its own height for network model training.
[0124] At the beginning of training, the warmup strategy is first adopted (the model is relatively unstable at the beginning of training, and the initial learning rate should be set very low, which can ensure that the network has good convergence. However, the low learning rate will make the training process very slow. Therefore, the "warm-up" stage of network training is achieved by gradually increasing the low learning rate to a higher learning rate, which is called warmup). The network is initially trained with a small learning rate and normal training starts after 20 epochs.
[0125] Specifically, the output information is compared with the answer annotation text, and based on the comparison result, the evaluation accuracy of the output vertebral body region fracture information is verified. The overlap degree between the detection box and the corresponding main region box in the answer annotation text can be judged, and at the same time, the percentage between the number of segmented pixels in the vertebral body sub-region and the corresponding pixel-level annotation result in the answer annotation text is calculated; the fracture classification information is compared with the classification information calibration structure in the answer annotation text, and at the same time, the distance between each vertebral body sub-region is statistically calculated based on the pixels of the body sub-region; the comparison result, the overlap degree judgment result, and the percentage calculation result are fused to analyze the output accuracy of the trained automatic recognition network structure.
[0126] When the loss stabilizes, select the model after 300 epochs and verify its performance on the test set (the data in this test set was not involved in model training). This includes the overlap between the vertebral body region bounding box and the calibration answer (i.e., the overlap between the detection box of the vertebral body region output by the model and the box of this region manually calibrated in the image, which is the IOU between the detection box predicted by the model and the manually calibrated answer box. The higher the overlap, the better the model performance; the lower the overlap, the worse the model performance). The overlap thresholds are 0.3, 0.4, 0.5, and 0.6 respectively. Test the accuracy of the detection box under different overlap thresholds, and check the effect of the segmentation branch. This is mainly to calculate the percentage between the number of pixel values effectively segmented within each sub-network structure and the calibration answer (for example, there are N1 pixel values in the current region, and these N1 pixel points are the manually calibrated answer. The values predicted by the network structure are compared with these 100. If the network structure predicts N2 values, then the percentage is N2 / N1). For the classification branch of whether there is a fracture, the network structure outputs two-dimensional information in this branch: there is a fracture and there is no fracture. Directly compare the result with the calibration answer. If the result output by the network structure is consistent with the calibrated answer, it is considered that the analysis of this image is correct; if not, it is considered that the network structure analysis is incorrect. Statistically calculate the final accuracy. Additionally, statistically calculate the distance between each sub-region and compare it with the accuracy of the calibration answer, and statistically calculate the percentage.
[0127] Convert the summarized text information into word vectors and send them into the built network structure for training. The text here includes information such as fracture location, fracture degree, treatment suggestions, patient age, weight, fracture history, fracture mechanism, healing time, stability of the fracture site, and risk of secondary fracture. After training the network structure, the network structure can output the severity of the fracture, healing time, risk of secondary fracture, and related treatment suggestions in the current CT image.
[0128] The fracture information recognition and judgment unit 3 is used to analyze the severity of the fracture based on the fracture information, evaluate the impact degree of the fracture using the severity analysis result, and transmit the severity analysis result and the impact degree of the fracture to the mobile terminal.
[0129] In this embodiment, the fracture information recognition and judgment unit 3 includes:
[0130] The fracture severity analysis module is used to analyze the size and location of the fracture based on the fracture information, and judge the severity of the vertebral fracture by analyzing the distance between each vertebral joint and adjacent sub-joints;
[0131] The fracture impact degree analysis module is used to input the severity of the vertebral fracture and the answer annotation text into the automatic recognition network structure to obtain the fracture impact degree;
[0132] The mobile display module is used to upload the severity analysis result and the fracture impact degree to the mobile end after converting the automatic recognition network structure into a mobile computing framework by using the quantization compression processing technology.
[0133] Specifically, in order to perform CT image recognition more conveniently and apply the network structure to the mobile end, the network structure has been designed for lightweight during construction, such as the number of channels and the number of network layers, etc. During use, the trained network model is converted into an int8 model of ncnn, which improves the running speed while ensuring the inference accuracy. Based on the Java language, component code based on the mobile end is developed, which can take pictures and display the effects on the mobile end, including whether there is a fracture, height loss, and treatment suggestions.
[0134] It should be explained that based on the relevant output information of the network structure, including the vertebral body area, the mask maps of each sub-area, and the corresponding attribute classification results, and processing the output information, the corresponding fracture degree can be obtained.
[0135] Specifically as follows:
[0136] Through the vertebral body area box, the vertebral body area image in the CT image can be obtained. During the training process, this area box not only helps the network better focus on the vertebral body area and prevent interference from other areas, but also helps the regression of the mask of each sub-joint. Secondly, on the obtained vertebral body area image and the mask output by the network model, the mask of each vertebral body sub-area is obtained, such as the complete area of the first vertebral body and the complete area of the second vertebral body. Here, the complete area refers to the pixel-level area of each vertebral body.
[0137] Based on the above results, all the segmentation results of each sub-joint in the whole vertebral body are obtained, that is, the segmentation maps of each sub-area of the vertebral body. The size of each sub-joint or vertebra of a normal person's vertebral body or the distance between adjacent vertebrae remains unchanged. According to the mask maps of each sub-joint in the network output, the distance between each sub-joint or vertebra and its adjacent sub-joint is calculated, such as the distance between the second and the first, and the distance between the second and the third. The distance is the height.
[0138] By analogy, calculate the front-to-back height between other sub-joints, and the front, middle, and back heights of each sub-joint in the entire vertebral body can be obtained. Here, the front height refers to the distance between the current sub-joint and the previous sub-joint, the middle height refers to the height of each sub-joint, and the back height refers to the distance between the current sub-joint and the next sub-joint. Then the fracture degree of the current vertebral body can be obtained. At the same time, the attribute classification information output by the network structure is used to further weight and confirm the fracture degree of the vertebral body. The classification entries include: unknown, mild fracture, moderate fracture, and severe fracture. Weight the attribute classification result and the determination result to output the final vertebral fracture situation. Here, the weighting means that if the score of the attribute classification result exceeds 0.95, the result is considered correct. At the same time, combined with the determination results of each sub-joint mask, the final result is given. Through the analysis of the vertebral fracture degree, corresponding treatment suggestions can be given.
[0139] Specifically, the vertebral fracture degree is divided into mild fracture, moderate fracture, and severe fracture, and the discrimination criteria are as follows:
[0140] Mild fracture: Compared with adjacent or the same vertebra, the front, middle, and back heights of the vertebra decrease by less than 25%; Moderate fracture: Compared with adjacent or the same vertebra, the front, middle, and back heights of the vertebra decrease by 25%-40%; Severe fracture: Compared with adjacent or the same vertebra, the front, middle, and back heights of the vertebra decrease by more than 40%.
[0141] Based on the fracture situation assessment and treatment suggestions, combined with the relevant information of the current patient himself: age, weight, fracture history, etc., all the information is sent into the network structure, and more comprehensive information can be obtained.
[0142] When it is a mild fracture, the healing period is 1-3 months, and the risk of secondary fracture is low. The treatment suggestion is: bed rest, conservative treatment.
[0143] When it is a moderate fracture, the healing period is 3-6 months, and the risk of secondary fracture is medium. The treatment suggestion is: vertebroplasty, minimally invasive surgery.
[0144] When it is a severe fracture, the healing period is 3-6 months, and the risk of secondary fracture is high. The treatment suggestion is: vertebral orthosis, open surgery.
[0145] In summary, by means of the above technical solutions of the present invention, the present invention proposes an automatic recognition and analysis system for vertebral fractures based on artificial intelligence through artificial intelligence technology. It analyzes the size, position, and type of fractures through artificial intelligence algorithms to judge the severity, evaluate the possible impact on patients, and give subsequent treatment suggestions. The actual operation is convenient and simple, and it has important promotion and application value.
[0146] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. The automatic identification and analysis system of vertebral fracture based on artificial intelligence is characterized by: include, A scanning image processing and annotation unit, used for performing a control and processing operation on the tomographic scanning image, and annotating the answer based on the tomographic scanning image after the processing; Vertebral fracture automatic recognition unit, used to construct an automatic recognition network structure, train the automatic recognition network structure using tomographic images and answer annotations, and recognize fracture information using the trained automatic recognition network structure; A fracture information recognition and judgment unit is used to analyze the severity of the fracture according to the fracture information, evaluate the impact of the fracture using the severity analysis result, and transmit the severity analysis result and the impact of the fracture to the mobile terminal; The scanned image processing and annotation unit comprises: An image rotation enhancement module is used to perform rotation processing on the tomographic image using image rotation technology and implement image enhancement operations after the rotation is completed; An image affine scaling module is used to perform affine transformation processing on the tomographic image and to scale the tomographic image using bilinear interpolation technology after the affine transformation; A scanning image annotation module is used to annotate the processed tomographic scan image with answers according to the input form of the deep learning network; The method of performing rotation processing on the tomographic image by using the image rotation technology and performing image enhancement operation after the rotation is completed includes: The tomographic image is rotated counterclockwise according to a predefined rotation point using an image rotation technology, and the rotation stops after reaching a preset angle; Combining the initial coordinate position of the tomographic image with the predefined rotation point to obtain the coordinate position of the rotated tomographic image; evaluating the image quality of the rotated tomographic image and fusing the evaluation result with a noise generation technique to add noise to the tomographic image; Performing an inverse transformation on the tomographic image after noise addition to complete the noise removal operation, thereby obtaining an enhanced tomographic image; The affine transformation of the tomographic image and the scaling of the tomographic image using bilinear interpolation technology after the affine transformation include: After determining the type and parameters of the affine transformation based on the tomographic image, a transformation matrix is constructed, and a mapping position of the tomographic image is defined by using the transformation matrix; Performing affine transformation on the tomographic image using the transformation matrix relationship to obtain a mapping coordinate matrix of the changed tomographic image; A tomographic image is magnified or scaled in the longitudinal and transverse directions based on a bilinear interpolation technique and the tomographic image; The answer labeling process for the processed tomographic image according to the input form of the deep learning network includes: Based on the rectangular annotation technology, the main area frame of the vertebra in the tomographic image is annotated to obtain the complete area of the vertebra, and the pixel-level annotation is used to perform secondary annotation in the complete area of the vertebra to obtain the vertebral sub-area; The vertebral sub-regions are digitally labeled in the order of labeling, and after the labeling is completed, the classification information of the vertebral sub-regions is calibrated to determine the fracture status of the vertebral body; Obtain personal information, injury mechanism and fracture history of the person corresponding to the tomographic image based on the image acquisition information, and combine with expert annotation technology to perform fracture status assessment and annotation on the corresponding tomographic image; The personal information, injury mechanism, fracture history and fracture status assessment annotation results of the corresponding person are integrated into the answer annotation text.
2. The artificial intelligence-based vertebral fracture automatic identification and analysis system according to claim 1, characterized in that: The vertebral fracture automatic identification unit comprises: The network structure building module is used to build a bimodal network and a large text semantic model structure, and automatically identify the network structure based on the construction results; Structural parameter configuration module, used to configure the hyper parameters of automatic identification network structure and training files; The recognition network structure training module is used to input the tomography image and the answer annotation into the automatic recognition network structure to perform network training and verify the evaluation accuracy of the output vertebral area fracture information; The network structure adjustment module is used to adjust the automatic recognition network structure based on the evaluation accuracy results.
3. The artificial intelligence-based vertebral fracture automatic identification and analysis system according to claim 2, characterized in that: The step of constructing a bimodal network and a text semantics large model structure and obtaining an automatic recognition network structure based on the construction result includes: The deep convolutional neural network model is selected as the basic network structure for feature extraction, the scale feature layer is set in combination with the pyramid structure, and convolution and deconvolution operations are performed on each scale feature layer; After the convolution is completed, a feature fusion operation is performed, and the branch functions of the feature layers at each scale are defined to obtain a bimodal network structure containing several groups of branch functions; Based on the output of the bimodal network structure and the text requirements, the text network infrastructure is selected, and the channel deletion operation is performed on the network infrastructure, and the long short-term memory network is spliced to obtain the text semantic network structure; The bimodal network structure is integrated with the text semantic network structure to obtain the automatic recognition network structure.
4. The artificial intelligence-based vertebral fracture automatic identification and analysis system according to claim 3, characterized in that: The tomographic image and the answer annotation are input into the automatic recognition network structure to perform network training, and the evaluation accuracy of the output vertebral area fracture information is verified, including: After the answer annotation text is converted into word vectors, the tomographic images are combined to divide the training set and the validation set according to a preset ratio, and the training set is input into the automatic recognition network structure; Use the learning rate adjustment strategy to set the initial learning rate to start the initial training of automatic recognition network structure, and increase the learning rate for normal training after reaching the number of traversals; During the training process of the automatic recognition network structure, the stability of the loss function is judged, and after the loss function is stable, the structure that meets the traversal times requirement is selected as the automatic recognition network structure that has completed the training; The validation set is input into the trained automatic recognition network structure, and the branching function is used to output the detection frame of the vertebral body area, the number of pixels in the vertebral sub-area segmentation, and the fracture classification information; The output information is compared with the answer annotation text, and the evaluation accuracy of the output vertebral area fracture information is verified based on the comparison results.
5. The artificial intelligence-based automatic identification and analysis system for vertebral fractures according to claim 4 is characterized in that: The output information is compared with the answer annotation text, and the evaluation accuracy of the output vertebral area fracture information is verified based on the comparison result, including: The overlap between the detection frame and the corresponding main area frame in the answer annotation text is judged, and the percentage between the number of segmented pixels of the vertebral sub-area and the corresponding pixel-level annotation result in the answer annotation text is calculated; Compare the fracture classification information with the classification information calibration structure in the answer annotation text, and count the distances between each vertebral sub-region based on the sub-region pixels; The comparison results, overlap judgment results and percentage calculation results are integrated to analyze the output accuracy of the trained automatic recognition network structure.
6. The artificial intelligence-based automatic identification and analysis system for vertebral fractures according to claim 5 is characterized in that: The fracture information identification and judgment unit comprises: The fracture severity analysis module is used to analyze the size and location of the fracture based on the fracture information, determine the distance between each vertebral joint and the adjacent sub-joints, and analyze the severity of the vertebral fracture; The fracture impact degree analysis module is used to input the severity of vertebral fractures and the answer annotation text into the automatic recognition network structure to obtain the fracture impact degree; The mobile terminal display module is used to convert the automatic recognition network structure into a mobile terminal computing framework using quantitative compression processing technology, and then upload the severity analysis results and the degree of fracture impact to the mobile terminal.
Citation Information
Patent Citations
Deep Learning-Based Assessment Method and System for Osteoporotic Vertebral Compression Fractures
CN114937502A
Deep learning-based thoracolumbar fracture recognition, segmentation, detection and positioning method
CN114494192A
Fracture auxiliary detection method and system, computer equipment and readable storage medium
CN115205238A