Spine segmentation method, device, equipment and medium based on deep learning

Through a two-stage deep learning method, combining global and local information, accurate segmentation of the cone segments in spinal images is achieved, solving the problem of low spinal segmentation accuracy in existing technologies and improving the accuracy and efficiency of clinical diagnosis and surgical planning.

CN119624998BActive Publication Date: 2025-09-30PERCEPTION VISION MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411701593.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-09-30
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

In existing technologies, two-dimensional deep learning networks based on patches and fully convolutional networks have difficulty accurately identifying specific cone segments in spinal images, resulting in low spinal segmentation accuracy and an inability to meet clinical needs.

Method used

A two-stage deep learning method is adopted. First, global spine segmentation is performed based on three-dimensional image blocks, and then precise cone segmentation is performed based on local information. The first deep learning model is used to obtain global information, and the second deep learning model is used for local refinement. The model training is optimized by combining multi-scale loss function.

Benefits of technology

It improves the accuracy and efficiency of spinal segmentation, conforms to the clinicians' reading habits, and provides more accurate spinal disease diagnosis and surgical planning support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119624998B_ABST
    Figure CN119624998B_ABST
Patent Text Reader

Abstract

The present application discloses a spinal segmentation method, apparatus, equipment and medium based on deep learning, which belongs to the field of artificial intelligence technology. In the embodiment of the present application, two deep learning models are used to perform two-stage image segmentation on the original image containing the spine, and the spine in the medical image is fully automatically identified, and each cone segment is automatically segmented and the position of each cone segment is identified, providing a new method for clinicians when performing spinal diseases and surgical planning. It is faster and more accurate, and can take into account global information or a part of local information, which is more in line with the clinician's reading habits, thereby achieving a more accurate segmentation result. In addition, the local segmentation process emphasizes the order of cone segment arrangement between the spine and the cervical vertebrae, so that the position of each cone segment is obtained by fine segmentation and identification, further improving the model accuracy, thereby improving the image segmentation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a spine segmentation method, apparatus, device and medium based on deep learning. Background Art

[0002] Spinal diseases are one of the main causes of chronic pain and functional impairment worldwide, including lumbar disc herniation, scoliosis and fractures. Among them, degenerative disc disease (DDD) such as lumbar spinal stenosis, lumbar disc herniation and cervical spondylosis is a common and frequently occurring disease that can seriously affect the quality of life of patients [1,2]. With the aging of the population, its incidence rate is gradually increasing, and the disease burden is becoming increasingly serious. Because the spinal structure is complex and adjacent to important nerves and vascular tissues, in addition, developmental variations, deformities or degeneration of the spine are also common, so how to improve the safety and accuracy of spinal surgery is crucial. In recent years, "precision medicine" has emerged, and studies [3] have shown that computer-guided spinal surgery can significantly improve the safety and efficiency of surgery. As the first step in diagnosis and computer guidance, an accurate and efficient spinal segmentation and labeling method will help subsequent work to be carried out more efficiently.

[0003] Spinal imaging examinations (such as CT and MRI) are important tools for diagnosing these diseases, and accurate spinal segmentation and cone identification are crucial for disease assessment and treatment plan formulation. Traditionally, imaging examinations rely on manual labeling and analysis by doctors, which is not only time-consuming and labor-intensive, but also easily affected by subjective factors, resulting in inconsistencies in labeling. In addition, with the increase in the amount of imaging data, the efficiency of manual analysis can no longer meet clinical needs. The application of automatic spinal segmentation and cone classification technology can significantly improve the efficiency and accuracy of image analysis, thereby promoting the formulation of individualized treatment plans. For example, in scoliosis surgery, accurate segmentation and labeling help identify affected vertebrae, thereby formulating a more accurate surgical path.

[0004] In the existing technology, a two-dimensional deep learning network is constructed based on patches (pre-dividing the image into a certain number of small image blocks) and a fully convolutional network to identify the spinal area in each image. However, it does not identify which specific cone segment it is. Therefore, the segmentation accuracy of this spinal segmentation method is low and the auxiliary effect for spinal imaging examination is poor. Summary of the Invention

[0005] The present invention provides a method, apparatus, device, and medium for spine segmentation based on deep learning, which can improve segmentation accuracy and achieve rapid and accurate results. The technical solution is as follows:

[0006] In one aspect, a spine segmentation method based on deep learning is provided, the method comprising:

[0007] Based on the original image containing the spine, multiple slice images are acquired;

[0008] For a first image among the multiple slice images, a three-dimensional image block consisting of the first image and its adjacent slice images is input into a first deep learning model, and the first deep learning model segments the first image based on the three-dimensional image block to obtain a first prediction result for the first image, where the first prediction result is used to indicate whether each pixel in the first image is a spinal region and the corresponding cone segment position;

[0009] In order from one end to the other end of the spine, the first prediction result of each target number of adjacent cone segments in the first prediction result and the second image are input into a three-dimensional second deep learning model. The second deep learning model segments the second image based on the first prediction result of each target number of cone segments, and outputs a second prediction result of the second image. The second image is the local image area corresponding to each target number of cone segments in the first image. The second prediction result is used to indicate whether each pixel in the second image is a spinal area and the corresponding cone segment position.

[0010] In some embodiments, acquiring multiple slice images based on an original image containing the spine includes:

[0011] The original image is preprocessed to obtain a first image, wherein the preprocessing includes cropping and / or normalization, wherein the cropping is to crop the background area other than the body in the original image according to grayscale information and / or position information.

[0012] In some embodiments, the process of obtaining the second prediction result of the second image based on the first prediction result and the second image is performed serially in an order from one end to the other end of the spine; or, the process of obtaining the second prediction result of the second image based on the first prediction result and the second image is performed in parallel through multiple processes.

[0013] In some embodiments, the training process of the first deep learning model is implemented based on a first loss function, and the first loss function is obtained by integrating multiple loss functions using multiple scales, and the weights of multiple loss functions in the first loss function are adaptively adjusted according to the training situation.

[0014] In some embodiments, the training process of the second deep learning model is based on at least one of the false positive rate of the segmentation result and the boundary loss function, and the first loss function is implemented, and the first loss function is obtained based on the integration of multiple loss functions using multi-scale, and the weights of multiple loss functions in the first loss function are adaptively adjusted according to the training situation.

[0015] In some embodiments, the method further comprises:

[0016] The second prediction result of the second image is added to the original image according to the position information of the pixel points to obtain a third prediction result of the original image, where the third prediction result of the original image includes a segmentation result and a labeling result of the original image.

[0017] In some embodiments, adding the second prediction result of the second image to the original image according to the pixel position information to obtain the third prediction result of the original image includes:

[0018] In response to the presence of overlapping pixels in the second prediction result of the second image, and the overlapping pixels are different in the second prediction results of different second images, the pixels around the overlapping pixels are clustered according to the second prediction results of the pixels around the overlapping pixels to obtain a third prediction result for the overlapping pixels.

[0019] In one aspect, a deep learning-based spinal segmentation apparatus is provided, comprising:

[0020] An acquisition module, configured to acquire a plurality of slice images based on an original image containing a spine;

[0021] a first segmentation module, configured to input, for a first image among the plurality of slice images, a three-dimensional image block consisting of the first image and its adjacent slice images into a first deep learning model, and have the first deep learning model segment the first image based on the three-dimensional image block to obtain a first prediction result for the first image, the first prediction result being used to indicate whether each pixel in the first image is a spinal region and a corresponding cone segment position;

[0022] The second segmentation module is used to input the first prediction results of each target number of adjacent cone segments in the first prediction results and the second image into a three-dimensional second deep learning model in order from one end to the other end of the spine. The second deep learning model segments the second image based on the first prediction results of each target number of cone segments, and outputs a second prediction result of the second image. The second image is the local image area corresponding to each target number of cone segments in the first image. The second prediction result is used to indicate whether each pixel in the second image is a spinal region and the corresponding cone segment position.

[0023] In some embodiments, the acquisition module is used to preprocess the original image to obtain a first image, and the preprocessing includes cropping and / or standardization, and the cropping is to crop the background area outside the body in the original image according to grayscale information and / or position information.

[0024] In some embodiments, the process of obtaining the second prediction result of the second image based on the first prediction result and the second image is performed serially in an order from one end to the other end of the spine; or, the process of obtaining the second prediction result of the second image based on the first prediction result and the second image is performed in parallel through multiple processes.

[0025] In some embodiments, the training process of the first deep learning model is implemented based on a first loss function, and the first loss function is obtained by integrating multiple loss functions using multiple scales, and the weights of multiple loss functions in the first loss function are adaptively adjusted according to the training situation.

[0026] In some embodiments, the training process of the second deep learning model is based on at least one of the false positive rate of the segmentation result and the boundary loss function, and the first loss function is implemented, and the first loss function is obtained based on the integration of multiple loss functions using multi-scale, and the weights of multiple loss functions in the first loss function are adaptively adjusted according to the training situation.

[0027] In some embodiments, the apparatus further comprises:

[0028] An adding module is used to add the second prediction result of the second image to the original image according to the position information of the pixel points to obtain a third prediction result of the original image, and the third prediction result of the original image includes the segmentation result and the labeling result of the original image.

[0029] In some embodiments, the adding module is used to:

[0030] In response to the presence of overlapping pixels in the second prediction result of the second image, and the overlapping pixels are different in the second prediction results of different second images, the pixels around the overlapping pixels are clustered according to the second prediction results of the pixels around the overlapping pixels to obtain a third prediction result for the overlapping pixels.

[0031] On the one hand, an electronic device is provided, which includes one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, and the at least one computer program is loaded and executed by the one or more processors to implement various optional implementations of the above-mentioned deep learning-based spine segmentation method.

[0032] On the one hand, a computer-readable storage medium is provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement various optional implementations of the above-mentioned deep learning-based spine segmentation method.

[0033] In one aspect, a computer program product or computer program is provided, comprising one or more program codes stored in a computer-readable storage medium. One or more processors of an electronic device are capable of reading the one or more program codes from the computer-readable storage medium, and executing the one or more program codes by the one or more processors, so that the electronic device can perform any of the above-described possible embodiments of the deep learning-based spine segmentation method.

[0034] In an embodiment of the present application, two deep learning models are used to perform two-stage image segmentation on the original image containing the spine, fully automatically identify the spine in the medical image, and automatically segment each cone segment and identify the position of each cone segment, providing clinicians with a new method when performing spinal diseases and surgical planning, which is more accurate and faster. The first stage is based on global information segmentation, and the second stage is based on local information segmentation. It can take into account both global information and a part of local information, which is more in line with the clinician's reading habits, thereby achieving a more accurate segmentation result. In the local segmentation process, the order of cone segment arrangement between the spine and the cervical vertebrae is emphasized, so that the position of each cone segment is obtained by fine segmentation and identification, further improving the model accuracy and image segmentation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0036] Figure 1 Schematic diagram of an implementation environment of a spine segmentation method based on deep learning provided in an embodiment of the present application;

[0037] Figure 2 This is a flowchart of a spine segmentation method based on deep learning provided in an embodiment of the present application;

[0038] Figure 3 This is a flowchart of a spine segmentation method based on deep learning provided in an embodiment of the present application;

[0039] Figure 4 1 is a schematic structural diagram of a spinal segmentation device based on deep learning provided in an embodiment of the present application;

[0040] Figure 5 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0041] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0042] In this application, the terms "first", "second", etc. are used to distinguish between identical or similar items that have substantially the same role and function. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is there any limitation on quantity or execution order. It should also be understood that although the following description uses the terms first, second, etc. to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the various described examples, a first image can be referred to as a second image, and similarly, a second image can be referred to as a first image. Both the first image and the second image can be images, and in some cases, can be separate and different images.

[0043] In this application, the term "at least one" means one or more, and the term "plurality" means two or more. For example, a plurality of data packets means two or more data packets.

[0044] It should be understood that the terminology used in the description of the various examples herein is for the purpose of describing particular examples only and is not intended to be limiting. As used in the description of the various examples and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0045] It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. The term "and / or" describes an association relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this application generally indicates that the associated objects are in an "or" relationship.

[0046] It should also be understood that in the various embodiments of the present application, the size of the serial number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0047] It should also be understood that determining B based on A does not mean determining B solely based on A. B can also be determined based on A and / or other information.

[0048] It will also be understood that the term “comprise” (also known as “inCludes,” “inCluding,” “Comprises,” and / or “Comprising”) when used in this specification specifies the presence of stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0049] It should also be understood that the term "if" may be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting." Similarly, the phrase "if it is determined that..." or "if [stated condition or event] is detected" may be interpreted to mean "upon determining that..." or "in response to determining that..." or "upon detecting [stated condition or event]" or "in response to detecting [stated condition or event]," depending on the context.

[0050] At least some of the functions of the apparatus or electronic device provided in the embodiments of the present application may be implemented using an AI model. For example, at least one of the multiple modules of the apparatus or electronic device may be implemented using an AI model. Functions associated with the AI ​​may be performed using non-volatile memory, volatile memory, and a processor.

[0051] The processor may include one or more processors. In this case, the one or more processors may be general-purpose processors, such as a central processing unit (CPU), an application processor (AP), etc., or pure graphics processing units, such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-specific processor, such as a neural processing unit (NPU).

[0052] The one or more processors control processing of input data according to predefined operating rules or artificial intelligence (AI) models stored in non-volatile memory and volatile memory. The predefined operating rules or artificial intelligence models are provided by training or learning.

[0053] Here, providing by learning means obtaining predefined operating rules or an AI model with desired characteristics by applying a learning algorithm to a plurality of learning data. This learning can be performed in the device or electronic device itself in which the AI ​​according to the embodiment is executed, and / or can be implemented by a separate server / system.

[0054] An AI model can include multiple neural network layers. Each layer has multiple weight values, and each layer performs neural network calculations by calculating between the input data of the layer (such as the calculation results of the previous layer and / or the input data of the AI ​​model) and the multiple weight values ​​of the current layer. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q networks.

[0055] A learning algorithm is a method for training a predetermined target device (e.g., a robot) using multiple learning data to enable, allow, or control the target device to make a determination or prediction. Examples of such learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0056] According to the present application, at least one step in the method performed in the electronic device, such as identifying architectural elements, classifying floor plan areas, etc., can be implemented using an artificial intelligence model. The processor of the electronic device can perform preprocessing operations on the data to convert it into a form suitable for use as an input to the artificial intelligence model. The artificial intelligence model can be obtained through training. Here, "obtained through training" means that a basic artificial intelligence model is trained with multiple training data through a training algorithm to obtain a predefined operating rule or artificial intelligence model configured to perform the desired feature (or purpose).

[0057] The following is an explanation of the nouns involved in this application.

[0058] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0059] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0060] Computer vision (CV) is the science of making machines "see." Specifically, it refers to machine vision techniques such as using cameras and computers to replace the human eye in identifying, tracking, and measuring objects. This involves further processing the images, transforming them into images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0061] Machine Learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI.

[0062] Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0063] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, robots, smart medical care, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0064] The solutions provided in the embodiments of this application involve artificial intelligence computer vision technology, machine learning and other technologies, which are specifically illustrated by the following embodiments.

[0065] The implementation environment of this application is described below.

[0066] Figure 1 1 is a schematic diagram of an implementation environment for a deep learning-based spine segmentation method provided in an embodiment of the present application. The implementation environment includes a terminal 101, or the implementation environment includes the terminal 101 and a deep learning-based spine segmentation platform 102. The terminal 101 is connected to the deep learning-based spine segmentation platform 102 via a wireless network or a wired network.

[0067] Terminal 101 is at least one of a smartphone, a game console, a desktop computer, a tablet computer, an e-book reader, an MP3 (Moving Picture Experts Group Audio Layer III) player, an MP4 (Moving Picture Experts Group Audio Layer IV) player, or a laptop computer. Terminal 101 has an application installed and running that supports deep learning-based spine segmentation.

[0068] For example, the terminal 101 has an image segmentation function. After acquiring an original image containing the spine, it can use two deep learning models to perform a two-stage image segmentation to obtain the spinal region and the location of each cone segment in the image. The terminal 101 can complete this task independently, or after acquiring the original image, the terminal 101 can transmit it to the deep learning-based spine segmentation platform 102 to provide image segmentation services. This embodiment of the present application is not limited to this.

[0069] The deep learning-based spine segmentation platform 102 includes at least one of a server, multiple servers, a cloud computing platform, and a virtualization center. The deep learning-based spine segmentation platform 102 is used to provide background services for applications that support deep learning-based spine segmentation. Optionally, the deep learning-based spine segmentation platform 102 undertakes the main processing work, and the terminal 101 undertakes the secondary processing work; or, the deep learning-based spine segmentation platform 102 undertakes the secondary processing work, and the terminal 101 undertakes the main processing work; or, the deep learning-based spine segmentation platform 102 or the terminal 101 undertakes the processing work separately. Alternatively, the deep learning-based spine segmentation platform 102 and the terminal 101 adopt a distributed computing architecture for collaborative computing.

[0070] Optionally, the deep learning-based spine segmentation platform 102 includes at least one server 1021 and a database 1022. The database 1022 is used to store data. In an embodiment of the present application, the database 1022 stores sample input data and sample annotation data, and can also store user input data to provide data services for at least one server 1021.

[0071] A server can be a standalone physical server, a server cluster composed of multiple physical servers, or a distributed system. It can also be a cloud server that provides basic cloud computing services, including cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. A terminal can be, but is not limited to, a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, and so on.

[0072] Those skilled in the art will appreciate that the number of terminals 101 and servers 1021 may be greater or lesser. For example, there may be only one terminal 101 or server 1021, or there may be dozens or hundreds of terminals 101 or servers 1021, or a greater number. The embodiments of the present application do not limit the number or device types of terminals or servers.

[0073] Most existing technologies segment the spinal region within an image, but fail to accurately identify each spinal segment and its corresponding location. In a precise surgical setting, surgeons need to clearly identify which spinal segments are problematic or require surgery. Therefore, simply segmenting the spine is insufficient; each segmented segment needs to be automatically labeled with its location and location. From a neural network training perspective, end-to-end segmentation and identification of each segment is possible. However, given the complex situation of seven cervical vertebrae, 12 thoracic vertebrae, five lumbar vertebrae, five sacral segments, and four coccygeal segments, where adjacent segments are very similar, end-to-end prediction struggles to accurately identify each segment. Furthermore, in actual clinical spinal scans, spinal regions often have unclear boundaries and are subject to significant noise due to the adjacent structural features of the spine and the imaging principles. This results in many algorithms segmenting spinal regions that appear contiguous, making it difficult to distinguish individual segments. For these reasons, simple 2D neural networks struggle to clearly identify which adjacent regions of different cone segments belong to which cone segment. 3D neural networks can make this determination based on information from adjacent images. Furthermore, clinical data primarily deals with diseased conditions, not conventional 3D spinal morphology. Therefore, a simple end-to-end neural network is unlikely to be effective in clinical application.

[0074] In response to the above problems, this application proposes a two-stage automatic spine segmentation and labeling method based on deep learning. The first stage segments the spine of the cervical spine, thoracic spine, lumbar spine, sacrum, and coccyx according to the global information of the image. This stage mainly distinguishes different spinal regions based on the different body tissues and organs around them. Because it relies on global information, some segmentation and recognition errors may occur locally. The second stage is based on the results of the first stage and performs local 3D neural network segmentation, aiming to obtain complete and independent segmentation results and correct labels for each cone segment in the local area. Although the running time will be slower than the end-to-end one-stage neural network model, it can reduce the time for subsequent modifications by doctors, and the segmentation results are more in line with clinical needs, which improves the doctor's work efficiency as a whole. The following is through Figure 2 The illustrated embodiment provides a detailed description of the above-mentioned two-stage image segmentation and labeling method based on deep learning.

[0075] Figure 2 This is a flowchart of a spine segmentation method based on deep learning provided by an embodiment of the present application. The method is applied to an electronic device, which is a server or a terminal. Figure 2 , the method includes the following steps.

[0076] 201. The electronic device obtains a plurality of slice images based on an original image including a spine.

[0077] This invention proposes a two-stage automatic spine segmentation and labeling method based on deep learning, hoping to provide an automatic spine segmentation and labeling method, providing clinicians with a new method for spinal disease and surgical planning, which is more accurate and faster. The embodiment of this application aims to fully automatically identify the spine in medical images, automatically segment each cone segment, and automatically identify the location and number of each cone segment. Based on this, a three-dimensional structure can be generated, providing important imaging information support for subsequent diagnosis and other work.

[0078] In the embodiments of the present application, the original image containing the spine may be a medical image containing the spine. A medical image containing the spine may be taken when a patient is undergoing medical consultation, so that the doctor can diagnose the patient's condition based on the medical image. The embodiments of the present application do not limit the specific imaging device used to capture the original image containing the spine.

[0079] In some embodiments, the original image containing the spine is typically a three-dimensional (3D) image. Therefore, after the electronic device acquires the original image containing the spine, it can first process it to obtain multiple slice images, and then perform image segmentation on the slice images. It should be noted that the original image containing the spine is typically a 3D image. Each image is generally referred to as a slice in medicine.

[0080] In some embodiments, the electronic device can preprocess the original image to obtain a first image, where the preprocessing includes cropping and / or standardization, and the cropping is to crop the background area outside the body in the original image according to grayscale information and / or position information.

[0081] Regarding cropping, the original image containing the spine may contain parts other than the user's body, such as the environment in which the user took the photo. Considering that the purpose of image segmentation is to identify the user's spine region and the locations of each cone segment within the spine region, parts other than the user's body are noise interference for image segmentation. Therefore, in this embodiment of the application, the electronic device can first crop the original image, cutting out parts other than the user's body, leaving only the image region of the user's body.

[0082] Specifically, the electronic device captures raw image data containing the patient's spine and crops the background area outside the body. This is because during the image acquisition phase, the background area outside the body is essentially air and is not imaged. Furthermore, the scanning bed, which may appear in the image, is a fixed mechanical device, and these can be removed using simple grayscale and position information.

[0083] For standardization processing, image data containing only body areas is standardized. The image acquisition may also include some boundary points, or the differences between the pixel points corresponding to different targets in the image are not too large. In this case, standardization can improve the accuracy of subsequent image segmentation.

[0084] 202. For the first image among the multiple slice images, the electronic device inputs a three-dimensional image block consisting of the first image and its adjacent slice images into a first deep learning model, and the first deep learning model segments the first image based on the three-dimensional image block to obtain a first prediction result of the first image, where the first prediction result is used to indicate whether each pixel in the first image is a spinal region and the corresponding cone position.

[0085] After obtaining multiple slice images, the electronic device can input each first image, along with the image blocks formed by the slice image and adjacent slice images, into the first deep learning model. To fully extract the characteristics of the tissues and organs surrounding the spine to assist in locating the cone segment, the first-stage model (also known as the first deep learning model) takes the processed slice image and its adjacent slices as input, forming a three-dimensional image block. This allows for global image information to be obtained while avoiding missed detections due to missing important information at certain levels.

[0086] In step 202, not only is each pixel identified as belonging to the spinal region, but also the specific cone segment to which the pixel belongs is identified. This first deep learning model performs segmentation based on the sliced ​​images derived from the original image, taking into account the global information of the image and avoiding missed detections due to missing important information at certain levels.

[0087] In some embodiments, the first deep learning model can segment the first image and output a first prediction result for the first image. In other embodiments, the first deep learning model can segment the first image and output whether each pixel in the first image is a spinal region, and then, through image post-processing methods, obtain the cone segment location corresponding to each pixel in the first image. The following description uses the first prediction result obtained by direct segmentation using the first deep learning model as an example.

[0088] The first deep learning model can be a pre-trained model for image segmentation, and the first deep learning model can be trained based on a large number of sample original images and target prediction results of the sample original images.

[0089] In some embodiments, the training process of the first deep learning model can be: obtaining a sample original image and a target prediction result of the sample original image; obtaining multiple sample slice images based on the sample original image, and then inputting an image block consisting of each sample slice image and adjacent sample slice images into the initial first deep learning model, and the initial first deep learning model segments the sample slice image based on the input image block, outputs the prediction result of the sample slice image, and then obtains a first loss function based on the prediction result of the sample slice image and its target prediction result, and then adjusts the model parameters of the initial first deep learning model based on the first loss function, and then repeats the above steps of inputting the sample original image, image segmentation and adjusting the model parameters until the first target condition is reached, and the first deep learning model is obtained.

[0090] Among them, the first target condition can be set by relevant technical personnel based on experience or needs. For example, the first target condition can be that the number of iterations reaches a first number. The specific value of the first number can be set by relevant technical personnel based on needs or experience, or the first target condition can be that the first loss function converges. The embodiments of the present application do not limit this.

[0091] For the first deep learning model, in step 202, the electronic device inputs the standardized data into the model of the first stage. The neural network of the first deep learning model can select various suitable deep neural network models, such as Unet, TransUnet, SwinUNETR, etc., or the model structure can be designed by relevant technical personnel, or any model with image segmentation function can be directly used. Of course, it is also possible to further increase or reduce the intermediate operation layer based on personal experience, such as introducing other advanced technologies such as attention mechanisms to improve the prediction accuracy and generalization ability of the model. When the model architecture design is large enough, the learning ability of the model is able to learn accurate spinal segmentation and labeling, so the embodiment of the present application does not constrain the first deep learning model to a certain model architecture.

[0092] For the first loss function, the first loss function can also be called an accuracy evaluation function, and the first loss function can be designed by relevant technical personnel according to needs or experience. In some embodiments, the training process of the first deep learning model is implemented based on the first loss function, and the first loss function can be obtained based on the integration of multiple loss functions using multiple scales, and the weights of multiple loss functions in the first loss function are adaptively adjusted according to the training situation. That is, the first loss function can use a multi-scale, multiple loss function integration structure, and adaptively adjust the weights of multiple loss functions according to the training situation. Exemplarily, the multiple loss functions can consider but are not limited to loss functions such as MSE (mean-square error), Dice coefficient, and boundary loss.

[0093] Among them, the mean squared error (MSE) and the Dice coefficient are metrics used to compare the similarity between two samples. They measure the degree of overlap between two sets and are commonly used in fields such as text similarity calculation, image processing, and bioinformatics. The Dice coefficient ranges from [value](value) to [value](value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value))))(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value)(value) ( ​​...

[0094] 203. The electronic device inputs the first prediction result of each target number of adjacent cone segments in the first prediction result and the second image into a three-dimensional second deep learning model in order from one end to the other end of the spine. The second deep learning model segments the second image based on the first prediction result of each target number of cone segments and outputs a second prediction result of the second image. The second image is a local image area corresponding to each target number of cone segments in the first image. The second prediction result is used to indicate whether each pixel in the second image is a spinal region and the corresponding cone segment position.

[0095] After the image segmentation in the previous stage is performed, considering that the image segmentation in the first stage is based on global information segmentation, which may lead to the local segmentation being not fine enough, for a simple end-to-end neural network, and when the first stage is based on the whole slice segmentation, the accuracy of the segmentation, especially at the boundary, may not be accurate enough. Therefore, in an embodiment of the present application, a second deep learning model is provided, and the result obtained based on the image segmentation in the first stage can be input into the second deep learning model together with the original image, and the local image input is segmented in a certain order, so that both global information and local information can be taken into account to obtain a more accurate image segmentation result.

[0096] The input of the second-stage model includes both the original image and the prediction results of the first stage. The advantage is that it allows the model to correct the results of the first stage based on local image details. This optimization method is also in line with the clinicians' own reading habits, from overall to local analysis.

[0097] In some embodiments, the two ends of the spine are the cervical vertebrae and the coccyx, respectively. The order from one end of the spine to the other end can be from the cervical vertebrae to the coccyx (that is, from top to bottom), or from the coccyx to the cervical vertebrae (that is, from bottom to top). This embodiment of the present application is not limited to this.

[0098] In some embodiments, the target number can be set by relevant technical personnel based on demand or experience, and the embodiments of the present application do not limit this. For example, the target number can be 3 or 4, that is, in step 203, the results of the first stage prediction are sequentially (this order can be from the cervical vertebra to the coccyx, or from the coccyx to the cervical vertebra, that is, from top to bottom or from bottom to top) taken as the input of the second stage each time for the adjacent 2 cone sections or 3 cone section areas predicted in the first stage, where the input includes the output of the first stage and the original image information of the corresponding local area. The output of the second stage model is the segmentation result of the new adjacent 2 cone sections or 3 cone sections and the corresponding cone section labels.

[0099] In some embodiments, the process of obtaining the second prediction result of the second image based on the first prediction result and the second image is performed serially from one end of the spine to the other; alternatively, the process of obtaining the second prediction result of the second image based on the first prediction result and the second image is performed in parallel via multiple processes. In other words, the image segmentation process for the target number of cone segments in step 203 can be performed serially or in parallel to improve processing efficiency. In other words, the inference of the second-stage model can be performed serially or in parallel to increase speed.

[0100] In some embodiments, the training process of the second deep learning model is based on at least one of the false positive rate of the segmentation result and the boundary loss function, and the first loss function is implemented, and the first loss function is obtained based on the integration of multiple loss functions using multiple scales, and the weights of multiple loss functions in the first loss function are adaptively adjusted according to the training situation. That is, the loss function (or accuracy evaluation function) of the second stage, in addition to the loss function consistent with the first stage, also needs to increase the constraints of the false positive rate of the classification label and the boundary loss. Thereby strengthening the model's constraints on the correctness of boundaries and classifications, and obtaining an accurate segmentation result.

[0101] In some embodiments, the training process of the second deep learning model may include: obtaining multiple sample slice images of a sample original image, the prediction results of the sample original image obtained by the first deep learning model, and the target prediction results of the sample original image; then, in order from one end of the spine to the other, inputting the prediction results of each target number of adjacent cone segments in the prediction results and the image area corresponding to the prediction results in the sample slice images into a three-dimensional initial second deep learning model; the initial second deep learning model segments the image area based on the prediction results of each target number of cone segments and outputs the prediction results of the image area. Then, based on the prediction results of the image area and its target prediction results, a second loss function is obtained, and the model parameters of the initial second deep learning model are adjusted based on the second loss function. The above steps of inputting the sample original image, image segmentation, and adjusting the model parameters are repeated until the second target condition is met, thereby obtaining the second deep learning model.

[0102] Among them, the second target condition can be set by relevant technical personnel based on experience or needs. For example, the second target condition can be that the number of iterations reaches a second number. The specific value of the second number can be set by relevant technical personnel based on needs or experience, or the second target condition can be that the second loss function converges. The embodiments of the present application are not limited to this.

[0103] In some embodiments, the first deep learning model and the second deep learning model can be jointly trained, so that the joint training of the two-stage models can enable the two models to be combined to segment the image more accurately. Specifically, the prediction results output by the initial first deep learning model can be input into the initial second deep learning model together with the sample slice image in the above order, and this is performed in each iteration. After the second deep learning model outputs the prediction result, the second loss function is obtained by the prediction result, and the first loss function and the second loss function are integrated into the target loss function. For example, the integration method can be weighted summation or other combination methods, which is not limited in the embodiments of the present application.

[0104] In some embodiments, after step 203, the electronic device may also add the second prediction result of the second image to the original image according to the position information of the pixel points to obtain a third prediction result of the original image, where the third prediction result of the original image includes the segmentation result and the labeling result of the original image.

[0105] In some embodiments, the results predicted by the second-stage model are placed back into the original image according to the position information to obtain the final segmentation and labeling results. In this process, if there is an inconsistency in the classification of overlapping pixels, local clustering of the surrounding pixels can be used to correct it (because the cone nodes are arranged in a top-down order). Therefore, in response to the presence of overlapping pixels in the second prediction result of the second image, and the second prediction results of the overlapping pixels in different second images are different, the electronic device can cluster the pixels around the overlapping pixels based on the second prediction results of the pixels around the overlapping pixels to obtain a third prediction result for the overlapping pixels.

[0106] In a specific example, Figure 3 As shown, raw image data of the patient's spine can be obtained, and then background removal and normalization are performed on the original image. This normalized data is then fed into the first-stage neural network to obtain preliminary segmentation results and labels for each conical segment of the spine. The local image information and segmentation results of the conical segments segmented in the first stage are then fed into the second-stage neural network, following the order of the spine, to obtain more accurate segmentation results and labels. The second-stage results are then placed back into the corresponding position in the original image, and the final spine segmentation and labeling results corresponding to the original image are output.

[0107] Through the above embodiments, the embodiments of the present application clearly provide an automatic method for automatic segmentation and labeling of the spine, that is, the image obtained by clinical scanning is segmented into the spinal area and marked as which part of the cone segment it is. It can save the time of clinicians to judge which cone segment a certain cone segment is, thereby improving work efficiency. Furthermore, we propose a two-stage method based on deep learning, because it is very easy to use the surrounding tissues and organs to judge the position of the cone segment, but it is not conducive to the fine segmentation of the cone segment; if you only focus on local spinal segmentation, it is easy to misclassify the cone segment. Therefore, the two-stage segmentation and labeling method proposed here can take advantage of each other's strengths and obtain the best results. Furthermore, according to the prior information that the spine is arranged in order, the corresponding loss function in the model training is strengthened, so that the model can more accurately segment each independent cone segment with the correct upper and lower order.

[0108] The advantages of the embodiments of this application are as follows:

[0109] 1) This method can automatically segment and label each cone segment of the spine. Compared with existing methods that only perform binary classification of the spine and segment the spinal region in the image, this method not only segments the spine but also accurately identifies the corresponding cone segment of each segment.

[0110] 2) Two-stage, more accurate segmentation: Compared to most end-to-end algorithms, which only consider global or partial local information, this two-stage approach is more in line with clinicians' reading habits. It first determines the approximate area and category of each cone segment, and then finely segments each cone segment boundary, thereby achieving a more accurate segmentation result.

[0111] 3) This method emphasizes the prior knowledge of the sequential arrangement of the spine, starting from the cervical vertebrae, continuing to the thoracic vertebrae, lumbar vertebrae, sacrum, and coccyx. During the model training phase, this prior information is used to impose targeted constraints on the loss function, further improving model accuracy. Compared to other methods that use general model training loss functions, this method introduces prior knowledge.

[0112] In an embodiment of the present application, two deep learning models are used to perform two-stage image segmentation on the original image containing the spine, fully automatically identify the spine in the medical image, and automatically segment each cone segment and identify the position of each cone segment, providing clinicians with a new method when performing spinal diseases and surgical planning, which is more accurate and faster. The first stage is based on global information segmentation, and the second stage is based on local information segmentation. It can take into account both global information and a part of local information, which is more in line with the clinician's reading habits, thereby achieving a more accurate segmentation result. In the local segmentation process, the order of cone segment arrangement between the spine and the cervical vertebrae is emphasized, so that the position of each cone segment is obtained by fine segmentation and identification, further improving the model accuracy and image segmentation accuracy.

[0113] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.

[0114] Figure 4 This is a schematic diagram of the structure of a spinal segmentation device based on deep learning provided in an embodiment of the present application. Figure 4 , the device comprises:

[0115] An acquisition module 401 is configured to acquire a plurality of slice images based on an original image containing the spine;

[0116] A first segmentation module 402 is configured to input, for a first image among the plurality of slice images, a three-dimensional image block consisting of the first image and its adjacent slice images into a first deep learning model, and have the first deep learning model segment the first image based on the three-dimensional image block to obtain a first prediction result for the first image, where the first prediction result indicates whether each pixel in the first image is a spinal region and the corresponding cone segment position;

[0117] The second segmentation module 403 is used to input the first prediction results of each target number of adjacent cone segments in the first prediction results and the second image into a three-dimensional second deep learning model in order from one end to the other end of the spine. The second deep learning model segments the second image based on the first prediction results of each target number of cone segments, and outputs a second prediction result of the second image. The second image is the local image area corresponding to each target number of cone segments in the first image. The second prediction result is used to indicate whether each pixel in the second image is a spinal region and the corresponding cone segment position.

[0118] In some embodiments, the acquisition module 401 is used to preprocess the original image to obtain a first image, and the preprocessing includes cropping and / or standardization, and the cropping is to crop the background area outside the body in the original image according to grayscale information and / or position information.

[0119] In some embodiments, the process of obtaining the second prediction result of the second image based on the first prediction result and the second image is performed serially in an order from one end to the other end of the spine; or, the process of obtaining the second prediction result of the second image based on the first prediction result and the second image is performed in parallel through multiple processes.

[0120] In some embodiments, the training process of the first deep learning model is implemented based on a first loss function, and the first loss function is obtained by integrating multiple loss functions using multiple scales, and the weights of multiple loss functions in the first loss function are adaptively adjusted according to the training situation.

[0121] In some embodiments, the training process of the second deep learning model is based on at least one of the false positive rate of the segmentation result and the boundary loss function, and the first loss function is implemented, and the first loss function is obtained based on the integration of multiple loss functions using multi-scale, and the weights of multiple loss functions in the first loss function are adaptively adjusted according to the training situation.

[0122] In some embodiments, the apparatus further comprises:

[0123] An adding module is used to add the second prediction result of the second image to the original image according to the position information of the pixel points to obtain a third prediction result of the original image, and the third prediction result of the original image includes the segmentation result and the labeling result of the original image.

[0124] In some embodiments, the adding module is used to:

[0125] In response to the presence of overlapping pixels in the second prediction result of the second image, and the overlapping pixels are different in the second prediction results of different second images, the pixels around the overlapping pixels are clustered according to the second prediction results of the pixels around the overlapping pixels to obtain a third prediction result for the overlapping pixels.

[0126] The device provided in the embodiment of the present application uses two deep learning models to perform two-stage image segmentation on the original image containing the spine, fully automatically identifies the spine in the medical image, and automatically segments each cone segment and identifies the position of each cone segment, providing clinicians with a new method when performing spinal diseases and surgical planning, which is more accurate and faster. The first stage is based on global information segmentation, and the second stage is based on local information segmentation. It can take into account global information or a part of local information, which is more in line with the clinician's reading habits, thereby achieving a more accurate segmentation result. In the local segmentation process, the order of cone segment arrangement between the spine and the cervical vertebrae is emphasized, so that the position of each cone segment is obtained by fine segmentation and identification, further improving the model accuracy and image segmentation accuracy.

[0127] It should be noted that the deep learning-based spine segmentation device provided in the above embodiment is only used as an example to illustrate the division of the above functional modules when trading stocks based on deep neural networks. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the deep learning-based spine segmentation device can be divided into different functional modules to complete all or part of the functions described above. In addition, the deep learning-based spine segmentation device provided in the above embodiment and the deep learning-based spine segmentation method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0128] Figure 5 It is a structural diagram of an electronic device provided in an embodiment of the present application. The electronic device 500 may have relatively large differences due to different configurations or performances, and can include one or more processors (Central Processing Units, CPU) 501 and one or more memories 502, wherein the memory 502 stores at least one computer program, and the at least one computer program is loaded and executed by the processor 501 to implement the deep learning-based spine segmentation method provided in the above-mentioned various method embodiments. The electronic device can also include other components for realizing the functions of the device. For example, the electronic device can also have components such as wired or wireless network interfaces and input and output interfaces for input and output. The embodiments of the present application are not described in detail here.

[0129] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including at least one computer program, wherein the at least one computer program is executable by a processor to perform the deep learning-based spine segmentation method described in the above embodiment. For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, or an optical data storage device.

[0130] In an exemplary embodiment, a computer program product or computer program is also provided, comprising one or more program codes stored in a computer-readable storage medium. One or more processors of an electronic device can read the one or more program codes from the computer-readable storage medium, and the one or more processors can execute the one or more program codes, so that the electronic device can perform the above-described deep learning-based spine segmentation method.

[0131] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0132] It should be understood that determining B based on A does not mean determining B based solely on A. B can also be determined based on A and / or other information.

[0133] Those skilled in the art will understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a disk or an optical disk, etc.

[0134] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A spine segmentation method based on deep learning, characterized in that: The method comprises: Based on the original image containing the spine, multiple slice images are acquired; For a first image among the multiple slice images, a three-dimensional image block consisting of the first image and its adjacent slice images is input into a first deep learning model, and the first deep learning model segments the first image based on the three-dimensional image block to obtain a first prediction result for the first image, where the first prediction result is used to indicate whether each pixel in the first image is a spinal region and the corresponding cone segment position; In order from one end to the other end of the spine, the first prediction result of each target number of adjacent cone segments in the first prediction result and the second image are input into a three-dimensional second deep learning model. The second deep learning model segments the second image based on the first prediction result of each target number of cone segments, and outputs a second prediction result of the second image. The second image is the local image area corresponding to each target number of cone segments in the first image. The second prediction result is used to indicate whether each pixel in the second image is a spinal area and the corresponding cone segment position.

2. The spine segmentation method based on deep learning according to claim 1, characterized in that The method of obtaining a plurality of slice images based on the original image containing the spine includes: The original image is preprocessed to obtain a first image, wherein the preprocessing includes cropping and / or normalization, wherein the cropping is to crop the background area other than the body in the original image according to grayscale information and / or position information.

3. The spine segmentation method based on deep learning according to claim 1, characterized in that The process of obtaining the second prediction result of the second image based on the first prediction result and the second image is performed serially in an order from one end to the other end of the spine; or, the process of obtaining the second prediction result of the second image based on the first prediction result and the second image is performed in parallel through multiple processes.

4. The spine segmentation method based on deep learning according to claim 1, characterized in that The training process of the first deep learning model is implemented based on a first loss function, which is obtained by integrating multiple loss functions using multiple scales. The weights of multiple loss functions in the first loss function are adaptively adjusted according to the training situation.

5. The spine segmentation method based on deep learning according to claim 1, characterized in that The training process of the second deep learning model is based on at least one of the false positive rate of the segmentation result and the boundary loss function, and the first loss function. The first loss function is obtained by integrating multiple loss functions using multiple scales, and the weights of multiple loss functions in the first loss function are adaptively adjusted according to the training situation.

6. The spine segmentation method based on deep learning according to claim 1, characterized in that The method further comprises: The second prediction result of the second image is added to the original image according to the position information of the pixel points to obtain a third prediction result of the original image, where the third prediction result of the original image includes a segmentation result and a labeling result of the original image.

7. The spine segmentation method based on deep learning according to claim 6, characterized in that The adding the second prediction result of the second image to the original image according to the position information of the pixel points to obtain a third prediction result of the original image includes: In response to the presence of overlapping pixels in the second prediction result of the second image, and the second prediction results of the overlapping pixels are different in different second images, the pixels with the same second prediction results around the overlapping pixels are clustered according to the second prediction results of the pixels around the overlapping pixels to obtain a third prediction result for the overlapping pixels.

8. A spinal segmentation device based on deep learning, characterized in that: The device comprises: An acquisition module, configured to acquire a plurality of slice images based on an original image containing a spine; a first segmentation module, configured to input, for a first image among the plurality of slice images, a three-dimensional image block consisting of the first image and its adjacent slice images into a first deep learning model, and have the first deep learning model segment the first image based on the three-dimensional image block to obtain a first prediction result for the first image, the first prediction result being used to indicate whether each pixel in the first image is a spinal region and a corresponding cone segment position; The second segmentation module is used to input the first prediction results of each target number of adjacent cone segments in the first prediction results and the second image into a three-dimensional second deep learning model in order from one end to the other end of the spine. The second deep learning model segments the second image based on the first prediction results of each target number of cone segments, and outputs a second prediction result of the second image. The second image is the local image area corresponding to each target number of cone segments in the first image. The second prediction result is used to indicate whether each pixel in the second image is a spinal region and the corresponding cone segment position.

9. An electronic device, characterized in that: The electronic device includes one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, and the at least one computer program is loaded and executed by the one or more processors to implement the deep learning-based spine segmentation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The storage medium stores at least one computer program, which is loaded and executed by a processor to implement the deep learning-based spine segmentation method according to any one of claims 1 to 7.