Recognition method of spine segment sequence, training method of spine segment sequence recognition model and electronic equipment

By introducing the anatomical axial sequence constraint loss function and sequence dependency modeling, the low-precision problem in spinal segment identification is solved, and high-precision spinal segment identification is achieved, ensuring that the surgical robot can accurately locate the spinal segments and reducing surgical risks.

CN120656041APending Publication Date: 2025-09-16BEIJING TINAVI MEDICAL TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510877820.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing technologies have the problem of low recognition accuracy in spinal segment identification, especially in the difficulty of accurate classification in similar segment areas, which may lead to risks such as pedicle screw misplacement or nerve root damage during surgery.

Method used

A sequential constraint loss function based on anatomical axial order and a sequence dependency modeling method are adopted. The image data is serialized and input into the first module to ensure the correctness of the sequence information. The neural network of the second module is used to model the anatomical dependency relationship between spinal segments. Combined with the multi-scale feature refinement module, the recognition accuracy is improved.

Benefits of technology

It significantly improves the accuracy of spinal segment identification, reduces the probability of identification failure, provides more reliable spinal segment positioning support for surgical robots, and ensures the safety and success rate of surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656041A_ABST
    Figure CN120656041A_ABST
Patent Text Reader

Abstract

The invention provides a spine segment sequence identification method, a spine segment sequence identification model training method and electronic equipment. According to the spine segment sequence identification method provided by the invention, the problem of low precision in spine segment identification is effectively solved by adopting the sequence constraint loss function based on the dissection axial sequence and combining sequence-dependent modeling. The first module is used for carrying out serialized input construction on image data to ensure that the input data keeps correct sequence information in the processing process, and the second module is used for modeling an anatomical dependency relationship between spinal segments through a neural network, so that the segments can be accurately identified from a global perspective. According to the method, errors caused by similar forms of adjacent segments, lack of anatomical reference structures and the like in the prior art are overcome, and the accuracy of spinal segment recognition is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to a spinal segment sequence recognition method, a spinal segment sequence recognition model training method, and an electronic device. Background Art

[0002] Spinal segment recognition is a key technology in medical image analysis. Its goal is to automatically locate and identify the anatomical position of each segment in spinal images through computer algorithms.

[0003] In the scenario of spinal segment recognition, there is a technical problem of low recognition accuracy. Summary of the Invention

[0004] The purpose of the present invention is to overcome the above-mentioned technical deficiencies and provide a method for identifying a spinal segment sequence, a training method for a spinal segment sequence identification model, and an electronic device to solve the technical problem of low recognition accuracy in the scenario of spinal segment identification in related technologies.

[0005] In order to achieve the above technical objectives, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a method for identifying a spinal segment sequence, the method comprising: determining image data representing at least two spinal segments; The image data is input into a preset spinal segment sequence recognition model for recognition to obtain a spinal segment sequence corresponding to the image data; wherein the spinal segment sequence recognition model includes at least a first module and a second module; wherein the first module is a module for performing serialized input construction on the input image data, and the second module is a neural network representing sequence to sequence, and the neural network is used to extract the anatomical dependency between the at least two spinal segments; wherein the spinal segment sequence recognition model is pre-trained using a preset loss function; wherein the preset loss function includes at least a first loss function, and the first loss function is a loss function for characterizing the spinal segment sequence constraint, and the sequence is based on the axial sequence of the spinal anatomy.

[0006] In a second aspect, the present invention provides a method for training a spinal segment sequence recognition model, comprising: Acquire training data; wherein the training data includes a plurality of image data for representing at least two spinal segments and corresponding label data; The training data is input into a spinal segment sequence recognition model to be trained for training; wherein the spinal segment sequence recognition model to be trained includes at least a first module and a second module; wherein the first module is a module for performing serialized input construction on the input image data, and the second module is a neural network representing sequence to sequence, and the neural network is used to extract the anatomical dependency between the at least two spinal segments; wherein the spinal segment sequence recognition model is trained using a preset loss function; wherein the preset loss function includes at least a first loss function, and the first loss function is a loss function for characterizing spinal segment sequence constraints, and the sequence is based on the axial sequence of the spinal anatomy; After the preset conditions are met, the spinal segment sequence recognition model is obtained.

[0007] In a third aspect, the present invention provides an electronic device comprising: a memory, and one or more processors communicatively connected to the memory; the memory stores instructions executable by the one or more processors, and the instructions are executed by the one or more processors so that the one or more processors implement the above-mentioned method.

[0008] Beneficial effects: The spinal segment sequence recognition method provided by the present invention effectively solves the low-precision problem in spinal segment recognition by adopting a sequence-constrained loss function based on the anatomical axial order and combining it with sequence dependency modeling. Specifically, the first module ensures that the input data maintains the correct sequence information during the processing by serializing the input structure of the image data, while the second module models the anatomical dependency relationship between spinal segments through a neural network, which can accurately identify the segments from a global perspective. This method overcomes the errors caused by the similar morphology of adjacent segments and the lack of anatomical reference structures in traditional technologies, and greatly improves the accuracy of spinal segment recognition. By introducing a sequence-constrained loss function, the rationality of the segment order is ensured, making the recognition of the spinal segment sequence more accurate, significantly reducing the recognition failure caused by sequence errors, and ultimately achieving high-precision spinal segment recognition, providing more reliable spinal segment positioning support for surgical robots. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 1 is a flow chart of a spinal segment sequence recognition method provided by an embodiment of the present invention; Figure 2 This is a scene example diagram of a spinal segment sequence recognition method provided by an embodiment of the present invention; Figure 3 This is a scene example diagram of a spinal segment sequence recognition method provided by an embodiment of the present invention; Figure 4Schematic diagram of a spinal segment sequence recognition model provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0010] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0011] In the related technologies, in the scenario of spinal segment recognition, especially for the processing of three-dimensional medical imaging data, the related technologies face the problem of low recognition accuracy. With the rapid development of robot-assisted surgery, accurate spinal segment recognition has become particularly important, especially when it is necessary to accurately locate each spinal segment during surgery. However, current spinal segment recognition technology relies on traditional image processing methods, such as template matching or segment-by-segment classifiers. These methods have many problems that affect the final recognition accuracy.

[0012] Specifically, in the surgical robotic system, multiple key hardware and software components are integrated to provide spinal segment positioning and surgical navigation.

[0013] In related art, a surgical robot system may include: The imaging system, which can use equipment such as cone-beam CT or intraoperative MRI, is used to collect three-dimensional image data of the spinal segments. The three-dimensional image data can be a 512×512×512 voxel three-dimensional image that includes each spinal segment and its surrounding structures.

[0014] The control device can be used for image processing and analysis to ensure rapid and accurate processing of image data during surgery and provide timely feedback.

[0015] The actuator, which may be a multi-axis robotic arm, is used to perform spinal segment positioning and precise manipulation.

[0016] Understandably, in spinal surgeries such as lumbar fusion, spinal segment identification serves as the coordinate reference for surgical navigation. This identification directly impacts the safety and success rate of surgery. Inaccurate identification can lead to adverse consequences such as pedicle screw misplacement and nerve root damage. Because inaccurate spinal segment identification can pose significant clinical risks, improving its accuracy is a critical issue in the current medical field.

[0017] Several techniques have been used in the field of spinal segment identification. These include template matching, single-segment classifiers, and sliding window prediction. Template matching locates spinal segments through rigid registration of preoperative CT templates with intraoperative CBCT images. However, this method relies on a complete anatomical reference and cannot effectively match anatomical structures such as ribs or the sacrum if they are missing from the image.

[0018] Single-segment classifiers can use deep learning models such as 3D-ResNet to independently classify each segment, but this method ignores the sequential dependencies between spinal segments and often leads to misclassification of adjacent segments.

[0019] The sliding window prediction method identifies spinal segments by treating multiple consecutive segments as a classification unit. However, this method still suffers from the situation where local information is separated from the global anatomical sequence, which may lead to recognition errors.

[0020] More specifically, especially between thoracic vertebrae T10 and T12, the geometry, tissue density, and spatial distribution of the segments are very similar. Single-segment classifiers and template matching methods are prone to recognition errors in these similar areas. Furthermore, related technologies overly rely on anatomical landmarks of the spine (ribs, sacrum, etc.) to locate the segments, and fail when these landmarks are missing. For example, when the spinal segment is located in the lumbar sacral region where the sacrum is missing, the recognition failure rate increases significantly. Sliding window prediction and template matching methods rely on local information for prediction, ignoring the order and dependencies between segments, which can easily lead to errors when processing adjacent segments.

[0021] It is understandable that the anatomical structure of the spinal segments is complex, especially the thoracic spine, where there are at least three similarities between the segments, including: 1. Similarity in geometric shape: that is, the difference in pedicle width between adjacent thoracic vertebrae is often small, which makes segment classification difficult.

[0022] 2. Similar tissue density: The CT values ​​of the cortical bones of the thoracic spine are all within a small range, making the tissue density of different segments in the CT image almost the same, which increases the complexity of image processing.

[0023] 3. Uniform spatial distribution: The coefficient of variation of intervertebral disc height in the thoracic spine segment is also small, resulting in almost no obvious difference in spatial distribution between different segments, further increasing confusion between adjacent segments.

[0024] The above factors make it difficult for the methods in the related art to accurately classify these similar segments.

[0025] In summary, in the scenario of spinal segment recognition, there is a technical problem of low recognition accuracy.

[0026] like Figure 1 、 Figure 2 and Figure 3 As shown, this embodiment provides a method for identifying spinal segment sequences. The method can be executed by a control device of a surgical robot or a separate image processing server. The control device can be an embedded controller, such as an ARM-based processor or an FPGA (field programmable gate array)-based controller, which can be used to accelerate image processing tasks. It is understood that the embedded controller is not only responsible for the calculation of spinal segment sequence identification but also for coordinating and controlling the operation of the robot actuators. For example, the separate image processing server can be a high-performance workstation or server, or a cloud server, etc. A preset spinal segment sequence recognition model can be deployed in the server. The server receives image data from an imaging system (e.g., CT, MRI, CBCT, etc.) and then calls the preset spinal segment sequence recognition model to identify the spinal segment sequence in the image data.

[0027] The method comprises: Step S12: Determine image data for representing at least two spinal segments.

[0028] In this embodiment, the determining action may be represented by performing a preset process on the three-dimensional medical image data received from the imaging system.

[0029] Specifically, the pre-set processing can be an image pre-processing algorithm (e.g., denoising) or a 3D mask extraction algorithm. In other words, the input data for the pre-set spinal segment sequence recognition model can be 3D medical image data that has undergone image pre-processing (e.g., denoising) or 3D mask data corresponding to the 3D medical image data.

[0030] Specifically, the image preprocessing algorithm is an algorithm used to characterize and improve the quality of three-dimensional medical image data. For example, it can be an image denoising algorithm, a contrast enhancement algorithm, a smoothing algorithm, and the like. Denoising can be used to remove random noise in medical images. Specifically, Gaussian filtering, median filtering, wavelet denoising, and the like can be used to ensure that useful information in three-dimensional medical image data is not interfered with. Contrast enhancement can highlight the features of the spinal segment area by performing local or global contrast enhancement processing on the three-dimensional medical image data, making the segment boundaries more obvious. Image smoothing can smooth the three-dimensional medical image data to remove unnecessary texture details while retaining the morphological features of the spinal segments to help subsequent models better identify the spinal area.

[0031] In this embodiment, the three-dimensional mask extraction process can adopt a threshold segmentation algorithm. For example, by setting an appropriate threshold, different regions in the three-dimensional medical image data can be segmented to extract the spinal segments from the surrounding tissue regions. Alternatively, an adaptive threshold segmentation algorithm can be used. A morphologically based processing algorithm can also be used to extract the three-dimensional mask. For example, morphological operations such as dilation and erosion can be used to further refine the boundaries of the spinal segments to ensure that the mask is highly consistent with the actual spinal segment region.

[0032] A region growing algorithm can also be used to extract a 3D mask. For example, starting from a predetermined seed point, the region can be expanded based on pixel value similarity or other features to gradually cover the entire spinal segment area.

[0033] Therefore, in this embodiment, the image data used to represent the at least two spinal segments may be three-dimensional medical image data, or three-dimensional mask data corresponding to the three-dimensional medical image data. Specifically, the three-dimensional medical image data may be CBCT data (cone-beam CT), CT data (computed tomography), or MRI data (magnetic resonance imaging).

[0034] Step S14: Input the image data into a preset spinal segment sequence recognition model for recognition to obtain a spinal segment sequence corresponding to the image data; wherein the spinal segment sequence recognition model includes at least a first module and a second module; wherein the first module is a module for performing serialized input construction on the input image data, and the second module is a neural network representing sequence to sequence, and the neural network is used to extract the anatomical dependency between the at least two spinal segments; wherein the spinal segment sequence recognition model is pre-trained using a preset loss function; wherein the preset loss function includes at least a first loss function, and the first loss function is a loss function for characterizing the spinal segment sequence constraint, and the sequence is based on the axial sequence of the spinal anatomy.

[0035] In this embodiment, the preset spinal segment sequence recognition model is obtained by pre-training using a base model (model to be trained). The base model may include a first module and a second module in terms of model structure. When the base model is pre-trained, a preset loss function may be used to train the base model to obtain a spinal segment sequence recognition model. The spinal segment sequence recognition model may be deployed in a control device or a server. It is understandable that the training data for training the base model is also prepared in advance. For example, in a specific embodiment, the training data may include the following categories of training data: 1. Only partial cervical and thoracic vertebrae CBCT data are available.

[0036] 2. Only partial thoracic and lumbar CBCT data are available.

[0037] 3. CBCT data of only partial thoracic spine and complete lumbar spine.

[0038] 4. Only CBCT data of the cervical spine.

[0039] 5. Only CBCT data of the thoracic spine are available.

[0040] 6. Only CBCT data of the lumbar spine are available.

[0041] It is also understood that the training data may include not only the aforementioned 3D medical image data or the 3D mask data corresponding to the 3D medical image data, but also label data. Label data is indicative data that corresponds one-to-one with the 3D medical image data or 3D mask data. This label data can provide supervisory information for the training of the base model, assist in the calculation of the loss function during training, and guide the base model in learning how to correctly identify and classify spinal segments.

[0042] In this embodiment, the label data may include category labels and spatial position information of the spinal segments. Specifically, each segment has a unique category label that indicates its position in the spine (e.g., C1, T1, L1, etc.). These category labels may be pre-labeled by people with relevant knowledge (e.g., people with a medical background), reflecting the correct order and identity of each segment in the anatomical structure. The spatial position information may be the specific spatial position information of the spinal segment in three-dimensional medical imaging data or three-dimensional mask data, and may be represented by coordinate points in a three-dimensional coordinate system. This spatial position information is used to help the model understand the relative positions between the segments.

[0043] In a specific embodiment, the label data may include a one-hot encoding label. Specifically, for each spinal segment, the label data may be represented in a one-hot encoding format. For example, for segments such as cervical vertebrae C1, thoracic vertebrae T1, and lumbar vertebrae L1, one-hot encoding maps the category of each segment to a 24-dimensional vector (which can also be adjusted according to actual task requirements and is not limited in this embodiment), where only the index position corresponding to the segment category is 1, and the remaining positions are 0. The label data for each spinal segment is a 24-dimensional one-hot vector. For example, the label of the C1 segment may be: [1, 0, 0, ..., 0], indicating that it is the first segment (C1). It is understandable that for a specific image data (including multiple segments), its label data may be a concatenation of multiple 24-dimensional one-hot vectors. For example, in a specific embodiment, if the image data includes 6 segments, the data dimension of the label data of the image data may be 6*1*24. The 6 indicates that there are six segments in the image (which can be 3D medical imaging data or 3D mask data) that need to be identified (a correct sequence). The 1*24 indicates that there are 24 possible classifications for each segment (humans typically have 24 spinal segments, which translates to 24 categories). It is understood that the label data can include spatial location information, which can be provided in the form of 3D coordinates. For example, each segment can have three coordinate values ​​(x, y, z).

[0044] It is understood that the number of segments in the image data may be less than or equal to 6 due to the limitation of the field of view of the imaging system. Therefore, the image data used to represent at least two spinal segments may be image data of two, three, four or even six spinal segments.

[0045] More specifically, the input image data can contain K spinal segments (K ≤ 6), each corresponding to an independent classification task. The label tensor then has the dimensions K*1*24 (standard when K = 6). The first dimension, K, represents the index position of the spinal segment in the sequence (sorted by anatomical axis), the first dimension, 1, is reserved for single-segment labels, and the third dimension, 24, represents the probability distribution space for the complete spinal segment class.

[0046] In this embodiment, the first module is a module for performing serialized input construction on the input image data. It is understood that the core function of this first module is to generate image embeddings for each spinal segment in the image data (based on this, it can further generate embeddings for the position information of each spinal segment, also known as position embeddings). Specifically, this first module can be represented as a network module that performs serialized input construction. This module's input is 3D image data, but its output is a sequence of feature vectors sorted based on the axial direction of the spinal anatomy. This first module can convert the raw spinal image data (which can be 3D medical image data or 3D mask data) into an ordered sequence of anatomical feature vectors, providing spatially structured input for subsequent sequence modeling. It is understood that traditional methods, due to their neglect of 3D spatial relationships, can lead to confusion between adjacent segments. In this embodiment, the first module can first discretize the raw spinal image data (including discretizing vertebral masks), then perform spatial sorting, and then perform feature embedding on the discretized image blocks (one block corresponds to one segment) to obtain image block embeddings. Finally, position enhancement can be performed, also known as position embedding.

[0047] In one possible and specific embodiment, the first module may include a discretization layer, a spatial sorting layer, a feature extraction layer, a position encoding enhancement layer, and an output layer, all connected in series. Specifically, the discretization layer discretizes the image data based on a preset threshold to generate a three-dimensional mask of the spinal segment. During this process, each spinal segment in the image is segmented based on anatomical features (such as vertebral boundaries), generating mask data corresponding to each segment. This mask data can be used as input for subsequent processing to define the region of each segment. The spatial sorting layer analyzes the spatial coordinate information of each segment along the anatomical axis of the spine and sorts the mask data according to the normal anatomical order of the spine (from C1 to S1). This sorting ensures that each image block (i.e., spinal segment) is arranged in the correct order, avoiding confusion between adjacent segments caused by traditional methods that ignore three-dimensional spatial relationships. The feature extraction layer extracts spatial features from the sorted spinal segment images through three-dimensional convolution and pooling operations. Three-dimensional convolution can effectively capture the spatial relationship and depth information between segments, while the three-dimensional pooling operation can further compress the dimension of the image data and extract feature vectors with higher discriminative power. Through such processing, the feature vector can accurately represent the spatial structural information of each spinal segment. The position coding enhancement layer can add position coding information to the feature vector of each image block. This process can embed the sequential information of the position of each segment in the anatomical axis of the spine into the feature vector through a preset coding scheme, ensuring that the model can understand the relative position relationship between the segments during subsequent sequence modeling. The output layer can output a structured sequence of feature vectors, where each feature vector represents the spatial characteristics of a spinal segment and contains position coding information, providing accurate spatial input for subsequent sequence modeling.

[0048] In one possible and specific embodiment, based on the above embodiment, the first module can be equipped with a multimodal input unit at its head end. This multimodal input unit is used to receive image information of other modalities representing the same spinal segment. In other words, the first module can accept different types of input through the multimodal input unit, perform adaptive processing, and then proceed to the subsequent feature extraction and sorting process.

[0049] In one possible and specific embodiment, the first module may include an anatomical position encoding unit and a three-dimensional feature extraction unit. The anatomical position encoding unit is configured to sort each spinal segment based on its spatial coordinate information along the spinal anatomical axis; the three-dimensional feature extraction unit is configured to sequentially perform three-dimensional convolution and three-dimensional pooling operations on the sorted segment images, outputting a sequence of reduced-dimensional feature vectors.

[0050] In this embodiment, the second module is a sequence-to-sequence neural network that is used to extract the anatomical dependencies between the at least two spinal segments. It is understood that the core function of this second module is to model the anatomical functional links between spinal segments (e.g., the spatial topology of nerve roots), thereby breaking through the traditional method's reliance on local features.

[0051] In this embodiment, the second module can be a sequence-to-sequence model based on the Transformer architecture. It can include at least an encoder, a decoder and a self-attention mechanism. Specifically, the encoder can encode the input spinal segment feature vector sequence and convert it into a set of context-rich representation vectors. The decoder can generate the sequential dependency of the spinal segments based on the representation information output by the encoder. Through the self-attention mechanism, the model can capture the long-term dependency between the spinal segments. It can be understood that this long-term dependency is crucial to ensuring the anatomical order.

[0052] In this embodiment, the second module can be a sequence-to-sequence model based on a recurrent neural network (RNN) architecture. In this model, the feature vectors of the spinal segments are input into the network one by one according to the anatomical order of the spine. At each step, the network processes the current segment features and the hidden state at the previous moment through a recursive relationship to update the current segment state.

[0053] In this embodiment, the second module can be a sequence-to-sequence model based on a long short-term memory (LSTM) network architecture. It may include at least an encoder-decoder structure and a long short-term memory (LSTM). The encoder extracts a feature vector sequence for each spinal segment, which is then decoded by the decoder to output sequential information about the spinal segments. The LSTM can retain long-range dependencies between spinal segments.

[0054] In this embodiment, the second module may be a sequence-to-sequence model based on a gated recurrent unit (GRU).

[0055] In this embodiment, the second module may be a sequence-to-sequence model based on bidirectional LSTM.

[0056] In this embodiment, the second module can be a sequence-to-sequence model based on a graph neural network. It can be understood that a graph neural network is a neural network that processes graph-structured data and can capture dependencies between non-continuous nodes by transferring information through edges between nodes. In spinal segment recognition, each segment can be regarded as a node in a graph, and the relationships between segments (such as intervertebral discs, ligament connections, adjacent relationships, etc.) are used as edges of the graph, and feature learning is performed through the structure of the graph. That is, each spinal segment corresponds to a node in the graph, and the features of the node represent the anatomical information (size, shape) of the segment. The connection relationship between spinal segments is represented by edges, and the existence of edges ensures that the network can learn the spatial dependencies between segments. The features of each node are updated through edges, and through graph convolution operations, nodes receive information from adjacent nodes and update their features.

[0057] In one possible and specific embodiment, the second module may include a sequence feature extraction unit and an independent classification unit connected in series. The sequence feature extraction unit is represented as a sequence modeling network with a self-attention mechanism, which takes as input the sequence of feature vectors and is used to extract anatomical dependencies between spinal segments; the independent classification unit is used to independently map the hidden vector of each spinal segment to a probability distribution of predefined spinal segment categories.

[0058] In this embodiment, the spinal segment sequence recognition model is pre-trained using a preset loss function; wherein the preset loss function includes at least a first loss function, which is a loss function used to characterize the spinal segment sequence constraint, and the sequence is based on the axial sequence of the spinal anatomy. It is understandable that in many cases, the second module (for example, when using the VIT architecture) is essentially an unordered input modeler. Although the segments can be sorted from top to bottom according to the z coordinate, the model cannot guarantee that the output labels are strictly increasing or continuous. For example, the output order of the model is not necessarily L1, L2, L3, L4, L5, but may be L2, L1, L3, L4, L5, or even random. Therefore, it is necessary to introduce a loss function to supervise the predicted spinal segment sequence. In this embodiment, the first loss function is used, and the first loss function (that is, the sequential supervision loss) is used to ensure that the predicted label index is increasing.

[0059] In one possible and specific embodiment, the first loss function may be: In the formula, It represents the anatomical index of the predicted label of the i-th segment, and its value range can be 1-24. It is represented by the number of input segments, which can be 2-6, that is, the maximum number of segments covered by intraoperative CBCT is 6. This ensures that there is still a penalty even when adjacent segments are equal (e.g., L3→L3).

[0060] In one possible and specific embodiment, the first loss function may be: In the formula, It can be expressed as a sequential tolerance gap, that is, only when < A penalty is imposed when the index of the next segment is not significantly greater than that of the previous segment (that is, a loss is triggered when the index of the next segment is not significantly greater than that of the previous segment).

[0061] In a possible and specific implementation, the preset loss function (that is, the total loss function) can be: In the formula, N represents the total number of spinal segments contained in a single input image. It can be expressed as a classification loss function, can be expressed as a regularization term, where is the regularization coefficient, are all weight parameters of the neural network. CE is represented by the cross entropy loss function, which is used to calculate the difference between the predicted probability distribution and the true distribution. is the predicted value of the i-th segment, that is, the predicted category probability distribution of each spinal segment, is the true value of the i-th segment, that is, the true category of each spinal segment.

[0062] In this embodiment, the preset spinal segment sequence recognition model may further include a third module.

[0063] Specifically, the third module may be a multi-scale feature refining module, and the multi-scale feature refining module may be provided between the first module and the second module.

[0064] The multi-scale feature refinement module takes as input the feature sequence output by the first module. It first extracts multi-scale features of the spinal segments using multiple convolutional branches (e.g., vertebral keypoint detection, bone density distribution, and intervertebral disc space analysis). Each branch is responsible for a different anatomical feature, such as vertebral morphology, bone density distribution, and intervertebral disc space. The features output by these multiple convolutional branches are then fused to obtain a more refined feature representation. Finally, the refined feature vector is passed to the second module, providing enhanced features for subsequent anatomical dependency modeling and spinal segment sequence optimization.

[0065] The spinal segment sequence recognition method provided in this embodiment effectively solves the low-precision problem in spinal segment recognition by adopting a sequence-constrained loss function based on the anatomical axial order and combining it with sequence dependency modeling. Specifically, the first module ensures that the input data maintains the correct sequence information during the processing by serializing the input structure of the image data, while the second module models the anatomical dependency relationship between spinal segments through a neural network, which can accurately identify the segments from a global perspective. This method overcomes the errors caused by the similar morphology of adjacent segments and the lack of anatomical reference structures in traditional technologies, and greatly improves the accuracy of spinal segment recognition. By introducing a sequence-constrained loss function, the rationality of the segment order is ensured, making the recognition of the spinal segment sequence more accurate, significantly reducing the recognition failure caused by sequence errors, and ultimately achieving high-precision spinal segment recognition, providing more reliable spinal segment positioning support for surgical robots.

[0066] In some embodiments, the first module includes: The anatomical position encoding unit is used to sort each spinal segment based on the spatial coordinate information of each spinal segment in the spinal anatomical axis.

[0067] In this embodiment, the primary function of the anatomical position encoding unit is to sort the input spinal segments based on their spatial coordinate information along the anatomical axis of the spine. This unit ensures that the order of the segment input conforms to the anatomical order of the spine, avoids confusion between adjacent segments, and provides ordered input data for the subsequent feature extraction process. Specifically, the features of each spinal segment are first associated with its spatial coordinates (e.g., Z-axis position), which can be provided by an imaging system (e.g., CBCT, CT, MRI, etc.). Each spinal segment can be sorted based on its spatial coordinates (e.g., Z-axis coordinates). For example, the order of the spinal segments along the Z-axis can be from bottom to top (i.e., from sacral to cervical vertebrae) or from top to bottom (i.e., from cervical vertebrae to sacral vertebrae). Therefore, the anatomical position encoding unit can sort the segments in ascending order based on their Z-axis coordinates, ensuring that the input order of each image block conforms to the natural anatomical order of the spine.

[0068] It is understandable that in order to enhance the model's perception of spatial position information, the anatomical position encoding unit can also assign a position encoding vector to each segment. This vector can be combined with the segment's feature vector in a predefined manner (such as a linear encoding generated based on the Z-axis position or through position embedding technology). The role of position encoding is to provide the model with clear spatial position information to help the model understand the relative position of each spinal segment in space. Specifically, the anatomical position encoding unit can add a separate embedding vector for each segment. This embedding vector can be generated based on the spatial coordinates of the segment, and by splicing or adding it with the feature vector (image embedding of the segment image block), the spatial information and feature data are input into the subsequent network together.

[0069] The three-dimensional feature extraction unit is used to perform three-dimensional convolution operations and three-dimensional pooling operations on the sorted segment images in sequence, and output a feature vector sequence after dimensionality reduction.

[0070] In this embodiment, the three-dimensional feature extraction unit mainly functions to extract useful spatial features from the sorted spinal segment images and compress these features into vectors of lower dimensions to facilitate subsequent classification tasks.

[0071] In this embodiment, the three-dimensional feature extraction unit may include a three-dimensional convolution layer and a three-dimensional pooling layer. Specifically, on the sorted spinal segment image, the three-dimensional convolution operation can extract spatial features from the three dimensions of the image: depth, width, and height. The convolution kernel slides in the three-dimensional space and extracts the local spatial features of the segment by weighted sum. The three-dimensional pooling layer can be performed after the convolution operation. The three-dimensional pooling layer compresses the spatial resolution of the data and retains key features by downsampling the three-dimensional feature map. The three-dimensional pooling layer can be maximum pooling (Max Pooling) and average pooling (Average Pooling). After convolution and pooling processing, the three-dimensional feature extraction unit can output a sequence of feature vectors after dimensionality reduction. These feature vector sequences represent the spatial features of each spinal segment in the three-dimensional space, while also retaining the anatomical information of the segment.

[0072] This embodiment effectively improves the accuracy and efficiency of spinal segment sequence recognition by introducing an anatomical position encoding unit and a three-dimensional feature extraction unit in the first module. The anatomical position encoding unit sorts each spinal segment based on the spatial coordinate information of the spinal segment in the anatomical axis, ensuring that the order of the input data conforms to the actual anatomical structure of the spine and avoiding the problem of confusion between adjacent segments. The three-dimensional feature extraction unit deeply mines the spatial features of each spinal segment through three-dimensional convolution and pooling operations, reduces the data dimension, retains the key spatial information of the segment, and optimizes the subsequent classification and recognition process. This design not only improves the recognition accuracy of the model, but also enhances its learning ability of spinal segment sequence dependence, providing more accurate and reliable spinal positioning support for the surgical robot.

[0073] In some embodiments, the second module includes: A sequence feature extraction unit is represented as a sequence modeling network with a self-attention mechanism, and its input is the feature vector sequence, which is used to extract the anatomical dependency between spinal segments.

[0074] In this embodiment, the sequence feature extraction unit utilizes a self-attention mechanism to extract anatomical dependencies between spinal segments from the input feature vector sequence. This unit captures the global contextual information between spinal segments, ensuring the model accurately understands the spatial and sequential relationships between segments, avoiding segment order errors caused by insufficient modeling of local dependencies in traditional methods.

[0075] It can be understood that the self-attention mechanism is a method that allows the model to learn the relationship between each element and other elements. In the sequence feature extraction unit, the self-attention mechanism calculates the correlation between each spinal segment feature and the features of other segments in the input sequence. For example, by calculating the relationship between the query, key, and value, the model can assign a weighted attention value to each spinal segment, thereby emphasizing other segment information related to the current segment. This weighting mechanism enables the model to automatically adjust the attention weights during the learning process, thereby better capturing the spatial dependencies of anatomically adjacent segments.

[0076] In this embodiment, the output of the sequence feature extraction unit is a sequence of feature vectors weighted by the self-attention mechanism. Each feature vector sequence corresponds to the contextual information of a spinal segment and includes the dependencies between that segment and other segments. These feature vectors can serve as input for subsequent classification tasks.

[0077] In this embodiment, the sequence feature extraction unit may be a Transformer encoder.

[0078] In this embodiment, the sequence feature extraction unit may be a model of CNN+(RNN, LSTM or GRU) structure.

[0079] In this embodiment, the sequence feature extraction unit can be a graph neural network.

[0080] In this embodiment, the sequence feature extraction unit may be a bidirectional LSTM.

[0081] An independent classification unit is connected in series with the sequence feature extraction unit and is used to independently map the hidden vector of each spinal segment into a probability distribution of a predefined spinal segment category.

[0082] In this embodiment, the independent classification unit independently maps the latent vector for each spinal segment in the feature vector sequence output by the sequence feature extraction unit into a probability distribution for the corresponding spinal segment class. This unit enables the model to classify each spinal segment and determine its specific anatomical location (e.g., C1, T1, L1, etc.).

[0083] In this embodiment, the independent classification unit may be an MLP classification head.

[0084] In this embodiment, the independent classification unit may be a fully connected network.

[0085] In this embodiment, the independent classification unit can be a classification head based on the self-attention mechanism.

[0086] In this embodiment, the independent classification unit may be a classification head based on a Transformer.

[0087] This embodiment significantly improves the accuracy and efficiency of spinal segment recognition by introducing a sequence feature extraction unit and an independent classification unit in the second module. The sequence feature extraction unit can effectively capture the anatomical dependencies between spinal segments through a sequence modeling network that carries a self-attention mechanism, ensuring that the model can understand the spatial and sequential dependencies between segments, thereby overcoming the sequential confusion problems that may exist in related technical methods. At the same time, the independent classification unit independently maps the extracted hidden features into a probability distribution of predefined spinal segment categories, ensuring that each spinal segment can be accurately classified and consistent with the anatomical order. This structured two-stage processing method enables the model to accurately grasp the anatomical dependencies and classification information when processing spinal segments, improves the overall recognition performance, and provides the surgical robot with more accurate spinal positioning and navigation capabilities.

[0088] In some embodiments, the preset spinal segment sequence recognition model is a spinal segment sequence recognition model based on the Vision-Transformer architecture, and the second module is a Transformer encoder and an MLP classification head connected in series therewith.

[0089] This embodiment uses a spinal segment sequence recognition model based on the Vision-Transformer (ViT) architecture to achieve efficient and accurate spinal segment recognition. The Vision-Transformer architecture effectively solves the long-range dependency problem that is difficult for convolutional neural networks in related technologies to handle by splitting image data into several image blocks and using a self-attention mechanism to capture global contextual information between segments. The Transformer encoder further strengthens the modeling of sequential dependencies between spinal segments, enabling the model to accurately understand the anatomical order of spinal segments. The MLP classification head independently classifies the hidden vectors of each spinal segment and accurately outputs the category probability distribution of each spinal segment. This design greatly improves the performance of the model in complex spinal segment classification tasks. It can not only effectively capture the dependencies between spinal segments, but also ensure efficient and accurate classification. It is particularly suitable for medical application scenarios that require precise positioning and identification of spinal segments.

[0090] In some embodiments, the image data is three-dimensional mask image data carrying at least two spinal segments or three-dimensional medical imaging data carrying at least two spinal segments; when the image data is three-dimensional medical imaging data carrying at least two spinal segments, the step of determining image data for representing at least two spinal segments includes: Step S122: Receive at least one three-dimensional medical image data carrying at least two spinal segments.

[0091] In this embodiment, the receiving action may be receiving one or more three-dimensional medical image data sent by the imaging system from the imaging system. The three-dimensional medical image data may carry 2-6 spinal segments, or a larger number.

[0092] Step S124: performing preprocessing on the at least one three-dimensional medical image data carrying at least two spinal segments based on a preset image preprocessing algorithm to obtain preprocessed three-dimensional medical image data.

[0093] In this embodiment, the image preprocessing algorithm may be an image denoising algorithm, a contrast enhancement algorithm, a smoothing algorithm, and the like.

[0094] This embodiment effectively improves the accuracy and efficiency of spinal segment identification by preprocessing three-dimensional medical imaging data containing at least two spinal segments. By applying preset image preprocessing algorithms (such as denoising and contrast enhancement), the quality of the image data is improved, making subsequent feature extraction and classification more accurate. The preprocessed image data can remove noise, enhance key information, and reduce data redundancy, thereby providing clearer input for accurate spinal segment identification. This design significantly improves the robustness and accuracy of the model, especially during real-time intraoperative diagnosis and navigation, providing more reliable support for spinal positioning.

[0095] In some embodiments, after the step of inputting the image data into a preset spinal segment sequence recognition model for recognition to obtain a spinal segment sequence corresponding to the image data, the method further comprises: Step S16: Determine whether the spinal segment sequence output by the spinal segment sequence recognition model needs to be post-processed.

[0096] In this embodiment, the spinal segment sequence output by the spinal segment sequence recognition model is the result of model prediction. This prediction result may be erroneous in some cases, so verification and judgment need to be performed.

[0097] In this embodiment, determining whether the spinal segment sequence output by the spinal segment sequence recognition model requires post-processing can include performing a sequence continuity check on the spinal segment sequence (prediction result), that is, determining whether the anatomical index values ​​of adjacent segments are strictly increasing. Alternatively, the determination can include determining whether the order of the spinal segment sequence output by the model conforms to an increasing order from the cervical vertebra (C1) to the sacral vertebra (S1). For example, if at least one segment in a spinal segment sequence does not conform to an increasing order, then post-processing is required for the spinal segment sequence (prediction result) to correct the error.

[0098] In this embodiment, the determination of whether the spinal segment sequence output by the spinal segment sequence recognition model needs to be post-processed can be: performing anatomical step verification on the spinal segment sequence (prediction result), that is, the index difference between adjacent segments is always 1 (no missing segments).

[0099] Step S18: When it is determined that post-processing is required, post-processing is performed on the spinal segment sequence output by the spinal segment sequence recognition model based on a preset post-processing algorithm to obtain a final spinal segment sequence.

[0100] In this embodiment, the preset post-processing algorithm may be an algorithm for performing anatomical rationality correction on the spinal segment sequence output by the model. For example, the anatomical rationality correction may be an algorithm for ensuring that the segment index values ​​are strictly monotonically increasing, or the anatomical index difference between adjacent segments is always 1 (no jumps or gaps).

[0101] In this embodiment, the post-processing algorithm can be a dynamic programming algorithm, which performs post-processing on the spinal segment sequence output by the spinal segment sequence recognition model through a preset dynamic programming algorithm to obtain a final spinal segment sequence. Specifically, the dynamic programming algorithm can be used to check whether the order of the spinal segments conforms to the standard anatomical order, and then the minimum cost path is selected based on the cumulative cost to correct possible sequence errors. The backtracking process of dynamic programming can readjust the order of the spinal segments according to the minimum cost path to ensure that the output spinal segment sequence conforms to the preset anatomical order (from top to bottom, or from bottom to top).

[0102] In this embodiment, the post-processing algorithm can be a greedy algorithm, which performs post-processing on the spinal segment sequence output by the spinal segment sequence recognition model through a preset greedy algorithm to obtain a final spinal segment sequence. Specifically, the greedy algorithm can correct the spinal segment sequence through a local optimization strategy. The algorithm selects the current optimal correction step each time, such as exchanging adjacent segments (solving the inversion problem) or inserting missing segments (solving the jump problem), and gradually optimizes the segment sequence based on these operations. More specifically, for each pair of adjacent segments, the greedy algorithm checks whether they are in the correct order. If the segment order is found to be wrong (for example, the order of L1 and L2 does not meet the anatomical requirements), the positions of the two segments are swapped until there is no inversion. If it is detected that some segments in the spinal segment sequence are missing (for example, the segment between L3 and L5 is not recognized), the greedy algorithm will insert a suitable label at the missing segment position to ensure the integrity of the spinal segment sequence.

[0103] In this embodiment, the post-processing algorithm can be a heuristic search algorithm, which performs post-processing on the spinal segment sequence output by the spinal segment sequence recognition model through a preset heuristic search algorithm to obtain the final spinal segment sequence. Specifically, the heuristic search algorithm is based on the combination of a cost function and a heuristic function, and finds the optimal spinal segment sequence by exploring different paths in the solution space. In this process, the heuristic search algorithm uses a heuristic function to evaluate the adjustment cost of the remaining segments, and calculates the number of modifications performed through the cost function. More specifically, the heuristic function can estimate the cost of the corrected path based on the number of remaining segments and the confidence of the model. The model will select the path with the minimum cost to correct the segment sequence. The heuristic search algorithm will explore different modification operations, such as exchanging adjacent segments, inserting missing segments, deleting abnormal segments, etc., until the order of the spinal segments meets the anatomical requirements and the confidence reaches a certain standard.

[0104] In this embodiment, the post-processing algorithm can be a genetic algorithm, which performs post-processing on the spinal segment sequence output by the spinal segment sequence recognition model through a preset genetic algorithm to obtain a final spinal segment sequence. Specifically, the genetic algorithm can generate a better spinal segment sequence by simulating the natural selection process and using operations such as crossover, mutation, and selection. Each individual (or "chromosome") represents a label sequence of a spinal segment, and the genetic algorithm gradually improves the order of the spinal segments through multiple generations of evolution. More specifically, the label of each spinal segment (for example, C1, T1, L1, etc.) is encoded as a chromosome, and each chromosome represents a possible spinal segment order. Through the crossover operation, the sequence parts of the two parent chromosomes are exchanged to generate new offspring to increase the diversity of the segment order.

[0105] In this embodiment, the post-processing algorithm can be a particle swarm optimization algorithm, which performs post-processing on the spinal segment sequence output by the spinal segment sequence recognition model through a preset particle swarm optimization algorithm to obtain the final spinal segment sequence. Specifically, particle swarm optimization defines the position (segment sequence) and movement speed (sequence adjustment direction) for each candidate solution (particle). When the particle swarm is initialized, the particle position is randomly perturbed near the original sequence. During the iteration process, each particle records its own optimal solution (historical optimal sequence) and tracks the particle swarm optimal solution at the same time. The particle updates its speed according to three components: the inertial component that maintains the original direction, the cognitive component that tends to its own optimal solution, and the social component that tends to the group optimal solution. After the position is updated, the anatomical order is forced to be corrected (to ensure that the index of the posterior segment is not less than that of the anterior segment). When the particle swarm optimal solution remains unchanged for multiple generations or the maximum number of iterations is reached, it is terminated and the optimal sequence is output.

[0106] In one possible and specific embodiment, the post-processing algorithm may be an optimization algorithm that minimizes the cost of spinal segment modification. For example, based on a preset standard spinal segment sequence, the spinal segment sequence output by the spinal segment sequence recognition model may be mapped into a continuous anatomical segment sequence using the optimization algorithm that minimizes the cost of spinal segment modification. The continuous anatomical segment sequence serves as the final spinal segment sequence.

[0107] This embodiment significantly improves the accuracy and robustness of spinal segment sequence recognition by introducing a post-processing step. Step S16 determines whether the results output by the spinal segment sequence recognition model need further processing. This step ensures the reliability and accuracy of the model output. In some cases, there may be small errors or inconsistencies in the order or labels of the spinal segments, especially when the image quality is poor. In step S18, if it is determined that post-processing is required, the spinal segment sequence output by the model is corrected and optimized using a preset post-processing algorithm to ensure that the final output spinal segment sequence conforms to the anatomical order and avoids incorrect label order and dependencies. This post-processing process provides additional accuracy guarantees for the model, making the final spinal segment recognition results more accurate and suitable for the high-precision requirements of intraoperative real-time navigation and spinal surgery.

[0108] In some embodiments, the preset post-processing algorithm is an optimization algorithm that minimizes the cost of spinal segment modification; the step of performing post-processing on the spinal segment sequence based on the preset post-processing algorithm to obtain the final spinal segment sequence includes: Step S182: Based on a preset standard spinal segment sequence, the spinal segment sequence output by the spinal segment sequence recognition model is mapped into a continuous anatomical segment sequence through an optimization algorithm that minimizes the spinal segment modification cost. The continuous anatomical segment sequence is the final spinal segment sequence.

[0109] In this embodiment, the spinal segment modification cost can be expressed as the correction cost required to map the model's original prediction sequence (i.e., the spinal segment sequence output by the spinal segment sequence recognition model) to a continuous anatomical sequence. This means minimizing interference with the original prediction sequence (preserving the true prediction results) while ensuring that the output conforms to anatomical authenticity (satisfying order and continuity).

[0110] In this embodiment, the minimization of the spinal segment modification cost may include a sequence adjustment cost. The sequence adjustment cost is used to measure the modification of the spinal segment sequence. For example, if the sequence of the spinal segments does not meet anatomical requirements, adjustment is required. The calculation of the sequence adjustment cost can be evaluated based on the relative position differences between the segments.

[0111] In this embodiment, the minimization of the spinal segment modification cost may include a label modification cost. The label modification cost is used to measure the cost of modifying the spinal segment label (e.g., C1, T1, L1, etc.). For example, if the model incorrectly labels T1 as C1 or L1 as T3, the label needs to be modified. The further the modified label deviates from the correct label, the higher the cost.

[0112] In this embodiment, the cost of minimizing spinal segment modification can be calculated by only considering the cost of label modification. In this case, the cost is calculated only for modifying spinal segment labels, without considering the segment order adjustment. In other words, the model will prioritize maintaining the original segment order and only make necessary corrections to the labels.

[0113] In this embodiment, the minimization of the spinal segment modification cost can simultaneously consider the costs of sequence adjustment and label modification. In this case, the cost calculation considers both sequence adjustment and label modification. That is, when adjusting the segment order, the model also needs to consider whether the labels are correct and calculate the total cost based on the extent of the modification. In other words, total cost = sequence adjustment cost + label modification cost.

[0114] In this embodiment, the minimization of spinal segment modification costs may also include structural intervention costs, that is, a high cost is imposed on the operation of adding / deleting segments.

[0115] In this embodiment, the minimization of the spinal segment modification cost may also include a cost based on prior knowledge. That is, it is defined based on prior knowledge of spinal anatomy, taking into account the anatomical relationship and importance between different spinal segments. For example, the order error of the cervical segments (C1, C2) may be more serious than the error of the lumbar segments (such as L1, L2). Specifically, the cost can be adjusted according to its anatomical importance by assigning different modification cost weights to different segments. For example, the label error cost of C1 is higher than the label error cost of L1. This cost can also take into account the impact of the relative position of a specific spinal segment on the modification, and adjust those segments that have a strong dependency on other segments.

[0116] In this embodiment, the minimization of the spinal segment modification cost can also include a cost based on the model confidence. Specifically, for each spinal segment, the model outputs a label category and its corresponding confidence, and the label modification cost can be weighted and adjusted according to the confidence. If the confidence of the model is high, the cost of label modification is low; conversely, if the confidence is low, the modification cost is high. This form of cost takes into account the uncertainty of the model, so that in high-confidence predictions, it is more inclined to keep the original label, and when the confidence is low, it is more inclined to make corrections, which can effectively avoid unnecessary modifications.

[0117] In a specific embodiment, step S182 may include: Because adjacent segments of the spine are highly similar, adjusting the order is not counted as a penalty during post-processing sorting; only modifying the label name is counted as a penalty. The goal is to minimize label name modifications. The specific implementation plan is as follows: First, based on the number of existing labels in the sequence (denoted as N), all consecutive subsequences of length N in the legal total sequence are enumerated as target sequences. The legal total sequence is: C1, C2, C3, C4, C5, C6, C7, T1, T2, T3, T4, T5, T6, T7, T8, T9, T10, T11, T12, L1, L2, L3, L4, L5. All subsequences in the above steps are matched with the predicted sequence, and the total penalty between each subsequence and the predicted sequence is calculated. Note that the penalty calculation here involves calculating the distance between the current sequence and the subsequence, such as the distance between L1 and T10 is 3, and the distance between L1 and T12 is 1. The total penalty sum is then compared, and the subsequence with the smallest total penalty sum is used as the label of the final sequence.

[0118] This embodiment effectively solves the problem of inconsistency in the spinal segment sequence by introducing a post-processing step based on an optimization algorithm that minimizes the modification cost of the spinal segments, ensuring that the spinal segment sequence finally output meets the anatomical continuity requirements. In step S182, the optimization algorithm minimizes the modification cost of the spinal segment sequence to ensure that the spinal segment sequence output by the recognition model can be accurately mapped into a continuous anatomical order. This optimization process not only improves the accuracy of the spinal segment sequence, but also ensures the logical consistency between the segments, avoiding the problem of sequence errors or segment jumps. Through this optimization strategy, the spinal segment recognition results are more stable, which greatly improves the reliability of surgical navigation and spinal positioning. In particular, during real-time processing during surgery, it can provide more accurate guidance, which helps to improve the safety and success rate of the operation.

[0119] In some embodiments, the step of mapping the spinal segment sequence output by the spinal segment sequence recognition model into a continuous anatomical segment sequence based on a preset standard spinal segment sequence by using an optimization algorithm that minimizes the spinal segment modification cost includes: Step S1822: In a preset standard spinal segment sequence, enumerate all continuous subsequences of length N; wherein N is the number of segments in the spinal segment sequence output by the spinal segment sequence recognition model.

[0120] In this embodiment, the preset standard spinal segment sequence can be a preset anatomical standard sequence, comprising 24 spinal segments arranged in a natural anatomical order. By traversing the standard spinal segment sequence, all continuous subsequences of length N are constructed. Each subsequence is composed of N continuous spinal segments. In other words, each continuous subsequence represents a possible spinal segment sequence.

[0121] Step S1824: For each continuous subsequence, calculate the total label modification cost between it and the spinal segment sequence; wherein the total label modification cost is defined as the sum of the absolute values ​​of the difference between the anatomical order index value of each segment label in the continuous subsequence and the anatomical order index value of the segment label at the corresponding position in the spinal segment sequence.

[0122] In this embodiment, for each spinal segment subsequence, the label modification cost is calculated. The label modification cost is measured by the difference between the position index of each segment label in the standard spinal segment sequence and the label index of the corresponding position in the spinal segment sequence predicted by the model. Assume that the label index of the continuous subsequence is , and the predicted label index of the spinal segment sequence is , then the computational cost is: | - That is, by calculating the absolute difference of the label position index of each segment, the label modification cost of each subsequence is obtained.

[0123] For each continuous subsequence, the total label modification cost is the sum of the label modification costs of all segments in the subsequence. In other words, the total cost is obtained by adding up the label modification costs of each segment.

[0124] In this embodiment, the total label modification cost may be: candidate In the formula, Output the i-th segment label of the sequence for the model, is the i-th segment label of the continuous subsequence. is the anatomical position number of the query tag in the standard sequence.

[0125] Step S1826: Select the continuous subsequence with the minimum total label modification cost as the final spinal segment sequence.

[0126] This embodiment first enumerates all possible continuous subsequences through the optimization process of step S1822 to step S1826, and calculates the label modification cost of each subsequence compared with the standard anatomical order. By calculating the total label modification cost, the method can accurately quantify the difference between the spinal segment order and the standard order, ensuring that the subsequence with the lowest cost is selected in the subsequent optimization process. Ultimately, by selecting the continuous subsequence with the lowest total label modification cost, the model can accurately adjust the spinal segment sequence to conform to the standard anatomical order, significantly improving the accuracy and reliability of spinal segment identification, especially in complex or abnormal situations, and can effectively avoid sequence errors and label deviations, thereby improving the accuracy of surgical navigation and spinal positioning.

[0127] This embodiment provides a training method for a spinal segment sequence recognition model, which can be executed by a deep learning server. The method may include: Step S22: Acquire training data; wherein the training data includes a plurality of image data for representing at least two spinal segments and corresponding label data.

[0128] In this embodiment, the training data may include: The first type of training data is used to represent a connection sequence across anatomical regions, which includes a continuous segment between the caudal cervical segment and the cranial thoracic segment, a continuous segment between the caudal thoracic segment and the cranial lumbar segment, or a continuous segment between the caudal thoracic segment and the complete lumbar segment; The second type of training data is used to represent an incomplete sequence of a single anatomical region, which includes a partial continuous segment of a cervical vertebrae segment, a thoracic vertebrae segment, or a lumbar vertebrae segment.

[0129] In some embodiments, the training data may also include: The third type of training data consists of spinal segment sequences with anatomical variations. This type of training data includes sequences of spinal segments with anatomical variations or abnormal morphologies, such as scoliosis, herniated discs, or congenital spinal deformities. These variations can complicate the relationships between segments, increasing the difficulty of segment identification. Using this type of data, the model can learn how to identify spinal segments in non-standard anatomical structures, enhancing its ability to handle anatomical variations and abnormalities.

[0130] The fourth type of training data is multimodal image data of spinal segments. Specifically, this type of training data can include multimodal spinal imaging data, such as CT data, MRI data, X-rays, and CBCT data. These images from different modalities provide multi-view and multi-dimensional information about the spinal segments.

[0131] The fifth type of training data consists of spinal segment sequences of different age groups and genders. This type of training data covers spinal segment sequences of different age groups and genders. Spinal structure may differ between different age groups and genders, and these differences may affect the accuracy of spinal segment recognition.

[0132] The sixth type of training data consists of spinal imaging data captured under various scanning conditions. This data includes spinal imaging data captured under different scanning conditions, including varying resolutions, scanning angles, and image noise levels. This data allows the model to learn how to maintain high recognition accuracy under varying scanning conditions, improving its adaptability to low-quality or noisy data.

[0133] Step S24: Input the training data into the spinal segment sequence recognition model to be trained for training; wherein, the spinal segment sequence recognition model to be trained includes at least a first module and a second module; wherein, the first module is a module for performing serialized input construction on the input image data, and the second module is a neural network representing sequence to sequence, and the neural network is used to extract the anatomical dependency between the at least two spinal segments; wherein, the spinal segment sequence recognition model is trained using a preset loss function; wherein, the preset loss function includes at least a first loss function, and the first loss function is a loss function for characterizing the spinal segment sequence constraint, and the sequence is based on the axial sequence of the spinal anatomy.

[0134] Step S26: After the preset conditions are met, a spinal segment sequence recognition model is obtained.

[0135] In this embodiment, the preset condition may be that the number of training epochs reaches a preset number.

[0136] During training, the model undergoes multiple rounds of training. Each round involves either traversing the entire training dataset or performing random sampling with replacement within the dataset according to a preset number of samplings. This means that each example may be sampled multiple times. A preset condition can be that training is considered complete when the maximum number of training rounds is reached.

[0137] In this embodiment, the preset condition may be that the model performance reaches a preset standard.

[0138] In this embodiment, the preset condition may be that the loss function converges.

[0139] In some embodiments, the training data includes at least: The first type of training data is used to represent a connection sequence across anatomical regions, which includes a continuous segment between the caudal cervical segment and the cranial thoracic segment, a continuous segment between the caudal thoracic segment and the cranial lumbar segment, or a continuous segment between the caudal thoracic segment and the complete lumbar segment; The second type of training data is used to represent an incomplete sequence of a single anatomical region, which includes a partial continuous segment of a cervical vertebrae segment, a thoracic vertebrae segment, or a lumbar vertebrae segment.

[0140] This embodiment effectively enhances the generalization ability and accuracy of the spinal segment sequence recognition model by introducing the first type of training data and the second type of training data. The first type of training data includes cross-anatomical region connection sequences. This type of data covers the cross-regional relationship between spinal segments (the connection between the cervical vertebrae and the thoracic vertebrae, and the thoracic vertebrae and the lumbar vertebrae), enabling the model to learn the transition and connection between different anatomical regions, and improving the ability to understand complex anatomical structures. The second type of training data represents partial continuous segments of a single anatomical region. This type of data helps the model focus on the identification of local segments. Through these diverse training data, the model can better adapt to the conditions of different spinal segments, enhance its application capabilities in different clinical scenarios, and thereby improve the accuracy and robustness of spinal segment recognition.

[0141] In some embodiments, the preset loss function also includes a second loss function, which is used to supervise the classification accuracy of the spinal segments; wherein the preset loss function is expressed as a preset weight multiplied by the first loss function + the second loss function; wherein the preset weight is used to control the intensity of the influence of the anatomical order constraint on the total loss.

[0142] In this embodiment, the preset loss function may be: In the formula, N represents the total number of spinal segments contained in a single input image. It can be expressed as a classification loss function (the second loss function). CE is expressed as a cross entropy loss function, which is used to calculate the difference between the predicted probability distribution and the true distribution. is the predicted value of the i-th segment, that is, the predicted category probability distribution of each spinal segment, is the true value of the i-th segment, that is, the true category of each spinal segment. The preset weight. is the first loss function.

[0143] This embodiment further improves the classification accuracy and anatomical sequence accuracy of the spinal segment sequence recognition model by introducing a second loss function. The second loss function is used to supervise the classification accuracy of the spinal segments to ensure that the model can accurately identify the category of each spinal segment when classifying spinal segments. This loss function is combined with the first loss function (used to constrain the anatomical order of the spinal segments) to form a comprehensive optimization goal. During the training process, by adjusting the preset weights, the relationship between the anatomical order constraints and the classification accuracy can be balanced, thereby optimizing the model's performance in sequence correctness and classification accuracy. In this way, the model can not only ensure that the order of the spinal segments meets the anatomical requirements, but also accurately classify each segment, improve the overall recognition effect, and provide more reliable and accurate recognition results, especially in complex spinal anatomical structures.

[0144] In one specific embodiment, a method for identifying spinal segment names based on the Vision-Transformer (VIT) architecture is provided. The spine (also known as the vertebral column) is the core structure of the human axial skeleton. It is composed of 23-24 vertebrae connected by intervertebral discs, ligaments, and joints. From top to bottom, it is divided into the cervical vertebrae (7 segments), thoracic vertebrae (12 segments), lumbar vertebrae (5 segments), and sacral vertebrae. Anatomically, a typical vertebra consists of the vertebral body (the load-bearing body), the vertebral arch (forming the vertebral foramen), and the transverse and spinous processes (muscle attachment points). Vertebral morphology varies significantly across different segments. For example, the transverse processes of the cervical vertebrae contain vertebral arterial foramina, and the thoracic vertebral bodies form articulations with the ribs. The fibrocartilage structure between adjacent vertebrae consists of an annulus fibrosus (resistance to tension) and a central, gelatinous nucleus pulposus (pressure buffering). The spine exhibits three curves: cervical (lordosis), thoracic (kyphosis), lumbar (lordosis), and sacral (kyphosis), contributing to the mechanical balance of upright posture. The anterior longitudinal ligament, posterior longitudinal ligament, and ligamentum flavum maintain vertebral stability, while deep muscles such as the erector spinae and multifidus control spinal movement. The core functions of the spine include mechanical support, neuroprotection, motor function, and biomechanical coordination.

[0145] In clinical applications, clinicians often identify spinal segments based on the ribs, sacrum, or atlantoaxial joint (thoracic vertebrae are connected to the ribs). Robotic-assisted surgery relies on automatically identifying each vertebra (segment) and assigning standardized names (such as C1, T1, L5, etc.). Both preoperative CT images and intraoperative CBCT images may not include ribs or sacrum as reference, which greatly inconveniences surgical planning before or during surgery. However, existing methods, which are mostly based on image template matching or segment-by-segment classifiers, lack the ability to model the overall sequence structure and are susceptible to factors such as image quality, viewing angle, and missing segments.

[0146] The similarity in strength of adjacent spinal structures poses a significant challenge to single-segment segmentation, especially when the differences between adjacent thoracic vertebrae are minimal. Therefore, developing a high-precision spinal segment naming method is of great clinical significance. The Transformer architecture, due to its powerful performance in sequence modeling, is widely used in fields such as natural language processing and computer vision. It has a low dependency on input order and can model sequential relationships through positional encoding, providing new insights into modeling spinal segment structure.

[0147] The purpose of this implementation is to overcome the shortcomings of existing technologies and propose a neural network-based single-segment spinal segmentation method while ensuring speed and accuracy. This method reduces the workload of segmentation, saves time, and makes spinal segmentation more efficient and rapid. It should be noted that the implementation of this implementation requires that the masks of each spinal segment have been generated.

[0148] In this embodiment, the direction from one transverse process to the other transverse process can be recorded as the coordinate X-axis direction, the direction from the spinal process to the vertebral body can be recorded as the coordinate Y-axis direction, and the direction from the sacrum to the cervical vertebra can be recorded as the coordinate Z-axis direction.

[0149] Step 1-1: Data preparation.

[0150] Collect CBCT data, including 512×512×512 spinal CBCT volume data and the corresponding single-segment vertebral mask. Based on the CBCT volume data, N represents the number of CBCT volume data layers, and 512×512 represents the size of each CBCT layer. Since CT data from the cervical spine to the sacral spine is generally not collected in actual use, the training set must include at least the following categories of images: 1. Only partial cervical and thoracic vertebrae CBCT data are available.

[0151] 2. Only partial thoracic and lumbar CBCT data are available.

[0152] 3. CBCT data of only partial thoracic spine and complete lumbar spine.

[0153] 4. Only CBCT data of the cervical spine.

[0154] 5. Only CBCT data of the thoracic spine are available.

[0155] 6. Only CBCT data of the lumbar spine are available.

[0156] During the implementation of this implementation plan, a total of 480 sets of data and corresponding masks and segment names were collected.

[0157] Step 1-2: Build a neural network, model the input sequence, and classify the output sequence.

[0158] In Vision-Transformer, to process 2D images, the image Interpolate (reshape) into a flattened 2Dpatches sequence , where (H, W) is the original image resolution, C is the number of original image channels (RGB image C=3), and (P, P) is the resolution of each image patch. It is also the effective input sequence length of VisionTransformer. Transformer flattens the image patches using a constant latent vector size D in all its layers and uses a trainable linear projection (FC layer) to transform the dimensions Mapped to D dimension while keeping the number of image patches N unchanged.

[0159] In this embodiment, due to the limited imaging range of the CBCT imaging device, a single image can only contain a maximum of six vertebrae. Therefore, a total of six patches are divided. Because the morphology of the sacral vertebrae is significantly different from that of the cervical, thoracic, and lumbar vertebrae and is larger in size, it does not need to be identified using this method. Depending on the imaging device, the image size is not exactly the same. In this embodiment, CBCT images acquired by Siemens CIOS SPIN are used, and the image size is 512*512*512. The maximum size of the spinal segment (max_x, max_y, max_z) is determined by counting the AABB bounding boxes of all spinal segment masks in the statistical data set. With the center of mass of the mask as the center, the image is cropped in the positive and negative directions of max_x / 2, max_y / 2, and max_z / 2 in the x, y, and z directions respectively. If the number of segments in the image is less than 6, a matrix of all zeros of the same size is added at the end.

[0160] Due to the large size of the spine mask image, three-dimensional convolution and three-dimensional pooling calculations are performed before inputting it into the neural network to extract the mask features while reducing the data dimension. The specific neural network architecture is shown in the figure below. Figure 4 shown.

[0161] Steps 1-3: Loss function design.

[0162] VIT is essentially an unordered input modeler. Although we can sort the segments from top to bottom according to the z-coordinate, the model cannot guarantee that the output labels are strictly increasing or continuous. For example, the output order of the model is not necessarily L1, L2, L3, L4, L5, but may be L2, L1, L3, L4, L5, or even random order. Therefore, it is necessary to introduce a loss function to supervise the predicted spinal segment order. In this implementation, sequence_order_loss is used to ensure that the predicted label index is increasing through sequential supervision loss.

[0163] Among them: classification_loss is the loss of spinal segment classification; order_loss is the loss of continuous supervision of spinal segment naming results; penalty_weight is used to control the influence of the loss of continuous supervision.

[0164] penalty_weight=0: order is not considered, the model only considers classification accuracy; penalty_weight=0.1: order is important, but does not overwhelm classification; penalty_weight=1.0: order is as important as classification; penalty_weight=5.0+: strongly penalizes disorder and may overfit the order.

[0165] In this embodiment, penalty_weight=0.5.

[0166] Steps 1-4: Post-processing sorting.

[0167] Since adjacent segments of the spine are highly similar, when performing post-processing sorting, adjusting the order is not counted as a penalty, and only modifying the label is counted as a penalty. The goal is to minimize the modification of label names. The specific implementation plan is as follows: According to the number of existing labels in the sequence (denoted as N), all consecutive subsequences of length N in the legal total sequence are enumerated as the target sequence. The legal total sequence is: C1, C2, C3, C4, C5, C6, C7, T1, T2, T3, T4, T5, T6, T7, T8, T9, T10, T11, T12, L1, L2, L3, L4, L5.

[0168] Try to match all the subsequences in the above steps with the predicted sequence, and calculate the total penalty between each subsequence and the predicted sequence. Note that the penalty calculation here is to calculate the distance between the current sequence and the subsequence, such as the distance between L1 and T10 is 3, and the distance between L1 and T12 is 1. Compare the penalty sums and take the subsequence with the smallest penalty sum as the label of the final sequence.

[0169] In summary, due to the high similarity between adjacent spinal segments, accurate sequence name recognition is difficult. Therefore, during neural network training, sequence labeling is supervised to ensure continuous and increasing sequence labels. Even under these conditions, it is not possible to completely guarantee the continuity of sequence labels, so post-processing is required to generate the final continuous sequence. In this post-processing, a dynamic programming algorithm is employed, which penalizes modifying label names while excluding swapping adjacent segment labels. This ensures that the continuity of spinal segment sequence names is maintained while minimizing segment name modifications.

[0170] According to an embodiment of the present invention, an electronic device is provided. The electronic device in this embodiment may include one or more of the following components: a processor, a network interface, a memory, a non-volatile memory, and one or more application programs, wherein the one or more application programs may be stored in the non-volatile memory and configured to be executed by the one or more processors, and the one or more programs are configured to perform the method described in the aforementioned method embodiment.

[0171] According to an embodiment of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a computer, the computer executes the method described in any one of the above embodiments.

[0172] According to an embodiment of the present invention, a computer program product comprising instructions is further provided. When the instructions are executed by a computer, the computer is enabled to perform a method in any one of the above embodiments.

[0173] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0174] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, and this embodiment will not be described in detail here.

[0175] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0176] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0177] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for identifying spinal segment sequences, characterized in that: The method comprises: determining image data representing at least two spinal segments; The image data is input into a preset spinal segment sequence recognition model for recognition to obtain a spinal segment sequence corresponding to the image data; wherein the spinal segment sequence recognition model includes at least a first module and a second module; wherein the first module is a module for performing serialized input construction on the input image data, and the second module is a neural network representing sequence to sequence, and the neural network is used to extract the anatomical dependency between the at least two spinal segments; wherein the spinal segment sequence recognition model is pre-trained using a preset loss function; wherein the preset loss function includes at least a first loss function, and the first loss function is a loss function for characterizing the spinal segment sequence constraint, and the sequence is based on the axial sequence of the spinal anatomy.

2. The method according to claim 1, characterized in that The first module includes: an anatomical position encoding unit, configured to sort each spinal segment based on spatial coordinate information of each spinal segment along the spinal anatomical axis; The three-dimensional feature extraction unit is used to perform three-dimensional convolution operations and three-dimensional pooling operations on the sorted segment images in sequence, and output a feature vector sequence after dimensionality reduction.

3. The method according to claim 2, characterized in that The second module includes: a sequence feature extraction unit, which is represented by a sequence modeling network with a self-attention mechanism, and whose input is the feature vector sequence, and is used to extract the anatomical dependencies between spinal segments; An independent classification unit is connected in series with the sequence feature extraction unit and is used to independently map the hidden vector of each spinal segment into a probability distribution of a predefined spinal segment category.

4. The method according to claim 3, characterized in that The preset spinal segment sequence recognition model is a spinal segment sequence recognition model based on the Vision-Transformer architecture, and the second module is a Transformer encoder and an MLP classification head connected in series with it.

5. The method according to claim 1, wherein The image data is three-dimensional mask image data carrying at least two spinal segments or three-dimensional medical image data carrying at least two spinal segments; when the image data is three-dimensional medical image data carrying at least two spinal segments, the step of determining image data for representing at least two spinal segments includes: receiving at least one three-dimensional medical image data carrying at least two spinal segments; Preprocessing is performed on the at least one three-dimensional medical image data carrying at least two spinal segments based on a preset image preprocessing algorithm to obtain preprocessed three-dimensional medical image data.

6. The method according to claim 1, wherein After the step of inputting the image data into a preset spinal segment sequence recognition model for recognition to obtain a spinal segment sequence corresponding to the image data, the method further includes: Determine whether the spinal segment sequence output by the spinal segment sequence recognition model needs to be post-processed; When it is determined that post-processing is required, post-processing is performed on the spinal segment sequence output by the spinal segment sequence recognition model based on a preset post-processing algorithm to obtain a final spinal segment sequence.

7. The method according to claim 6, characterized in that The preset post-processing algorithm is an optimization algorithm that minimizes the cost of spinal segment modification; the step of performing post-processing on the spinal segment sequence based on the preset post-processing algorithm to obtain the final spinal segment sequence includes: Based on a preset standard spinal segment sequence, the spinal segment sequence output by the spinal segment sequence recognition model is mapped into a continuous anatomical segment sequence through an optimization algorithm that minimizes the spinal segment modification cost. The continuous anatomical segment sequence is the final spinal segment sequence.

8. The method according to claim 7, characterized in that The step of mapping the spinal segment sequence output by the spinal segment sequence recognition model into a continuous anatomical segment sequence based on a preset standard spinal segment sequence by using an optimization algorithm that minimizes the spinal segment modification cost comprises: In a preset standard spinal segment sequence, all continuous subsequences of length N are enumerated; where N is the number of segments in the spinal segment sequence output by the spinal segment sequence recognition model; For each continuous subsequence, the total label modification cost between it and the spinal segment sequence is calculated; wherein the total label modification cost is defined as the sum of the absolute values ​​of the difference between the anatomical order index value of each segment label in the continuous subsequence and the anatomical order index value of the segment label at the corresponding position in the spinal segment sequence; The continuous subsequence with the minimum total label modification cost is selected as the final spinal segment sequence.

9. A training method for a spinal segment sequence recognition model, characterized in that: include: Acquire training data; wherein the training data includes a plurality of image data for representing at least two spinal segments and corresponding label data; The training data is input into a spinal segment sequence recognition model to be trained for training; wherein the spinal segment sequence recognition model to be trained includes at least a first module and a second module; wherein the first module is a module for performing serialized input construction on the input image data, and the second module is a neural network representing sequence to sequence, and the neural network is used to extract the anatomical dependency between the at least two spinal segments; wherein the spinal segment sequence recognition model is trained using a preset loss function; wherein the preset loss function includes at least a first loss function, and the first loss function is a loss function for characterizing spinal segment sequence constraints, and the sequence is based on the axial sequence of the spinal anatomy; After the preset conditions are met, the spinal segment sequence recognition model is obtained.

10. The method according to claim 9, characterized in that The training data at least includes: The first type of training data is used to represent a connection sequence across anatomical regions, which includes a continuous segment between the caudal cervical segment and the cranial thoracic segment, a continuous segment between the caudal thoracic segment and the cranial lumbar segment, or a continuous segment between the caudal thoracic segment and the complete lumbar segment; The second type of training data is used to represent an incomplete sequence of a single anatomical region, which includes a partial continuous segment of a cervical vertebrae segment, a thoracic vertebrae segment, or a lumbar vertebrae segment.

11. The method according to claim 9, characterized in that The preset loss function also includes a second loss function, which is used to supervise the classification accuracy of the spinal segments; wherein the preset loss function is expressed as a preset weight multiplied by the first loss function + the second loss function; wherein the preset weight is used to control the intensity of the influence of the anatomical order constraint on the total loss.

12. An electronic device, characterized in that: include: a memory, and one or more processors communicatively coupled to the memory; The memory stores instructions that can be executed by the one or more processors, and the instructions are executed by the one or more processors to enable the one or more processors to implement the method according to any one of claims 1 to 8 or the method according to any one of claims 9 to 11.