Scoliosis classification model training method and scoliosis classification method and system

By constructing a class-balanced dataset and optimizing feature extraction methods, combined with the Mamba structure and a bidirectional state-space model, the problems of class imbalance and computational resource consumption in scoliosis detection are solved, achieving efficient scoliosis detection.

CN120913007APending Publication Date: 2025-11-07XIN HUA HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511096951.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing point cloud detection methods suffer from class imbalance in scoliosis datasets, resulting in insufficient detection accuracy and generalization ability. Furthermore, traditional methods are deficient in feature extraction capabilities and computational resource consumption.

Method used

By constructing a class-balanced meta-dataset and a basic dataset, and combining farthest point sampling and optimal transmission algorithms to optimize point cloud feature extraction, we adopt the Mamba structure and attention mechanism, use a bidirectional state-space model to optimize the computation process, and introduce a Bias model to adjust sample weights, thereby reducing computational resource requirements.

Benefits of technology

It improves the accuracy and generalization ability of scoliosis detection, enhances feature extraction capabilities, reduces computational resource requirements, and enables the detection method to run effectively on ordinary equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913007A_ABST
    Figure CN120913007A_ABST
Patent Text Reader

Abstract

The invention relates to the field of spine detection, and provides a scoliosis classification model training method, a scoliosis classification method and a scoliosis classification system in order to reduce computing resources and improve detection precision, and the scoliosis classification model is trained by constructing a class-balanced metadata set and a class-unbalanced basic data set. The model is guided to learn features of each category, the problem of category imbalance in a scoliosis data set is solved, and the generalization ability of the model under different scoliosis angle categories is effectively improved, so that the detection precision is improved; a Mama structure and an attention mechanism are introduced into the scoliosis classification model, and by optimizing a calculation process and a feature extraction mode, on the premise that the detection precision is guaranteed, the requirement for calculation resources is reduced, and the detection method can be effectively operated on common equipment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of spinal column detection, and in particular to a scoliosis classification model training method, a scoliosis classification method and system. BACKGROUND

[0002] Scoliosis is a common spinal deformity that often occurs in adolescents. Early detection and intervention are crucial to prevent the condition from worsening. According to medical statistics, the prevalence of scoliosis is about 2% to 3%, and most patients are mild (less than 10 degrees).

[0003] In recent years, three-dimensional point cloud-based spinal morphology detection has been widely explored, and representative methods include:

[0004] (1) Deep learning scheme based on PointNet++

[0005] This scheme first uses statistical outlier rejection (SOR) and voxel grid (Voxel Grid) downsampling to denoise and simplify the original back point cloud, then calculates the normal vector of each point to enhance the spatial geometric expression; then uses the hierarchical aggregation structure of PointNet++, to learn local-global features of the back point cloud, and finally outputs the "normal / mild / moderate / severe" four-classification results.

[0006] (2) Point cloud scheme based on Transformer

[0007] This scheme serializes the point cloud into a high-dimensional token of "coordinates + normal vectors + density", and introduces a sine-cosine position encoding, so that the Transformer can use the multi-head self-attention mechanism to describe the global correlation between any two points, thereby improving the modeling ability of complex three-dimensional curvature.

[0008] The existing point cloud methods have the following problems:

[0009] Class imbalance: existing point cloud detection methods are usually trained based on deep learning models, but in the scoliosis data set, most samples are concentrated in the "0-10 degree" normal range, resulting in class imbalance, affecting the detection accuracy and generalization ability of the model.

[0010] Limited feature extraction capability: traditional point cloud processing methods (such as PointNet) have limited ability to capture the relationship between points at a long distance, and cannot effectively identify the complex morphological features of the spine.

[0011] Large consumption of computing resources: when using high-performance models such as Transformer, the demand for computing and storage resources is high, making it difficult for the detection system to run efficiently on ordinary devices. SUMMARY

[0012] In order to reduce the computing resources while improving the detection accuracy, the application provides a scoliosis classification model training method, a scoliosis classification method and a system.

[0013] The technical scheme adopted by the application to solve the above problems is:

[0014] The scoliosis classification model training method comprises:

[0015] Step 1: Obtain training samples, the training samples comprising back point cloud data and corresponding classification result labels;

[0016] Step 2: Create a basic data set and a meta data set based on the training samples, the basic data set being generated by random sampling from the training samples, and the meta data set being generated by class-balanced sampling from the training samples;

[0017] Step 3: Construct a BaseMamba model, a MetaMamba model and a Bias model, the BaseMamba model and the MetaMamba model having the same structure and sharing parameters, the BaseMamba model being used to extract back point cloud data features in the basic data set, the extracted features being processed based on an encoder and classification being completed through a classification head; the MetaMamba model being used to extract back point cloud data features in the meta data set, the extracted features being processed based on an encoder and classification being completed through a classification head; and the Bias model being used to predict class imbalance bias;

[0018] Step 4: Input a batch of data in the basic data set into the MetaMamba model to calculate an initial loss of each sample; use the Bias model to assign a weight to each initial loss , and weight the initial losses based on the weights to calculate a pseudo loss; update the MetaMamba model parameters based on the pseudo loss; input a batch of data in the meta data set into the updated MetaMamba model to calculate a meta loss of each sample; and update the Bias model parameters based on the meta loss;

[0019] Use the BaseMamba model to calculate a true loss of the current batch of the basic data set, and input the true loss into the updated Bias model to obtain a weight ; weight the true loss based on the weight to calculate a total loss; and update the BaseMamba model parameters based on the total loss;

[0020] Repeat step 4 until the training completion requirement is met, and the trained BaseMamba model is the scoliosis classification model.

[0021] Further, the BaseMamba model and the MetaMamba model both include:

[0022] a feature embedding module for converting the back point cloud data into a high-dimensional feature vector;

[0023] a feature vector processing module for adding global information and position encoding to the high-dimensional feature vector;

[0024] an encoder stacked by T identical encoder layers, each layer including a local feature extraction unit for capturing local features and a global feature extraction unit for capturing long-range dependencies between features; and outputting a feature representation that integrates local features and global features;

[0025] a classification module for completing classification based on the output result of the encoder through a classification head.

[0026] Further, the operation steps of the feature embedding module are as follows: selecting a center point; constructing a local neighborhood based on the center point; and performing spatial pose alignment, local feature extraction, feature alignment, and global feature aggregation on the point cloud data of each local neighborhood through a PointNet network, and finally generating G high-dimensional feature vectors.

[0027] Further, the operation steps of the local feature extraction unit are as follows:

[0028] calculating the normalized difference of neighbor features relative to the center feature ;

[0029] concatenating and the center point feature ;

[0030] generating new features through a learnable linear transformation , wherein represents feature concatenation, and are linear transformation parameters;

[0031] the updated center point feature is calculated as follows: , where k is the number of neighbor features, is the i-th new feature, is the exponential function value of the i-th neighbor feature.

[0032] Further, the operation steps of the global feature extraction unit are as follows:

[0033] forward scanning: scanning the input sequence from head to tail to capture dependencies in one direction;

[0034] Reverse scan: flip the channel of the feature vector of the input sequence and then scan it;

[0035] Fusion: add the original input, the output of the forward scan and the output of the reverse scan to get the final result.

[0036] Further, step 2 further comprises: preprocessing the back point cloud data in the basic data set and the metadata set, wherein the preprocessing refers to feature aggregation and dimension reduction of the acquired point cloud data based on the farthest point sampling algorithm and the optimal transport algorithm.

[0037] Further, the specific steps of preprocessing are: generating an initial class center set based on the farthest point sampling algorithm; dynamically optimizing the class center through the optimal transport algorithm to minimize the transport cost of the point cloud distribution to the class center distribution; and outputting structured point cloud data containing class centers and nearest neighbor points.

[0038] Further, the training completion requirement is that the number of training reaches a preset training threshold, or the classification accuracy of the model on the validation set is not less than a preset target precision threshold, or the model shows a convergence trend in continuous several training periods.

[0039] The scoliosis classification method comprises:

[0040] Obtaining the back point cloud data to be classified and preprocessing the same, wherein the preprocessing refers to feature aggregation and dimension reduction of the acquired point cloud data based on the farthest point sampling algorithm and the optimal transport algorithm;

[0041] Classifying the preprocessed back point cloud data to be classified by using the trained scoliosis classification model, wherein the scoliosis classification model is trained by using the scoliosis classification model training method.

[0042] The scoliosis classification system comprises:

[0043] The point cloud acquisition module is configured to obtain the back point cloud data to be classified;

[0044] The OT sampling module is configured to preprocess the acquired back point cloud data to be classified, wherein the preprocessing refers to feature aggregation and dimension reduction of the acquired point cloud data based on the farthest point sampling algorithm and the optimal transport algorithm;

[0045] The scoliosis classification module is configured to classify the preprocessed back point cloud data to be classified by using the trained scoliosis classification model, wherein the scoliosis classification model is trained by using the scoliosis classification model training method.

[0046] The present application has the beneficial effects compared with the prior art: by constructing a class-balanced metadata set and a class-unbalanced basic data set to train the scoliosis classification model, guiding the model to learn the features of each class, solving the problem of class imbalance in the scoliosis data set, effectively improving the generalization ability of the model under different scoliosis angle classes, thereby improving the detection accuracy; the scoliosis classification model introduces the Mamba structure and attention mechanism, by optimizing the calculation process and feature extraction method, under the premise of ensuring the detection accuracy, reduces the demand for computing resources, so that the detection method can effectively run on ordinary devices; when preprocessing the point cloud data, combining the optimal transport to dynamically extract key information in the point cloud, enhancing the feature extraction capability, avoiding the problem of insufficient processing of long-distance point relationships in traditional methods, improving the accuracy and robustness of point cloud data processing, and enhancing the recognition ability of the model for different scoliosis shapes. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1 A scoliosis classification model training method flowchart is provided.

[0048] Figure 2 A lightweight PointNet model structure diagram is provided.

[0049] Figure 3 A T-Net network structure diagram is provided.

[0050] Figure 4 An encoding layer structure diagram is provided.

[0051] Figure 5 A bi-SSM module schematic diagram is provided.

[0052] Figure 6 A model training structure diagram is provided. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below with examples. It should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application.

[0054] As shown in Figure 1 The scoliosis classification model training method comprises:

[0055] Step 1: Obtain training samples, the training samples comprising back point cloud data and corresponding classification result labels.

[0056] Step 2: Create a basic data set and a metadata set based on the training samples, the basic data set being generated by randomly sampling the training samples, and the metadata set being generated by class-balanced sampling the training samples.

[0057] The base dataset is generated by random sampling from the training samples, and the sample category distribution is uneven, for example, in the setting of batchsize of 8 and total category number of 4, the sample category that can be extracted is The metadata set is generated by class-balanced sampling from the original training data, ensuring that the number of samples of each category is relatively balanced. For example, in the setting of batchsize of 8 and total category number of 4, the sample category that can be extracted is .

[0058] In order to enhance the feature extraction ability, the application proposes a method combining Farthest Point Sampling (FPS) and Optimal Transport (OT) for optimizing the selection of class centers in point cloud classification. The method is divided into two main stages: initial class center selection and dynamic class center optimization.

[0059] (1) Initial class center selection

[0060] Given a point cloud dataset , where P represents the entire point cloud dataset, containing N three-dimensional points, , represents the i-th point, which is a three-dimensional coordinate vector , first use the farthest point sampling algorithm to select the initial class center. FPS realizes it by gradually selecting the point farthest from the selected point set , its definition is:

[0061] ,

[0062] where represents the initialized class center set, S represents the current selected point set, P\S represents the point set that has not been selected, represents the Euclidean distance between point p and the selected point s, represents the shortest distance from point p to the current selected point set S, represents the point with the maximum distance from the current point set, that is, the point farthest from the current point set.

[0063] (2) Dynamic class center optimization

[0064] In the initial class center , where is the set of initial class centers, represents G initial class centers, represents the three-dimensional coordinates of the j-th initial class center, and the superscript 0 indicates that this is the initial state of the class center, and the subscript , After the determination of the class index, i.e., G target classes, the class centers are optimized using the optimal transport method. Let the point set of each class be with the weight distribution , where w j represents the weight distribution of each point in the jth class, all elements are non-negative and the sum is 1, indicating the relative importance of the point in the overall point set; the distribution of the initial class center is , representing the target distribution of the class center, the expected distribution of the point cloud mapping to the G class centers, also satisfying non-negative and sum of 1. The optimization goal is to minimize the transport cost of the point cloud distribution to the class center distribution:

[0065] ,

[0066] where represents the joint distribution set satisfying the transport boundary constraint, representing the quality of the assignment from the point to the class center, which needs to satisfy the row sum and the column sum .

[0067] ,

[0068] The class center is updated by the optimal transport mapping matrix :

[0069] ,

[0070] where represents the number of iterations.

[0071] After the initial class center is generated by the FPS, the class center is dynamically adjusted through several OT optimization iterations. Finally, the class center set is used for feature extraction and prediction of the point cloud classification model, and the K nearest points of the class center obtained according to the OT algorithm form the point cloud data information representing the class center, finally forming a point cloud data of size [B, G, K, 3].

[0072] Step 3: Build BaseMamba model, MetaMamba model and Bias model, the BaseMamba model and the MetaMamba model have the same structure and share parameters, the BaseMamba model is used to extract the feature of the back point cloud data in the basic data set, the extracted feature is processed based on the encoder and the classification is completed through the classification head; the MetaMamba model is used to extract the feature of the back point cloud data in the meta data set, the extracted feature is processed based on the encoder and the classification is completed through the classification head; the Bias model is used to predict the class imbalance bias, which is composed of three Fully Connected Layers, the input is the class imbalance loss value of the sample, and the output is the prediction of the weight of each sample.

[0073] The BaseMamba model and the MetaMamba model each include a feature embedding module, a feature vector processing module, an encoder, and a classification module.

[0074] Feature embedding module: This module is responsible for converting the input raw point cloud (a bunch of three-dimensional coordinate points) into a series of "features" or "tokens" with rich features, preparing for subsequent encoder processing. This process consists of three steps as follows:

[0075] Selecting a center point: Here, the class center set is directly used as the center point of each group;

[0076] Constructing a local neighborhood: Here, the point cloud data is directly used ;

[0077] Extracting features: From the previous steps, we have obtained point cloud data containing G groups. Here, we need to convert the geometric information of each point cloud data into a high-dimensional feature vector. A lightweight PointNet network is used here, such as Figure 2 , which first receives a three-dimensional point cloud as input and immediately learns a transformation matrix through a network called T-net to align the spatial pose of the entire point cloud, eliminating the effects of rotation and other factors. After alignment, the model will independently extract high-dimensional local features for each point through a multi-layer perceptron (MLP). Then, the model will again use a T-net network to align the extracted features themselves. Finally, the model uses a Max Pooling operation to aggregate the feature information of all points into a global feature vector. Input: G K × 3 point cloud position information; Output: G C-dimensional feature vector, which is the feature embedding (Patch Embeddings), representing the geometric features of each local region. The T-net network is a transformation network that mainly learns and generates a k × k transformation matrix, as shown in Figure 3 . At this point, the original point cloud has been converted into G high-dimensional feature vectors, which can be treated like "tokens" in text or images and sent to the subsequent encoder.

[0078] Feature vector processing module: used to add global information and position encoding to the high-dimensional feature vector.

[0079] Before sending the feature embedding to the encoder, two things need to be done:

[0080] (1) Add [CLS] Token: Like Vision Transformer (ViT) and BERT models, a learnable special token called [CLS] Token is added at the beginning of the G feature embedding sequence. This token will eventually gather the global information of the entire point cloud for downstream tasks such as classification. The sequence length is now G + 1.

[0081] (2) Add Positional Encoding: Point clouds do not have an order, but they are actually processed as a sequence. To let the model know the relative position information of each feature, a learnable position encoding is added to each of the G + 1 tokens. This position encoding will be added again in each layer of the encoder to continuously enhance the model's spatial perception ability.

[0082] The initial data sent to the encoder The calculation formula is: where, is the [CLS] Token, is the feature embedding, is the position encoding.

[0083] Encoder: The encoder is the core of the model, responsible for deep feature learning of feature embeddings and fusion of local and global information. The encoder is stacked by T identical encoder layers (as shown in Figure 4 The structure of each layer is similar to a Transformer, containing two core components (LNP module and bi-SSM module), and using residual connection (Residual Connection) and layer normalization (LayerNorm) to stabilize training.

[0084] The processing flow of an encoder layer is as follows:

[0085] (1) Input: Output of the previous layer .

[0086] (2) Local feature extraction (LNP module): .

[0087] The input first goes through layer normalization (LN) and position encoding, and then goes into the LNP (Local Norm Pooling) module. LNP is responsible for exchanging information between adjacent features to capture local geometric features. The calculation result is added to the original input through residual connection.

[0088] (3) Global feature extraction (bi-SSM module): .

[0089] Output of the previous step After layer normalization, the output is fed into the bi-SSM (bi-directional state space model) module. The bi-SSM has a global receptive field and is responsible for capturing long-range dependencies between features. The result is added to the input through a residual connection to obtain the final output of the layer .

[0090] The two key modules, LNP and bi-SSM, are described in detail below.

[0091] LNP (Local Norm Pooling) module:

[0092] The LNP module aims to efficiently extract local geometric features of point clouds. It does so through a two-step process: feature propagation and aggregation, which simulates the exchange of information within a local neighborhood.

[0093] (1) Feature propagation (K-norm): The purpose of this step is to compute a "relative feature representation" for each neighbor feature of a center point, relative to the center point. For a center point feature and its neighbor features , the computation process is as follows: first, calculate the normalized difference of neighbor features relative to the center feature : where represents the variance of the above difference vector, is a very small constant for numerical stability, then concatenate this relative feature with the center point feature , and generate the final "post-propagation feature" for aggregation through a learnable linear transformation (parameterized by and ). This new feature combines the topological relationship between neighbor points and the center point. where represents feature concatenation.

[0094] (2) Feature aggregation (K-pooling): After obtaining the post-propagation neighbor features , this step aggregates this information back to update the center point's feature. Here, a weighted summation mechanism inspired by Softmax is adopted, which assigns weights according to the importance of each post-propagation feature, thereby effectively aggregating while reducing information loss. The updated center point feature is calculated as follows: , where k is the number of neighbor features, For the i-th new feature, Exponential function value for the i-th neighbor feature.

[0095] Through these two steps, the feature of each point (as a center point) effectively absorbs the geometric and semantic information within its neighborhood, thus enhancing the local perceptual ability of the model.

[0096] bi-SSM (Bidirectional SSM) module:

[0097] The standard Mamba model is unidirectional, suitable for processing text and other data with a clear order. However, for unordered point clouds, unidirectional scanning will produce a dependence on pseudo-order (i.e., mistakenly learning the Token order randomly assigned). Here, the bi-directional state space model (bi-SSM) designed for unordered data is used.

[0098] The operation process is as shown in Figure 5 :

[0099] A. Forward scanning (Forward SSM, L+SSM): The input sequence F (also denoted as ) is scanned from beginning to end (e.g., Token 0 to 7), capturing dependencies in one direction.

[0100] B. Reverse scanning (C-SSM / Feature Reverse SSM): A traditional bidirectional model would horizontally flip the Token sequence and scan it again (e.g., from 7 to 0), but doing so would exacerbate the pseudo-order dependence problem in point clouds. The solution here: instead of flipping the order of Tokens, flip the channels of each Token feature vector. This is called "vertical flipping" or "Feature Reverse SSM (C-SSM)". For example, if a feature vector is [c1, c2,..., cC], it becomes [cC,..., c2, c1] after flipping.

[0101] C. Fusion: Finally, the original input F, the output of forward scanning L+SSM, and the output of reverse scanning C-SSM are added together to obtain the final result.

[0102] In this way, the bi-SSM module not only utilizes the ability of SSM to capture long-distance dependencies, but also successfully adapts to the unordered nature of point cloud data through the novel "feature flipping" strategy.

[0103] Classification module: Based on the output of the encoder, the classification is completed through the classification head.

[0104] Step 4: Input a batch of data in the base dataset into the MetaMamba model to calculate the initial loss of each sample; use the Bias model to assign weights to the initial loss , and calculate the pseudo-loss based on the weights ; update the MetaMamba model parameters based on the pseudo-loss; input a batch of data in the meta dataset into the updated MetaMamba model to calculate the meta-loss of each sample; update the parameters of the Bias model based on the meta-loss

[0105] Calculate the true loss of the current batch of the base dataset using the BaseMamba model and input it into the updated Bias model to obtain the weights ; weight the true loss based on the weights to calculate the total loss; update the parameters of the BaseMamba model based on the total loss

[0106] Repeat Step 4 until the training completion requirements are met, and the trained BaseMamba model is the scoliosis classification model.

[0107] The core idea of the scoliosis classification model training is to use a meta dataset with balanced categories to guide the Bias model on how to assign weights to the samples in the base dataset with unbalanced categories; this process is not simply alternating training of each model, but in each iteration, through a double-layer optimization strategy, the Bias model and the BaseMamba model are updated simultaneously.

[0108] As shown in Figure 6 , the entire training framework consists of two core modules: a base training module ( Figure 6 upper half) and a meta-learning module ( Figure 6 lower half), which work together to optimize the main model (Base Mamba).

[0109] The specific process is as follows:

[0110] First stage: update the Bias model

[0111] The goal of this stage is to update the Bias model so that it learns to assign appropriate weights to the samples in the base dataset to improve the model's generalization ability under an unbiased distribution.

[0112] Input a batch of data in the base dataset into the MetaMamba model to calculate the initial loss of each sample. Then, use the current Bias model V to assign weights to these losses , thereby obtaining the pseudo-loss : , , .

[0113] Based on , compute the gradient of the MetaMamba model parameters with respect to the MetaMamba model parameters, and perform a virtual parameter update on the MetaMamba model parameters based on the gradient, to obtain a set of hypothetical optimized MetaMamba model parameters .

[0114] Evaluate the updated MetaMamba model (with parameters ) on the meta-dataset, and compute the meta-loss : , which measures the impact of the weight strategy generated by the Bias model on the final generalization performance.

[0115] Backpropagate , and only update the parameters of the Bias model.

[0116] Second stage: training the BaseMamba model

[0117] After the Bias model is updated, use it to guide the training process of the BaseMamba model:

[0118] Compute the true loss of the current batch of base dataset using the BaseMamba model , and input it into the updated Bias model V to obtain the weights : , .

[0119] According to the weights , weight the loss to obtain the total loss used for backpropagation : .

[0120] Backpropagate the total loss to update the BaseMamba model parameters.

[0121] Through this mechanism, the BaseMamba model can focus on the samples considered "more valuable" by the Bias model, thereby alleviating the negative impact of class imbalance.

[0122] In this embodiment, the model training process will be automatically stopped when any of the following termination conditions is met: (1) the training round reaches a preset upper limit (e.g., 100 epochs); (2) the classification accuracy of the model on the validation set reaches or exceeds a set target threshold (e.g., 90.0%); (3) the model shows a convergence trend in a number of consecutive training cycles (e.g., 5 consecutive epochs), i.e., the standard deviation of the validation set accuracy is less than 0.1%, and the average change amplitude of the loss function value is less than 1e-4. When the training termination mechanism is triggered, the system will automatically select the set of weights that performs best on the validation set from the intermediate saved model snapshots as the final model output. The baseMamba model obtained at this time is the scoliosis classification model.

[0123] The application dynamically adjusts the accurate matching of class centers to the physiological distribution of the spine based on the point cloud optimization technology of optimal transport (OT) and farthest point sampling (FPS) fusion, eliminates point cloud disorder interference by combining the channel flipping mechanism of the bi-directional state space model (bi-SSM), and significantly enhances the feature extraction capability for complex spinal curvature; adopts a meta-learning driven double-layer optimization framework, dynamically punishes class bias in basic training by using an equilibrium meta-dataset to train a Bias model, and forces the model to focus on minority class samples from the loss function level; both solve the problem of insufficient long-distance modeling of the spine by traditional point cloud methods, and simultaneously overcome the generalization bottleneck caused by class imbalance through the lightweight bi-SSM sample weight distribution mechanism, and finally realize the reduction of computing resources while improving detection accuracy.

[0124] Correspondingly, the application also provides a scoliosis classification method, comprising:

[0125] Obtaining the to-be-classified back point cloud data and preprocessing, the preprocessing refers to feature aggregation and dimensionality reduction of the obtained point cloud data based on the farthest point sampling algorithm and the optimal transport algorithm;

[0126] Classifying the preprocessed to-be-classified back point cloud data by using the trained scoliosis classification model.

[0127] The scoliosis classification system comprises:

[0128] The point cloud acquisition module is used to obtain the to-be-classified back point cloud data;

[0129] The OT sampling module is used to preprocess the obtained to-be-classified back point cloud data, and the preprocessing refers to feature aggregation and dimensionality reduction of the obtained point cloud data based on the farthest point sampling algorithm and the optimal transport algorithm;

[0130] The scoliosis classification module classifies the preprocessed to-be-classified back point cloud data based on the trained scoliosis classification model.

Claims

1. A method for training a scoliosis classification model, the method comprising: The method comprises the following steps: Step 1: obtaining training samples, the training samples comprising back point cloud data and corresponding classification result labels; Step 2: creating a basic data set and a metadata set based on the training samples, the basic data set being generated by randomly sampling the training samples, and the metadata set being generated by class-balanced sampling the training samples; Step 3: constructing a BaseMamba model, a MetaMamba model and a Bias model, the BaseMamba model and the MetaMamba model having the same structure and sharing parameters, the BaseMamba model being used to extract features of the back point cloud data in the basic data set, the extracted features being processed based on an encoder and classification being completed through a classification head; the MetaMamba model being used to extract features of the back point cloud data in the metadata set, the extracted features being processed based on an encoder and classification being completed through a classification head; and the Bias model being used to predict class imbalance bias; Step 4: inputting a batch of data in the basic data set into the MetaMamba model, calculating an initial loss of each sample; Bias model is used to assign weights to each initial loss and based on the weights weighting the initial losses to calculate pseudo-losses; updating the MetaMamba model parameters based on the pseudo-losses; inputting a batch of data in the meta dataset into the updated MetaMamba model, calculating the meta-loss of each sample; updating the parameters of the Bias model based on the meta-loss; Compute the true loss for the current batch of the base dataset using the BaseMamba model and input it into the updated Bias model to get the weights ; according to the weights Weight the true loss to compute the total loss; updating parameters of the BaseMamba model based on the total loss; repeating step 4 until a training completion requirement is met, and the trained BaseMamba model being a scoliosis classification model.

2. The scoliosis classification model training method of claim 1, wherein, Both the BaseMamba model and the MetaMamba model comprise: a feature embedding module for converting the back point cloud data into a high-dimensional feature vector; a feature vector processing module for adding global information and position encoding to the high-dimensional feature vector; an encoder stacked by T identical encoder layers, each layer comprising a local feature extraction unit for capturing local features and a global feature extraction unit for capturing long-distance dependency between features; and outputting feature representations fused with the local features and the global features; a classification module for completing classification through a classification head based on the output of the encoder.

3. The scoliosis classification model training method of claim 2, wherein, The operation steps of the feature embedding module are: selecting a center point; constructing a local neighborhood based on the center point; and performing spatial pose alignment, local feature extraction, feature alignment and global feature aggregation on the point cloud data of each local neighborhood through a PointNet network, and finally generating G high-dimensional feature vectors.

4. The scoliosis classification model training method of claim 2, wherein, The operation steps of the local feature extraction unit are: Computing a normalized difference of neighbor features with respect to a center feature ; will be described below. with the center point feature stitching; Generating new features by learnable linear transformations , where represents feature concatenation, and are linear transformation parameters; Updated center point features is calculated as follows: where k is the number of neighbor features, is the i-th new feature, is the exponential function value of the i-th neighbor feature.

5. The scoliosis classification model training method of claim 2, wherein, The operation steps of the global feature extraction unit are: forward scanning: scanning the input sequence from head to tail to capture the dependency in one direction; backward scanning: flipping the channel of the feature vector of the input sequence and then scanning; fusion: adding the original input, the output of the forward scanning and the output of the backward scanning to obtain the final result.

6. The scoliosis classification model training method of claim 1, wherein, Step 2 further comprises: pre-processing the back point cloud data in the basic data set and the metadata set, the pre-processing referring to feature aggregation and dimension reduction on the obtained point cloud data based on a farthest point sampling algorithm and an optimal transport algorithm.

7. The scoliosis classification model training method of claim 6, wherein, The specific steps of the pre-processing are: generating an initial class center set based on the farthest point sampling algorithm; dynamically optimizing the class center through the optimal transport algorithm to minimize the transport cost of the point cloud distribution to the class center distribution; and outputting structured point cloud data comprising the class center and the nearest neighbor point.

8. The scoliosis classification model training method of claim 1, wherein, The training completion requirement is that the number of training reaches a preset training threshold or the classification accuracy of the model on the validation set is not lower than a preset target precision threshold or the model presents a convergence trend in continuous training cycles.

9. A method of classifying scoliosis, characterized by, The method comprises the following steps: Obtaining and preprocessing the back point cloud data to be classified, wherein the preprocessing refers to feature aggregation and dimension reduction of the obtained point cloud data based on a farthest point sampling algorithm and an optimal transport algorithm; Classifying the preprocessed back point cloud data to be classified by using the trained scoliosis classification model, wherein the scoliosis classification model is trained by using the scoliosis classification model training method in any one of claims 1-8.

10. A scoliosis classification system characterized by, The method comprises the following steps: A point cloud acquisition module is configured to obtain the back point cloud data to be classified; An OT sampling module is configured to preprocess the obtained back point cloud data to be classified, wherein the preprocessing refers to feature aggregation and dimension reduction of the obtained point cloud data based on a farthest point sampling algorithm and an optimal transport algorithm; A scoliosis classification module is configured to classify the preprocessed back point cloud data to be classified based on the trained scoliosis classification model, wherein the scoliosis classification model is trained by using the scoliosis classification model training method in any one of claims 1-8.

Citation Information

Cited By

  • Spinal curvature measuring method, device and equipment based on large visual model

    CN122156060A