Ankylosing spondylitis ossification classification method based on incomplete data

Through the feature selection algorithm based on self-representation learning, the characteristics related to AS ossification are identified and reconstructed, and the problem of incomplete biochemical index data is solved, which significantly improves the accuracy and stability of the AS ossification classification model.

CN120105044APending Publication Date: 2025-06-06BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510156550.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the classification of ankylosing spondylitis (AS), incomplete biochemical index data leads to underestimation of the importance of characteristics, affecting the generalization ability and effect of the model.

Method used

A feature selection algorithm based on self-representation learning is proposed. Through weighted self-representation learning and l2,1 constrain least squares regression, identify the features with the largest amount of information and reconstruct the feature space to reduce the impact of missing data.

Benefits of technology

It effectively solves the impact of incomplete data on model construction, improves the accuracy and stability of the AS ossification classification model, and greatly improves the performance compared with other classification models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105044A_ABST
    Figure CN120105044A_ABST
Patent Text Reader

Abstract

The invention discloses an ankylosing spondylitis ossification classification method based on incomplete data, and belongs to the field of machine learning and artificial intelligence. Firstly, a feature selection algorithm under missing data based on self-expression learning is provided, the algorithm can identify a part with the maximum information amount from incomplete indexes, a feature space is reconstructed by capturing high-order correlation between missing indexes, and then the feature space is reconstructed by adopting constrained least square regression. And the influence of missing data on the validity of feature selection is further reduced. The obtained feature subsets are sent to a classifier, an AS ossification classification model is constructed, the model can discriminate and classify samples by using missing biochemical indexes, and compared with other classification models, the performance of the AS ossification classification model is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of machine learning and artificial intelligence, and specifically relates to technologies such as feature selection and data mining. Background Art

[0002] Ankylosing Spondylitis (AS) can cause new bone movement, spinal ankylosis, and ligament calcification. Currently, the classification of AS ossification is mainly displayed on two types of data: biochemical data and imaging data. Imaging data, such as computed tomography (CT) and X-ray (X-radiation, X-ray), are not only expensive, but also deeply influenced by subjective concepts. In contrast, biochemical data obtained by biochemical analyzers are easier to obtain and are also objective and effective.

[0003] The biochemical index data used to build the AS ossification classification model are often incomplete, and data missing is common. Missing data will lead to the underestimation of the importance of certain features. Without these features, the generalization ability of the general model will be reduced, affecting the model effect.

[0004] To meet this challenge, this paper proposes an AS ossification classification method based on incomplete data. First, a feature selection algorithm under missing data based on self-representation learning is proposed, which can identify the most informative part from incomplete indicators, reconstruct the feature space by capturing the high-order correlation between missing indicators, and then use l 2,1 Constrained least squares regression was used to further mitigate the impact of missing data on the effectiveness of feature selection. Next, the obtained feature subset was sent to the classifier to construct an AS ossification classification model, which can use the missing biochemical indicators to discriminate and classify samples, and its performance is greatly improved compared with other classification models. Summary of the invention

[0005] The present invention proposes an AS ossification classification method based on incomplete data, which can still achieve accurate classification of AS ossification when some data are missing.

[0006] In order to achieve the above object, the present invention proposes the following technical solutions:

[0007] 1) A feature selection algorithm based on self-representation learning under data loss is proposed. This algorithm can identify the most relevant features related to AS ossification. Weighted self-representation learning can reduce the impact of missing data. 2,1 Norm-constrained least squares regression, and a simple and efficient optimization algorithm is proposed to reconstruct information and obtain the optimal subset;

[0008] 2) Using this subset, an efficient, robust and interpretable support vector machine classifier is trained to obtain the AS ossification classification model under incomplete data.

[0009] The scheme includes three steps: constructing an AS ossification classification dataset, proposing a feature selection algorithm based on self-representation learning under data loss, and constructing an AS ossification classification model. Each step is described in detail below.

[0010] Step 1: Construct AS ossification classification dataset

[0011] After an eight-hour fasting period, blood samples were drawn from the antecubital fossa vein and 40 biochemical indices were measured. The specific biochemical indices are shown in Table 1 .

[0012] Step 2: Propose a feature selection algorithm based on self-representation learning under data loss

[0013] In order to mitigate the impact of missing data, this paper proposes a novel feature selection algorithm to ensure that the estimated missing data remains close to the original data. A weighted mask matrix is ​​used to help identify the location of missing data, thereby reducing their impact on the effectiveness of the feature selection method. 2,1 The least squares regression of the norm is used to explore the correlation between the label space Y and the feature space X, and then the final objective function of the feature selection algorithm is obtained.

[0014] In order to solve the problem, the present invention proposes an alternating iterative updating method to solve the equation, and the matrix W and the matrix B are updated alternately until convergence.

[0015] Step 3: Build an ossification classification model

[0016] The optimal feature subset should be selected to train the support vector machine (SVM) to obtain the AS ossification warning model, which can predict whether AS patients will develop ossification.

[0017] Compared with the prior art, the present invention has the following obvious advantages and beneficial effects:

[0018] First, the present invention aims at the problem of missing data in biochemical indicators and proposes an AS ossification classification method with incomplete data. Most of the existing technologies are aimed at complete data samples, and the prediction performance is not ideal under missing data. The algorithm proposed in the present invention can obtain relevant information of the missing data through self-representation learning, which can solve the impact of incomplete data on model construction. In addition, the method has low time complexity, is easy to train and use, and has good practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1This is a flow chart of the ossification classification method proposed in the present invention. DETAILED DESCRIPTION

[0020] In combination with the above description and the accompanying diagrams, the process of the present invention is further described in detail, but the scope of protection of this patent is not limited to this implementation process.

[0021] Table 1 AS ossification dataset built by the present invention

[0022]

[0023] Step 1: Construct AS ossification classification dataset

[0024] The data samples were divided into two categories: ossified and non-ossified, and each sample contained a total of 40 indicators, such as gender, age, and four blood indicators, as shown in Table 1. Blood indicators include: (1) Complete blood count: This group includes 26 indicators. (2) Renal function: This group consists of two indicators: urea and creatinine, which are used to evaluate renal function. (3) Liver function: It includes two indicators: alanine aminotransferase and aspartate aminotransferase, which are used to evaluate liver health. (4) Hematology ratio: This group includes 8 indicators.

[0025] Step 2: Propose a feature selection algorithm based on self-representation learning under data loss

[0026] To mitigate the impact of missing data, this paper proposes a feature selection method for missing data to ensure that the estimated missing data remains close to the original data.

[0027] First, let me introduce that in this invention, italic capital letters represent matrices, and italic lowercase letters represent vectors. m×n ,s i Indicates the i-th row, s j represents the jth column, S ij Represents the value of the i-th row and j-th column. For the matrix S∈R m×n , ‖S‖ F represents its Frobenius norm, which is defined as:

[0028]

[0029] The matrix l 2,1 The norm can be used to rotate the invariant l 1 The present invention introduces l 2,1 Loss is used to reduce the impact of missing data and improve the robustness of the model. It is defined as follows:

[0030]

[0031] Therefore, we can deduce that l2,p norm, which is introduced to select discriminative and representative features, where r and p are parameters in the norm, respectively controlling the weighting of the intra-row elements and inter-row norms. By adjusting r and p, different matrix norms can be obtained to measure different characteristics of the matrix. In the present invention, r is 2, and p is in the range of 0.5, which can be expressed as:

[0032]

[0033] (1) Self-representation learning

[0034] In order to mitigate the impact of missing data, this paper proposes a WSRL method to ensure that the estimated missing data remains close to the original data. To this end, the matrix B∈R d×d Used as a coefficient matrix to recover missing data, where d represents the total number of features. The coefficient matrix B is constrained by the Frobenius norm to promote row sparsity.

[0035] It can be expressed as follows:

[0036]

[0037] where X∈R n×d A matrix representing data, with each row x i ∈R 1×d represents a sample. n represents the number of samples. Next, the weighted mask matrix A∈R is used n×d To help identify the location of missing data, thereby reducing their impact on the effectiveness of feature selection methods. The weighted mask matrix A is defined as follows:

[0038]

[0039] By combining the above equations, the following formula can be derived:

[0040]

[0041] Through the above WSRL, the most informative AS data can be identified in the presence of incomplete indicators.

[0042] (2) Using l 2,1 Least Squares Regression with Norm

[0043] Least squares regression is used to explore the correlation between label space Y and feature space X. 2,1 The least squares regression of the norm can be formulated as follows:

[0044]

[0045] where Y∈R n×crepresents the AS ossification label matrix. c represents the number of label types. I represents the identity matrix. b∈R 1 ×c is the bias vector. W∈R d×c is the projection matrix. To better minimize the impact of missing data, use l 2,1 Norm constrained LSR, the following formula is obtained:

[0046]

[0047] Then, the last item is provided with l 2,p The definition of the norm is given by:

[0048]

[0049] Reducing p can increase the sparsity of W and control the complexity of the optimization framework. Therefore, the problem is stated as follows:

[0050]

[0051] where λ 2 is the regularization parameter.

[0052] (3) The final objective function of FS-WSRL

[0053] Finally, combining equations (6) and (10), we can obtain the objective function of FS-WSRL, which can be expressed as:

[0054]

[0055] where λ 1 and λ 2 is a regularization parameter, ranging from 0 to 1.

[0056] Optimization algorithm:

[0057] The present invention introduces an optimization scheme for the FS-WSRL objective function. There are three variables, including W, B, and b in the objective function. The objective equation (11) is equivalent to:

[0058]

[0059] l 2,1 The diagonal matrix of the norm can be expressed as where x i ,y i 、b i Indicates the values ​​of X, Y, and b corresponding to the i-th row.

[0060] Next, to get the derivative with respect to i, simplify the equation. Setting the derivative to 0, the solution for i is:

[0061]

[0062] Where m = I T G 0 I.

[0063] Therefore, by substituting equation (14) into equation (13), the latter can be expressed as follows:

[0064]

[0065] in m has been defined above and I is the identity matrix.

[0066] In the following, an alternating iterative updating method is proposed to solve equation (15).

[0067] Step 1: Fix B and update W

[0068] When B is fixed, equation (15) can be equivalently written as:

[0069]

[0070] To optimize W, the iteratively reweighted least squares (IRLS) method is used. The diagonal weight matrix G of W is 1 can be written as:

[0071]

[0072] in represents the j-th diagonal element.

[0073] If W ≥ 0, introduce the Lagrange multiplier ψ to obtain its Lagrangian function:

[0074]

[0075] The above equation can be obtained by taking the partial derivative of L(W):

[0076]

[0077] According to the Karush-Kuhn-Tucker (KKT) complementarity condition ψ ij W ij =0, where ψ ij represents the ψ value of the i-th row and j-th column, and the update expression of W is:

[0078]

[0079] Step 2: Fix W and update B

[0080] When W is fixed, equation (15) becomes:

[0081]

[0082] The premise is that B ≥ 0, and the Lagrangian multiplier σ is introduced to obtain its Lagrangian function:

[0083]

[0084] Here, tr represents the trace of the matrix.

[0085] The above equation can be obtained by taking the partial derivative of L(B):

[0086]

[0087] According to the Karush-Kuhn-Tucker (KKT) complementarity condition σ ij B ij =0, where σ ij represents the σ value of the i-th row and j-th column, and the update expression of B is:

[0088]

[0089] The algorithm of FS-WSRL is described in Algorithm 1. In this algorithm, the regression matrix W and the coefficient matrix B are updated alternately until convergence. Finally, by calculating ‖W i ‖ 2 And sort them in descending order to perform feature selection for missing data.

[0090] Algorithm 1. Weighted Self-Representation Learning (FS-WSRL) Feature Selection Algorithm

[0091]

[0092]

[0093] Step 3: Build an ossification classification model

[0094] The present invention selects linear SVM to design the AS ossification classification method. Linear SVM has low computational complexity, fast training speed and strong interpretability. During training, all data are randomly divided into training set and test set in the ratio of 80% and 20%. The training set data is sent to the model for training, iterated 500 times, and the convergence condition is that the change of the objective function between two consecutive times is less than 10 -5 Optimize parameters and improve model performance. 1 ,λ 2The three parameters p are continuously adjusted to find the most suitable data of 0.4, 0.6, and 0.5 respectively. Finally, the test set is used for testing to verify the model performance. Compared with other methods, the accuracy of the method proposed in the present invention is significantly improved, and the average accuracy under different missing degrees reaches 93.67%.

[0095] In response to the problem of missing data, the present invention designs an AS ossification classification method for incomplete data. Compared with the traditional model, this method effectively solves the missing problem, achieves higher accuracy, and has better stability and interpretability.

Claims

1. A classification method for ankylosing spondylitis ossification based on incomplete data, characterized in that The following steps are involved: Step 1: Construct AS ossification classification dataset Each sample in the AS ossification dataset contains a total of 40 indicators, as shown above; Blood indices include: (1) complete blood count: this group includes 26 indices; (2) renal function: this group consists of two indices: urea and creatinine, which are used to evaluate renal function; (3) liver function: including two indices: alanine aminotransferase and aspartate aminotransferase, which are used to evaluate liver health; (4) hematological ratios: this group includes 8 indices; Step 2: Propose a feature selection algorithm based on self-representation learning under data loss In the following content, italic capital letters represent matrices, italic lowercase letters represent vectors; for any matrix S∈R m×n ,s i Indicates the i-th row, s j represents the jth column, S ij Represents the value of the i-th row and j-th column; for the matrix S∈R m×n , ‖S‖ F represents its Frobenius norm, which is defined as: The matrix l 2,1 The norm is used to rotate the invariant l1 norm; l is introduced 2,1 Loss, which is defined as follows: Therefore, it is deduced that 2,p Norm, which is introduced to select discriminative and representative features, where r and p are parameters in the norm, respectively controlling the weighting of the norm of the elements within the row and between rows. The value of r is 2, and the value of p is in the range of 0.

5. It is expressed as: (1) Self-representation learning Coefficient matrix B∈R d×d The coefficient matrix used to restore missing data, where d represents the total number of features; the coefficient matrix B is constrained by the Frobenius norm and is expressed as follows: where X∈R n×d A matrix representing data, with each row x i ∈R 1×d represents a sample; n represents the number of samples; next, the weighted mask matrix A∈R is used n×d To help identify the location of missing data; the weighted mask matrix A is defined as follows: By combining the above equations, the following formula is derived: Through the above constraints, the most informative AS data is identified in the presence of incomplete indicators; (2) Using l 2,1 Least Squares Regression with Norm Least squares regression is used to explore the correlation between label space Y and feature space X; with l 2,1 The least squares regression formulation of the norm is as follows: where Y∈R n×c represents the AS ossification label matrix; c represents the number of label types; I represents the identity matrix; b∈R 1×c is the bias vector; W∈R d×c is the projection matrix; use e 2,1 Norm constrained LSR, the following formula is obtained: Then, the last item is provided with e 2,p The definition of the norm is given by: where reducing p increases the sparsity of W and controls the complexity of the optimization framework; therefore, the problem is stated as follows: Where λ2 is the regularization parameter; (3) The final objective function of FS-WSRL Combining equations (6) and (10), we can obtain the objective function of FS-WSRL, which is expressed as: Among them, λ1 and λ2 are regularization parameters, ranging from 0 to 1; Optimization scheme of FS-WSRL objective function; there are three variables, including W, B, and b in the objective function; the objective equation (11) is equivalent to: l 2,1 The diagonal matrix of the norm is expressed as where x i ,y i , b i Indicates the values ​​of X, Y, and b corresponding to the i-th row; Next, to get the derivative with respect to b, simplify the equation; setting the derivative to 0, the solution for b is: Where m = I T G0I; Therefore, by substituting equation (14) into equation (13), the latter is expressed as follows: in m has been defined above, I is the identity matrix; In the following, an alternating iterative updating method is proposed to solve equation (15); Step 1: Fix B and update W When B is fixed, equation (15) is equivalently written as: To optimize W, an iterative weighted least squares method is used; the diagonal weight matrix G1 of W is written as: in represents the jth diagonal element; If W ≥ 0, introduce the Lagrange multiplier ψ to obtain its Lagrangian function: By taking the partial derivative of L(W) we obtain: According to the Karush-Kuhn-Tucker (KKT) complementarity condition ψ ij W ij =0, where ψ ij represents the ψ value of the i-th row and j-th column, and the update expression of W is: Step 2: Fix W and update B When W is fixed, equation (15) becomes: The premise is that B ≥ 0, and the Lagrangian multiplier σ is introduced to obtain its Lagrangian function: Where tr represents the trace of the matrix; The above equation is obtained by taking the partial derivative of L(B): According to the Karush-Kuhn-Tucker (KKT) complementarity condition σ ij B ij =0, where σ ij represents the σ value of the i-th row and j-th column, and the update expression of B is: The algorithm of FS-WSRL is described in Algorithm 1. In this algorithm, the regression matrix W and the coefficient matrix B are updated alternately until convergence. Finally, by calculating ‖W i ‖2 and sort them in descending order to perform feature selection on missing data; Algorithm 1. Weighted Self-Representation Learning FS-WSRL feature selection algorithm is as follows enter AS indicator feature matrix X∈R n×d AS ossification label matrix Y∈R n×c Mask matrix A∈R n×d Regularization parameters λ1, λ2, p Set t = 0 and initialize W (0) ∈R d×c , B (0) ∈R d×d ; cycle Update the diagonal matrices G0 and G1; Update b (t+1) According to formula (14); Update W (t+1) According to formula (20); According to B (t+1) According to formula (24); Update t = t + 1; Until the convergence standard condition is met, the objective function changes less than 10 -5 ; Output Feature selection matrix W∈R d×c Step 3: Build an ossification classification model Linear SVM was selected to design the AS ossification classification method. During training, all data were randomly divided into training set and test set in the ratio of 80% and 20%, respectively. The training set data was sent to the model for training, and iterated for more than 500 rounds. The convergence condition was that the change of the objective function between two consecutive times was less than 10 -5 , get the trained model; apply the trained model for classification.