A method, device and storage medium for clinical subtype classification of basal cell carcinoma
By using the multi-source data set constructed using the k-means algorithm and knowledge graph, basal cell carcinoma is fine-grained, which solves the problem of inaccurate classification in the existing technology, and achieves higher accuracy and the formulation of personalized treatment plans.
Patent Information
- Application Number
- CN202411611775.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-11-12
AI Technical Summary
The prior art has problems with inaccurateness in the clinical subtype classification of basal cell carcinoma, which is mainly due to the reliance on human experience and knowledge, resulting in subjectivity and inconsistency in the classification results.
The pretreated basal cell carcinoma case data sets were clustered by the k-means algorithm, fine-grained typing was mined, and the data quality and consistency were constructed through knowledge graphs and multi-source and multi-modal data sets were ensured.
It improves the accuracy and resolution of clinical subtype classification of basal cell carcinoma, provides more accurate disease assessment and personalized treatment options, and reduces unnecessary side effects.
Smart Images

Figure CN119560131B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer-aided medical diagnosis, and in particular to a method, device and storage medium for classifying clinical subtypes of basal cell carcinoma. Background Art
[0002] Basal cell carcinoma (BCC) is one of the most common malignant skin tumors. Due to its diverse clinical manifestations, it is easy to be misdiagnosed and mistreated, often causing local organ damage or even death. Therefore, early diagnosis and precise treatment are very important.
[0003] At present, the traditional clinical skin lesions of BCC are divided into 5 types (i.e., nodular, superficial, pigmented, infiltrative, and fibroepithelioma), and the histopathological manifestations are divided into 7 types (i.e., solid, palisade, superficial, nodular, morphea-like, infiltrative, and fibroepithelioma). Although these methods have been used in clinical practice, the traditional classification method mainly relies on the experience of clinical and pathological experts, the classification is rough, and is limited by human subjectivity and limited knowledge. There is a lack of unified and objective classification standards. Different experts may have different classification results for the same case based on different experience and knowledge backgrounds, and the classification accuracy and resolution cannot be ensured, which cannot help accurately evaluate the progression and prognosis of the disease and formulate personalized treatment plans. Summary of the invention
[0004] The present invention provides a method, device and storage medium for classifying clinical subtypes of basal cell carcinoma, so as to solve the technical problem that the existing classification methods are not accurate enough in classifying clinical subtypes of basal cell carcinoma.
[0005] According to one aspect of the present invention, a method for classifying clinical subtypes of basal cell carcinoma is provided, comprising:
[0006] Acquire a basal cell carcinoma case data set, and preprocess the basal cell carcinoma case data set;
[0007] Based on the preprocessed basal cell carcinoma case data set, the k-means algorithm was used to mine the fine-grained classification of basal cell carcinoma, and the reliability and adequacy of the classification were verified;
[0008] Acquire patient data, match the patient data with the fine-grained typing results, and obtain the clinical subtype classification of basal cell carcinoma corresponding to the patient.
[0009] Furthermore, obtaining a basal cell carcinoma case data set includes:
[0010] A data resource integration platform is constructed based on the knowledge graph, and a basal cell carcinoma case subject database is established, wherein the data resource integration platform is connected to N clients for obtaining basal cell carcinoma case data from the clients, and the clients are used to manage case data of corresponding hospitals;
[0011] The basal cell carcinoma case data are obtained through the basal cell carcinoma case subject database to construct a basal cell carcinoma case data set.
[0012] Further, obtaining basal cell carcinoma case data through the basal cell carcinoma case subject database, and constructing a basal cell carcinoma case data set includes:
[0013] Using random sampling without replacement to obtain basal cell carcinoma case data from the basal cell carcinoma case subject database, constructing a data set D;
[0014] Let the model layer of the basal cell carcinoma knowledge graph be MLbcc and the concept model of the dataset D be C D , QQ plot was used to test the uniformity of sampling data, ANOVA was used to test the balance of sampling data, and Levene's test was used to test the homogeneity of variance of sampling data;
[0015] Determine whether the data set D meets the following requirements:
[0016] Conceptual model and model layer mapping f: C D —>MLbcc is a surjection, that is, C D Each concept in can find a corresponding mapping in MLbcc;
[0017] The data in dataset D are balanced, uniform, and have homogeneous variance;
[0018] The size of the dataset D |D| ≥ F((VC+ln(1 / d)) / e), where |D| is the size of the dataset, d is the failure rate, e is the error rate, and VC is the VC dimension;
[0019] The dataset D consists of case data and medical images corresponding to the cases, and it is a one-to-many (1:m) relationship;
[0020] If the dataset D meets the requirements, the constructed dataset D is considered to meet the requirements, and the dataset D that meets the requirements is used as the basal cell carcinoma case dataset.
[0021] Further, preprocessing the basal cell carcinoma case data set includes:
[0022] Mean value filling and regression algorithms were used to process missing values in the basal cell carcinoma case dataset, and normalization methods were used for feature scaling;
[0023] The spatial transformation algorithm is used for image data enhancement, where the image is transformed by 180° rotation, horizontal flipping and shearing transformation, with a transformation coefficient of 0.2;
[0024] The enhanced basal cell carcinoma case dataset was split into a training set and a test set according to a 80:20 split ratio.
[0025] Furthermore, the k-means algorithm is used to mine fine-grained typing based on the preprocessed basal cell carcinoma case data set, including:
[0026] The data features in the preprocessed basal cell carcinoma case data set are divided into logical type, classification type and quantitative type, and standardized. The standardization process includes converting the logical type data into 0, 1; converting the classification type data into 0, 1, 2, ...; and mapping the quantitative type data into the range of [0, 1].
[0027] Based on the standardized data features, the basal cell carcinoma feature vector X is constructed, where X = {k 1 x 1 , k 2 x 2 , k 3 x 3 , …, k i x i}, where x 1 , x 2 , x 3 , …, x i is the standardized data feature, k 1 , k 2 , k 3 , …, k i is the corresponding weight value;
[0028] Based on the Euclidean distance, the k-means algorithm is used to cluster the feature vector X to mine and summarize the statistical distribution of basal cell carcinoma features. The objective function of the k-means algorithm is as follows:
[0029]
[0030] C k represents the kth cluster, K represents the total number of clusters, i.e. the number of clusters, x (i) Represents cluster C k The i-th data point in (k) Represents cluster C k The centroid of the cluster is the average value of all data points in the cluster. The objective function J represents the sum of the squares of the distances from all data points in the cluster to their centroids. The goal of the algorithm is to minimize J.
[0031] Furthermore, clustering the feature vector X using the k-means algorithm includes:
[0032] Randomly select K samples as the initial cluster centers, denoted as μ(1), μ(2), …μ(k), where K is a hyperparameter representing the number of clusters, i.e. the number of categories, μ(1) is the first cluster center, μ(2) is the second cluster center, and μ(k) is the kth cluster center;
[0033] Iteration step: For each sample x in the dataset i , calculate its Euclidean distance to the K cluster centers, and assign it to the class corresponding to the cluster center with the smallest distance, completing a cluster division; for each cluster, recalculate its cluster center position, if the geometric center does not coincide with the cluster center, use the geometric center as the new cluster center and re-divide the cluster;
[0034] Repeat the above iterative steps until the change in the cluster center position is less than the preset threshold, then stop and output the clustering results;
[0035] Plot the clustering performance indicators under different numbers of clusters into line graphs;
[0036] Observe the line chart and find an inflection point, the Elbow point;
[0037] The cluster number K corresponding to the Elbow point is the optimal cluster center.
[0038] Further, after using the k-means algorithm to mine the fine-grained typing of basal cell carcinoma, the method further includes:
[0039] The principal component analysis of clinical and pathological characteristics was performed to summarize the correlation of various factors affecting fine-grained typing. The principal component analysis included:
[0040] Standardize the data: subtract the mean from each feature of the original data and then divide by the standard deviation to ensure that each feature has the same importance;
[0041] Calculate the covariance matrix: Calculate the covariance matrix of the standardized data set to understand the correlation between different features;
[0042] Calculate eigenvalues and eigenvectors: Perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors; the eigenvectors represent the principal component directions in the data set, and the eigenvalues represent the variance of the data in these principal component directions;
[0043] Select the number of principal components: Select the number of principal components to be retained based on the size of the eigenvalues;
[0044] Constructing a projection matrix: Select the number of principal components selected before and construct a projection matrix; the projection matrix is a matrix composed of eigenvectors and is used to project the original data into the new principal component space;
[0045] Dimensionality reduction transformation: Multiply the standardized data set by the projection matrix to transform the data set from the original high-dimensional space to a new low-dimensional principal component space.
[0046] Further, after using the k-means algorithm to mine the fine-grained typing of basal cell carcinoma, the method further includes:
[0047] Factor analysis of clinical and pathological features was performed to reveal the underlying factor structure, including:
[0048] Standardize data: Use z-score standardization or minimum and maximum scaling methods to standardize data to ensure that all variables have the same scale;
[0049] Extract factors: Use statistical software to calculate and obtain the results of factor loading matrix and commonality matrix; these results reveal the relationship between factors and variables and the internal structure of factors;
[0050] Determine the number of factors: Kaiser criterion, Scree plot and explained cumulative variance were used to explain the variance contribution and cumulative variance contribution of each factor to determine the number of factors to be retained;
[0051] Factor rotation: Use maximization Varimax and maximum likelihood estimation Promax to rotate the extracted factors to make the factors easier to interpret;
[0052] Factor interpretation: Analyze the rotated factor loading matrix, determine the relationship between each variable and factor based on the absolute value and sign of the loading, and assign factor interpretations;
[0053] Factor score calculation: Calculate the score of each sample on each factor to further analyze the position and distribution of the sample in the factor space.
[0054] According to another aspect of the present invention, there is also provided a device for classifying clinical subtypes of basal cell carcinoma, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method for classifying clinical subtypes of basal cell carcinoma as described above when executing the computer program.
[0055] According to another aspect of the present invention, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for classifying clinical subtypes of basal cell carcinoma as described above are implemented.
[0056] The present invention provides a clustering-based basal cell carcinoma clinical subtype classification method, device and storage medium, wherein the k-means algorithm adopted is an unsupervised clustering algorithm that can divide data into k non-overlapping subsets (i.e., clusters) through iterative calculations so that the sum of the distances between each data point and the center of the cluster to which it belongs is minimized. The k-means algorithm is simple and efficient, has clear spatial division, and is suitable for large-scale data sets.
[0057] The present invention adopts the BCC knowledge graph and constructs BCC data resources under the guidance of the knowledge graph, thereby minimizing human selection bias.
[0058] The present invention adopts a multi-source, multi-modal, and continuously updated BCC normative dataset. The dataset comes from a wide range of sources and integrates data from multiple medical institutions, making it more representative and diverse. It also ensures the quality and consistency of the dataset by formulating unified data collection and processing standards. The diverse data modalities help to more comprehensively and accurately describe the biological characteristics of BCC, and the classification results are more accurate.
[0059] The present invention is based on a clustering fine-grained basal cell carcinoma classification method, which improves the accuracy and resolution of clinical classification. The results of this clinical subtype classification are crucial for personalized treatment and precision medicine. By understanding the specific subtype to which each patient belongs, doctors can understand the patient's disease status more accurately, thereby formulating more targeted treatment plans, improving treatment effects and reducing unnecessary side effects. It also helps to more accurately evaluate the progression and prognosis of the disease, monitor the recurrence and metastasis of the disease, and promptly discover and take corresponding intervention measures. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0061] Figure 1 is a schematic flow chart of a method for classifying clinical subtypes of basal cell carcinoma in an embodiment of the present invention;
[0062] Figure 2 It is a schematic diagram of a process flow of an embodiment of the present invention in an implementation scenario;
[0063] Figure 3 This is a schematic diagram of the BBC clinical fine-grained classification mining process according to an embodiment of the present invention;
[0064] Figure 4 This is a visualization diagram of the clinical fine-grained classification of BCC screened out by cluster analysis in an embodiment of the present invention. DETAILED DESCRIPTION
[0065] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only embodiments of a part of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the present application. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0066] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0067] In addition, the terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover a non-exclusive inclusion, for example, a process, method, product, or apparatus comprising a series of steps or elements is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product, or apparatus.
[0068] Example 1
[0069] In this embodiment, a method for classifying clinical subtypes of basal cell carcinoma is provided.
[0070] Reference Figure 1 , Figure 1 is a flow chart of a method for classifying clinical subtypes of basal cell carcinoma in an embodiment of the present invention, such as Figure 1 As shown, the method comprises the following steps:
[0071] S1, obtaining a basal cell carcinoma case data set, and preprocessing the basal cell carcinoma case data set;
[0072] Basal cell carcinoma case datasets can be obtained through databases of hospitals, research institutes, or other institutions.
[0073] Preprocessing includes denoising, supplementing missing values, and normalizing the data.
[0074] Furthermore, in this embodiment, the obtaining of the basal cell carcinoma case data set includes: constructing a data resource integration platform based on the knowledge graph, and establishing a basal cell carcinoma case subject database, wherein the data resource integration platform is data-connected with N clients to obtain basal cell carcinoma case data from the clients, and the clients are used to manage case data of corresponding hospitals; basal cell carcinoma case data is obtained through the basal cell carcinoma case subject database to construct a basal cell carcinoma case data set.
[0075] Under the guidance of the knowledge graph, this embodiment adopts a 1+N mode to deploy the platform and the client's multi-center BCC data resource integration platform to obtain BCC data resources, where "1" represents the data resource integration platform and "N" represents multiple clients (such as hospitals). Through the guidance of the knowledge graph, the platform can more effectively integrate and manage data resources from different hospitals. The client (hospital) is responsible for managing and providing its own case data, which are integrated and analyzed through the data resource integration platform.
[0076] Further, obtaining basal cell carcinoma case data through the basal cell carcinoma case subject database to construct a basal cell carcinoma case data set includes: obtaining basal cell carcinoma case data from the basal cell carcinoma case subject database using a random sampling method without replacement to construct a data set D;
[0077] Random sampling without replacement ensures that each case has the same probability of being selected, and once selected, it will not be put back into the database again, avoiding repeated sampling;
[0078] Let the pattern layer of the basal cell carcinoma knowledge graph be MLbcc. The MLbcc pattern layer mainly defines the core concepts and attributes related to basal cell carcinoma, such as the site of disease, symptoms, histopathological characteristics, treatment methods, etc. The conceptual pattern of the dataset D is CD. The conceptual pattern mainly defines the data elements and data structures related to basal cell carcinoma, such as patient information, lesion information, diagnosis information, treatment information, etc.
[0079] QQ plot was used to test the uniformity of sampling data, ANOVA was used to test the balance of sampling data, and Levene's test was used to test the homogeneity of variance of sampling data;
[0080] Among them, the QQ plot (Quantile-Quantile Plot) is used to test the uniformity of sampled data. It is a graphical statistical method used to compare whether the shapes of two data distributions are the same. If the data points are roughly distributed along the diagonal line, it indicates that the data distribution is uniform.
[0081] ANOVA (Analysis of Variance) is used to test the balance of sampled data. It is a statistical method used to compare whether there are significant differences in the means of three or more groups of data. In a data set, balance means that the number of data in different categories or groups is similar and there is no significant deviation.
[0082] Levene is used to test the homogeneity of variance of sampling data. It is a robust test method for homogeneity of variance. Homogeneity of variance means that data of different groups or categories have similar variances, that is, the degree of discreteness of the data is similar.
[0083] Determine whether the data set D meets the following requirements:
[0084] (a) Conceptual model and model layer mapping f: C D —>MLbcc is a surjection, that is, C D Each concept in can find a corresponding mapping in MLbcc;
[0085] Condition a ensures that the information in dataset D can be fully mapped into the knowledge graph;
[0086] (b) The data classification in dataset D is balanced, uniform, and has homogeneous variance;
[0087] Condition b ensures the accuracy of subsequent analysis and avoids bias caused by data imbalance or large variance differences;
[0088] (c) The size of the dataset D |D| ≥ F((VC + ln(1 / d)) / e), where |D| is the size of the dataset, i.e., the number of samples in the dataset; d is the failure rate, i.e., the upper limit of the probability of model prediction errors; e is the error rate, i.e., the expected value of the model prediction error; and VC is the VC dimension (Vapnik-Chervonenkis Dimension), which is an indicator of model complexity. The larger the VC dimension, the more complex the data distribution that the model can fit, but it may also lead to overfitting.
[0089] Condition c ensures that the dataset D is of sufficient size to support subsequent statistical analysis;
[0090] (d) Dataset D consists of case data and medical images corresponding to the cases, and it is a one-to-many (1:m) relationship;
[0091] Condition d means that each case may contain multiple related medical images, which helps to provide more comprehensive data information to support analysis and modeling;
[0092] If the dataset D meets the above requirements a, b, c, and d, the constructed dataset D is considered to meet the requirements, and the dataset D that meets the requirements is used as the basal cell carcinoma case dataset.
[0093] It should be noted that the constructed basal cell carcinoma case dataset is a multimodal normative dataset for basal cell carcinoma. Multimodality means that the dataset contains multiple types of data, such as text (case records), values (test result values), and images (medical pictures).
[0094] Furthermore, preprocessing the basal cell carcinoma case data set includes: using mean value filling and regression algorithms to process missing values of the basal cell carcinoma case data set, and using a normalization method to perform feature scaling; using a spatial transformation algorithm to enhance image data, wherein the image is transformed by 180° rotation, horizontal flipping and shearing transformation, and the transformation coefficient is 0.2; and splitting the enhanced basal cell carcinoma case data set into a training set and a test set according to a split ratio of 80:20.
[0095] By performing preprocessing steps such as missing value processing, feature scaling, image data enhancement, and dataset splitting on the basal cell carcinoma case dataset, the quality and diversity of the data can be further improved, providing reliable data support for the follow-up and helping to improve the accuracy and robustness of the model.
[0096] S2, based on the preprocessed basal cell carcinoma case data set, the k-means algorithm was used to mine the fine-grained typing of basal cell carcinoma and verify the reliability and adequacy of the typing results;
[0097] Referring to the current clinical and pathological classification, combined with the clinical and pathological characteristics of BCC, the data features in the pre-processed basal cell carcinoma case data set are divided into logical type, classification type and quantitative type, and standardized. The standardization process includes converting the logical type data into 0, 1; converting the classification type data into 0, 1, 2, ...; and standardizing and mapping the quantitative type data into the range of [0, 1].
[0098] Logical types are usually expressed as yes / no or true / false, such as whether a person has a certain complication, etc. Through standardization, logical type data is converted into 0 or 1, such as no = 0, yes = 1;
[0099] Categorical types are usually represented as multiple categories, such as cancer stage, pathological type, etc. Through standardization, the categorical type data is converted into integers such as 0, 1, 2, ..., usually using One-Hot Encoding or Label Encoding for conversion, but here, for simplicity, it is directly mapped to continuous integers;
[0100] Quantitative types are usually expressed as continuous values, such as tumor size, age, etc. Standardization processing maps quantitative type data to the range of [0, 1]. The Min-Max normalization method can be used:
[0101]
[0102] Among them, x' represents the normalized value, x represents the original value, min(x) represents the minimum value in the data set, and max(x) represents the maximum value in the data set;
[0103] Based on the standardized data features, the basal cell carcinoma feature vector X is constructed, where X = {k 1 x 1 , k 2 x 2 , k 3 x 3 , …, k i x i}, where x 1 , x 2 , x 3 , …, x i is the standardized data feature, k 1 , k 2 , k 3 , …, k i is the corresponding weight value;
[0104] Based on the Euclidean distance, the k-means algorithm is used to cluster the feature vector X, and the statistical distribution of basal cell carcinoma characteristics is mined and summarized, which is the fine-grained classification of BCC. The objective function of the k-means algorithm is as follows:
[0105]
[0106] Among them, C k represents the kth cluster, K represents the total number of clusters, i.e. the number of clusters, x (i) Represents cluster C k The i-th data point in (k) Represents cluster C k The centroid of the cluster is the average value of all data points in the cluster. The objective function J represents the sum of the squares of the distances from all data points in the cluster to their centroids. The goal of the algorithm is to minimize J.
[0107] The specific process of clustering feature vector X using the k-means algorithm is as follows:
[0108] Randomly select K samples as the initial cluster centers, denoted as μ(1), μ(2), …μ(k), where K is a hyperparameter representing the number of clusters, i.e. the number of categories, μ(1) is the first cluster center, μ(2) is the second cluster center, and μ(k) is the kth cluster center;
[0109] Iteration step: For each sample x in the dataset i , calculate its Euclidean distance to the K cluster centers, and assign it to the class corresponding to the cluster center with the smallest distance, completing a cluster division; for each cluster, recalculate its cluster center position, if the geometric center does not coincide with the cluster center, use the geometric center as the new cluster center and re-divide the cluster;
[0110] Repeat the above iterative steps until the change in the cluster center position is less than the preset threshold, then stop and output the clustering results;
[0111] Plot the clustering performance indicators under different numbers of clusters into line graphs;
[0112] Observe the line graph and find an inflection point, namely the Elbow point. The Elbow point refers to the point in the line graph where the performance indicator value begins to stabilize or the change is no longer significant. This point usually corresponds to an "inflection point", that is, after this point, adding more cluster centers will no longer significantly improve the performance;
[0113] The cluster number K corresponding to the Elbow point is the optimal cluster center.
[0114] Furthermore, verifying the reliability and adequacy of the classification results includes: evaluating the reliability of the classification by calculating the internal evaluation indicators of the clustering results (such as the silhouette coefficient, Calinski-Harabasz index, etc.), verifying the adequacy of the classification by comparing with known clinical or pathological classifications, and using methods such as cross-validation to further verify the stability and accuracy of the classification results.
[0115] By performing feature classification, standardization, feature vector construction, and k-means clustering on the preprocessed basal cell carcinoma case data set, the fine-grained typing of basal cell carcinoma can be discovered. These typing results help to deeply understand the biological characteristics and clinical characteristics of basal cell carcinoma and provide a scientific basis for formulating personalized treatment plans. At the same time, by verifying the reliability and adequacy of the typing results, the accuracy and practicality of the typing results can be ensured.
[0116] Furthermore, after using the k-means algorithm to mine the fine-grained classification of basal cell carcinoma, the method further includes: relying on Psych, performing principal component analysis on clinical and pathological characteristics to summarize the correlation of various factors affecting the fine-grained classification, wherein the principal component analysis includes:
[0117] Standardize the data: subtract the mean from each feature of the original data and then divide by the standard deviation to ensure that each feature has the same importance;
[0118] Calculate the covariance matrix: Calculate the covariance matrix of the standardized data set to understand the correlation between different features;
[0119] Calculate eigenvalues and eigenvectors: Perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors; the eigenvectors represent the principal component directions in the data set, and the eigenvalues represent the variance of the data in these principal component directions;
[0120] Select the number of principal components: select the number of principal components to be retained based on the size of the eigenvalues; usually the number of principal components to be retained is the number of principal components that can explain a larger part of the data variance; this can be achieved by drawing a scree plot of the eigenvalues or setting a threshold;
[0121] Constructing a projection matrix: Select the number of principal components selected before and construct a projection matrix; the projection matrix is a matrix composed of eigenvectors and is used to project the original data into the new principal component space;
[0122] Dimensionality reduction transformation: Multiplying the standardized data set by the projection matrix can transform the data set from the original high-dimensional space to a new low-dimensional principal component space while retaining most of the information of the original data.
[0123] Furthermore, after using the k-means algorithm to mine the fine-grained classification of basal cell carcinoma, the method further includes: performing factor analysis on the clinical and pathological characteristics to reveal the potential factor structure, wherein the factor analysis includes:
[0124] Standardize data: Use z-score standardization or minimum and maximum scaling methods to standardize data to ensure that all variables have the same scale;
[0125] Extract factors: Use statistical software (such as SPSS, etc.) to calculate and obtain the results of factor loading matrix and commonality matrix; these results reveal the relationship between factors and variables and the internal structure of factors;
[0126] Determine the number of factors: Kaiser criterion, Scree plot and explained cumulative variance were used to explain the variance contribution and cumulative variance contribution of each factor to determine the number of factors to be retained;
[0127] Factor rotation: Use maximization Varimax and maximum likelihood estimation Promax to rotate the extracted factors to make the factors easier to interpret;
[0128] Factor interpretation: Analyze the rotated factor loading matrix, determine the relationship between each variable and factor based on the absolute value and sign of the loading, and give the factor a reasonable interpretation;
[0129] Factor score calculation: Calculate the score of each sample on each factor to further analyze the position and distribution of the sample in the factor space.
[0130] Through principal component analysis and factor analysis, we can have a deeper understanding of the correlation between clinical and pathological characteristics and their potential factor structure behind the fine-grained classification of basal cell carcinoma. These analysis results can provide more accurate diagnostic basis and treatment recommendations.
[0131] S3, obtaining patient data, matching the patient data with the fine-grained typing results, and obtaining the clinical subtype classification of basal cell carcinoma corresponding to the patient.
[0132] The fine-grained classification is matched with the platform BCC patient data, and key features related to basal cell carcinoma, such as tumor size, shape, color, and infiltration depth, are extracted from the patient data. The extracted patient features are matched with the features in the fine-grained classification results. According to the matching results, the patients are classified into the most similar clinical subtypes of basal cell carcinoma, and the clinical pathological manifestations, treatment plans, and prognosis of patients with different classifications are evaluated. Among them, the clinical subtype classification of basal cell carcinoma includes solid type, palisade type, superficial type, nodular type, morphea-like type, invasive type, fibroepithelioma type, etc.
[0133] Furthermore, after obtaining the clinical subtype classification of basal cell carcinoma corresponding to the patient, this embodiment also includes verifying the reliability and adequacy of the fine-grained classification, such as conducting expert sampling review, and relying on the reviewed sample labels to test the confidence of accurate labeling. When the confidence is greater than a preset threshold, it indicates that the reliability is good, otherwise re-labeling is performed.
[0134] The present invention improves the accuracy and resolution of clinical classification based on clustering fine-grained basal cell carcinoma classification. By understanding the specific subtype of each patient, doctors can understand the patient's disease status more accurately, thereby formulating more targeted treatment plans, improving treatment effects and reducing unnecessary side effects.
[0135] Example 2
[0136] The technical solution of the present invention is further described in detail below in conjunction with a specific embodiment and the accompanying drawings. It should be understood that the following embodiment is only used to explain the present invention, and is not used to limit the present invention.
[0137] Reference Figure 2 , Figure 2 FIG. 1 is a flow chart of an embodiment of the present invention in an implementation scenario. Figure 2 As shown, a method for classifying clinical subtypes of basal cell carcinoma in an embodiment of the present invention comprises the following steps:
[0138] Step 1: Through the multi-center sharing platform, under the guidance of the knowledge graph, 17,000 case data and 59,000 image data of BCC patients from 87 hospitals across the country (clinical case management system) were collected to obtain BCC multimodal original data, hierarchical data and thematic data, which were integrated into the multi-center BCC data resource integration platform to build a BCC case thematic database. The clinical data of BCC patients collected by the multi-center sharing platform were approved by the ethics committee, and all patients had signed informed consent forms;
[0139] Among them, the inclusion criteria for the integrated cases were: all cases were confirmed as BCC by histopathology, clinical images of skin lesions were routinely taken for each patient, and detailed information of the patients and skin lesions was complete; the exclusion criteria were: the diagnosis in the pathology report was unclear or controversial; the clinical information was missing, wrong, or incomplete; the imaging data or clinical images did not meet the requirements, such as low image clarity, out-of-focus or blurred images, etc.
[0140] Step 2: Extract BCC subject data from the BCC case subject database. Take the BCC subject data as the population and repeatedly use random sampling without replacement to form a multimodal normative data set D. Use QQ plot to test uniformity, ANOVA to test balance, and Levene's test for homogeneity of variance.
[0141] Let the model layer of the BCC knowledge graph be MLbcc and the concept model of the dataset D be C D , if D satisfies the following conditions:
[0142] (a)f:C D —>MLbcc is a full shot;
[0143] (b) Classification balance, uniformity, and homogeneity of variance;
[0144] (c) The scale of D |D| ≥ F((VC+ln(1 / d)) / e), where d is the failure rate, e is the error rate, and VC is the VC dimension;
[0145] (d) It consists of a one-to-many (1:m) pair of case data and medical images corresponding to the case;
[0146] Then D is called the BCC multimodal canonical dataset.
[0147] Step three: For metadata, the missing values are processed by using mean filling and regression algorithms, and the normalization method is used for feature scaling. The spatial transformation algorithm is used for image data enhancement. The image is transformed by 180° rotation, horizontal flipping and shearing transformation. The transformation coefficient is 0.2, and the 80:20 split rule is applied to construct multimodal standard training sets and test sets.
[0148] Step 4: Use the k-means algorithm to mine the BCC fine-grained classification and verify the reliability and adequacy of the obtained clinical classification.
[0149] Among them, refer to Figure 3 , Figure 3 FIG. 1 is a schematic diagram of the BBC clinical fine-grained classification mining process according to an embodiment of the present invention. Figure 3 As shown in the figure, the workflow of using k-means algorithm to mine BCC fine-grained typing specifically includes:
[0150] Sub-step 1: Referring to the current clinical and pathological classification, combined with the clinical and pathological characteristics of BCC, the characteristics are divided into three types: logical, categorical and quantitative. The logical type is standardized to 0, 1, the categorical type is standardized to 0, 1, 2, 3, ..., and the quantitative type is standardized to R + , R + ∈[0,1], forming the BCC clinical feature vector X = {x 1 , x 2 , …};
[0151] For example, "Whether to smoke" is described by logical type, yes: 1, no: 0;
[0152] "Blood type" is described using classification types, A: 0, B: 1, AB: 2, O: 3;
[0153] "Age" is described quantitatively using the standardized formula Map the data to between [0, 1];
[0154] As shown in Table 1, Table 1 is a table of the names, descriptions and characteristics of factors related to the classification of clinical subtypes of basal cell carcinoma:
[0155] Table 1
[0156]
[0157]
[0158] Sub-step 2: construct the BCC feature vector X = {k 1 x 1 , k 2 x 2 , k 3 x 3, …}, where k 1 , k 2 , k 3 , …, k i >0,∈I + , which is the weight value set artificially based on the doctor's clinical experience.
[0159] Sub-step 3, based on the Euclidean distance, the k-means algorithm is used to cluster X, and the statistical distribution of BCC features is mined and summarized, which is the BCC fine-grained classification. The objective function of the k-means algorithm is as follows:
[0160]
[0161] The specific process of clustering X using the k-means algorithm is as follows:
[0162] ① Randomly select K samples as the initial cluster centers, denoted as μ(1), μ(2), …μ(k); K is a hyperparameter, representing the number of clusters, that is, the number of categories;
[0163] ②For each sample x in the data set (i) , calculate its Euclidean distance to the K cluster centers, and assign it to the class corresponding to the cluster center with the smallest distance, completing a cluster division;
[0164] ③ For each cluster, recalculate its cluster center position. If the geometric center does not coincide with the cluster center, use the geometric center as the new cluster center and re-divide the clusters.
[0165] ④ Repeat the above steps ② and ③ until the center position of the cluster remains unchanged, stop and output the clustering results.
[0166] ⑤ Draw the clustering performance indicators under different cluster numbers into a line graph;
[0167] ⑥Observe the line chart and find an inflection point, namely the Elbow point;
[0168] ⑦The cluster number K corresponding to the Elbow point is the best choice.
[0169] Sub-step 4, relying on Psych, principal component analysis (PCA) and exploratory factor analysis (EFA) were performed on the clinical and pathological characteristics to summarize the correlation of various factors affecting the fine-grained typing;
[0170] Among them, the specific steps of principal component analysis PCA are as follows:
[0171] Standardize the data: subtract the mean from each feature of the original data and then divide by the standard deviation to ensure that each feature has the same importance;
[0172] Calculate the covariance matrix: Calculate the covariance matrix of the standardized data set to understand the correlation between different features;
[0173] Calculate eigenvalues and eigenvectors: Perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors; the eigenvectors represent the principal component directions in the data set, and the eigenvalues represent the variance of the data in these principal component directions;
[0174] Select the number of principal components: select the number of principal components to be retained based on the size of the eigenvalues; usually the number of principal components selected to be retained is the number of principal components that can explain a larger part of the data variance;
[0175] Construct a projection matrix: Select the number of principal components selected in the previous step and construct a projection matrix; the projection matrix is a matrix composed of eigenvectors and is used to project the original data into the new principal component space;
[0176] Dimensionality reduction transformation: Multiply the standardized data set by the projection matrix to transform the data set from the original high-dimensional space to a new low-dimensional principal component space;
[0177] The specific steps of factor analysis EFA are as follows:
[0178] Standardize data: Use z-score standardization or minimum and maximum scaling methods to standardize data to ensure that all variables have the same scale;
[0179] Extract factors: Use statistical software to calculate and obtain factor loading matrices, commonality matrices and other results; these results can reveal the relationship between factors and variables and the internal structure of factors;
[0180] Determine the number of factors: Kaiser criterion, Scree plot and explained cumulative variance were used to explain the variance contribution and cumulative variance contribution of each factor to determine the number of factors to be retained;
[0181] Factor rotation: Use maximization Varimax and maximum likelihood estimation Promax to rotate the extracted factors to make the factors easier to interpret;
[0182] Factor interpretation: Analyze the rotated factor loading matrix, determine the relationship between each variable and factor based on the absolute value and sign of the loading, and give the factor a reasonable interpretation;
[0183] Factor score calculation: Calculate the score of each sample on each factor to further analyze the position and distribution of the sample in the factor space.
[0184] Sub-step 5: Match the fine-grained classification results with the platform BCC patient data, evaluate the clinical pathological manifestations, treatment plans and prognosis of patients with different classifications, and verify the reliability and adequacy of the fine-grained classification.
[0185] Step five: manually annotate the fine-grained typing labels of 17,000 case data and 59,000 images, conduct expert sampling review, and use the reviewed sample labels to test the confidence level of accurate labeling to ensure that the p-value is ≥ 0.05, otherwise re-label.
[0186] like Figure 4 As shown, Figure 4 This is a visualization diagram of the clinical fine-grained typing of BCC screened out by cluster analysis in an embodiment of the present invention, in which different colors represent different fine-grained typings. Such a diagram helps to intuitively understand the differences and relationships between various clusters (i.e., fine-grained typing). In the diagram, each data point may represent a patient or a clinical sample of a patient. These data points are divided into different groups by the clustering algorithm according to their clinical information (such as patient demographic information, clinical electronic medical record data, auxiliary examination results, etc.). The data points in each group are relatively close in the feature space, while the data points between different groups are relatively far away. Different colors are used to distinguish these clusters. For example, red may represent a specific BCC subtype, blue represents another subtype, and so on. The choice of color is arbitrary, and it is important to keep the color consistent with the corresponding cluster so that different fine-grained typings can be clearly represented in the figure.
[0187] Example 3
[0188] In this embodiment, a basal cell carcinoma clinical subtype classification device is also provided, which is used to implement the above embodiments and preferred implementations, and the descriptions that have been made will not be repeated. The basal cell carcinoma clinical subtype classification device includes a memory and a processor, the memory stores a computer program, and the processor implements the steps of the basal cell carcinoma clinical subtype classification method as described above when executing the computer program. The specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and this embodiment will not be repeated here.
[0189] An embodiment of the present invention further provides a storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.
[0190] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store computer programs.
[0191] An embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0192] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0193] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0194] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0195] In the several embodiments provided in this application, it should be understood that the disclosed technical contents can be implemented in other ways.
[0196] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for classifying clinical subtypes of basal cell carcinoma, characterized in that: include: Acquire a basal cell carcinoma case data set, and preprocess the basal cell carcinoma case data set; Based on the preprocessed basal cell carcinoma case data set, the k-means algorithm was used to mine the fine-grained classification of basal cell carcinoma, and the reliability and adequacy of the classification were verified; Acquiring patient data, matching the patient data with the fine-grained typing results, and obtaining a clinical subtype classification of basal cell carcinoma corresponding to the patient; The obtaining of the basal cell carcinoma case data set comprises: A data resource integration platform is constructed based on the knowledge graph, and a basal cell carcinoma case subject database is established, wherein the data resource integration platform is connected to N clients for obtaining basal cell carcinoma case data from the clients, and the clients are used to manage case data of corresponding hospitals; Acquire basal cell carcinoma case data through the basal cell carcinoma case subject database to construct a basal cell carcinoma case data set; The k-means algorithm is used to mine fine-grained typing based on the preprocessed basal cell carcinoma case data set, including: The data features in the preprocessed basal cell carcinoma case data set are divided into logical type, classification type and quantitative type, and standardized. The standardization process includes converting the logical type data into 0, 1; converting the classification type data into 0, 1, 2, ...; and mapping the quantitative type data into the range of [0, 1]. Based on the standardized data features, the basal cell carcinoma feature vector X is constructed, where X={ k 1 x 1 , k 2 x 2 , k 3 x 3 , …, k i x i },in, x 1 , x 2 , x 3 , …, x i is the standardized data feature, k 1 , k 2 , k 3 , …, k i is the corresponding weight value; Based on the Euclidean distance, the k-means algorithm is used to cluster the feature vector X to mine and summarize the statistical distribution of basal cell carcinoma features. The objective function of the k-means algorithm is as follows: ; C k Indicates k clusters, K represents the total number of clusters, that is, the number of clusters, x (i) Representation Cluster C k The i Data points, μ (k) Representation Cluster C k The centroid of the cluster is the average value of all data points in the cluster. The objective function J Represents the sum of the squares of the distances from all data points in the cluster to their centroids. The goal of the algorithm is to minimize J .
2. The method for classifying clinical subtypes of basal cell carcinoma according to claim 1, characterized in that: Acquiring basal cell carcinoma case data through the basal cell carcinoma case subject database, and constructing a basal cell carcinoma case data set includes: Using random sampling without replacement to obtain basal cell carcinoma case data from the basal cell carcinoma case subject database, constructing a data set D; Let the model layer of the basal cell carcinoma knowledge graph be MLbcc and the concept model of the dataset D be C D , QQ plot was used to test the uniformity of sampling data, ANOVA was used to test the balance of sampling data, and Levene's test was used to test the homogeneity of variance of sampling data; Determine whether the data set D meets the following requirements: Conceptual model and model layer mapping f: C D —>MLbcc is a surjection, that is, C D Each concept in can find a corresponding mapping in MLbcc; The data in dataset D are balanced, uniform, and have homogeneous variance; The size of the dataset D |D| ≥ F((VC+ln(1 / d)) / e), where |D| is the size of the dataset, d is the failure rate, e is the error rate, and VC is the VC dimension; The dataset D consists of case data and medical images corresponding to the cases, and it is a one-to-many (1:m) relationship; If the dataset D meets the requirements, the constructed dataset D is considered to meet the requirements, and the dataset D that meets the requirements is used as the basal cell carcinoma case dataset.
3. The method for classifying clinical subtypes of basal cell carcinoma according to claim 1, characterized in that: Preprocessing the basal cell carcinoma case data set includes: Mean value filling and regression algorithms were used to process missing values in the basal cell carcinoma case dataset, and normalization methods were used for feature scaling; The spatial transformation algorithm is used for image data enhancement, where the image is transformed by 180° rotation, horizontal flipping and shearing transformation, with a transformation coefficient of 0.2; The enhanced basal cell carcinoma case dataset was split into a training set and a test set according to a 80:20 split ratio.
4. The method for classifying clinical subtypes of basal cell carcinoma according to claim 1, characterized in that: Clustering feature vector X using the k-means algorithm includes: Randomly select K samples as the initial cluster centers, denoted as μ(1) , μ(2) ,… μ(k) , where K is a hyperparameter, representing the number of clusters, that is, the number of categories, μ(1) is the first cluster center, μ(2) is the second cluster center, μ(k) For the k Cluster centers; Iteration step: for each sample in the dataset x i , calculate its Euclidean distance to the K cluster centers, and assign it to the class corresponding to the cluster center with the smallest distance, completing a cluster division; for each cluster, recalculate its cluster center position, if the geometric center does not coincide with the cluster center, use the geometric center as the new cluster center and re-divide the cluster; Repeat the above iterative steps until the change in the cluster center position is less than the preset threshold, then stop and output the clustering results; Plot the clustering performance indicators under different numbers of clusters into line graphs; Observe the line chart and find an inflection point, the Elbow point; The cluster number K corresponding to the Elbow point is the optimal cluster center.
5. The method for classifying clinical subtypes of basal cell carcinoma according to claim 1, characterized in that: After using the k-means algorithm to mine the fine-grained typing of basal cell carcinoma, the method further includes: The principal component analysis of clinical and pathological characteristics was performed to summarize the correlation of various factors affecting fine-grained typing. The principal component analysis included: Standardize the data: subtract the mean from each feature of the original data and then divide by the standard deviation to ensure that each feature has the same importance; Calculate the covariance matrix: Calculate the covariance matrix of the standardized data set to understand the correlation between different features; Calculate eigenvalues and eigenvectors: Perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors; the eigenvectors represent the principal component directions in the data set, and the eigenvalues represent the variance of the data in these principal component directions; Select the number of principal components: Select the number of principal components to be retained based on the size of the eigenvalues; Constructing a projection matrix: Select the number of principal components selected before and construct a projection matrix; the projection matrix is a matrix composed of eigenvectors and is used to project the original data into the new principal component space; Dimensionality reduction transformation: Multiply the standardized data set by the projection matrix to transform the data set from the original high-dimensional space to a new low-dimensional principal component space.
6. The method for classifying clinical subtypes of basal cell carcinoma according to claim 5, characterized in that: After using the k-means algorithm to mine the fine-grained typing of basal cell carcinoma, the method further includes: Factor analysis of clinical and pathological features was performed to reveal the underlying factor structure, including: Standardize data: Use z-score standardization or minimum and maximum scaling methods to standardize data to ensure that all variables have the same scale; Extract factors: Use statistical software to calculate and obtain the results of factor loading matrix and commonality matrix; these results reveal the relationship between factors and variables and the internal structure of factors; Determine the number of factors: Kaiser criterion, Scree plot and explained cumulative variance were used to explain the variance contribution and cumulative variance contribution of each factor to determine the number of factors to be retained; Factor rotation: Use maximization Varimax and maximum likelihood estimation Promax to rotate the extracted factors to make the factors easier to interpret; Factor interpretation: Analyze the rotated factor loading matrix, determine the relationship between each variable and factor based on the absolute value and sign of the loading, and assign factor interpretations; Factor score calculation: Calculate the score of each sample on each factor to further analyze the position and distribution of the sample in the factor space.
7. A device for classifying clinical subtypes of basal cell carcinoma, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method for classifying clinical subtypes of basal cell carcinoma according to any one of claims 1 to 6 are implemented.
8. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for classifying clinical subtypes of basal cell carcinoma according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-modal probability distribution adaptive primary liver cancer pathology grading prediction method
CN114898872A
Skin basal cell carcinoma fine-grained typing method based on clustering analysis algorithm
CN118379730A