A blood vessel data classification and identification method based on a large language model and a clustering algorithm
By extracting vascular lesion features from literature using a large language model and clustering algorithm, and combining them with a weighted K-means model for automatic annotation, the problem of low efficiency of manual annotation in existing technologies is solved, and efficient and accurate vascular lesion risk prediction is achieved.
Patent Information
- Application Number
- CN202510751644.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2025-05-26
- Filing Date
- 2025-06-06
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-06-06
AI Technical Summary
Existing technologies for predicting disease risk using venous and arterial vascular data require extensive manual annotation, resulting in low efficiency in data collection and annotation, and consuming significant human and financial resources.
We use a large language model to extract vascular lesion-related features from literature, use clustering algorithms for data filtering, filling and clustering, and combine a weighted K-means model for automatic annotation to reduce manual annotation work.
It reduced the cost of data collection and labeling, improved efficiency, and achieved more accurate prediction of vascular lesion risk through the application of weight calculation and clustering models.
Smart Images

Figure CN120611221B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence (large language models, clustering algorithms) technology, and in particular to a method and system for classifying and identifying vascular data using large language models and clustering algorithms. This invention originates from a research project: the National Science and Technology Major Project for the Prevention and Treatment of Cancer, Cardiovascular and Cerebrovascular Diseases, and Respiratory and Metabolic Diseases. Background Technology
[0002] Currently, when using venous vascular data for venous vascular disease risk prediction, a venous vascular disease risk prediction model is first built and trained using historical venous vascular data. Then, new venous vascular data to be predicted is input into the trained model for prediction. This process requires manual annotation of each historical venous vascular data point. This manual annotation is done by experts and doctors using their professional knowledge to annotate the data, resulting in labeled data. The same applies to arterial vascular data for arterial vascular disease risk prediction. We know that to obtain a relatively accurate venous / arterial vascular disease risk prediction model, a large amount of historical venous / arterial vascular data needs to be collected and annotated to obtain a large number of model training samples. In other words, a large amount of data collection and annotation work is required, which consumes significant human and financial resources and is inefficient. Given the limited time and energy of experts and doctors, how to improve the efficiency of vascular data collection and annotation, and achieve vascular data classification and identification, has become an urgent technical problem to be solved. Summary of the Invention
[0003] This invention addresses the problems and shortcomings of existing technologies by providing a method and system for classifying and recognizing vascular data based on a large language model and clustering algorithm.
[0004] The present invention solves the above-mentioned technical problems through the following technical solution:
[0005] This invention provides a method for classifying and recognizing vascular data based on a large language model and clustering algorithm, characterized by comprising:
[0006] S1. Use a large language model to find N literatures related to vascular lesions, and extract different features and data related to vascular lesions from each literature as a corresponding literature data. Select A different features from the N literature data and sort them from high to low frequency of occurrence. Discard A1 features with too low frequency of occurrence and keep A2 features. A = A1 + A2. The occurrence of a feature in a literature data is counted as one occurrence.
[0007] S2. Select from N document data that a certain document data contains all A2 features or is missing only one of the A2 features or one feature. The number of selected documents is denoted as N1.
[0008] S3. For the missing features or data in N1 document data, construct a regression model using the document data that is not missing to fill in the missing feature data;
[0009] S4. Calculate the weight value of each of the A2 features by using the occurrence frequency of the A2 features in the N literature data;
[0010] S5. Obtain the data of each feature in the A2 feature of each blood vessel in M clinical patients to obtain M clinical data. M clinical patients include patients with various types of vascular lesions without vascular lesions. Actively label M1 of the clinical data as various vascular lesion types, M=M1+M2, M1<M2.
[0011] S6. Using the labeled M1 clinical data and the weight values of each feature, the weighted K-means clustering model is clustered into clusters of various vascular lesion types.
[0012] S7. Input the unlabeled M2 clinical data points and N1 literature data points into the clustering model to obtain each clustering result, and label the M2 clinical data points and N1 literature data points according to the clustering results.
[0013] This invention also provides a vascular data classification and recognition system based on a large language model and clustering algorithm, characterized in that it includes:
[0014] A data extraction module is used to find N vascular lesion-related literature using a large language model, and extract different features and data related to vascular lesions from each literature as a corresponding literature data. A different features are selected from the N literature data and sorted from high to low frequency of occurrence. A1 features with too low frequency of occurrence are discarded, and A2 features are kept. A = A1 + A2. The occurrence of a feature in a literature data is counted as one occurrence.
[0015] A data filtering module is used to filter out data from N documents that contain all A2 features or are missing only one of the A2 features. The number of documents filtered out is denoted as N1.
[0016] A data imputation module is used to construct a regression model using the literature data that is not missing to fill in the missing feature data for N1 literature data that only lacks features or data.
[0017] A weight calculation module is used to calculate the weight value of each of the A2 features by using the occurrence frequency of A2 features in N literature data;
[0018] A data acquisition module is used to acquire data of each feature in the A2 feature of each blood vessel in M clinical patients, and obtain M clinical data. The M clinical patients include patients with various types of vascular lesions without vascular lesions. Active type labeling is performed on M1 of the clinical data, and the labels are marked as various types of vascular lesions. M = M1 + M2, M1 < M2.
[0019] A data clustering module is used to cluster the weighted K-means clustering model using M1 labeled clinical data and the weight values of each feature, and the clusters are divided into clusters of various vascular lesion types.
[0020] A data annotation module is used to input the unannotated M2 clinical data entries and N1 literature data entries into the clustering model to obtain each cluster result, and to annotate the M2 clinical data entries and N1 literature data entries with their types based on the clustering results.
[0021] The positive and progressive effects of this invention are as follows:
[0022] In this invention, a portion of the feature data related to vascular lesions is extracted from existing literature as training samples for the vascular lesion risk prediction model. This can greatly reduce data collection work, save manpower and financial resources, and improve collection efficiency. Furthermore, if the extracted feature data related to vascular lesions corresponds to a vascular lesion type, it can further reduce data annotation work, further save manpower and financial resources, and further improve collection efficiency.
[0023] In this invention, feature data from multiple clinical patients are collected, and only a portion (less than half) of the feature data is labeled. The labeled feature data is used to train a clustering model. The remaining unlabeled feature data is then clustered using the clustering model to obtain corresponding clustering results. These clustering results are then automatically labeled one by one. By using a portion of the labeled feature data to train the clustering model, an accurate clustering model can be obtained. This clustering model is then used to cluster and label the unlabeled feature data, which can greatly reduce the data labeling work, save manpower and financial resources, improve labeling efficiency, and achieve vascular data classification and recognition.
[0024] In this invention, features that appear more frequently are given a larger weight, while features that appear less frequently are given a smaller weight. This results in more accurate clustering results, which in turn makes the annotation results more accurate and is beneficial for the training of subsequent vascular lesion risk prediction models.
[0025] In this invention, filling in the missing features or data in the literature data can expand the number of training samples, which is beneficial for the training of subsequent vascular lesion risk prediction models. Attached Figure Description
[0026] Figure 1 This is a flowchart of a preferred embodiment of the vascular data classification and recognition method of the present invention.
[0027] Figure 2 This is a structural block diagram of a vascular data classification and recognition system according to a preferred embodiment of the present invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example 1
[0029] like Figure 1 As shown, this embodiment of the invention provides a method for classifying and recognizing vascular data based on a large language model and clustering algorithm, including:
[0030] Step 101: Using a large language model, find N relevant documents on vascular (venous or arterial) lesions by searching for keywords related to vascular lesions. Extract different features and their data related to vascular lesions from each document as a corresponding document entry. Each document entry is associated with its source document for easy tracing of the original source. Select A different features from the N document entries and sort them by frequency of occurrence from highest to lowest. Discard A1 features with too low an occurrence frequency (e.g., discard a feature whose occurrence frequency is below a preset value), leaving A2 features. A = A1 + A2, where each occurrence of a feature in a document entry is counted as one occurrence. Here, N, A, A1, and A2 are all positive integers.
[0031] In this step, the selected A different features are sorted in descending order of frequency of occurrence. If a feature is ranked higher, it means that the feature is mentioned more often, proving that the feature is more important for characterizing vascular lesions. If a feature is ranked lower, it means that the feature is mentioned less often, proving that the feature is relatively less important for characterizing vascular lesions and can be discarded.
[0032] In step 101, the sources of the literature include CNKI, VIP, Wangfang, PubMed, Sino Med, or the Web of Science database, etc.
[0033] Step 102: Select from N document data that all contain all A2 features or are missing only one of the A2 features. The number of selected documents is denoted as N1.
[0034] In this step, if a document contains all A2 features and their data, it indicates that the document is complete and is retained. If a document is missing only one of the A2 features or its data, it indicates that the document has a defect but can be remedied and is also retained. All other documents are not retained. This process can be used to select N1 documents.
[0035] Furthermore, to obtain more literature data entries that meet the screening criteria, this embodiment of the invention designs a manual completion-assisted screening method. The manual completion-assisted screening method is as follows: For any literature data entry lacking two or three features or any data of the A2 features among N literature data entries, a reminder message is sent to staff. The staff confirms whether the data of at least one of the missing two or three features can be calculated based on the data provided in the literature traced back to that literature data entry. If so, the missing feature is filled in, thus obtaining N literature data entries after manual completion. From the N literature data entries after manual completion, a literature data entry containing all of the remaining A2 features or missing only one of the remaining A2 features or any missing feature is selected. The number of selected entries is denoted as N1. In this method, for a literature data entry lacking two or three features, or for a literature data entry lacking two or three features, staff participate. The staff confirms whether the data of at least one missing feature can be calculated based on the data of that literature data traced back to the literature provided in the literature. If so, the missing feature is filled in, thus obtaining N literature data entries after manual completion. Then, literature data that meets the screening requirements is selected from the N literature data entries after manual completion.
[0036] Furthermore, in order to obtain more literature data entries that meet the screening criteria, this embodiment of the invention designs an automatic filling-assisted screening method. The automatic filling-assisted screening method is as follows: In step 101, a calculation formula is predefined for each of the remaining A2 features; In step 102, for any literature data entry in the N literature data entries that is missing a feature or its data among the A2 features, it is determined whether the data of the missing feature can be calculated based on the data provided in the literature traced back to that literature data entry and the calculation formula of each missing feature. If so, the missing feature is automatically calculated and filled in, thereby obtaining N literature data entries after automatic filling; From the N literature data entries after automatic filling, a certain literature data entry is selected that contains all of the remaining A2 features or only lacks one of the remaining A2 features or its feature. The number of selected entries is denoted as N1. In this method, for a document data that is missing two or three features, or for a document data that is missing two or three features, the system automatically determines whether the missing features can be calculated based on the data provided in the document and the calculation formula for each missing feature. If so, the missing features are automatically calculated and filled in, thereby obtaining N documents data with automatic filling. Then, documents data that meet the screening requirements are selected from the N documents data with automatic filling.
[0037] Step 103: For the missing features or data in the N1 literature data, construct a regression model using the literature data that is not missing to fill in the missing feature data.
[0038] The specific implementation method of this step is as follows: The dataset with no missing features in the N1 document data is designated as the normal dataset, and the dataset with missing features is designated as the abnormal dataset. For a specific document data in the abnormal dataset, the feature with missing features is used as the target quantity, and the feature with no missing features is used as the variable. For each document data in the normal dataset, a regression model is built using the feature data of the target quantity as the output and the feature data of the variable as the input, resulting in a trained regression model. For the document data in the abnormal dataset, the feature data of the variable is substituted into the trained regression model to calculate the value, and the calculated value fills in the missing feature data.
[0039] To more quickly fill in the missing feature data, the embodiments of the present invention have been further optimized:
[0040] Construct the matrix: Take each of the N1 document data as a row and leave each of the A2 features as a column to construct an N1-row A2-column matrix.
[0041] Row swapping: Swap rows with no missing feature data to the top row of the matrix, swap rows with missing feature data to the bottom row of the matrix, and arrange rows with missing the same feature data into adjacent rows for easier subsequent processing.
[0042] For the top row of each row that lacks the same feature data, the feature that lacks the feature data is used as the target quantity, and the feature that does not lack the feature data is used as the variable.
[0043] For each row with complete feature data, a training regression model is constructed using the feature data of the target quantity as output and the feature data of the variable as input, resulting in a trained regression model.
[0044] For the top row of each row that lacks the same feature data, substitute the feature data of the variable into the trained regression model to calculate the value, and fill the missing feature data with the calculated value.
[0045] For rows other than the top row that lack the same feature data, the feature data of the variable is substituted into the regression model after training to calculate the value. The calculated value fills the missing feature data in this row. This can quickly calculate the data that lacks the same feature, saving time.
[0046] Step 104: Calculate the weight value of each of the A2 features by using the occurrence counts of the A2 features in the N literature data. The weight value of each feature is equal to the occurrence count of each feature divided by the total occurrence count of the A2 features.
[0047] Step 105: Obtain the data of each feature in the A2 feature of the blood vessels of each of the M clinical patients to obtain M clinical data. The M clinical patients include patients with various types of vascular lesions (including those without vascular lesions). Actively label the M1 clinical data as various types of vascular lesions containing those without vascular lesions. M = M1 + M2, M1 < M2, where M, M1, and M2 are all positive integers.
[0048] In this step, the M clinical data points are divided into M1 and M2 clinical data points. The M1 clinical data points are actively labeled with vascular lesion types, while the M2 clinical data points are left unlabeled. This step only labels a portion of the collected M clinical data points. The labeled feature data is then used to train the clustering model. The remaining unlabeled feature data is clustered using the clustering model to obtain the corresponding clustering results. The clustering results are then used for automatic labeling, which can significantly reduce data collection work, save manpower and financial costs, and improve collection efficiency.
[0049] Step 106: Use the labeled M1 clinical data and the weight values of each feature to perform clustering on the weighted K-means clustering model. The clusters are divided into clusters containing various types of vascular lesions with no vascular lesions. Each cluster is represented by a number, with the cluster with no vascular lesions represented by 0.
[0050] If the blood vessel is a vein, 0 indicates no venous vascular lesions, 1 indicates deep vein valve insufficiency, and 2 indicates deep vein thrombosis, etc.
[0051] If the blood vessel is an artery, 0 represents a cluster without arterial vascular lesions, 1 represents a cluster of atherosclerosis, 2 represents a cluster of aneurysms, 3 represents a cluster of thromboangiitis obliterans, and 4 represents a cluster of arterial dissections, etc.
[0052] Step 107: Input the unlabeled M2 clinical data points and N1 literature data points into the weighted K-means clustering model to obtain each clustering result, and label the M2 clinical data points and N1 literature data points according to the clustering results.
[0053] In this step, each feature data point in the M2 clinical datasets can be clustered, and the feature data point is labeled based on this clustering result. Similarly, each feature data point in the N1 literature datasets can be clustered, and the feature data point is labeled based on this clustering result. After all the feature data points are labeled, they can be used as training samples to train the vascular lesion risk prediction model, resulting in a more accurate and optimized vascular lesion risk prediction model.
[0054] like Figure 2 As shown, this embodiment of the invention also provides a vascular data classification and recognition system based on a large language model and clustering algorithm, including a data extraction module 1, a data filtering module 2, a data filling module 3, a weight calculation module 4, a data acquisition module 5, a data clustering module 6, and a data annotation module 7.
[0055] Data extraction module 1 is used to find N vascular lesion-related literature using a large language model, and extract different features and data related to vascular lesions from each literature as a corresponding literature data. A different features are selected from the N literature data and sorted in descending order of frequency of occurrence. A1 features with too low frequency of occurrence are discarded, and A2 features are kept. A = A1 + A2. The occurrence of a feature in a literature data is counted as one occurrence.
[0056] Data filtering module 2 is used to filter out data from N documents that contain all of the A2 features or that are missing only one of the A2 features. The number of filtered documents is denoted as N1.
[0057] The data imputation module 3 is used to construct a regression model using the literature data that is not missing to fill in the missing feature data for the N1 literature data that are only missing features or their data.
[0058] The weight calculation module 4 is used to calculate the weight value of each of the A2 features by using the occurrence frequency of the A2 features in N literature data.
[0059] The data acquisition module 5 is used to acquire the data of each feature in the A2 feature of each blood vessel in M clinical patients, and obtain M clinical data. The M clinical patients include patients with various types of vascular lesions without vascular lesions. Active type labeling is performed on M1 of the clinical data, and the labels are marked as various types of vascular lesions with non-vascular lesions. M = M1 + M2, M1 < M2.
[0060] The data clustering module 6 is used to cluster the weighted K-means clustering model using the labeled M1 clinical data and the weight values of each feature, and the clusters are divided into clusters containing various types of vascular lesions without vascular lesions.
[0061] The data annotation module 7 is used to input the unannotated M2 clinical data and N1 literature data into the weighted K-means clustering model to obtain each clustering result, and to annotate the M2 clinical data and N1 literature data according to the clustering results.
[0062] This invention also provides an electronic device, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to execute the aforementioned method.
[0063] This invention also provides a computer-readable storage medium storing computer program instructions thereon, which, when executed by a processor, implement the aforementioned method.
[0064] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention. Example 2
[0065] This embodiment optimizes the technical solution based on embodiment 1. When the corresponding vascular lesion type can be extracted at the same time as the features and data, it can be extracted together, so that the literature data does not need to be labeled again, further reducing the labeling work.
[0066] Step 201: Using a large language model, find N relevant documents on vascular lesions by searching for keywords related to vascular lesions. For any of the N relevant documents, if the document contains at least one type of vascular lesion, use the large language model to extract each type of vascular lesion and its corresponding different features and data as a corresponding document data. If the document does not contain any type of vascular lesion, use the large language model to extract different features related to vascular lesions and their data as a corresponding document data. Based on this, a total of N' document data are obtained, where N' ≥ N. Each document data is associated with the source document, facilitating the tracing of the original source. Select A different features from the N' document data and sort them from highest to lowest frequency. Discard A1 features with too low a frequency (e.g., if the frequency of a feature is below a preset value, discard it), leaving A2 features, where A = A1 + A2. The occurrence of a feature in a document data is counted as one occurrence.
[0067] Step 202: Select from N' document data that all document data contain all of the A2 features or that are missing only one of the A2 features or one of the features. The number of selected documents is denoted as N1.
[0068] Step 203: For the missing features or data in the N1 literature data, construct a regression model using the literature data that is not missing to fill in the missing feature data.
[0069] Step 204: Calculate the weight value of each of the A2 features by using the occurrence frequency of the A2 features in the N' literature data.
[0070] Step 205: Obtain the data of each feature in the A2 feature of each blood vessel in M clinical patients to obtain M clinical data. M clinical patients include patients with various types of vascular lesions without vascular lesions. Actively label M1 of the clinical data as various types of vascular lesions without vascular lesions, M = M1 + M2, M1 < M2.
[0071] Step 206: Using the labeled M1 clinical data and N1 literature data with vascular lesion types and the weight values of each feature, the weighted K-means clustering model is clustered into clusters containing various vascular lesion types without vascular lesions. Each cluster is represented by a number, where the cluster without vascular lesions is represented by 0.
[0072] Step 107: Input the unlabeled M2 clinical data and N1 literature data without vascular lesion types into the weighted K-means clustering model to obtain each cluster result, and label the M2 clinical data and N1 literature data according to the clustering results.
[0073] While specific embodiments of the present invention have been described above, those skilled in the art should understand that these are merely illustrative examples, and the scope of protection of the present invention is defined by the appended claims. Those skilled in the art can make various changes or modifications to these embodiments without departing from the principles and essence of the present invention, but all such changes and modifications fall within the scope of protection of the present invention.
Claims
1. A method for classifying and recognizing vascular data based on a large language model and clustering algorithm, characterized in that, include: S1. Use a large language model to find N literatures related to vascular lesions. For any one of the N literatures related to vascular lesions, if the literature contains at least one type of vascular lesion, use the large language model to extract each type of vascular lesion and its corresponding different features and data as a corresponding literature data. If the literature does not contain any type of vascular lesion, use the large language model to extract different features related to vascular lesions and their data as a corresponding literature data. A total of N' literature data are obtained, N'≥N. Select A different features from the N' literature data and sort them from high to low frequency. Discard A1 features with too low frequency and keep A2 features. A=A1+A2. The occurrence of a feature in a literature data is counted as one occurrence. S2. Select from N' document data a document that contains all A2 features or is missing only one of the A2 features or one of the features. The number of selected documents is denoted as N1. S3. For the missing features or data in N1 document data, construct a regression model using the document data that is not missing to fill in the missing feature data; S4. Calculate the weight value of each of the A2 features by using the occurrence frequency of the A2 features in the N' literature data; S5. Obtain the data of each feature in the A2 feature of each blood vessel in M clinical patients to obtain M clinical data. M clinical patients include patients with various types of vascular lesions without vascular lesions. Actively label M1 of the clinical data as various vascular lesion types, M=M1+M2, M1<M2. S6. Using the labeled M1 clinical data and N1 literature data with vascular lesion types and the weight values of each feature, the weighted K-means clustering model is used to cluster the data into clusters of various vascular lesion types. S7. Input the unlabeled M2 clinical data and N1 literature data that do not contain vascular lesion types into the weighted K-means clustering model to obtain each clustering result, and label the M2 clinical data and N1 literature data according to the clustering results.
2. The vascular data classification and recognition method based on a large language model and clustering algorithm as described in claim 1, characterized in that, In S1, each piece of literature data is associated with the literature that served as the source; In S2, for any document data in N' document data that is missing two or three features or data of A2 features, a reminder message is sent to the staff. The staff then confirms whether the data of at least one of the missing two or three features can be calculated based on the data provided in the document traced by that document data. If so, the missing features are filled in, thus obtaining N' document data after manual filling. From the N' manually filled literature data, select a literature data that contains all of the A2 features or is missing only one of the A2 features or one feature. The number of selected literature data is denoted as N2. In S3, for the features or data that are missing in the N2 document data, a regression model is constructed using the document data that is not missing to fill in the missing feature data; In S6, the weighted K-means clustering model is used to cluster the M1 labeled clinical data and N2 literature data with vascular lesion types and the weight values of each feature to form clusters of various vascular lesion types. S7. Input the M2 unlabeled clinical data and N2 literature data that do not contain vascular lesion types into a weighted K-means clustering model to obtain each clustering result. Label the M2 clinical data and N2 literature data according to the clustering results.
3. The vascular data classification and recognition method based on a large language model and clustering algorithm as described in claim 1, characterized in that, In S1, each piece of literature data is associated with the literature that is the source, and the calculation formula for each of the remaining A2 features is predefined. In S2, for any document data in N' document data that lacks a feature or its data among the A2 features, it is determined whether the data of the missing feature can be calculated based on the data provided in the document traced by that document data and the calculation formula of each missing feature. If it can, the missing feature data is automatically calculated and filled in, thereby obtaining N' document data after automatic filling. From the N' automatically filled document data, select a document data that contains all of the A2 features or is missing only one of the A2 features or one feature. The number of selected documents is denoted as N3. In S3, for the features or data that are missing in the N3 literature data, a regression model is constructed using the literature data that are not missing to fill in the missing feature data; In S6, the weighted K-means clustering model is used to cluster the M1 labeled clinical data and N3 literature data with vascular lesion types and the weight values of each feature to form clusters of various vascular lesion types. S7. Input the unlabeled M2 clinical data and N3 literature data that do not contain vascular lesion types into the weighted K-means clustering model to obtain each cluster result, and label the M2 clinical data and N3 literature data according to the clustering results.
4. The method for classifying and recognizing vascular data based on a large language model and clustering algorithm as described in claim 1, characterized in that, In S3, N1 document data entries without missing feature data are recorded as normal datasets, and those missing feature data are recorded as abnormal datasets. For a specific document data in an anomalous dataset, the feature with missing features is used as the target quantity, and the feature with not missing features is used as the variable. For each document data in the normal dataset, a regression model is built using the feature data of the target quantity as output and the feature data of the variable as input, and the trained regression model is obtained. In the abnormal dataset, the feature data of the document is substituted into the regression model after training to calculate the value, and the calculated value fills the missing feature data.
5. The vascular data classification and recognition method based on a large language model and clustering algorithm as described in claim 4, characterized in that, Construct the matrix: Take each of the N1 document data as a row and leave each of the A2 features as a column to construct an N1-row A2-column matrix; Row swapping: Swap rows with no missing feature data to the top row of the matrix, swap rows with missing feature data to the bottom row of the matrix, and arrange rows with missing the same feature data into adjacent rows; For the top row among all rows that lack the same feature data, the feature that lacks the feature data is used as the target quantity, and the feature that does not lack the feature data is used as the variable. For each row that does not lack feature data, a training regression model is constructed using the feature data of the target quantity as output and the feature data of the variable as input, and the trained regression model is obtained. For the top row of each row that is missing the same feature data, substitute the feature data of the variable into the trained regression model to calculate the value, and fill the missing feature data with the calculated value. For each row lacking the same feature data (excluding the top row), substitute the variable's feature data into the trained regression model to calculate the value, and fill the missing feature data in that row with the calculated value.
6. The method for classifying and recognizing vascular data based on a large language model and clustering algorithm as described in claim 1, characterized in that, In S1, the large language model uses vascular lesions and related keywords to find N literatures related to vascular lesions, where the vascular lesions are either veins or arteries. In S6, each cluster is represented by a number, with the cluster of non-vascular lesions represented by 0.
7. A vascular data classification and recognition system based on a large language model and clustering algorithm, characterized in that, include: A data extraction module is used to find N vascular lesion-related literature using a large language model. For any one of the N vascular lesion-related literature, if the literature contains at least one type of vascular lesion, the large language model is used to extract each type of vascular lesion and its corresponding different features and data as a corresponding literature data. If the literature does not contain any vascular lesion type, the large language model is used to extract different features related to vascular lesions and their data as a corresponding literature data. A total of N' literature data are obtained, where N'≥N. From the N' literature data, A different features are selected and sorted from high to low frequency. A1 features with too low frequency are discarded, and A2 features are kept, where A=A1+A2. The occurrence of a feature in a literature data is counted as one occurrence. A data filtering module is used to filter out data from N' literature data that contains all of the A2 features or only lacks one of the A2 features or one feature from the A2 features. The number of filtered data is denoted as N1. A data imputation module is used to construct a regression model using the literature data that is not missing to fill in the missing feature data for N1 literature data that only lacks features or data. A weight calculation module is used to calculate the weight value of each of the A2 features by using the occurrence frequency of the A2 features in N' literature data; A data acquisition module is used to acquire data of each feature in the A2 feature of each blood vessel in M clinical patients, and obtain M clinical data. The M clinical patients include patients with various types of vascular lesions without vascular lesions. Active type labeling is performed on M1 of the clinical data, and the labels are marked as various types of vascular lesions. M = M1 + M2, M1 < M2. A data clustering module is used to cluster the weighted K-means clustering model using the labeled M1 clinical data and N1 literature data with vascular lesion types and the weight values of each feature, and the clusters are divided into clusters of various vascular lesion types. A data annotation module is used to input the unannotated M2 clinical data and N1 literature data without vascular lesion types into a weighted K-means clustering model to obtain each cluster result, and to annotate the M2 clinical data and N1 literature data with type based on the clustering results.
8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in memory to execute the method of any one of claims 1-6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When computer program instructions are executed by a processor, they implement the method of any one of claims 1-6.
Citation Information
Patent Citations
Method and device for classifying documents
CN101853250A
Medical literature classification model training method and device, and medical literature classification method and device
CN108959236A
Cited By
Cardiovascular early warning method and system based on millimeter wave radar and physiological signal decoupling
CN121910352A