Lymphangioleiomyomatosis misdiagnosis risk assessment system and method based on big data analysis

Through a system based on big data analysis, machine learning algorithms are used to analyze the data of patients with lymphangiofibroid disease, and the risk of misdiagnosis is automatically evaluated, which solves the problem of high misdiagnosis rate in the existing technology, and achieves more accurate and objective diagnostic assistance.

CN120183687APending Publication Date: 2025-06-20梁笔帆
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510239551.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Due to the incomplete clarity of the pathogenesis of lymphangiofibroid disease and the complex clinical manifestations, the incidence of misdiagnosis and misdiagnosis is high. The existing technology relies on experts' experience and judgment, which is subjective and lacks objectivity and accuracy.

Method used

A system based on big data analysis is adopted to analyze patient data through data acquisition, preprocessing, analysis and visualization modules, machine learning algorithms are used to analyze patient data, automatically evaluate misdiagnosis risks, and display results in intuitive graphical form.

Benefits of technology

It realizes automatic identification of the risk of misdiagnosis in patients and assists doctors in making correct diagnostic decisions, improves the objectivity and accuracy of diagnosis, and reduces the incidence of misdiagnosis and misdiagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183687A_ABST
    Figure CN120183687A_ABST
Patent Text Reader

Abstract

The invention discloses a lymphangioleiomyomatosis misdiagnosis risk assessment system based on big data analysis, and the system comprises a data collection module, a data preprocessing module connected with the data collection module, a big data analysis module connected with the data preprocessing module, and a misdiagnosis risk assessment module connected with the big data analysis module. And the visualization module is connected with the misdiagnosis risk assessment module. The method has the advantages that the misdiagnosis risk of the patient can be automatically recognized, and therefore a doctor is assisted in making a correct diagnosis decision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to medical diagnosis technology, and particularly to a misdiagnosis risk assessment system and method for lymphangioleiomyomatosis based on big data analysis. Background Art

[0002] Lymphangioleiomyomatosis (LAM), also known as lymphangioleiomyomatosis, lymphangiosarcoma, lymphangiomyomatous hyperplasia, and diffuse benign interstitial lymphangioleiomyomatosis, is a rare multi-system disease that mainly affects the lungs and uterus of women, with less frequent involvement of other organs. In recent years, the incidence has shown an upward trend and a tendency towards younger age, and there is currently no radical cure. Since the pathogenesis of this disease has not been fully elucidated and the clinical manifestations are complex and diverse, it is easy to be confused with many other diseases, resulting in a relatively high incidence of misdiagnosis and missed diagnosis, which seriously affects the treatment effect and prognosis. Therefore, establishing an effective misdiagnosis risk assessment system is crucial for early detection and correct diagnosis of LAM.

[0003] In related technologies, the misdiagnosis risk of patients is usually determined by the experience judgment of experts. This method is highly subjective, lacks objectivity and accuracy, and cannot meet the requirements of practical applications. Summary of the Invention

[0004] In view of the above problems, the purpose of the present invention is to provide a misdiagnosis risk assessment system and method for lymphangioleiomyomatosis based on big data analysis, which can automatically identify the misdiagnosis risk of patients and thus assist doctors in making correct diagnostic decisions.

[0005] To achieve the above purpose, the present invention provides a misdiagnosis risk assessment system for lymphangioleiomyomatosis based on big data analysis. The system includes a data acquisition module, a data preprocessing module connected to the data acquisition module, a big data analysis module connected to the data preprocessing module, a misdiagnosis risk assessment module respectively connected to the big data analysis module, and a visualization module connected to the misdiagnosis risk assessment module. The data acquisition module is used to collect data related to lymphangioleiomyomatosis from multiple channels. The data preprocessing module is used to clean, standardize, and extract features from the collected data. The big data analysis module is used to analyze the preprocessed data using machine learning algorithms. The misdiagnosis risk assessment module is used to input the data of new patients into the trained model, and the model will output the misdiagnosis risk probability of the patient. The visualization module is used to display the misdiagnosis risk assessment results in an intuitive chart form.

[0006] In some embodiments, the data acquisition module includes the patient's basic information in the hospital information system (HIS), the clinical symptom descriptions, laboratory test results, imaging examination results, and pathological examination reports in the electronic medical record (EMR). The above are not limited to the patient's basic information in the hospital information system (HIS), the clinical symptom descriptions, laboratory test results, imaging examination results (such as CT, MRI images and reports), pathological examination reports, etc.;

[0007] In some embodiments, there is also an update module separately connected to the big data analysis module; the update module is used to monitor the latest medical literature in real time, extract new knowledge and new evidence related to lymphangioleiomyomatosis, and update the data and models in the system.

[0008] In some embodiments, a training set and a use test set are separately connected to the big data analysis module. The training set is used to train the model so that the model learns the mapping relationship from input features to misdiagnosis. The use test set is used to evaluate the trained model, continuously adjust the parameters of the model, and improve the accuracy and generalization ability of the model.

[0009] In some embodiments, the data preprocessing module is used to extract features from the imaging data, convert information such as the size, quantity, and distribution of pulmonary cystic lesions in the image into quantifiable feature vectors; and perform encoding processing on the clinical symptoms, converting symptoms such as "dyspnea" and "cough" into corresponding numerical values or category codes; then, perform standardization processing on the laboratory test results and convert data with different units and reference ranges.

[0010] In some embodiments, the big data analysis module uses algorithms such as support vector machine (SVM), random forest (RF), or deep learning neural network (such as convolutional neural network CNN combined with long short-term memory network LSTM) to analyze the preprocessed data.

[0011] In some embodiments, the misdiagnosis risk assessment module is provided with an input of new patient data and an output of a misdiagnosis risk probability value; the output misdiagnosis risk probability value is between 0 and 1. When inputting the data of a patient with specific clinical symptoms, laboratory test results, and imaging manifestations, the model will output a value between 0 and 1, representing the possibility of misdiagnosis of the patient. For example, 0.2 means that the patient has a 20% possibility of being misdiagnosed.

[0012] Another object is to provide a method for assessing the misdiagnosis risk of lymphangioleiomyomatosis based on big data analysis, which is characterized in that the specific steps are as follows:

[0013] (1) Collect data related to lymphangioleiomyomatosis from multiple channels, including but not limited to the basic information of patients in the hospital information system (HIS), the description of clinical symptoms, laboratory test results, imaging examination results (such as CT, MRI images and reports), and pathological examination reports in the electronic medical record (EMR);

[0014] (2) Clean, standardize and extract features from the collected data;

[0015] (3) Use machine learning algorithms to analyze the preprocessed data;

[0016] (4) Input the data of new patients into the trained model, and the model will output the misdiagnosis risk probability of the patient;

[0017] (5) Display the misdiagnosis risk assessment results in an intuitive chart form.

[0018] The beneficial effect of the present invention is that by collecting, analyzing, and using machine learning on lymphangioleiomyomatosis data, outputting a trained model, and finally generating a visual evaluation result content, it realizes the automatic identification of the misdiagnosis risk of patients, thereby assisting doctors in making correct diagnostic decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is the block schematic diagram of the system of the present invention;

[0020] Figure 2 is the block schematic diagram of the data preprocessing module in the system of the present invention;

[0021] Figure 3 is the block schematic diagram of the big data analysis module in the system of the present invention.

[0022] Figure 4 is the block schematic diagram of the misdiagnosis risk assessment module in the system of the present invention.

[0023] Figure 5 is the block schematic diagram of the visualization module in the system of the present invention.

[0024] Figure 6 is the block flow chart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0025] The invention will be further described in detail below with reference to the accompanying drawings.

[0026] As Figures 1 to 5As shown, a misdiagnosis risk assessment system and method for lymphangioleiomyomatosis based on big data analysis. The system includes: a data collection module 101, which is used to collect data related to lymphangioleiomyomatosis from multiple channels, including but not limited to the basic information of patients in the hospital information system (HIS), the description of clinical symptoms in the electronic medical record (EMR), laboratory test results, imaging examination results (such as CT, MRI images and reports), pathological examination reports, etc.; a data preprocessing module 102, which is used to clean, standardize and extract features from the collected data; a big data analysis module 103, which is used to analyze the preprocessed data by using machine learning algorithms; a misdiagnosis risk assessment module 104, which is used to input the data of new patients into the trained model, and the model will output the misdiagnosis risk probability of the patient; a visualization module 105, which is used to display the misdiagnosis risk assessment results in an intuitive chart form.

[0027] The functions of the above - mentioned modules are as follows: The data acquisition module 101 is used to collect data related to lymphangioleiomyomatosis from multiple channels, including but not limited to the basic patient information in the hospital information system (HIS), the clinical symptom descriptions, laboratory test results, imaging examination results (such as CT, MRI images and reports), pathological examination reports, etc. in the electronic medical record (EMR). Among them, HIS is an integrated management system composed of computer hardware and software. It completes the management of patient information resources through functions such as input, storage, processing, sorting, retrieval, transmission, and printing of various information generated during the patient's diagnosis and treatment process. EMR refers to the digital information in various forms and carriers such as text, symbols, pictures, images, audio, and video generated by medical institutions in their daily work, reflecting the patient's health status, the process of medical activities, and their results. In this way, valuable information can be obtained from a large number of historical cases to help us better understand the clinical characteristics and diagnostic difficulties of LAM. The data pre - processing module 102 is used to clean, standardize, and extract features from the collected data. Among them, data cleaning is mainly to remove duplicate, missing, incorrect, or abnormal data to ensure the quality and reliability of the data. Data standardization is to eliminate the differences between data from different sources, formats, or units, making it unified and standardized. For example, different hospitals may use different inspection instruments and techniques, resulting in large deviations in the results of the same indicator. Through standardization processing, they can be converted into the same unit and reference range for easy comparison and analysis. Feature extraction is to extract useful features from the original data for subsequent analysis and modeling. In the misdiagnosis risk assessment of LAM, many factors need to be considered, such as age, gender, smoking history, family history, clinical manifestations, laboratory tests, imaging findings, etc. There may be complex interactions and influences among these factors, which are difficult to directly use as inputs for the model. Through feature extraction, they can be transformed into a series of numerical values or variables, facilitating the processing of machine learning algorithms. The big data analysis module 103 is used to analyze the pre - processed data using machine learning algorithms. Among them, machine learning is a research direction in the field of artificial intelligence, aiming to enable computers to have the ability of self - learning and improvement. It can be applied to many fields, including natural language processing, image recognition, recommendation systems, bioinformatics, etc. In the misdiagnosis risk assessment of LAM, we can use supervised learning or unsupervised learning methods to construct prediction models or clustering models.

[0028] Specifically, it can be achieved in one of the following ways: Divide the dataset into a training set and a test set. The training set is used to train the model so that the model learns the mapping relationship from input features to misdiagnosis or not. Use the test set to evaluate the model and continuously adjust the model's parameters to improve the model's accuracy and generalization ability. This is one of the most commonly used machine learning methods and also the basic idea of supervised learning. First, we need to prepare a set of data with known labels (i.e., misdiagnosis or not) as the training set, and then select appropriate algorithms and hyperparameters to train a classifier or regressor so that its prediction results are as close as possible to the true values. Then, another set of unseen data can be used as the test set to evaluate the performance and stability of the model. Finally, the model architecture and parameters can be adjusted according to the actual situation to optimize its performance. Algorithms such as Support Vector Machine (SVM), Random Forest (RF), or deep learning neural networks (such as Convolutional Neural Network CNN combined with Long Short-Term Memory Network LSTM) are used to analyze the preprocessed data. SVM is a binary classification algorithm that attempts to find an optimal hyperplane to separate positive and negative samples. RF is an ensemble learning method that obtains the final prediction result by constructing multiple decision trees and taking the average vote. CNN is a feedforward neural network commonly used in image processing tasks. It can capture local features and gradually form global representations. LSTM is a recurrent neural network mainly used for learning sequence data. It can retain long-term dependencies and avoid the problem of gradient vanishing. Each of the above algorithms has its own advantages and disadvantages and is suitable for different types of problems and scenarios. The misdiagnosis risk assessment module 104 is used to input the data of a new patient into the trained model, and the model will output the misdiagnosis risk probability of this patient. When inputting the data of a patient with specific clinical symptoms, laboratory test results, and imaging findings, the model will output a value between 0 and 1, representing the possibility of this patient being misdiagnosed. For example, 0.2 means that this patient has a 20% possibility of being misdiagnosed. The visualization module 105 is used to display the misdiagnosis risk assessment results in an intuitive chart form. This can help doctors quickly understand and interpret the output of the model and also enhance the user's interaction experience and satisfaction. Common visualization tools include tables, line charts, bar charts, pie charts, scatter plots, heat maps, and so on. The update module 106 is used to monitor the latest medical literature in real time and extract new knowledge and new evidence related to lymphangioleiomyomatosis to update the data and model in the system. The big data analysis module 103 is specifically used for: Dividing the dataset into a training set and a test set. The training set is used to train the model so that the model learns the mapping relationship from input features to misdiagnosis or not. Using the test set to evaluate the model and continuously adjust the model's parameters to improve the model's accuracy and generalization ability. The data preprocessing module 102 is also used for: Extracting features from the imaging data and converting information such as the size, quantity, and distribution of pulmonary cystic lesions in the image into quantifiable feature vectors;

[0029] Encode clinical symptoms, converting symptoms such as "dyspnea" and "cough" into corresponding numerical or categorical codes; standardize laboratory test results, converting data with different units and reference ranges. The big data analysis module 103 is specifically used for: analyzing the preprocessed data using algorithms such as support vector machine (SVM), random forest (RF), or deep learning neural network (such as convolutional neural network CNN combined with long short-term memory network LSTM). The misdiagnosis risk assessment module 104 is also used for: when inputting the data of a patient with specific clinical symptoms, laboratory test results, and imaging findings, the model will output a value between 0 and 1, representing the possibility of misdiagnosis of the patient. For example, 0.2 means that the patient has a 20% possibility of being misdiagnosed.

[0030] Such as Figure 6 As shown, a method for assessing the risk of misdiagnosis of lymphangioleiomyomatosis based on big data analysis, the specific steps are as follows: (1) Collect data related to lymphangioleiomyomatosis from multiple channels, including but not limited to the basic information of patients in the hospital information system (HIS), the description of clinical symptoms in the electronic medical record (EMR), laboratory test results, imaging test results (such as CT, MRI images and reports), pathological examination reports, etc.; (2) Clean, standardize, and extract features from the collected data; (3) Use machine learning algorithms to analyze the preprocessed data; (4) Input the data of a new patient into the trained model, and the model will output the misdiagnosis risk probability of the patient; (5) Display the misdiagnosis risk assessment results in an intuitive chart form. The above steps continuously monitor the latest medical literature and extract new knowledge and new evidence related to lymphangioleiomyomatosis to update the data and models in the system. Using machine learning algorithms to analyze the preprocessed data includes: dividing the data set into a training set and a test set. The training set is used to train the model so that the model learns the mapping relationship from input features to misdiagnosis or not; using the test set to evaluate the model and continuously adjusting the parameters of the model to improve the accuracy and generalization ability of the model. Cleaning, standardizing, and extracting features from the collected data includes: extracting features from imaging data, converting information such as the size, quantity, and distribution of pulmonary cystic lesions in the image into quantifiable feature vectors; encoding clinical symptoms, converting symptoms such as "dyspnea" and "cough" into corresponding numerical or categorical codes;

[0031] Standardize laboratory test results, converting data with different units and reference ranges.

[0032] The above are only some embodiments of the present invention. For those of ordinary skill in the art, without departing from the inventive concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the invention.

Claims

1. A risk assessment system for misdiagnosis of lymphangioleiomyomatosis based on big data analysis, characterized in that ,The system includes a data acquisition module, a data preprocessing module connected to the data acquisition module, a big data analysis module connected to the data preprocessing module, a misdiagnosis risk assessment module respectively connected to the big data analysis module, and a visualization module connected to the misdiagnosis risk assessment module; The data collection module is used to collect data related to lymphangioleiomyomatosis from multiple channels; The data preprocessing module is used to clean, standardize and extract features from the collected data; The big data analysis module is used to analyze the pre-processed data using a machine learning algorithm; The misdiagnosis risk assessment module is used to input the data of a new patient into the trained model, and the model will output the misdiagnosis risk probability of the patient; The visualization module is used to display the misdiagnosis risk assessment results in the form of intuitive charts.

2. According to claim 1, a lymphangioleiomyomatosis misdiagnosis risk assessment system based on big data analysis is characterized in that ,The data collection module includes the basic information of patients in the hospital information system HIS, ,the clinical symptom description, laboratory test results, imaging test results, and pathological examination ,reports in the electronic medical record EMR.

3. A lymphangioleiomyomatosis misdiagnosis risk assessment system based on big data analysis according to claim 1 or 2, characterized in that , also includes an update module connected to the big data analysis module respectively; The updating module is used to monitor the latest medical literature in real time, extract new knowledge and new evidence related to lymphangioleiomyomatosis, and update the data and models in the system.

4. The risk assessment system for misdiagnosis of lymphangioleiomyomatosis based on big data analysis according to claim 1 is characterized in that ,The big data analysis module is respectively connected with a training set and a test set; The training set is used to train the model so that the model learns the mapping relationship from input features to whether a misdiagnosis occurs; The test set is used to evaluate the training model, continuously adjust the parameters of the model, and improve the accuracy and generalization ability of the model.

5. The interpretable multimodal information system for postoperative chronic pain assessment according to claim 1, characterized in that: The data preprocessing module is used to extract features from imaging data, converting information such as the size, number, and distribution of lung cystic lesions in the image into quantifiable feature vectors; and to encode clinical symptoms, converting symptoms into corresponding numerical values ​​or category codes; then, standardizing laboratory test results and converting data in different units and reference ranges.

6. The interpretable multimodal information system for postoperative chronic pain assessment according to claim 1, characterized in that: The big data analysis module uses support vector machine SVM, random forest RF or deep learning neural network algorithm to analyze the preprocessed data.

7. The interpretable multimodal information system for postoperative chronic pain assessment according to claim 1, characterized in that: The misdiagnosis risk assessment module is provided with input of new patient data and output of misdiagnosis risk probability value; the output misdiagnosis risk probability value is between 0 and 1.

8. A method for assessing the risk of misdiagnosis of lymphangioleiomyomatosis based on big data analysis, characterized in that , the specific steps are as follows: (1) Collect data related to lymphangioleiomyomatosis from multiple channels, including but not limited to basic patient information in the hospital information system (HIS), clinical symptom descriptions in the electronic medical record (EMR), laboratory test results, imaging test results (such as CT, MRI images and reports), and pathological examination reports; (2) Clean, standardize and extract features from the collected data; (3) Use machine learning algorithms to analyze the preprocessed data; (4) Input the new patient's data into the trained model, and the model will output the patient's misdiagnosis risk probability; (5) The misdiagnosis risk assessment results are presented in an intuitive graphical form.