Chest surgery examination data arrangement system based on related data characteristics
By introducing data acquisition, classification empowerment, related load, labeling addition and data adjustment modules into the thoracic surgery examination data sorting system, the problem of poor data relationship analysis in the existing system is solved, high-precision processing and analysis of data is achieved, and the efficiency of the system and the accuracy of diagnosis and treatment decisions are improved.
Patent Information
- Application Number
- CN202510181921.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing thoracic surgery data sorting system based on relevant data characteristics cannot effectively analyze the relationship between the data, resulting in low data accuracy and unsatisfactory use efficiency.
A system including data acquisition, classification empowerment, related load, labeling addition and data adjustment modules were designed. Through these modules, the thoracic surgery examination data were cleaned, classified, empowered, correlation analysis and labeled to ensure the accuracy and consistency of the data.
Through meticulous data processing processes, the quality and availability of data are improved, the accuracy and relevance of data analysis are ensured, and resource allocation and diagnosis and treatment decisions are optimized.
Smart Images

Figure CN120108700A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data management, and in particular to a thoracic surgery examination data sorting system based on relevant data features. Background Art
[0002] Thoracic surgery examination is a series of examinations of the chest and the various organs it contains to evaluate the health of the organs and diagnose possible diseases. The compilation of thoracic surgery examination data is crucial to ensure that patients receive accurate and effective medical care. By systematically compiling and analyzing thoracic surgery examination data, doctors can accurately identify symptoms and causes, helping doctors better determine the nature of lung lesions or heart problems. Therefore, the compilation of thoracic surgery examination data not only plays a key role for clinicians in providing individualized treatment plans, but also has important significance for the efficiency and quality assurance of the entire medical system.
[0003] In the thoracic surgery examination data collation system, relevant data features refer to those attributes or indicators that are important for the analysis, diagnosis and treatment of thoracic surgery diseases. The features are described from multiple dimensions, including patients' personal information, clinical indicators, imaging features, laboratory test results, etc., and data features are acquired, classified, weighted, analyzed and labeled in the thoracic surgery examination data collation system so that doctors and researchers can better understand disease patterns, evaluate treatment effects and develop personalized treatment plans.
[0004] However, when the existing thoracic surgery examination data collation system based on relevant data features is used, the relationship between the data is not analyzed, resulting in the inability to guarantee the accuracy of the data when the thoracic surgery examination data collation system is used. In addition, when the existing thoracic surgery examination data collation system based on relevant data features is used, the load generated by the data is not considered, resulting in the inability of the thoracic surgery examination data collation system based on relevant data features to clarify the potential connection between different data, resulting in the unsatisfactory use efficiency of the thoracic surgery examination data collation system based on relevant data features.
[0005] Currently, no effective solution has been proposed for the problems in the related technologies. Summary of the invention
[0006] In view of the problems in the related technology, the present invention proposes a thoracic surgery examination data sorting system based on relevant data features to overcome the above-mentioned technical problems existing in the existing related technology.
[0007] To this end, the specific technical solution adopted by the present invention is as follows: A thoracic surgery examination data sorting system based on relevant data features, the thoracic surgery examination data sorting system based on relevant data features comprises: a data acquisition module, a classification weighting module, a relevant load module, a labeling adding module and a data adjustment module; The data acquisition module is used to acquire thoracic surgery examination data and clean the thoracic surgery examination data; A classification and weighting module is used to classify the thoracic surgery examination data after data cleaning, and to weight the thoracic surgery examination data based on the classification results; A correlation load module is used to perform correlation analysis on the weighted results of thoracic surgery examination data and calculate the thoracic surgery examination data load value based on the correlation analysis results; An annotation adding module is used to add annotations to the thoracic surgery examination data according to the thoracic surgery examination data classification results, thoracic surgery examination data weighting results, correlation analysis results and thoracic surgery examination data load values; A data adjustment module, used for arranging and storing thoracic surgery examination data based on the annotations added to the thoracic surgery examination data; The data acquisition module, the classification weighting module, the related load module, the annotation adding module and the data adjustment module are connected in sequence.
[0008] As a preferred solution, the classification weighting module includes: a data processing module, a feature extraction module, a feature classification module, a weight allocation module and an optimization verification module; Among them, the data processing module is used for standardizing the thoracic surgery examination data; A feature extraction module is used to preset feature extraction rules and extract features from the standardized thoracic surgery examination data based on the feature extraction rules; A feature classification module is used to classify thoracic surgery examination data according to feature extraction results; A weight allocation module is used to set classification weight rules and assign weights to the classification results of thoracic surgery examination data based on the classification weight rules; An optimization verification module is used to verify the classification and weighting results of thoracic surgery examination data, and optimize the classification and weighting results of thoracic surgery examination data based on the verification results; The data processing module, feature extraction module, feature classification module, weight allocation module and optimization verification module are connected in sequence.
[0009] As a preferred solution, setting classification weight rules and assigning weights to the classification results of thoracic surgery examination data based on the classification weight rules include: Set the factor proportion weights and time sequence weights in the thoracic surgery examination data, and combine the factor proportion weights with the time sequence weights to obtain the classification weight rules; Extract data proportion factors and data time information from surgical examination data classification results; Based on the classification weight rules, weight the data proportion factor and data time information, and calculate the weight value of the data proportion factor and the weight value of the data time information; Summarize the weight values of the data proportion factors and the weight values of the data time information, and verify the summary results.
[0010] As a preferred solution, extracting data proportion factors and data time information from surgical examination data classification results includes: Set data segmentation rules, and segment the surgical examination data classification results by data proportion based on the data segmentation rules; Calculate each data proportion segmentation value based on the proportion segmentation results, and sort the proportion segmentation results according to the data proportion segmentation values; Set timestamp extraction rules, extract data acquisition time information in surgical examination data classification results based on the timestamp extraction rules, and unify the format of the data acquisition time information.
[0011] As a preferred solution, the relevant load module includes: a data related module, a data load module, a data adjustment module and a relevant impact module; Among them, the data correlation module is used to evaluate the correlation value between each data weighted result in the thoracic surgery examination data weighted result; A data load module is used to extract the load characteristics of thoracic surgery examination data and calculate the load value of the test data based on the load characteristics; A data adjustment module, used to optimize and adjust the detection data load value according to the correlation value between the data weighting results; A related impact module is used to perform load impact analysis based on the optimized and adjusted detection data load value, and output the optimized and adjusted detection data load value and load impact analysis; The data related module, the data load module, the data adjustment module and the related impact modules are connected in sequence.
[0012] As a preferred solution, the data related module includes: a related linkage module, a related calculation module and a related evaluation module; Among them, the relevant linkage module is used to preset data linkage rules and add relevant annotations between data weighting results based on the data linkage rules; A correlation calculation module is used to calculate the correlation value between the data weighting results according to the relevant annotations; A correlation evaluation module is used to classify the correlation between the data weighting results according to the correlation value and add labels based on the classification results; The related linkage modules, the related calculation modules and the related evaluation modules are connected in sequence.
[0013] As a preferred solution, the calculation formula for the correlation value between the weighted data results calculated based on the relevant annotations is: ; in, r Assign correlation values between data weight results; n is the total number of data weighting results; The ranking difference between the results of weighting the data; i is the index value between the data weighting results.
[0014] As a preferred solution, extracting the load characteristics of thoracic surgery examination data, and calculating the load value of the test data based on the load characteristics includes: Setting inspection data load rules, and extracting detection load characteristic values and data load characteristic values of thoracic surgery inspection data based on the inspection data load rules; Construct a load model, and substitute the detected load characteristic value into the load model to calculate the detected data load value; The detection load value and the detection data load value are summarized, and the summary results are verified, and the summary results of the verified detection data load value and data load characteristic value are output.
[0015] As a preferred solution, a load model is constructed, and the detected load characteristic value is substituted into the load model to calculate the detected data load value, including: Perform data cleaning on the detected load characteristic values, and normalize the load characteristic values after data cleaning; Setting load matching rules and a load model library, and matching the normalized load characteristic value with the load model library based on the load matching rules; The normalized load characteristic values are divided into a training set and a test set, the matching load model is trained by the training set, and the trained load model is tested by the test set; Substitute the load characteristic value into the load model after the test, calculate the test data load value, and cross-validate the test data load value.
[0016] As a preferred solution, the annotation adding module includes: a data analysis module, a result annotation module, an annotation verification module and an annotation comprehensive module; Among them, the data analysis module is used to analyze the data characteristic factors of thoracic surgery examination data classification results, thoracic surgery examination data weighting results, correlation analysis results and thoracic surgery examination data load value; A result annotation module, used to generate data factor annotations based on data characteristic factors; The annotation verification module is used to compare and verify the data factor annotation and analysis of the thoracic surgery examination data classification results, thoracic surgery examination data weighting results, correlation analysis results and thoracic surgery examination data load values; The annotation comprehensive module is used to summarize the annotations of the data factors after comparison and verification, and to add annotations to the thoracic surgery examination data based on the data factor annotation summary results.
[0017] The beneficial effects of the present invention are: 1. The present invention divides data processing into data acquisition, classification and weighting, related load, annotation addition and data adjustment modules, processes complex medical data in a meticulous and orderly manner, optimizes the data processing flow, ensures the accuracy and consistency of the data, and removes inconsistent and erroneous data through data cleaning, improves the quality and availability of data, and ensures the accuracy and relevance of data analysis by accurately classifying and weighting the data.
[0018] 2. The present invention conducts correlation analysis, calculates the data load value based on the analysis results, understands the data characteristics and internal correlations, optimizes resource allocation and improves the key to diagnosis and treatment decisions, builds a load model and calculates the test data load value, deeply understands the load characteristics of the data, and calculates the data load value through correlation analysis to identify the potential connections between different data.
[0019] 3. The present invention uses a weight distribution module to distribute weights to data according to different factors and time sequence, thereby increasing the depth and sophistication of data analysis, and in the relevant load module, performs data cleaning and normalization on the load characteristic values, thereby improving the quality and consistency of data analysis, identifying and extracting key features, and classifying the data accordingly, thereby achieving accurate understanding and analysis of the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0021] Figure 1 The system block diagram of a thoracic surgery examination data sorting system based on relevant data features according to an embodiment of the present invention.
[0022] In the figure: 1. Data acquisition module; 2. Classification weighting module; 3. Related load module; 4. Label adding module; 5. Data adjustment module. DETAILED DESCRIPTION
[0023] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0024] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0025] According to an embodiment of the present invention, a thoracic surgery examination data sorting system based on relevant data features is provided.
[0026] The present invention will now be further described with reference to the accompanying drawings and specific embodiments. Figure 1 As shown, according to a thoracic surgery examination data sorting system based on relevant data features according to an embodiment of the present invention, the thoracic surgery examination data sorting system based on relevant data features includes: a data acquisition module 1, a classification weighting module 2, a relevant load module 3, a labeling adding module 4 and a data adjustment module 5; The data acquisition module 1 is used to acquire thoracic surgery examination data and clean the thoracic surgery examination data; Specifically, determine the source of thoracic surgery examination data, which may include hospital information systems, electronic medical records, radiology information systems, laboratory information management systems, etc., and then extract the required data from the above systems through direct database queries, etc., and ensure that data extraction complies with relevant regulations on medical information security and privacy protection. Design the format and scope of data extraction based on the required data type and purpose. For example, for thoracic surgery examination data, include basic patient information, examination results, imaging data, biomarkers, etc.
[0027] Delete duplicate records, ensure that each record in the data set is unique, and handle missing values in the data, including data interpolation, deletion of missing data, or estimation of missing values using statistical methods. Identify possible outliers or erroneous data through statistical analysis methods, and correct or delete them. Standardize or normalize the data, for example, unify all date formats or standardize continuous variables to the same scale. Ensure that the type of each column of data meets expectations, such as numeric type, text type, etc., and ensure that the data remains intact during the extraction and cleaning process without losing important information. Perform data quality assessments regularly to ensure the accuracy and availability of the data during data cleaning and processing.
[0028] Classification and weighting module 2, used to classify the thoracic surgery examination data after data cleaning, and weight the thoracic surgery examination data based on the classification results; Specifically, the classification weighting module 2 includes: a data processing module, a feature extraction module, a feature classification module, a weight allocation module and an optimization verification module; Among them, the data processing module is used for standardizing the thoracic surgery examination data; Specifically, ensure that all date and time data is in the same format, and adjust all numeric data to the same measurement unit and range, such as converting all weight data to kilograms, and expand abbreviations in text or convert all text to the same uppercase and lowercase format.
[0029] Scale the data so that the final data range falls within a specific interval, correct obvious entry errors, such as entering a non-standard value in the gender column, use box plots or standard deviation methods to identify and handle outliers for numerical data, for example, replace data points that are out of the normal range with the median or mean, convert text category data to numerical values, perform mathematical and statistical analysis, such as creating a new binary column for each category, suitable for category data without a sequential relationship, and then merge data from different sources into a unified data set to handle key-value correspondence and data duplication issues.
[0030] A feature extraction module is used to preset feature extraction rules and extract features from the standardized thoracic surgery examination data based on the feature extraction rules; Specifically, the purpose and goal of feature extraction should be clarified, including understanding which features are most critical for the diagnosis, treatment or disease monitoring of thoracic surgery. For example, features that need attention include the patient's age, gender, pathological results, surgical records, radiological imaging features, etc., and determining which clinical parameters and biomarkers are most critical for disease diagnosis and treatment.
[0031] Conduct exploratory data analysis, such as correlation analysis and principal component analysis, to identify key features and hidden patterns in the data, use machine learning algorithms such as feature selection or feature importance assessment to identify the most effective features, and use programming languages and related libraries to extract features according to defined rules, and ensure that all extracted features follow a unified format and standard, such as standardizing or normalizing all continuous variables and appropriately encoding categorical variables, and test the effects of feature extraction rules on different data subsets to ensure that the features have good generalization capabilities, and continuously adjust and optimize feature extraction rules based on model performance and clinical feedback.
[0032] A feature classification module is used to classify thoracic surgery examination data according to feature extraction results; Specifically, evaluate the relationship between the extracted features and the target you want to classify. For example, if the classification goal is to predict the success rate of surgical outcomes, focus on those features that are significantly correlated with surgical outcomes, such as the patient's age, surgery type, pre-existing conditions, etc. According to the characteristics of the data and the classification goal, select an appropriate classification algorithm, including decision trees, random forests, support vectors, and logistic regression. Before classification, ensure that all selected features are properly preprocessed, such as standardization, normalization, and handling of missing values and outliers. Use the selected algorithm and preprocessed data to train the classification model, select a training set to train the model, and optimize the model parameters through cross-validation and other techniques, and use the test set to evaluate the performance of the model. The indicators usually concerned include accuracy, recall, F1 score, etc. Based on the feedback from these indicators, further adjust the model parameters or reselect features.
[0033] A weight allocation module is used to set classification weight rules and assign weights to the classification results of thoracic surgery examination data based on the classification weight rules; Specifically, setting classification weight rules and assigning weights to the classification results of thoracic surgery examination data based on the classification weight rules includes: Set the factor proportion weights and time sequence weights in the thoracic surgery examination data, and combine the factor proportion weights with the time sequence weights to obtain the classification weight rules; Specifically, data analysis techniques such as correlation analysis and regression analysis are used to determine the influence of each factor on the result, so as to assign corresponding weights and adopt models such as exponential decay or linear decay. For example, the weight decreases according to the number of months from the current time, and the closer the data is, the greater the weight. A time window is set, such as the last three months. The data in the window is given a higher weight, and the weight of the data outside the window gradually decreases. Combining these two weights, it is usually necessary to establish a comprehensive weight model. The final weight of each data point is determined by the factor proportion weight and the time sequence weight. A factor proportion weight is set for each factor, and then the weight is adjusted according to the time sequence of the data. For example, assuming that age and surgery type are two key factors, and the data in the last six months is more important, the following rules are set: the age weight is 0.3, the surgery type weight is 0.7, and the time weight is set to start from the current month, and the weight decreases by 10% for each month forward, that is, the weight of the last month is 1, the previous month is 0.9, and so on.
[0034] Extract data proportion factors and data time information from surgical examination data classification results; Specifically, the data proportion factors and data time information extracted from the surgical examination data classification results include: Set data segmentation rules, and segment the surgical examination data classification results by data proportion based on the data segmentation rules; Specifically, clarify the purpose of data segmentation, and determine which factors or attributes will be used to classify data in order to train machine learning models, conduct statistical analysis, or simply understand the proportion of data in each category, including the patient's age, gender, diagnosis type, treatment outcomes, etc. Classification is usually based on project requirements or data analysis goals. According to the analysis needs, set the data proportion target for each classification group. For example, if the goal is to ensure the representativeness of each age group in the data, the data of each age group should account for an equal proportion of the total data or be adjusted according to the actual population distribution. If the data needs to reflect a specific structure in the population, use a stratified sampling method to divide the population into different layers or subsets. Each layer represents a classification standard, such as age group, and then samples are drawn from each layer according to the set proportion.
[0035] If the classified data needs to be subjected to machine learning or statistical analysis, a random segmentation method is used to ensure that the samples in each category are randomly selected to reduce bias. In practical applications, the segmentation rules need to be adjusted based on the preliminary data analysis results. For example, if a category exhibits high variance or bias due to insufficient data, the data proportion rule for that category may need to be readjusted. After setting and implementing the data segmentation rules, regularly monitor and evaluate whether the segmentation effect meets project requirements. The evaluation is based on data coverage, model performance, or uniformity of data segmentation.
[0036] Calculate each data proportion segmentation value based on the proportion segmentation results, and sort the proportion segmentation results according to the data proportion segmentation values; Specifically, the data proportion is calculated from each segmented category, the data is counted for each segmented category, the total number of data items in each category is determined, and the total number of data items in the overall data set is calculated. For each category, its data proportion is obtained by dividing the number of data items in the category by the total data volume, and the proportion represents the proportion of the category in the entire data set. The proportion value is usually expressed as a percentage to intuitively understand the proportion of each category in the whole. For example, if a category has 50 data items and the total data volume is 1000, the proportion of the category is 5%.
[0037] Calculate the proportion of each category, sort the classification results according to these proportions, and then sort the categories from high to low or from low to high according to the proportion. In order to understand the data distribution more intuitively, use charts such as bar charts or pie charts to show the proportion ranking of different categories. The sorted data can be used for a variety of applications. For example, when resources are limited, give priority to categories with high proportions. In risk management, understand which categories have a high proportion and formulate targeted strategies. In market analysis, understand the proportion of different customer groups to optimize marketing strategies.
[0038] Set timestamp extraction rules, extract data acquisition time information in surgical examination data classification results based on the timestamp extraction rules, and unify the format of the data acquisition time information.
[0039] Specifically, it is necessary to clarify what time information is extracted from the data. In surgical examination data, this includes examination time, operation time, admission and discharge time, etc. Determining these details helps design effective extraction rules. According to the source and format of the data, rules are defined to identify and extract timestamps. For unstructured text data, regular expressions are used to match and extract standard date and time formats, such as YYYY-MM-DD or DD / MM / YYYY, to extract time information from the data set. During the extraction process, ensure that the rules or functions used can accurately match the time format in the target data to avoid incorrect extraction, and check whether the extracted time data is complete and confirm that no data is omitted or extracted incorrectly. Unifying the format of the extracted time information is an important step to ensure data consistency. Select a standard time format and convert all time data to this format.
[0040] Based on the classification weight rules, weight the data proportion factor and data time information, and calculate the weight value of the data proportion factor and the weight value of the data time information; Specifically, define specific weight rules based on the importance and time sensitivity of the data. For example, set weights based on the proportion of different categories in the overall data set, and categories with higher proportions get higher weights. Then set weights based on the temporal proximity of the data, and newer data get higher weights. Calculate weights based on the proportion of data in each category. For example, use the proportion directly as the weight or adjust the proportion value through a logarithmic function to make the weight distribution more reasonable. Set a decay function, such as exponential decay, so that the weight of the data gradually decreases over time. If the amount of data in each category is known, divide the amount of data in each category by the total amount of data to get the proportion, and then calculate the weight value according to the weighting rule. For each data record, calculate the weight based on the time difference between its timestamp and the current date. The time difference can be calculated in days, months or years. According to specific needs, the calculated weight value is applied to the data analysis or machine learning model, and the influence of each data point will be adjusted according to its weight value.
[0041] Summarize the weight values of the data proportion factors and the weight values of the data time information, and verify the summary results.
[0042] Specifically, determine how to summarize the weights of data proportion factors and data time information, for example, whether it is necessary to calculate the comprehensive weight value of the data points or generate summary statistics of the weight values; for the proportion weight of each category, perform a simple summation to obtain the overall proportion weight; and check the weight distribution of different categories to identify whether there are significant differences or deviations; summarize the time weights of all data points according to the calculation method of time weight, for example, use a weighted average to obtain the overall time weight; check the time distribution of the data to ensure that the distribution of time weights on the time axis is reasonable and consistent with expectations.
[0043] Combine the two weight values to get a comprehensive weight indicator, for example, use the geometric mean to ensure that the weights are not overly affected by extreme values when combined. After combining the weights, verifying the results is a key step to ensure the accuracy and effectiveness of data processing and weight combination, and to ensure that the combined weight values are within a reasonable range, such as between 0 and 1 or other appropriate ranges. At the same time, check whether the weight distribution meets expectations and ensure that there are no unreasonable deviations or outliers. Test the summarized weight results in actual applications, such as using them in classification models, to ensure that the model performance meets expectations.
[0044] An optimization verification module is used to verify the classification and weighting results of thoracic surgery examination data, and optimize the classification and weighting results of thoracic surgery examination data based on the verification results; Specifically, determine the key indicators for verifying the weighting results, including checking the accuracy of matching the classified data with the actual data, evaluating the internal consistency of data classification, ensuring that the same type of data is consistently classified, and using indicators such as recall, precision, and F1 score. Verify by applying the classification weighting results to a known test set or through cross-validation, including creating a confusion matrix to view details such as true positive examples and false positive examples of different classifications, and analyze the data collected during the verification process in detail to identify any significant patterns or anomalies, including checking for misclassification and determining whether there are systematic errors, such as certain specific types of data that are always misclassified, evaluating the specific impact of different weights on the classification results, and determining whether the weight settings need to be adjusted. According to the validation results, the weighting rules are optimized, and based on the error analysis results, the weight coefficients of different classifications and time information are adjusted to accurately reflect their importance, and whether it is necessary to introduce new weighting factors, such as patient prognostic indicators or specific treatment responses, to improve the accuracy and relevance of classification. The optimized weighting results are revalidated to ensure that the changes actually improve the classification effect. This process is repeated until satisfactory accuracy and effectiveness are achieved. The performance of weighted classification is continuously monitored in actual applications, and as time goes by or data accumulates, the weighting rules need to be regularly reviewed and adjusted to adapt to new trends or discoveries.
[0045] The data processing module, feature extraction module, feature classification module, weight allocation module and optimization verification module are connected in sequence.
[0046] The correlation load module 3 is used to perform correlation analysis on the weighted results of thoracic surgery examination data, and calculate the thoracic surgery examination data load value based on the correlation analysis results; Specifically, the related load module 3 includes: a data related module, a data load module, a data adjustment module and a related impact module; Among them, the data correlation module is used to evaluate the correlation value between each data weighted result in the thoracic surgery examination data weighted result; Specifically, the calculation formula for calculating the correlation value between the data weighting results according to the relevant annotations is: ; in, r Assign correlation values between data weight results; n is the total number of data weighting results; The ranking difference between the results of weighting the data; i is the index value between the data weighting results.
[0047] Specifically, the data-related modules include: a related linkage module, a related calculation module, and a related evaluation module; Among them, the relevant linkage module is used to preset data linkage rules and add relevant annotations between data weighting results based on the data linkage rules; Specifically, if the data comes from the same time period or meets a specific time series pattern, it will be automatically linked. The data will be linked based on shared attributes or features such as patient ID, treatment type, etc., and when a specific event occurs, for example, a patient receives a specific treatment, the data linkage is triggered. The system is set up to automatically identify and execute linkage rules, for example, using database triggers or data stream processing platforms to achieve this, and merge or associate data that meets the linkage conditions for comprehensive analysis. This may include merging data records, associating tables, etc.
[0048] In data presentations or reports, use visual or text annotations to emphasize key connections between data, such as using arrows, links, or highlights. Provide descriptive text or annotations to explain the relationship between data, such as explaining how a weighted result depends on the specific characteristics of another data set. It is crucial to verify the effectiveness and accuracy of data linkage rules. It is necessary to ensure that the linked data is correct and conforms to the expected logical relationship. Test the linkage rules on different data sets to check their accuracy and responsiveness. Collect feedback from users or data analysts on the effectiveness of data linkage and its annotations, make adjustments based on the feedback, and update the rules based on new data patterns or analysis needs. Explore new annotation techniques and tools to more effectively convey the relationship between data.
[0049] A correlation calculation module is used to calculate the correlation value between the data weighting results according to the relevant annotations; Specifically, all data points with relevant annotations are collected, including the weight value of each data point and the annotation information between them, such as the annotation type, such as positive correlation, negative correlation, irrelevant, etc. and strength, to measure the strength and direction of the linear relationship between the two variables, and organize the data according to the relevant annotations. For example, if there is a positive correlation annotation between data point A and data point B, their weight values should be analyzed together, and then the correlation between each pair of data points should be calculated using the selected correlation measurement method.
[0050] The obtained correlation values are analyzed to determine the strength of the relationship between the data points. A high positive correlation value indicates a strong positive correlation, and a high negative correlation value indicates a strong negative correlation. Based on the correlation results, the weighting or processing of the data is further adjusted. For example, if two variables are found to be highly correlated, it is decided to merge these variables or downgrade one of the variables to reduce redundancy. The results of the correlation analysis are verified to be in line with expectations, and necessary adjustments are made according to the needs of the actual application to ensure the accuracy and applicability of the results.
[0051] A correlation evaluation module is used to classify the correlation between the data weighting results according to the correlation value and add labels based on the classification results; Specifically, the standard or threshold for determining the correlation value classification is set based on the nature of the data and the analysis objectives. For example, the correlation values are divided into the following levels: strong correlation, the correlation value is greater than a certain high threshold, medium correlation, the correlation value is within the medium threshold range, weak correlation, the correlation value is within the low threshold range, and no correlation, the correlation value is lower than a set threshold. The correlation values between the data weighting results are calculated using the method described previously, and the correlation values are divided into different levels according to the set threshold.
[0052] Once the correlation between the data weighting results is graded, add corresponding labels for each level. For example, label A indicates strong correlation, label B indicates medium correlation, label C indicates weak correlation, and label D indicates no correlation. Then apply the calculated labels to the data weighting results. Consider this correlation information in the subsequent data analysis and decision-making process, verify whether the application of grading and labels is reasonable, and make adjustments according to the actual application needs.
[0053] The related linkage modules, the related calculation modules and the related evaluation modules are connected in sequence.
[0054] A data load module is used to extract the load characteristics of thoracic surgery examination data and calculate the load value of the test data based on the load characteristics; Specifically, extracting the load characteristics of thoracic surgery examination data and calculating the load value of the test data based on the load characteristics includes: Setting inspection data load rules, and extracting detection load characteristic values and data load characteristic values of thoracic surgery inspection data based on the inspection data load rules; Specifically, load rules are defined. These rules will determine the load level based on different characteristics of the data, including setting rules based on factors such as the type, difficulty, and time required for detection, and setting rules based on factors such as the amount of data, processing difficulty, and the number of data sources. Characteristic value indicators used to quantify the load are determined, including the number of detections, the type of detection, the detection difficulty coefficient, and the like, such as the amount of data, the number of data sources, and the data processing time. According to the defined rules, thoracic surgery examination data are classified and evaluated to determine their load characteristic values. According to the defined rules, the data are classified into different load levels, and the load characteristic values of each classification are calculated, such as by calculating the mean, median, or other statistics, and then characteristic values such as the detection type and difficulty coefficient are extracted from the detection data, and characteristic values such as the amount of data and the number of data sources are extracted from the data set. It is verified whether the extracted characteristic values meet the expected load rules and are adjusted according to the needs of the actual application.
[0055] Construct a load model, and substitute the detected load characteristic value into the load model to calculate the detected data load value; Specifically, constructing a load model and substituting the detected load characteristic value into the load model to calculate the detected data load value includes: Perform data cleaning on the detected load characteristic values, and normalize the load characteristic values after data cleaning; Specifically, identify and handle missing data, choose to fill missing values, for example, use the median or mean or delete records containing missing values as appropriate, identify outliers in the data through statistical methods such as box plots or standard deviation ranges, handle outliers according to specific circumstances, such as adjustment or deletion, and ensure that the data type of each feature is suitable for subsequent analysis, for example, convert dates from strings to date types, or convert numerical classification labels to integers. After the cleaning process, perform data validation to ensure data accuracy and consistency, and ensure that all data is logically consistent, such as checking whether the data range and data format comply with the defined business rules, checking and deleting duplicate records unless the duplication has specific business significance, and normalization is to standardize the data within a unified range. Setting load matching rules and a load model library, and matching the normalized load characteristic value with the load model library based on the load matching rules; Specifically, the matching model is determined by comparing the size of the load characteristic value with a preset threshold, and the matching model is determined by comparing the load characteristic value with a predefined load range. The load characteristic value is matched with the model in the model library according to its level, such as low, medium, and high. The load model library is a data set containing multiple load models, each model represents a specific load state or working mode, and ensures that the model library contains a wide range of load models covering different types of load characteristic values, verifies the validity and accuracy of each model, ensures that it can correctly reflect the load state, and regularly updates the model library to reflect new load characteristics or improved load models.
[0056] Use the established load matching rules to match the normalized load characteristic values with the models in the load model library, and use algorithms or automation tools to automatically match the characteristic values with the models in the model library according to the matching rules, verify the validity of the load matching results, and make adjustments based on the needs of actual applications. Use different test data sets to test the effectiveness of the matching rules, collect feedback from users or experts, and adjust the matching rules and model library based on the feedback.
[0057] The normalized load characteristic values are divided into a training set and a test set, the matching load model is trained by the training set, and the trained load model is tested by the test set; Substitute the load characteristic value into the load model after the test, calculate the test data load value, and cross-validate the test data load value.
[0058] The detection load value and the detection data load value are summarized, and the summary results are verified, and the summary results of the verified detection data load value and data load characteristic value are output.
[0059] A data adjustment module, used to optimize and adjust the detection data load value according to the correlation value between the data weighting results; Specifically, by optimizing the data load value, reducing unnecessary repeated detection or excessive load, and ensuring that the detection tasks are evenly distributed, avoiding excessive concentration of a certain part of the detection, reducing errors and deviations, and improving the reliability of the detection results, by calculating the correlation value between the weighted results, identifying which data have strong correlations, such as using statistical methods or data visualization to determine the correlation, and identifying data weighting results with higher correlation values. For example, if there is a high correlation between certain detection data, it may indicate that they share the same source or characteristics. For highly correlated data, consider merging or reducing repeated detections to reduce the load. If the importance of certain detection data changes due to correlation, consider adjusting its weight to ensure a more accurate load value. Reallocate detection tasks based on correlation to avoid over-concentration in a specific area or time period.
[0060] A related impact module is used to perform load impact analysis based on the optimized and adjusted detection data load value, and output the optimized and adjusted detection data load value and load impact analysis; Specifically, collect and organize the load values of the optimized and adjusted detection data to ensure that the data is up-to-date and accurately reflects all recent adjustments. Perform a detailed analysis of the adjusted load data to evaluate its impact on system performance, efficiency and reliability, such as comparing it with the data before optimization, evaluating the impact of load adjustment on detection time, processing speed and error rate, analyzing whether there are new or unresolved bottlenecks, and evaluating whether resource utilization has improved, such as hardware utilization, human resource allocation, etc. Use statistical methods and data visualization techniques to quantify the load impact. For example, use scatter plots, bar charts or line charts to display performance indicators before and after load changes. Based on the analysis results, write a detailed report, including a clear display of the data comparison before and after load adjustment, a detailed description of the specific impact of load adjustment on system performance and resource utilization, pointing out current problems and suggestions for further optimization, and outputting the optimized and adjusted detection data load values and load impact analysis results in an appropriate format for further review and decision-making. The output format can be a spreadsheet, database or interactive dashboard, depending on the user's needs and technical platform.
[0061] The data related module, the data load module, the data adjustment module and the related impact modules are connected in sequence.
[0062] Annotation adding module 4, used for adding annotations to the thoracic surgery examination data according to the thoracic surgery examination data classification results, thoracic surgery examination data weighting results, correlation analysis results and thoracic surgery examination data load values; Specifically, the annotation adding module 4 includes: a data analysis module, a result annotation module, an annotation verification module and an annotation synthesis module; Among them, the data analysis module is used to analyze the data characteristic factors of thoracic surgery examination data classification results, thoracic surgery examination data weighting results, correlation analysis results and thoracic surgery examination data load value; Specifically, use indicators such as classification accuracy, recall rate, F1 score, etc. to evaluate the performance of the classification model, use machine learning models to evaluate which features have the greatest impact on the classification results, analyze cases of classification errors, find possible reasons, such as data quality issues or insufficient features, study how the weights of different data points affect the final decision or model output, statistically analyze the distribution of weights, identify whether the weight distribution is uniform or whether there is over-concentration, and use methods such as Pearson correlation coefficient and Spearman rank correlation coefficient to measure the correlation between variables.
[0063] Visualize the correlation results through heat maps or scatter plots to intuitively display the relationship strength and pattern between variables, analyze how high correlation affects the diagnosis or treatment decisions of thoracic surgery examinations, review the calculation method of the load value to ensure that it reflects the actual data processing or detection complexity, analyze how the high and low load values affect the system's response time and processing capabilities, and based on the load value analysis results, put forward suggestions for system optimization or data processing process improvement.
[0064] A result annotation module, used to generate data factor annotations based on data characteristic factors; The annotation verification module is used to compare and verify the data factor annotation and analysis of the thoracic surgery examination data classification results, thoracic surgery examination data weighting results, correlation analysis results and thoracic surgery examination data load values; Specifically, clarify the standards or rules for data factor annotation, including data type, importance level, data source, etc., annotate each data point according to the defined standards, check the consistency between different analysis results, for example, whether the key features in the classification results have received corresponding attention in the weighting results, analyze how the annotated data factors affect the classification results, weighting results and correlation analysis results, verify whether the load value matches the data factor annotation, for example, whether high load values are associated with complex or important data features, use statistical tests to evaluate the significant impact of data factor annotation on the analysis results, and then use charts to visualize the relationship between data factor annotation and analysis results, evaluate the results of comparison and verification, determine the consistency and difference between data factor annotation and analysis results, and adjust and optimize the data factor annotation standards or analysis methods based on the evaluation results to improve the accuracy of the analysis.
[0065] The annotation comprehensive module is used to summarize the annotations of the data factors after comparison and verification, and to add annotations to the thoracic surgery examination data based on the data factor annotation summary results.
[0066] Specifically, all comparison and verification results are collected and organized, including the matching between data factor annotations and analysis results, inconsistencies, and any observed patterns or trends, to identify and emphasize key findings, especially those that have a significant impact on thoracic surgery examination data analysis and decision-making, and to identify any existing problems or challenges that may be caused by inaccurate data factor annotations or improper interpretation of analysis results.
[0067] Based on the summarized key findings and identified issues, update or modify the standards or rules for data factor annotation, create or update annotation guidelines, ensure that all relevant personnel can understand and follow the new or updated annotation standards, train the personnel responsible for annotation to ensure that they understand the new annotation standards and can apply them consistently, re-annotate the thoracic surgery examination dataset or add new annotations according to the updated annotation standards, test on the updated annotated data subset to verify whether the new annotation improves data quality and analysis accuracy, and use the same or updated evaluation indicators as before to measure the performance of the updated annotated data, update all relevant data documents and metadata to reflect the new annotation standards and any data changes, and record all information about the annotation update and verification process, including decision reasons, execution steps, and evaluation results.
[0068] A data adjustment module 5, used for arranging and storing the thoracic surgery examination data based on the annotations added to the thoracic surgery examination data; Specifically, determine how to organize data based on annotations, such as grouping by disease type, examination type, or patient characteristics, select a suitable storage format, ensure data accessibility and scalability, classify data into different groups or categories based on annotations, and sort data based on relevance, importance, or other criteria, select a suitable storage system, such as a local file system, cloud storage service, or database management system, import the organized data into the selected storage system, ensure data integrity and consistency, create metadata for each dataset or data file, and record data source, annotation information, data format, and other relevant information.
[0069] Ensure that data storage complies with relevant data protection regulations and standards, set up appropriate access controls to ensure that only authorized personnel can access sensitive data, create detailed documentation for the data collation and storage process, including operating steps, tools and technologies used, and record all changes to data collation and storage, including dates, executors and reasons for changes.
[0070] The data acquisition module 1, the classification weighting module 2, the related load module 3, the annotation adding module 4 and the data adjustment module 5 are connected in sequence.
[0071] To summarize, with the aid of the above-mentioned technical solutions of the present invention, the present invention divides data processing into data acquisition, classification empowerment, related load, annotation addition and data adjustment modules, so as to process complex medical data in a meticulous and orderly manner, optimize the data processing flow, ensure the accuracy and consistency of the data, and the data cleaning function removes inconsistent and erroneous data, improves the quality and availability of the data, and ensures the accuracy and relevance of data analysis by accurately classifying and weighting the data.
[0072] In addition, the present invention conducts correlation analysis, calculates the data load value through the analysis results, understands the data characteristics and internal correlations, optimizes resource allocation and improves the key to diagnosis and treatment decisions, builds a load model and calculates the test data load value, deeply understands the load characteristics of the data, and calculates the data load value through correlation analysis to identify the potential connections between different data.
[0073] In addition, the present invention uses a weight allocation module to allocate weights to data according to different factors and time sequences, thereby increasing the depth and sophistication of data analysis, and in related load modules, performs data cleaning and normalization on load characteristic values, thereby improving the quality and consistency of data analysis, identifying and extracting key features, and classifying the data accordingly, thereby achieving accurate understanding and analysis of the data.
[0074] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A thoracic surgery examination data sorting system based on relevant data features, characterized in that: The thoracic surgery examination data sorting system based on relevant data features includes: a data acquisition module, a classification weighting module, a relevant load module, a labeling adding module and a data adjustment module; Wherein, the data acquisition module is used to acquire thoracic surgery examination data and perform data cleaning on the thoracic surgery examination data; The classification and weighting module is used to classify the thoracic surgery examination data after data cleaning, and to weight the thoracic surgery examination data based on the classification results; The correlation load module is used to perform correlation analysis on the weighted results of thoracic surgery examination data, and calculate the thoracic surgery examination data load value based on the correlation analysis results; The annotation adding module is used to add annotations to the thoracic surgery examination data according to the thoracic surgery examination data classification results, thoracic surgery examination data weighting results, correlation analysis results and thoracic surgery examination data load values; The data adjustment module is used to organize and store the thoracic surgery examination data based on the annotations added to the thoracic surgery examination data; The data acquisition module, the classification weighting module, the related load module, the annotation adding module and the data adjustment module are connected in sequence.
2. A thoracic surgery examination data sorting system based on relevant data features according to claim 1, characterized in that: The classification weighting module includes: a data processing module, a feature extraction module, a feature classification module, a weight allocation module and an optimization verification module; Wherein, the data processing module is used for standardizing the thoracic surgery examination data; The feature extraction module is used to preset feature extraction rules and perform feature extraction on the standardized thoracic surgery examination data based on the feature extraction rules; The feature classification module is used to classify the thoracic surgery examination data according to the feature extraction results; The weight allocation module is used to set classification weight rules and assign weights to the classification results of thoracic surgery examination data based on the classification weight rules; The optimization verification module is used to verify the classification and weighting results of thoracic surgery examination data, and optimize the classification and weighting results of thoracic surgery examination data based on the verification results; The data processing module, the feature extraction module, the feature classification module, the weight allocation module and the optimization verification module are connected in sequence.
3. A thoracic surgery examination data sorting system based on relevant data features according to claim 2, characterized in that: The step of setting classification weight rules and weighting the classification results of thoracic surgery examination data based on the classification weight rules includes: Set the factor proportion weights and time sequence weights in the thoracic surgery examination data, and combine the factor proportion weights with the time sequence weights to obtain the classification weight rules; Extract data proportion factors and data time information from surgical examination data classification results; Based on the classification weight rules, weight the data proportion factor and data time information, and calculate the weight value of the data proportion factor and the weight value of the data time information; Summarize the weight values of the data proportion factors and the weight values of the data time information, and verify the summary results.
4. A thoracic surgery examination data sorting system based on relevant data features according to claim 3, characterized in that: The extraction of data proportion factors and data time information in the surgical examination data classification results includes: Set data segmentation rules, and segment the surgical examination data classification results by data proportion based on the data segmentation rules; Calculate each data proportion segmentation value based on the proportion segmentation results, and sort the proportion segmentation results according to the data proportion segmentation values; Set timestamp extraction rules, extract data acquisition time information in surgical examination data classification results based on the timestamp extraction rules, and unify the format of the data acquisition time information.
5. A thoracic surgery examination data sorting system based on relevant data features according to claim 1, characterized in that: The related load module includes: a data related module, a data load module, a data adjustment module and a related impact module; Wherein, the data correlation module is used to evaluate the correlation value between each data weighted result in the thoracic surgery examination data weighted result; The data load module is used to extract the load characteristics of thoracic surgery examination data and calculate the detection data load value based on the load characteristics; The data adjustment module is used to optimize and adjust the detection data load value according to the correlation value between the data weighting results; The related impact module is used to perform load impact analysis according to the optimized and adjusted detection data load value, and output the optimized and adjusted detection data load value and load impact analysis; The data related module, the data load module, the data adjustment module and the related impact module are connected in sequence.
6. A thoracic surgery examination data sorting system based on relevant data features according to claim 5, characterized in that: The data related module includes: a related linkage module, a related calculation module and a related evaluation module; The related linkage module is used to preset data linkage rules and add related annotations between data weighting results based on the data linkage rules; The correlation calculation module is used to calculate the correlation value between the data weighting results according to the relevant annotations; The correlation evaluation module is used to classify the correlation between the data weighting results according to the correlation values, and add labels based on the classification results; The related linkage module, the related calculation module and the related evaluation module are connected in sequence.
7. A thoracic surgery examination data sorting system based on relevant data features according to claim 6, characterized in that: The calculation formula for the correlation value between the weighted results of the data calculated based on the relevant annotations is: ; in, r Assign correlation values between data weight results; n is the total number of data weighting results; The ranking difference between the results of weighting the data; i is the index value between the data weighting results.
8. A thoracic surgery examination data sorting system based on relevant data features according to claim 5, characterized in that: The step of extracting the load characteristics of the thoracic surgery examination data and calculating the load value of the examination data based on the load characteristics comprises: Setting inspection data load rules, and extracting detection load characteristic values and data load characteristic values of thoracic surgery inspection data based on the inspection data load rules; Construct a load model, and substitute the detected load characteristic value into the load model to calculate the detected data load value; The detection load value and the detection data load value are summarized, and the summary results are verified, and the summary results of the verified detection data load value and data load characteristic value are output.
9. A thoracic surgery examination data sorting system based on relevant data features according to claim 8, characterized in that: The construction of the load model and substituting the detected load characteristic value into the load model to calculate the detected data load value comprises: Perform data cleaning on the detected load characteristic values, and normalize the load characteristic values after data cleaning; Setting load matching rules and a load model library, and matching the normalized load characteristic value with the load model library based on the load matching rules; The normalized load characteristic values are divided into a training set and a test set, the matching load model is trained by the training set, and the trained load model is tested by the test set; Substitute the load characteristic value into the load model after the test, calculate the test data load value, and cross-validate the test data load value.
10. A thoracic surgery examination data sorting system based on relevant data features according to claim 1, characterized in that: The annotation adding module includes: a data analysis module, a result annotation module, an annotation verification module and an annotation synthesis module; The data analysis module is used to analyze the data characteristic factors of the thoracic surgery examination data classification results, thoracic surgery examination data weighting results, correlation analysis results and thoracic surgery examination data load value; The result annotation module is used to generate data factor annotations based on data characteristic factors; The annotation verification module is used to compare and verify the data factor annotation and analysis of the thoracic surgery examination data classification results, thoracic surgery examination data weighting results, correlation analysis results and thoracic surgery examination data load values; The annotation and synthesis module is used to summarize the annotations of the data factors after comparison and verification, and to add annotations to the thoracic surgery examination data according to the data factor annotation and summary results.