Tumor disease analysis system based on electronic cases
By segmented annotation of electronic case data and multi-level time interval index construction, combined with distribution comparison and aggregation statistical techniques and evaluation of time impact index and characteristic deviation index, the problems of heterogeneity and inconsistency of electronic case data are solved, and the accuracy and credibility of tumor disease trend prediction are significantly improved.
Patent Information
- Application Number
- CN202510443483.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-05-23
AI Technical Summary
In long-term follow-up management based on electronic cases, early or irregular records may have outliers or missing fields, and direct inclusion in the analysis will lead to deviations in follow-up results. The existing technology lacks a refined distinction strategy for different stages and attributes, resulting in systematic errors.
It provides a tumor disease analysis system based on electronic cases. Through segmented annotation of historical records and follow-up data and multi-level time interval index construction, it uses distribution comparison and aggregation statistical technology to screen abnormal data, and realizes the deweighting or elimination of data through dynamic evaluation of time impact index and characteristic deviation index, combining cross-validation of multiple sources and cross-time periods to ensure the reliability and integrity of the data.
It effectively reduces the analysis error caused by heterogeneity and inconsistency of electronic case data, improves the accuracy and credibility of tumor disease trend prediction, and ensures the effectiveness and reliability of the analysis.
Smart Images

Figure CN120032784A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic medical record data processing, and more specifically, to a tumor disease analysis system based on electronic medical records. Background Art
[0002] In long-term follow-up management based on electronic medical records, information is usually extracted from historical records to provide reference for subsequent follow-up stages. Some early or irregular records may contain outliers or missing fields. If they are directly included in the overall analysis, they are likely to interfere with the follow-up results, resulting in deviations in trend assessment or statistical inference. How to remove noise data that is not helpful or highly interfering to follow-up analysis while retaining key disease evolution information has become a core challenge that is difficult to avoid in the process of data cleaning and aggregation. Existing solutions lack refined differentiation strategies for different stages and attributes. Inconsistent records from different sources or across time periods are often regarded as scattered anomalies, making it difficult to perform unified cross-validation and weight reduction elimination, which ultimately leads to systematic errors in follow-up trend analysis. In order to solve the above problems, a technical solution is now provided. Summary of the invention
[0003] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a tumor disease analysis system based on electronic medical records, which constructs a multi-level time interval index through segmented annotation of historical records and follow-up data; uses distribution comparison and aggregation statistical technology to screen out abnormal data that deviates from the normal distribution, and comprehensively quantifies its impact through dynamic evaluation of the time impact index and the feature deviation index, thereby achieving the downgrading or elimination of abnormal data; combines multi-source and cross-time period cross-validation to identify data consistency and potential deviations, and ensure the reliability and integrity of data fusion; thereby effectively reducing the analysis errors caused by the heterogeneity and inconsistency of electronic medical record data, improving the accuracy and credibility of tumor disease trend prediction, and ensuring the effectiveness and reliability of tumor disease analysis in practical applications, so as to solve the problems raised in the above-mentioned background technology.
[0004] To achieve the above object, the present invention provides the following technical solutions: The tumor disease analysis system based on electronic medical records includes: time period marking module, distribution screening module, weight evaluation module and consistency verification module; Time period annotation module: segmentally annotate historical records and follow-up data, establish multi-level time intervals and indexes according to different periods and record types, and output segmented annotated data sets and multi-level time interval indexes to the distribution screening module; Distribution screening module: performs distribution comparison and aggregate statistics on each key indicator field within the divided time interval, screens out suspicious data items that deviate from the normal distribution, summarizes them into a candidate list, and outputs the candidate list containing suspicious data items to the weight evaluation module; Weight evaluation module: For each record in the candidate list, the time decay and feature weight mechanism is used to re-evaluate its impact on the follow-up analysis results according to the distance from the current period and the importance of the field, make a downgrade or elimination decision, and output the downgraded or eliminated records to the consistency verification module; Consistent Verification Module: Cross-verify the same data entries that appear repeatedly from multiple sources or across time periods.
[0005] In a preferred embodiment, the time period marking module includes the following contents: First, the historical information is preliminarily segmented with reference to the occurrence date and category of the record to form several adjacent or continuous time blocks; then, the records of the follow-up stage are divided into another time series separately according to the preset start and end times in the follow-up plan; after the segmentation is completed, multi-dimensional tags are attached to the records of each time period; on this basis, index fields are generated for each time period block and the record objects within it; all index fields are stored in a unified index structure.
[0006] In a preferred embodiment, the distribution screening module includes the following contents: S2.1, the set of divided time intervals is , each time interval Include records, the record set is ; Each record The key indicator field set is ,in Representation record Medium Index The specific value of; the global reference benchmark distribution is ,in is the indicator number; S2.2, for each indicator in each time interval, a local distribution feature model is constructed, and the distribution density function in the interval is generated by the non-parametric kernel density estimation method. ; The calculation formula is as follows: ;in is the kernel smoothing bandwidth, which controls the smoothness of the distribution curve; As the kernel function, Gaussian kernel is used ; Refer to the global data and generate a global benchmark distribution density function for each indicator , calculated using the same kernel density estimation method as the local distribution; to improve the dynamic responsiveness of the distribution, a time-weighted function is used ,in Indicates time interval The interval from the current time, is the time attenuation coefficient; the weighted local distribution after superposition is: .
[0007] In a preferred embodiment, S2.3, the difference between the local distribution and the global distribution of a single record is calculated by integration to capture the abnormal deviation characteristics; the characteristic deviation value of each record on the indicator in the current time interval is calculated. : ;in is the local distribution kernel width; calculate the deviation index for each record , through the deviation index, multiple indicator information is aggregated into a single indicator for screening; the comprehensive deviation of all fields is calculated for each record: ;in For global indicators The distribution span of S2.4, mark the records whose deviation index is greater than the deviation threshold as suspicious entries, and generate a candidate list.
[0008] In a preferred embodiment, the weight evaluation module includes the following contents: First, the records in the candidate list are sorted according to the time interval and feature deviation To perform dynamic segmentation: Time dimension segmentation: based on time interval Divide records into recent segments , mid-term period and forward period , each segment represents a different time impact range; and It is the threshold used for time dimension division in the dynamic segmentation process; Feature dimension segmentation: based on feature deviation value The discrete degree is divided into segments, which are divided into high deviation segments , mid-deviation segment and low deviation segment ,in is the global median of the feature, is the deviation range.
[0009] In a preferred embodiment, the weight evaluation module further includes the following contents: The time impact index measures the dynamic impact of the recorded time position on the follow-up analysis based on the consistency of the time distance and the data trend within the interval; formula: ;in It is a record Time impact index; It is the time interval The proportion of records in the time period indicates the importance of the corresponding time period; is the time decay rate, reflecting the weight of the influence of time distance; It is the time interval The time interval from the current period; It is the time interval The trend consistency is defined as: ;in For time interval The standard deviation of the characteristic deviation values recorded in, is the mean of the characteristic deviation values in the corresponding time interval.
[0010] In a preferred embodiment, the weight evaluation module further includes the following contents: The feature deviation index quantifies the abnormality of records by building a dynamic segmentation model, combining the deviation degree of candidate records on key indicators with their global distribution characteristics; formula: ;in It is a record The characteristic deviation index of is recorded in the indicator The deviation on Is an indicator The global median of ; Is an indicator The deviation range; index The significance score in the high deviation segment is defined as: ;in Indicates that in all records, the indicator The characteristic deviation value of Exceeding the threshold The number of records; Indicator The total number of occurrences in all records, that is, the number of all records including high deviation segments, medium deviation segments, and low deviation segments.
[0011] In a preferred embodiment, the weight evaluation module further includes the following contents: The time impact index and the feature deviation index are calculated based on the comparison increment to obtain the decision coefficient; based on the calculated decision coefficient , classify the records in the candidate list: ;in and represent the first decision threshold and the second decision threshold respectively.
[0012] In a preferred embodiment, the consistency verification module includes the following: In the candidate list after decision making, according to the time interval of each record and source identification Group and generate unique identifiers for duplicate entries ; Representation record A unique identifier for tracking its relevance across sources and time; For each unique ID, extract the set of records from different sources ; The consistency score measures the similarity of feature deviation values of data from multiple sources. By comparing the differences in feature values from different sources, it evaluates whether there are significant inconsistencies between records from different sources; Calculate the consistency score : ;in is the number of sources; is the number of features; is recorded in the source Medium Features The deviation value of It is a characteristic that uniquely identifies a record across all sources. The mean of is a small positive number that prevents the denominator from being zero.
[0013] In a preferred embodiment, the consistency verification module further includes the following: For each unique identifier, check whether its change trend across time intervals is consistent; the cross-period deviation score measures the characteristic change trend of the record in multiple time intervals, and evaluates whether the record shows a stable trend by comparing the change in the deviation value of adjacent time intervals; calculate the cross-period deviation score : ;in Is a unique identifier The number of time intervals that occurred; Is recorded in the time interval The characteristic deviation value in ; is a tiny value that prevents the denominator from being zero; is the cross-period consistency score; the comprehensive verification score combines the consistency score and the cross-period deviation score to quantify the global stability of records across multiple sources and time intervals through nonlinear combinations; calculate the comprehensive verification score : ; Decision rules: ;in It is the elimination threshold. Records below this value are considered to be inconsistent and eliminated; records above this value are considered to be consistent and retained directly.
[0014] The technical effects and advantages of the tumor disease analysis system based on electronic medical records of the present invention are as follows: The present invention provides a tumor disease analysis method based on electronic medical records. By segmenting and labeling historical records and follow-up data and constructing a multi-level time interval index, accurate management and organization of different time periods and data sources are achieved. The use of distribution comparison and aggregation statistical technology can effectively screen out potential abnormal data that deviates from the normal distribution, and through the dynamic evaluation mechanism of the time impact index and the feature deviation index, the degree of influence of candidate records can be comprehensively quantified, thereby realizing the refined downgrading or elimination decision of abnormal data. Combined with cross-validation of multiple sources and repeated records across time periods, the consistency and potential deviations of the data can be accurately identified to ensure the reliability and integrity of data fusion. This method significantly reduces the analysis errors caused by the heterogeneity and inconsistency of electronic medical record data, improves the accuracy and credibility of tumor disease trend prediction, optimizes the data processing process, provides more reliable and accurate disease management support, and ensures the effectiveness and reliability of tumor disease analysis based on electronic medical records in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic diagram of the structure of the tumor disease analysis system based on electronic medical records of the present invention. DETAILED DESCRIPTION
[0016] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0017] Embodiment 1: Figure 1 The invention provides a tumor disease analysis system based on electronic medical records, including: a time period marking module, a distribution screening module, a weight evaluation module and a consistency verification module; Time period annotation module: segmentally annotate historical records and follow-up data, establish multi-level time intervals and indexes according to different periods and record types, and output segmented annotated data sets and multi-level time interval indexes to the distribution screening module; Distribution screening module: performs distribution comparison and aggregate statistics on each key indicator field within the divided time interval, screens out suspicious data items that deviate from the normal distribution, summarizes them into a candidate list, and outputs the candidate list containing suspicious data items to the weight evaluation module; Weight evaluation module: For each record in the candidate list, the time decay and feature weight mechanism is used to re-evaluate its impact on the follow-up analysis results according to the distance from the current period and the importance of the field, make a downgrade or elimination decision, and output the downgraded or eliminated records to the consistency verification module; Consistency verification module: Cross-verify the same data items that appear repeatedly from multiple sources or across time periods. If multiple sources show consistency anomalies, the corresponding items will be eliminated; if there are differences in some sources, the corresponding items will be retained to avoid overall deviations caused by distortion of a single source.
[0018] In the follow-up management process of electronic medical records, historical records often contain heterogeneous data from different periods and sources, and there are large differences in their input formats and data quality. If they are directly compared with follow-up data, the inconsistency of time span and data attributes may lead to deviations in analysis results. By segmenting and labeling historical records and follow-up data, and establishing a multidimensional index structure for each time interval and record entry, different data types and sources can be effectively separated, and the location and management of information at each stage can be clarified, providing a basis for subsequent precise analysis and abnormal screening.
[0019] The time period marking module includes the following: First, the historical information is segmented initially according to the date and category of the record, forming several adjacent or continuous time blocks; then, the records of the follow-up phase are divided into another time series separately according to the preset start and end time in the follow-up plan. After the segmentation is completed, the records of each time period are attached with multi-dimensional tags, including data attributes (such as surgical information, test information, imaging information, etc.), input sources (such as different departments or different information systems) and data quality levels (comprehensively evaluated by indicators such as record completeness and internal consistency). On this basis, index fields are generated for each time period block and the record objects within it, including unique record identifiers, time period identifiers, data attribute labels and quality rating labels. All index fields are stored in a unified index structure, allowing subsequent processes to quickly locate the target time period or specific types of records based on the index.
[0020] Through segmented annotation and the establishment of multi-level indexes, historical records and follow-up data are systematically organized into structured data sets with clear time intervals and type dimensions, forming a foundation for cross-period, multi-source data management. The multi-dimensional index structure provides efficient data retrieval capabilities for subsequent steps, ensures the consistency of entries during aggregate analysis and anomaly screening, and lays a rigorous and clear technical framework for the entire data processing chain.
[0021] In the historical records and follow-up data of electronic medical records, the distribution characteristics of key indicators may show complex dynamic changes due to different time periods or data sources. Abnormal deviations of some indicators may be hidden in the local distribution and not covered by the global trend, making it difficult for subsequent analysis to accurately identify potential abnormal items. By dynamically modeling and weighted analysis of indicator distribution and capturing the deviation between the local distribution and the global benchmark, potential interference data can be effectively screened out, providing a basis for data cleaning for subsequent decision-making and trend modeling.
[0022] First, the time interval division and multi-dimensional index annotation of the data have been completed in the time period annotation module. On this basis, the goal of this step is to identify suspicious record entries that deviate from the normal distribution. To achieve this goal, it is necessary to build an indicator distribution feature model and analyze the difference between local distribution and global distribution.
[0023] The distribution screening module includes the following: S2.1, the set of divided time intervals is , each time interval Include records, the record set is .
[0024] Each record The key indicator field set is ,in Representation record Medium Indicator The specific value of .
[0025] The global reference base distribution is ,in The indicator number.
[0026] S2.2, for each indicator in each time interval, a local distribution feature model is constructed, and the distribution density function in the interval is generated by the non-parametric kernel density estimation method. The calculation formula is as follows: ;in is the kernel smoothing bandwidth, which controls the smoothness of the distribution curve; As the kernel function, Gaussian kernel is used .
[0027] Refer to the global data (i.e. the data set of all time intervals) to generate a global benchmark distribution density function for each indicator , calculated using the same kernel density estimation method as for the local distribution.
[0028] To improve the dynamic responsiveness of the distribution, a time weighting function is used ,in Indicates time interval The interval from the current time, is the time decay coefficient.
[0029] The weighted local distribution after superposition is: .
[0030] S2.3, by integrating and calculating the difference between the local distribution and the global distribution of a single record, the abnormal deviation characteristics are captured. The greater the difference, the more likely the record is abnormal data. Calculate the characteristic deviation value of each record on the indicator in the current time interval : ;in is the local distribution kernel width, which is used to capture the distribution deviation of records in a small range.
[0031] Calculate the deviation index for each record , comprehensively consider the deviation of all key indicators. Through the deviation index, multiple indicator information is aggregated into a single indicator for screening. Calculate the comprehensive deviation of all fields for each record: ;in For global indicators The distribution span (difference range) of .
[0032] S2.4, set a deviation threshold, and mark the records with a deviation index greater than the deviation threshold as suspicious entries. For each suspicious record, attach its source time interval, deviation information of specific indicators, and comprehensive deviation of all fields.
[0033] Generate a candidate list, where each record contains the following structured content: record identifier (time period tag and index ID), deviation indicator list (including field name and deviation value) and time interval information.
[0034] The distribution screening module successfully achieved dynamic capture of key indicator distribution characteristics and quantitative analysis of deviations through multi-layer distribution modeling and time weighting mechanism. Through the calculation of difference integral and comprehensive deviation index, suspicious data items that deviate from the normal distribution are effectively identified, and a candidate list is generated, laying an accurate data foundation for the next step of time decay and feature weight evaluation. The processing results have the advantages of clear distribution characteristics, traceable data items, and quantifiable abnormal screening, providing key support for the accuracy and credibility of the final follow-up analysis.
[0035] In the distribution screening module, candidate abnormal records that deviate from the normal distribution are screened out and a list is generated through distribution comparison and aggregate statistics. These records may interfere with the analysis results of the follow-up data, but relying solely on the degree of deviation is not enough to fully evaluate their impact. The weight evaluation module further quantifies the comprehensive impact of the records by constructing the time impact index and feature deviation index and comprehensively generating the decision coefficient, providing an accurate basis for the decision to exclude or downgrade, and laying the foundation for subsequent verification and trend modeling.
[0036] The weight evaluation module includes the following: First, the records in the candidate list are sorted according to the time interval and feature deviation Perform dynamic segmentation.
[0037] Time dimension segmentation: based on time interval Divide records into recent segments , mid-term period and forward period , each segment represents a different time impact range. and It is the threshold used for time dimension segmentation during dynamic segmentation.
[0038] Feature dimension segmentation: based on feature deviation value The discrete degree is divided into segments, which are divided into high deviation segments , mid-deviation segment and low deviation segment ,in is the global median of the feature, is the deviation range.
[0039] The calculation idea of the time impact index is based on the consistency of time distance and data trend within the interval, and comprehensively measures the dynamic impact of the time position of the record on the follow-up analysis. By introducing the time decay model, the time distance between the record and the current period is converted into an exponential form. As the time interval increases, the weight of the record gradually decreases. In addition, considering that the data within different time periods may be volatile, in order to avoid the instability caused by data fluctuations, the interval trend consistency parameter is introduced into the index to quantify the distribution stability within the interval. Trend consistency is calculated by the ratio of the standard deviation of the deviation within the interval to the mean, reflecting the internal discreteness of the interval in which the record is located, and finally realizing a comprehensive evaluation of time distance and interval stability. The design of the time impact index ensures that records with closer time and more stable internal distribution have greater weight in the analysis, thereby enhancing the attention and utilization of recent data. Formula: ;in It is a record Time impact index; It is the time interval The proportion of internal records, indicating the importance in the corresponding time period; is the time decay rate, reflecting the weight of time distance on the impact; is the time interval The time interval from the current period; is the time interval The trend consistency of, defined as: ; where is the time interval The standard deviation of the feature deviation values recorded in, is the mean of the feature deviation values for the corresponding time interval.
[0040] The feature deviation index calculates the abnormality recorded on key features by segments, integrating the significance of high-deviation segments and the distribution concentration of medium- and low-deviation segments. The feature deviation index constructs a dynamic segmentation model, combines the deviation degree of candidate records on key indicators with their global distribution characteristics, and quantifies the abnormality of records. In the calculation, the deviation features are first divided into high-deviation, medium-deviation, and low-deviation segments to capture the significance of different deviation degrees in the global context. Subsequently, through the squared processing of relative deviation values, the significant contribution of records with larger deviation degrees to the abnormality assessment is emphasized. At the same time, the significance score of high-deviation segments is introduced into the index to measure the rarity or importance of candidate records in the overall dataset. Finally, the feature deviation index realizes the effective aggregation of multi-feature deviation information and dynamically adjusts the influence weight of deviation characteristics on the abnormality assessment, enabling the index to comprehensively and accurately capture the abnormal behavior of data records in the feature dimension. Formula: ; where is the record 's feature deviation index; is the deviation degree of the record on the indicator ; is the global median of the indicator ; is the deviation range of the indicator ; The significance score of the indicator in the high-deviation segment, defined as: ; where represents that among all records, the feature deviation value of the indicator exceeds the threshold The number of records; ; represents the total number of times the indicator appears among all records, that is, the total number of records including high-deviation segments, medium-deviation segments, and low-deviation segments.
[0041] The decision coefficient is calculated based on the time impact index and the feature deviation index in a comparative increment manner. For example, it can be calculated in the following way: ; Calculation logic: Part I By calculating the sum of the squares of the time impact index and the feature deviation index, the combined weight of the two is quantified and the synergy between the two is amplified. The denominator combines the difference between the two indexes to avoid the dominance of a single index and ensure a smoother distinction between abnormal records and normal records.
[0042] Part 2 The inverse tangent function is introduced to handle the nonlinear changes of the time impact index, emphasizing the decreasing effect of time distance on abnormal data. The denominator is combined with the square of the characteristic deviation exponent to further reduce the suppression effect of extreme deviation records on the overall decision coefficient.
[0043] Based on the calculated decision coefficient , classify the records in the candidate list: ; Logical explanation: and represent the first decision threshold and the second decision threshold respectively.
[0044] Elimination: When the decision coefficient is greater than the first decision threshold, it means that the combined weight of the recorded time influence and characteristic deviation is too high, and it is regarded as an interference item for the follow-up analysis and is eliminated.
[0045] De-weighting: When the decision coefficient is within the first decision threshold and the second decision threshold of the tolerance interval, it means that the impact of the record is still within the controllable range, but its analysis weight needs to be reduced.
[0046] Retention: When the decision coefficient is less than the second decision threshold, the recorded time and features have low impact and can be directly retained for subsequent analysis.
[0047] For records that need to be downgraded, adjust their time impact and feature deviation weights: , ;in, ∈(0,1) is the weight reduction coefficient, which is used to weaken the weight influence of the record.
[0048] The decision coefficient calculation method based on the comparison increment effectively captures the abnormal characteristics and comprehensive weights of the candidate records by integrating the time impact index and the feature deviation index. The threshold judgment and weight reduction mechanism further refine the record processing strategy to ensure that the cleaning and optimization of the follow-up data are both rigorous and flexible. Ultimately, the processed records provide high-quality data support for multi-source cross-validation and trend model construction.
[0049] The weight assessment module calculates the time impact index and feature deviation index, generates decision coefficients in combination with the comparative increment method, and comprehensively evaluates the potential interference of candidate abnormal records on follow-up analysis. Through the classification processing methods of elimination, downgrading and retention, the quality and credibility of the data are optimized, providing efficient and reliable data input for the next step of multi-source cross-validation and the construction of the final trend model.
[0050] In the weight evaluation module, the decision coefficient of each candidate record is generated through the comprehensive calculation of the time impact index and the feature deviation index, and preliminary elimination is performed based on the threshold. However, some items in the candidate list may involve multiple sources or repeated records across time intervals. The abnormal properties of such records may be due to deviations from a single source rather than global consistency issues. The consistency verification module further refines the processing strategy for these records through multi-source and cross-time cross-validation to ensure the accuracy and comprehensiveness of the final results.
[0051] The consistent verification module includes the following: In the candidate list after decision making, according to the time interval of each record and source identification Group and generate unique identifiers for duplicate entries ; Representation record A unique identifier for a resource, used to track its relevance across sources and time.
[0052] The grouping is based on time interval and source, ensuring that the same entries from different sources are logically aggregated together to facilitate subsequent verification.
[0053] For each unique ID, extract the set of records from different sources The consistency score measures the similarity of feature deviation values of data from multiple sources. By comparing the differences in feature values from different sources, it evaluates whether there are significant inconsistencies between records from different sources. Calculating the consistency score : ;in is the number of sources; is the number of features; is recorded in the source Medium Features The deviation value of It is a characteristic that uniquely identifies a record across all sources. The mean of is a small positive number that prevents the denominator from being zero.
[0054] The closer the consistency score is to 1, the smaller the difference in the deviation values of record characteristics between different sources is; conversely, the greater the difference is.
[0055] For each unique identifier, check whether its change trend across time intervals is consistent. The cross-period deviation score measures the change trend of the characteristics of the record in multiple time intervals. By comparing the change in the deviation value of adjacent time intervals, it evaluates whether the record shows a stable trend. Calculate the cross-period deviation score : ;in Is a unique identifier The number of time intervals that occurred; Is recorded in the time interval The characteristic deviation value in ; is a tiny value that prevents the denominator from being zero; It is the consistency score across time periods. The closer the value is to 1, the more stable the change across time periods. The deviation score across time periods is used to judge the consistency of records in time and identify abnormal deviations caused by time period changes.
[0056] The composite validation score combines the consistency score and the cross-period deviation score to quantify the global stability of the records across multiple sources and time intervals through a nonlinear combination. Calculating the composite validation score : ; Decision rules: ;in The threshold for elimination. Records with a value lower than this value are considered to be inconsistent and eliminated. Records with a value higher than this value are considered to be consistent and retained directly. The consistency verification module further optimizes the processing results of candidate records by cross-validating from multiple sources and across time periods, comprehensively considering the consistency and trend stability of records. The multi-source consistency score emphasizes the consistency of feature deviations between different sources, and the cross-time deviation score captures the trend changes in the time dimension. The combined comprehensive verification score of the two provides a rigorous decision-making basis. The final output processing results provide a clearer and higher-quality data foundation for the construction of the follow-up trend model.
[0057] The present invention provides a tumor disease analysis system based on electronic medical records, which realizes the precise management of historical records and follow-up data through systematic data segmentation annotation and the establishment of multi-level time interval index. On this basis, the system performs distribution comparison and aggregation statistics on key indicator fields in each time interval, screens out suspicious data entries that deviate from the normal distribution and generates a candidate elimination list. Subsequently, the candidate records are re-evaluated using an innovative time impact index and feature deviation index mechanism, and a decision is made whether to downgrade and eliminate them based on their time distance from the current time period and the importance of key indicators. Finally, the system performs rigorous cross-validation on the same data entries that appear repeatedly from multiple sources or across time periods to ensure the consistency and reliability of the data and avoid affecting the overall analysis results due to distortion of a single source. Through the above steps, the present invention effectively solves the problems of data heterogeneity, deviation and inconsistency in electronic medical records, and significantly improves the accuracy and credibility of tumor disease analysis.
[0058] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.
[0059] The above description is only by way of illustration of certain exemplary embodiments of the present invention. It is undoubted that those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
[0060] It should be noted that, in this article, if there are relational terms such as first and second, etc., they are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "including a..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0061] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A tumor disease analysis system based on electronic medical records, characterized in that: include: Time period labeling module, distribution screening module, weight evaluation module and consistency verification module; Time period annotation module: segmentally annotate historical records and follow-up data, establish multi-level time intervals and indexes according to different periods and record types, and output segmented annotated data sets and multi-level time interval indexes to the distribution screening module; Distribution screening module: performs distribution comparison and aggregate statistics on each key indicator field within the divided time interval, screens out suspicious data items that deviate from the normal distribution, summarizes them into a candidate list, and outputs the candidate list containing suspicious data items to the weight evaluation module; Weight evaluation module: For each record in the candidate list, the time decay and feature weight mechanism is used to re-evaluate its impact on the follow-up analysis results according to the distance from the current period and the importance of the field, make a downgrade or elimination decision, and output the downgraded or eliminated records to the consistency verification module; Consistent Verification Module: Cross-verify the same data entries that appear repeatedly from multiple sources or across time periods.
2. The tumor disease analysis system based on electronic medical records according to claim 1, characterized in that: The time period marking module includes the following: First, the historical information is preliminarily segmented with reference to the occurrence date and category of the record to form several adjacent or continuous time blocks; then, the records of the follow-up stage are divided into another time series separately according to the preset start and end times in the follow-up plan; after the segmentation is completed, multi-dimensional tags are attached to the records of each time period; on this basis, index fields are generated for each time period block and the record objects within it; all index fields are stored in a unified index structure.
3. The tumor disease analysis system based on electronic medical records according to claim 2 is characterized in that: The distribution screening module includes the following: S2.1, the set of divided time intervals is , each time interval Include records, the record set is ; Each record The key indicator field set is ,in Representation record Medium Indicator The specific value of; the global reference benchmark distribution is ,in is the indicator number; S2.2, for each indicator in each time interval, a local distribution feature model is constructed, and the distribution density function in the interval is generated by the non-parametric kernel density estimation method. ; The calculation formula is as follows: ;in is the kernel smoothing bandwidth, which controls the smoothness of the distribution curve; As the kernel function, Gaussian kernel is used ; Refer to the global data and generate a global benchmark distribution density function for each indicator , calculated using the same kernel density estimation method as the local distribution; to improve the dynamic responsiveness of the distribution, a time-weighted function is used ,in Indicates time interval The interval from the current time, is the time attenuation coefficient; The weighted local distribution after superposition is: .
4. The tumor disease analysis system based on electronic medical records according to claim 3 is characterized in that: S2.3, by integrating the difference between the local distribution and the global distribution of a single record, capture the abnormal deviation characteristics; calculate the characteristic deviation value of each record on the indicator in the current time interval : ;in is the local distribution kernel width; Calculate the deviation index for each record , through the deviation index, multiple indicator information is aggregated into a single indicator for screening; the comprehensive deviation of all fields is calculated for each record: ;in For global indicators The distribution span of S2.4, mark the records whose deviation index is greater than the deviation threshold as suspicious entries, and generate a candidate list.
5. The tumor disease analysis system based on electronic medical records according to claim 4 is characterized in that: The weight evaluation module includes the following: First, the records in the candidate list are sorted according to the time interval and feature deviation To perform dynamic segmentation: Time dimension segmentation: based on time interval Divide records into recent segments , mid-term period and forward period , each segment represents a different time impact range; and It is the threshold used for time dimension division in the dynamic segmentation process; Feature dimension segmentation: According to the characteristic deviation value The discrete degree is divided into high deviation segments , mid-deviation segment and low deviation segment ,in is the global median of the feature, is the deviation range.
6. The tumor disease analysis system based on electronic medical records according to claim 5, characterized in that: The weight evaluation module also includes the following: The temporal impact index measures the dynamic impact of the temporal position of the record on the follow-up analysis based on the consistency of temporal distance and data trends within the interval; formula: ;in It is a record Time impact index; It is the time interval The proportion of records in the time period indicates the importance of the corresponding time period; is the time decay rate, reflecting the weight of the influence of time distance; It is the time interval The time interval from the current period; It is the time interval The trend consistency is defined as: ;in For time interval The standard deviation of the characteristic deviation values recorded in, is the mean of the characteristic deviation values in the corresponding time interval.
7. The tumor disease analysis system based on electronic medical records according to claim 6, characterized in that: The weight evaluation module also includes the following: The feature deviation index quantifies the abnormality of records by building a dynamic segmentation model, combining the deviation degree of candidate records on key indicators with their global distribution characteristics; formula: ;in It is a record The characteristic deviation index of is recorded in the indicator The deviation on Is an indicator The global median of ; Is an indicator The deviation range; index The significance score in the high deviation segment is defined as: ;in Indicates that in all records, the indicator The characteristic deviation value of Exceeding the threshold The number of records; Indicator The total number of occurrences in all records, that is, the number of all records including high deviation segments, medium deviation segments, and low deviation segments.
8. The tumor disease analysis system based on electronic medical records according to claim 7, characterized in that: The weight evaluation module also includes the following: The time impact index and the feature deviation index are calculated based on the comparison increment to obtain the decision coefficient; based on the calculated decision coefficient , classify the records in the candidate list: ;in and represent the first decision threshold and the second decision threshold respectively.
9. The tumor disease analysis system based on electronic medical records according to claim 8, characterized in that: The consistent verification module includes the following: In the candidate list after decision making, according to the time interval of each record and source identification Group and generate unique identifiers for duplicate entries ; Representation record A unique identifier for tracking its relevance across sources and time; For each unique ID, extract the set of records from different sources The consistency score measures the similarity of feature deviation values of data from multiple sources. By comparing the differences in feature values from different sources, it assesses whether there are significant inconsistencies between records from different sources. Calculating the consistency score : ;in is the number of sources; is the number of features; is recorded in the source Medium Features The deviation value of It is a characteristic that uniquely identifies a record across all sources. The mean of is a small positive number that prevents the denominator from being zero.
10. The tumor disease analysis system based on electronic medical records according to claim 9, characterized in that: The consistent verification module also includes the following: For each unique identifier, check whether its change trend across time intervals is consistent; The cross-period deviation score measures the characteristic change trend of the record in multiple time intervals. By comparing the change in the deviation value of adjacent time intervals, it evaluates whether the record shows a stable trend; Calculating the cross-period deviation score : ;in Is a unique identifier The number of time intervals that occurred; Is recorded in the time interval The characteristic deviation value in ; is a tiny value that prevents the denominator from being zero; is the cross-period consistency score; the comprehensive verification score combines the consistency score and the cross-period deviation score to quantify the global stability of records across multiple sources and time intervals through nonlinear combinations; Calculating the overall validation score : ; Decision rules: ;in It is the elimination threshold. Records below this value are considered to be inconsistent and eliminated; records above this value are considered to be consistent and retained directly.
Citation Information
Cited By
Novel early prediction method for coronavirus myocardial injury based on mitochondrial function
CN120766978A
Automobile extended insurance multi-source heterogeneous data cleaning evaluation system
CN120929453A