Large-span regional tumor early screening, early warning, and testing data management system and methods

By standardizing and individualizing baseline calculations of cross-institutional cancer early screening data, combined with two-dimensional trend analysis and instrument bias correction, the standardization problem of heterogeneous cross-institutional data was solved, improving the accuracy of early cancer screening and the effectiveness of the early warning system.

CN122135948APending Publication Date: 2026-06-02HENAN ZHONGKANG MEDICAL LAB CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HENAN ZHONGKANG MEDICAL LAB CO LTD
Filing Date
2026-04-20
Publication Date
2026-06-02

Smart Images

  • Figure CN122135948A_ABST
    Figure CN122135948A_ABST
Patent Text Reader

Abstract

This invention relates to tumor early warning and data processing, and more particularly to a tumor early screening and early warning data management system and method spanning a large region. The method includes: converting cross-institutional raw test data into standardized records with uniform dimensions and merging them into individual longitudinal files; extracting historical test data to calculate the individual baseline mean and fluctuation range; calculating the trend slope based on recent continuous sequences and combining it with the current baseline to calculate the standardized drift; determining the trigger for graded early warning based on the two-dimensional combination of slope and drift, and dynamically calibrating the threshold based on a follow-up closed-loop system. This invention can solve or at least mitigate the problems of early subtle trend missed diagnoses and systemic bias accumulation caused by cross-domain data heterogeneity barriers and single-point static thresholds, providing a tumor early screening and early warning data management system and method spanning a large region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to tumor early warning and data processing, and more particularly to a tumor early screening, early warning and testing data management system and method spanning a large area. Background Technology

[0002] In modern regional public health and early cancer prevention and control systems, large-scale regional cancer early screening and early warning is a core component for reducing malignant tumor mortality and optimizing medical resource allocation. This places extremely high engineering demands on the homogeneity and longitudinal tracking capabilities of cross-institutional testing data. With the continuous advancement of regional medical informatization, tumor marker screening networks are evolving towards "multi-level institutional interconnection," massive data aggregation, and high automation. This not only requires the underlying data management system to maintain extremely high dimensional consistency when processing multi-source heterogeneous data, but also demands that it possess the ability to highly fit the actual pathological evolution of examinees and provide anomaly feedback when faced with perturbation factors such as individual physiological differences. The temporal completeness of screening data and individual baseline heterogeneity are related to the quantitative benchmark for tumor early warning judgment. However, in automated continuous monitoring, regardless of whether it involves laboratory testing errors or data degradation during data preprocessing, significant biases in trend calculations will occur in early warning analysis, leading to missed detections of early tumors or misjudgments of physiological fluctuations.

[0003] Current technologies include several methods for compiling and alerting medical data for early cancer risk screening. One approach is based on single-point threshold monitoring of cross-sectional data, such as the commonly used population reference interval comparison. This involves establishing a comparison logic between a single test measurement and a preset upper limit of normal reference value to determine abnormal conditions. This approach can detect significant exceedances of biochemical indicators in a single medical visit, and the filtering of basic patients is improved through a backend rule engine. Another approach starts with comparing the differences in test results, monitoring the absolute change between two adjacent test results using built-in range rules. This design attempts to achieve preliminary identification of the examinee's condition by comparing one test result longitudinally with another, utilizing the dispersion of the numbers.

[0004] The coupling of the working principles of the aforementioned existing technologies with complex real-world clinical screening reveals that, under the requirements of new-generation regional quality control standards for high predictive accuracy and low over-intervention rates, these data management and early warning logics reflect underlying technological deficiencies. 1. The traditional regional direct upload mechanism is essentially a simple aggregation of heterogeneous data. The macro and micro environments naturally lack unified standardized data conversion rules to account for differences in laboratory brand and dimensionality. When a cross-laboratory data query is initiated, a single combination of numbers is insufficient to break down the heterogeneous barriers of coding and units between different laboratories. Relying solely on direct data comparison not only leads to the fragmentation of individual longitudinal medical records but also destroys the horizontal comparability boundaries of test results from different laboratories during the aggregation period. It is also impossible to eliminate quality degradation noise caused by sample hemolysis and collection time lags at the same time sequence.

[0005] 2. Traditional single-point threshold monitoring can screen for advanced lesions to some extent, but in the early stages of tumor markers, the internal score changes are significantly influenced by individual physiological baselines during the slow initial rise. Most current early warning models use a one-size-fits-all approach with indiscriminate judgment across population reference intervals, failing to effectively isolate individual basal metabolic fluctuations. The lack of support from individual historical baselines for physiological load can easily lead to missed detections and monitoring failures before indicators exceed the upper limit of the population reference interval, resulting in the loss of time-series trend characteristics and worsening of early warnings.

[0006] 3. Due to the inevitable large-scale screening across different medical institutions, systemic errors in the instruments used by these institutions are unavoidable. Current early warning systems cannot impose quantitative constraints on the biases of different medical institutions, similar to those imposed by standard quality assessment samples, and lack closed-loop correction channels based on actual epidemiological diagnostic results. Over time, this leads to false trend slopes, false alarms, and a deterioration in the adaptability of early warning thresholds. These systemic errors among different medical institutions result in the specificity of early cancer warnings remaining consistently at a low level.

[0007] Therefore, standardization of heterogeneous data across institutions, damping of individual historical baselines and two-dimensional trend analysis, and early warning systems for laboratory data with cross-institutional bias correction and threshold adaptive closed-loop compensation capabilities are technical challenges that urgently need to be addressed in medical informatics and automated evaluation of early cancer screening. Summary of the Invention

[0008] To achieve the above-mentioned objectives, this invention provides a large-span regional tumor early screening, warning, and testing data management system and method, aiming to solve or at least mitigate the technical problems in the prior art, such as the failure of horizontal comparison caused by heterogeneous cross-institutional testing data, the missed diagnosis of weak early abnormal signals of tumors caused by traditional single-point static thresholds, and the false trend false alarms and deterioration of early warning effectiveness caused by deviations in cross-institutional testing instrument systems.

[0009] To achieve the above objectives, the present invention provides the following technical solution: a method for early cancer screening and early warning and test data management across a large geographical area, comprising the following steps: Receive raw test data from screening institutions in various regions and convert the raw test data into standardized records with unified dimensions and unified item codes; The standardized records are matched according to the examinee's identity information, and cross-institutional standardized records belonging to the same examinee are merged into the corresponding individual longitudinal file; Historical testing data are extracted from the individual's longitudinal archives to calculate the individual's historical baseline mean and individual historical fluctuation range for the test items; The trend slope is calculated based on recent continuous detection data in the individual's longitudinal profile, and the standardized drift is calculated based on the current detection value, the individual's historical baseline mean, and the individual's historical fluctuation range. Based on the comparison results of the trend slope and the preset trend slope threshold, and the comparison results of the standardized drift amount and the preset drift amount threshold, a two-dimensional combined judgment is performed to trigger the corresponding level of warning.

[0010] To further realize the present invention, the following technical solutions may be preferred: Preferably, the specific method for converting the original test data into standardized records is as follows: conversion is performed based on the unit conversion coefficient and project code lookup table pre-set by each institution; and the source institution identifier field, conversion coefficient version number field, and data receipt timestamp field are written into the standardized records.

[0011] Preferably, the step of performing a two-dimensional combined judgment to trigger a warning of the corresponding level specifically includes: When both the trend slope and the standardized drift exceed their respective thresholds, a Level 1 warning is triggered, and a clinical referral recommendation is output. A second-level warning is triggered when only the trend slope or only the standardized drift exceeds its respective threshold, and a re-examination suggestion is output. When neither the trend slope nor the standardized drift exceeds the threshold, but the difference between either of them and the threshold is less than the preset threshold, a third-level warning is triggered, and a suggestion to shorten the follow-up interval is output.

[0012] Preferably, the standardized record further includes a sample quality score field, which is calculated based on the sample collection time interval, the hemolysis index field, and the sample type field. When calculating the trend slope, data points with sample quality scores below a preset quality threshold are assigned low weight coefficients, and data points with sample quality scores that meet the preset requirements are assigned normal weight coefficients, and the trend slope is calculated in a weighted manner.

[0013] Preferably, the step of extracting historical testing data to calculate the individual's historical baseline mean and individual historical fluctuation range specifically involves: The stable period dataset is defined as the continuous record segment in the first batch of continuous detection data after the establishment of an individual's file, in which the change range between two adjacent detection values ​​does not exceed the preset stability judgment range. When the adjacent change range of a newly added record in the file meets the stability judgment range condition, the new record is included in the stable period dataset and the individual's historical baseline mean and individual historical fluctuation range are recalculated.

[0014] Preferably, the matching of standardized records according to the examinee's identity information specifically includes: The matching key is a combination of the examinee's name, ID number, and date of birth. If all three fields match, the examinee is identified as the same person. If the ID number matches but the name or date of birth differs, the record is marked as pending verification and is not merged into the existing file. If the match fails, a new individual file is created for the examinee.

[0015] Preferably, it also includes a cross-agency inspection deviation correction step: The system periodically pushes standard quality assessment sample inspection tasks to each access institution, collects the inspection results of each institution on the quality assessment samples, calculates the deviation between the inspection results of each institution and the reference value, and stores it as the system deviation coefficient of each institution; in the conversion of the data standardization record, the system deviation coefficient corresponding to each institution is used to correct the deviation of the transmitted data of that institution.

[0016] Preferably, the preset trend slope threshold and the preset drift threshold are differentiated threshold parameters configured based on different geographical regions and different test items; The method also includes: writing the follow-up diagnosis results of subjects who have triggered the warning into the corresponding individual files; calculating the sensitivity and specificity of each warning level according to a preset period; and recalibrating the threshold parameters of the corresponding region based on the updated regional positive detection rate when the sensitivity or specificity is lower than the preset lower limit.

[0017] Preferably, when the number of data points in the time series with consecutive sample quality scores below a preset quality threshold exceeds a preset number, a sample quality early warning notification is sent to the corresponding screening institution, and the early warning judgment of the institution's data is suspended until the quality problem is resolved and confirmed.

[0018] A large-scale regional tumor early screening, warning, and testing data management system includes: The data standardization subsystem is equipped with a data access interface, format conversion module, unit conversion module, item mapping module, and quality labeling module. It is used to convert the original test data of screening institutions in various regions into standardized records with unified units and unified item codes and add quality scores. The individual longitudinal record management subsystem is equipped with an identity recognition module, a record database, and a baseline parameter calculation module. It is used to merge standardized records into the individual longitudinal records of the corresponding examinees and calculate the individual's historical baseline parameters. The trend anomaly detection subsystem is equipped with a time series construction module, a trend slope calculation module, a drift calculation module, and an anomaly judgment module. It is used to calculate the trend slope and standardized drift of time series data in individual longitudinal archives and perform two-dimensional early warning judgment. The graded early warning response subsystem is equipped with an early warning rule configuration module, an early warning level classification module, and an early warning information push module. It is used to match response measures according to the early warning level and push early warning information to the corresponding institution or clinical system. The four subsystems are connected by a data bus. Data flows sequentially through the data standardization subsystem, the individual longitudinal file management subsystem, and the trend anomaly detection subsystem. Early warning signals are pushed to the corresponding institutions by the hierarchical early warning response subsystem.

[0019] The beneficial effects of this invention are: First, this invention transforms heterogeneous cross-institutional test data into standardized records with unified dimensions and unified item codes, constructing strict data normalization constraints and quality control benchmarks using unit conversion coefficients and quality scoring rules. Faced with complex disturbances from multi-source data aggregation, the data standardization subsystem achieves deep nesting of test values ​​and quality scoring weights. Compared to the cross-institutional comparison biases generated by traditional direct data connection systems, this invention successfully eliminates heterogeneous noise caused by system differences or sample preprocessing defects within the same time series, stabilizing the merging success rate of multi-institutional longitudinal archives within a high range. Simultaneously, by triggering commands for extracting stable-period datasets and calculating baseline parameters in individual longitudinal archives, the system applies baseline mean and fluctuation range support tailored to the individual's historical physiological level. This individualized parameter effectively offsets background noise introduced by population physiological differences, maintaining the assessment benchmark for early tumor markers within the individualized normal fluctuation range, and avoiding the omission of early weak abnormal signals due to rigid static threshold determination mechanisms.

[0020] Faced with the quantitative measurement requirements of large-scale early cancer screening, early warning, and subsequent intervention, this system's early warning judgment results, relying on its anti-artifact bias correction mechanism and highly continuous time series network, establish an excellent benchmark for early anomaly transmission. Upon receiving newly added test data, the underlying trend anomaly detection algorithm can highly sensitively extract the trend slope of continuous test data and simultaneously calculate the standardized drift relative to the individual baseline. This invention extracts and quantifies the two-dimensional evolutionary load data of the time series, comprehensively evaluates trend anomaly characteristics, and triggers a tiered early warning response. Compared to traditional single-point early warning, this two-dimensional combined judgment mechanism, after offsetting cross-institutional instrument system biases, improves the accuracy of predicting early tumor trend changes, avoids unnecessary excessive clinical intervention, and drives adaptive correction of regional thresholds through continuous follow-up feedback. Attached Figure Description

[0021] Figure 1 This is the overall system architecture diagram of the present invention.

[0022] Figure 2 This is a flowchart of the data standardization process of the present invention.

[0023] Figure 3 This is a flowchart of the individual longitudinal profile identity matching process of the present invention.

[0024] Figure 4 This is a schematic diagram of the dual-dimensional early warning judgment of trend slope and drift amount of the present invention.

[0025] Figure 5 This is a schematic diagram illustrating the fitting effect of individual AFP time series and trend in this invention.

[0026] Figure 6 This is a comparison chart of the regional differential threshold calibration effect of the present invention.

[0027] Figure 7 This is a comparison chart of data before and after the cross-mechanism deviation correction of the present invention.

[0028] Figure 8 This is a flowchart of the follow-up feedback iteration process of the early warning model of the present invention. Detailed Implementation

[0029] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example 1

[0031] like Figure 1 As shown, the large-span regional tumor early screening and early warning system and laboratory data management system is deployed on the central server of the regional health information platform, establishing encrypted data transmission links with the laboratory information systems of screening institutions in each region via dedicated internet lines or VPN channels. Combined with... Figure 1 As shown in the overall architecture diagram, the system consists of four parts: a data standardization subsystem, an individual longitudinal record management subsystem, a trend anomaly detection subsystem, and a tiered early warning response subsystem. These four parts achieve asynchronous and decoupled communication through a message middleware as a data bus, ensuring unidirectional data processing flow and avoiding interference from reverse data backflow into existing processing links. Specifically, the testing systems of each regional screening institution are connected to the data access interface of the data standardization subsystem as external input terminals; the tiered early warning response subsystem connects to the workstation systems of the screening institutions and the HIS systems of clinical departments, serving as the output terminal for early warning information.

[0032] The data standardization subsystem is used to eliminate format barriers between multi-source heterogeneous data. For example... Figure 2 The data standardization process shown involves the data access interface receiving raw test data and loading the institution's configuration file. This file pre-configured parameters include the institution's geographic region code, a list of commonly used test items, and conversion coefficients. The format conversion module maps fields according to the configuration file, the unit conversion module performs unit conversion, and the item mapping module performs standard code replacement. Additionally, the quality labeling module reads parameters such as the time interval between sample collection and testing completion, hemolysis index, and sample type to calculate the sample quality score. After standardization, the system generates a standardized record containing the system standard code, unified unit value, source identifier, and quality score, and pushes it to the message queue.

[0033] The individual longitudinal record management subsystem is responsible for collecting the examinee's data. For example... Figure 3The individual longitudinal profile identity matching process shown involves the identity recognition module extracting the ID card number from standardized records and performing a query in the profile database. If the primary key matches and all three fields (name, date of birth, etc.) are verified to be consistent, the record is appended to the existing profile; if the primary key matches but auxiliary fields are questionable, it is written to the pending verification queue; if no match is found, a new profile is created and subsequent baseline calculations are triggered. The profile database uses the examinee's unique identifier as the primary key and stores the examination records and the calculated individual historical baseline mean and fluctuation range in chronological order.

[0034] The trend anomaly detection subsystem performs time-series-based risk assessment. The time-series construction module extracts continuous, valid detection data from the archive to construct a dataset. The trend slope calculation module performs a weighted least squares linear fit on the data point set to obtain the slope. The drift calculation module calculates the degree of deviation of the current value from the baseline. Figure 4 As shown, the anomaly detection module compares the points on the horizontal axis (absolute value of trend slope) and the vertical axis (absolute value of standardized drift) with the threshold lines (T_slope and T_drift). Falling into the upper right quadrant triggers a first-level warning, falling into the lower right or upper left quadrant triggers a second-level warning, and falling into the near-boundary zone (within the T_near range) inside the threshold lines triggers a third-level warning.

[0035] The tiered early warning response subsystem is responsible for distributing early warnings and providing closed-loop feedback. The early warning rule configuration module stores threshold parameters differentiated by geographical region. Based on the judgment results, the early warning information push module generates and pushes early warning reports, re-inspection recommendations, or follow-up plan adjustment notices to the corresponding institutions. Figure 8 The illustrated early warning model follow-up feedback iteration process includes a follow-up result feedback interface that receives clinical diagnoses or follow-up conclusions and writes them into the result field. The system periodically calculates various sensitivities and specificities, generates a threshold calibration report, and updates it after administrator review, forming a complete closed loop from data acquisition to parameter correction. Example 2

[0036] In this embodiment, the method for managing early cancer screening and testing data across large regions is implemented in four sequential stages: data standardization, individual record merging and baseline establishment, trend anomaly detection, and graded early warning response.

[0037] During the data standardization process, the testing information systems of screening institutions in each region push the test results data to the data access interface after testing is completed, according to a preset transmission protocol. Alternatively, administrators can upload batch data files to the import interface. After receiving the raw test data, the data access interface reads the institution identifier carried in the data and queries the corresponding institution configuration file, loading it into the context of this processing task. Subsequently, the format conversion module parses and extracts fields from the raw test data according to the field mapping rules in the institution configuration file, mapping each field of the raw test data structure to the corresponding field of the system's internal standard data structure, generating structured intermediate records. The dimension conversion module further performs multiplication conversion on the values ​​of each test item in the intermediate records according to the unit conversion coefficients of each item in the configuration file, uniformly writing them into the standard dimension value field, while retaining the original values ​​and units for subsequent traceability verification. The item mapping module then replaces the local item codes in the intermediate records with the system's standard item codes according to a preset item code lookup table. To ensure the reliability of data entering the analysis phase, the quality labeling module reads the sample collection timestamp, test completion timestamp, hemolysis index field (if present in the report), and sample type field from the intermediate records, and calculates the sample quality score according to preset quality scoring rules. For example, a base score is awarded if the time interval between collection and testing is within a specified number of hours, and deductions are made for any excess; if the hemolysis index exceeds the limit or the sample type is incorrect, corresponding deductions are executed. The quality score and all calculation parameters are written to the additional field area of ​​the standardized record. Finally, the standardized record is written to the data buffer along with a processing batch identifier and pushed to the message queue of the individual longitudinal record management subsystem via the message middleware.

[0038] During the individual file merging and baseline establishment phase, the logical processing unit primarily handles the identification of individuals across different institutions and the automatic establishment of historical benchmarks. The identity recognition module consumes standardized records from the message queue, retrieving three identity identifiers: name, ID number, and date of birth. The logical processing unit performs a precise query in the file database using the ID number as the primary key. If the query matches, the name and date of birth are further compared. If all three identifiers match, it is considered continuous monitoring data belonging to the same examinee, and the record is directly appended to the examinee's file's verification record sequence, with the file's last activity time updated. If the query matches, but one or both the name and date of birth are inconsistent, an error is considered. To prevent baseline contamination due to examinee merging errors, the record is written to the pending verification task queue, automatic merging of this data is stopped, and a notification to the data management personnel's terminal to verify the data is pushed, awaiting manual processing by the data management personnel. When the primary key query fails, it is identified as a new member of the group, a unique examinee identifier is assigned to this examinee within the logical system, a new file is created, and the file status is set to "baseline pending establishment." After the file is updated, the baseline parameter calculation module determines the baseline establishment status of the current examinee's file and whether the quality score of newly added samples meets the initiation criteria. (For files with baselines yet to be established, it checks whether the number of valid records for the current test item (records whose quality scores meet quality control requirements) is at the initiation threshold (e.g., 3). When the threshold is reached, the module sorts the valid records for that test item in the file by test time from most recent to oldest, and concatenates those records where the absolute value of two test values ​​changes within the baseline range into a subset, which is the stable period dataset. The module calculates the arithmetic mean of all records in this stable period dataset, which is the individual's historical baseline mean; then it calculates the average of the absolute deviations of each test value from the mean, which is the individual's historical fluctuation range. The module then updates the file status.) The baseline is now established. When a baseline is established and a qualified new record arrives, the module compares the absolute value of the difference between the new record's test value and the individual's historical baseline mean to see if it exceeds 1.5 times the current fluctuation range. If so, the new record is rejected; otherwise, the new record is added to the stable period dataset, and the mean and fluctuation range of the dataset are recalculated, with the new mean and fluctuation range replacing the old ones. After baseline maintenance is complete, the module informs the trend anomaly detection subsystem via the internal bus that a sample arrival event has occurred for a certain test item for a certain examinee. The information provided includes the examinee's code and the test item code.

[0039] During the trend anomaly detection phase, the system focuses on capturing the slow, gradual increase in tumor markers at individual physiological levels. Upon receiving a trigger notification, the trend anomaly detection subsystem retrieves the subject's archival metadata and test result sequences from the archival database. For example... Figure 5As shown, the time series construction module extracts inspection records with satisfactory quality scores from each inspection item listed in the notification, within a specified time window (e.g., the most recent 365 days). Figure 5 As shown in the scatter plot, each actual detection value is assigned a different weight based on its quality score (visually represented by the point diameter in the figure; the higher the score, the larger the point diameter). The extracted data points are arranged in ascending order of inspection time, with the number of days since the first record on the horizontal axis and the detection value on the vertical axis to construct the data point set. If the number of valid records is less than the minimum calculation requirement, the item is marked as insufficient data and skipped; if the requirement is met, the trend slope calculation module performs a weighted least squares linear fit on the data point set, outputting... Figure 5 The trend line shown by the solid line represents the rate of change of the indicator (in standard units per day). Simultaneously, the drift calculation module extracts the most recent detection value and calculates its correlation with... Figure 5 The difference between the individual's historical baseline mean (shown by the dashed line) and the individual's historical fluctuation range (marked by the wavy line) yields the dimensionless standardized drift. If the historical fluctuation range is calculated as zero due to identical values, it is substituted into the system's preset minimum fluctuation range for calculation. Subsequently, the anomaly detection module retrieves multi-level threshold parameters applicable to the examinee's geographical region and the current test item. The system logically compares the calculated absolute value of the trend slope and the absolute value of the standardized drift with the corresponding thresholds. If both the slope and drift exceed the limits, the system determines the anomaly signal is significant and issues a Level 1 warning; if only one dimension exceeds the limit, it issues a Level 2 warning; if neither dimension exceeds the limit but the value falls within the set range approaching the threshold, it issues a Level 3 warning; if both are within the safe range, it is recorded as normal. The system combines the highest warning level of each indicator to generate a warning record containing a calculation timestamp, parameter details, and level conclusion, writes it to the warning database, and pushes it to the next stage. This method can intuitively identify situations where an early warning is triggered because the trend slope has significantly exceeded the individual threshold, even when the AFP test value has not yet exceeded the upper limit of the population reference interval (20 ng / mL).

[0040] During the tiered early warning response phase, the system executes differentiated intervention actions based on the early warning level and establishes a dynamic feedback mechanism. Upon receiving an early warning record, the tiered early warning response subsystem, if determined to be a Level 1 warning, extracts the examinee's basic information, time-series data from the past 12 months, and current trend calculation parameters. It then generates a standardized early warning report containing data tables and intervention recommendations using a pre-set template and pushes it concurrently to the corresponding screening institution's workstation and the designated clinical department's HIS system via the system interface, ensuring immediate clinical response. For Level 2 warnings, the system focuses on laboratory verification, generating only a re-examination recommendation notification containing current trend data and distributing it to the screening institution. For Level 3 warnings, the system automatically generates a follow-up plan adjustment notification by updating the recommended screening date field in the record (e.g., 60 days in advance). In addition to positive information distribution, the system simultaneously runs a feedback closed-loop task. Medical institutions transmit the examinee's final pathological diagnosis or follow-up screening conclusions through the feedback interface. The system aggregates early warning records with follow-up results at preset intervals. By comparing actual clinical diagnoses with previous early warning judgments, it calculates the sensitivity and specificity performance indicators of each region and each test item under the current threshold parameters. When an indicator deviates from the set lower limit, the system automatically generates a threshold calibration report containing suggested adjustment ranges. After confirmation by the authorized administrator, the updated threshold parameters take effect immediately and are applied to the calculation of the next batch of data. Example 3

[0041] This implementation method uses a cohort of high-risk individuals for liver cancer from seven regional screening centers across three cities in a province as the validation scenario, spanning 18 months. A total of 4820 participants were included, aged 35 to 75 years, all of whom were high-risk individuals with positive hepatitis B surface antigen or a history of cirrhosis. The seven screening centers were located in city A (3 centers), city B (2 centers), and city C (2 centers). The testing environment exhibited heterogeneity: AFP reporting units existed in both ng / mL and μg / L formats, and local coding for the test items existed in three independent systems.

[0042] In data standardization and identity matching verification, the system processed 14,630 original inspection records in the first month, with 100% accuracy in both dimension conversion and item coding mapping. In the identity matching stage, the three-field accurate matching success rate was 98.6% (14,405 / 14,630), and most of the records to be verified were manually confirmed to belong to the same examinee, indicating that the identity merging strategy based on multi-dimensional fields has high engineering reliability. Regarding baseline establishment, by the end of the sixth month, 81.2% (3,916 people) of examinees had completed the initial individual baseline establishment containing three or more valid records, confirming that the stable period dataset extraction method can operate effectively within the actual business cycle.

[0043] like Figure 7As shown, the system effectively corrected for cross-institutional testing bias. After sending the same batch of AFP standard quality assessment samples (reference value 9.50 ng / mL) to 7 screening centers, as shown... Figure 7 As shown in the left-hand bar chart, there is a significant positive deviation between center 3 (10.10 ng / mL) and center 6 (10.30 ng / mL). The system automatically calculates and applies deviation coefficients (0.9406 and 0.9223, respectively). Taking a subject with consecutive tests at centers 3 and 6 as an example, the trend slope calculated after correcting for the first two test data (10.10 and 10.30 ng / mL) is +0.033 ng / mL / day; Figure 7 As shown on the right, after applying bias correction, the two values ​​were corrected to a stable 9.50 ng / mL, and the trend slope returned to zero (0.000 ng / mL / day). This validation demonstrates that the system can effectively eliminate spurious trend interference introduced by the testing instrument system error and avoid false alarms caused by hardware differences.

[0044] In the performance validation of the early warning system, 43 newly diagnosed liver cancer cases were included within 18 months. Using the traditional cross-sectional threshold comparison method (AFP ≥ 20 ng / mL), the recognition rate within 6 months prior to diagnosis was 65.1% (28 / 43), while the number of false triggers in the normal population was 93. Using the dual-dimensional trend drift detection method of this invention, 37 out of the 43 cases triggered the first or second level early warning, increasing the recognition rate to 86.0%; simultaneously, false triggers in normal subjects decreased by 24.6%. The few cases not identified by the system in advance were mainly attributed to insufficient data due to follow-up dropout or belonging to a low AFP expression subtype.

[0045] Figure 6 The effectiveness of configuring independent thresholds for different regions was verified. Combined with epidemiological data, the historical AFP positivity rate in city A (8.3%) was significantly higher than that in cities B and C. The system assigned a stricter trend slope threshold (0.08 ng / mL / day) to city A, and a more lenient threshold (0.12 and 0.11 ng / mL / day) to cities B and C, respectively. Figure 6 As shown in the bar chart, compared to the control group using a uniform threshold (0.10 ng / mL / day) across the province, the differentiated approach reduced the false trigger rate of Level 1 warnings in City A by 18.3%, while the missed diagnosis rates in City B and City C increased by 9.2% and 7.8%, respectively, confirming the feasibility and necessity of dynamic parameter adjustment based on regional baselines. Based on a closed-loop feedback mechanism, at the end of the second quarter, the system automatically generated suggestions based on 24 clinical follow-up data, adjusting the drift threshold for City B from 1.8 to 2.0. Subsequently, the false trigger rate of Level 2 warnings in this region further decreased by 11.3%, verifying the effectiveness of the adaptive parameter evolution design.

[0046] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for early cancer screening and early warning and test data management across a large geographical area, characterized in that, Includes the following steps: Receive raw test data from screening institutions in various regions and convert the raw test data into standardized records with unified dimensions and unified item codes; The standardized records are matched according to the examinee's identity information, and cross-institutional standardized records belonging to the same examinee are merged into the corresponding individual longitudinal file; Historical testing data are extracted from the individual's longitudinal archives to calculate the individual's historical baseline mean and individual historical fluctuation range for the test items; The trend slope is calculated based on recent continuous detection data in the individual's longitudinal profile, and the standardized drift is calculated based on the current detection value, the individual's historical baseline mean, and the individual's historical fluctuation range. Based on the comparison results of the trend slope and the preset trend slope threshold, and the comparison results of the standardized drift amount and the preset drift amount threshold, a two-dimensional combined judgment is performed to trigger the corresponding level of warning.

2. The method according to claim 1, characterized in that, The specific method for converting the original test data into standardized records is as follows: conversion is performed based on the unit conversion coefficient and project code comparison table pre-set by each institution; and the source institution identifier field, conversion coefficient version number field, and data receipt timestamp field are written into the standardized records.

3. The method according to claim 1, characterized in that, The process of performing a two-dimensional combined judgment to trigger a warning of the corresponding level specifically includes: When both the trend slope and the standardized drift exceed their respective thresholds, a Level 1 warning is triggered, and a clinical referral recommendation is output. A second-level warning is triggered when only the trend slope or only the standardized drift exceeds its respective threshold, and a re-examination suggestion is output. When neither the trend slope nor the standardized drift exceeds the threshold, but the difference between either of them and the threshold is less than the preset threshold, a third-level warning is triggered, and a suggestion to shorten the follow-up interval is output.

4. The method according to claim 3, characterized in that, The standardized record also includes a sample quality score field, which is calculated based on the sample collection time interval, hemolysis index field, and sample type field. When calculating the trend slope, data points with sample quality scores below a preset quality threshold are assigned low weight coefficients, while data points with sample quality scores that meet preset requirements are assigned normal weight coefficients, and the trend slope is calculated in a weighted manner.

5. The method according to claim 1, characterized in that, The extraction of historical testing data to calculate the individual's historical baseline mean and individual historical fluctuation range specifically involves: The stable period dataset is defined as the continuous record segment in the first batch of continuous detection data after the establishment of an individual's file, in which the change range between two adjacent detection values ​​does not exceed the preset stability judgment range. When the adjacent change range of a newly added record in the file meets the stability judgment range condition, the new record is included in the stable period dataset and the individual's historical baseline mean and individual historical fluctuation range are recalculated.

6. The method according to claim 1, characterized in that, The matching of standardized records according to the examinee's identity information specifically involves: The matching key is a combination of the examinee's name, ID number, and date of birth. If all three fields match, the examinee is identified as the same person. If the ID number matches but the name or date of birth differs, the record is marked as pending verification and is not merged into the existing file. If the match fails, a new individual file is created for the examinee.

7. The method according to claim 1, characterized in that, It also includes cross-agency inspection deviation correction steps: The system periodically pushes standard quality assessment sample inspection tasks to each access institution, collects the inspection results of each institution on the quality assessment samples, calculates the deviation between the inspection results of each institution and the reference value, and stores it as the system deviation coefficient of each institution; in the conversion of the data standardization record, the system deviation coefficient corresponding to each institution is used to correct the deviation of the transmitted data of that institution.

8. The method according to claim 1, characterized in that, The preset trend slope threshold and preset drift threshold are differentiated threshold parameters configured based on different geographical regions and different test items; The method also includes: writing the follow-up diagnosis results of subjects who have triggered the warning into their corresponding individual files; The sensitivity and specificity of each warning level are statistically analyzed according to a preset cycle. When the sensitivity or specificity is lower than the preset lower limit, the threshold parameters of the corresponding area are recalibrated based on the updated regional positive detection rate.

9. The method according to claim 4, characterized in that, When the number of consecutive data points with sample quality scores below the preset quality threshold in the time series exceeds the preset number, a sample quality early warning notification is sent to the corresponding screening institution, and the early warning judgment for the institution's data is suspended until the quality problem is resolved and confirmed.

10. A large-span regional tumor early screening, warning, and testing data management system, characterized in that, include: The data standardization subsystem is equipped with a data access interface, format conversion module, unit conversion module, item mapping module, and quality labeling module. It is used to convert the original test data of screening institutions in various regions into standardized records with unified units and unified item codes and add quality scores. The individual longitudinal record management subsystem is equipped with an identity recognition module, a record database, and a baseline parameter calculation module. It is used to merge standardized records into the individual longitudinal records of the corresponding examinees and calculate the individual's historical baseline parameters. The trend anomaly detection subsystem is equipped with a time series construction module, a trend slope calculation module, a drift calculation module, and an anomaly judgment module. It is used to calculate the trend slope and standardized drift of time series data in individual longitudinal archives and perform two-dimensional early warning judgment. The graded early warning response subsystem is equipped with an early warning rule configuration module, an early warning level classification module, and an early warning information push module. It is used to match response measures according to the early warning level and push early warning information to the corresponding institution or clinical system. The four subsystems are connected by a data bus. Data flows sequentially through the data standardization subsystem, the individual longitudinal file management subsystem, and the trend anomaly detection subsystem. Early warning signals are pushed to the corresponding institutions by the hierarchical early warning response subsystem.