Tumor follow-up data intelligent processing service system

By designing an intelligent processing service system for tumor follow-up data, the problem of insufficient data quality control in the existing technology is solved, and the efficient, accurate processing and quality control of data is achieved, which significantly improves the management efficiency and accuracy of tumor follow-up.

CN120183590APending Publication Date: 2025-06-20HANGZHOU SHENGWU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510254542.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing tumor follow-up data processing methods have problems with insufficient data quality control, resulting in errors, missing or inconsistent data, affecting clinical decision-making and treatment plan formulation.

Method used

A tumor follow-up data intelligent processing service system is designed, including a data input module, a data preprocessing module, a data quality evaluation module and a data response module. The system receives multi-channel data through a unified interface, performs standardized processing, automatically detects missing values ​​and abnormal data, evaluates data quality, and triggers an adaptive correction early warning mechanism based on abnormal detection scores.

Benefits of technology

It improves the control efficiency of data quality, significantly reduces the risk of human errors and omissions, ensures the efficiency and accuracy of tumor follow-up work, and thus improves the follow-up effect and patient management quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183590A_ABST
    Figure CN120183590A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent processing service system for tumor follow-up visit data, which relates to the technical field of data processing, and is characterized in that when the system runs, multi-channel data from medical equipment, patient self-reports, experimental detection reports and the like are received through a uniform interface, and the data are subjected to standardized processing, so that the data format is uniform, and subsequent analysis is facilitated. Through the data preprocessing module, the system can automatically detect missing values and abnormal data, so that the quality problem possibly existing in follow-up visit data is eliminated, and the reliability of the data is ensured. And then, the data quality evaluation module comprehensively evaluates the quality of the data by calculating key indexes such as a data missing rate Dmiss, a consistency score Dcons and an accuracy score Dac, and generates a quality evaluation vector Q, thereby providing a basis for subsequent anomaly detection and correction. The risk of human errors and omission is remarkably reduced, and the high efficiency and accuracy of tumor follow-up visit work are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and specifically to an intelligent processing service system for tumor follow-up data. Background Art

[0002] Tumor treatment and management is a complex and crucial direction in the medical field. With the rapid development of medical technology and big data, the follow-up management of tumor patients has gradually shifted from traditional manual records to an intelligent, data-driven precision medicine model. Tumor follow-up not only involves the treatment process of patients, but also includes dynamic monitoring of various aspects such as the progression of their condition, treatment response, and recurrence risk. Therefore, the efficient collection and accurate processing of tumor follow-up data are important components of the clinical decision support system.

[0003] Although tumor follow-up management has received increasing attention in the medical system, the existing methods for processing follow-up data still have many deficiencies, especially in data quality control. Currently, many hospitals and clinics rely on manual methods or simple electronic health record systems to collect and archive follow-up data. Such systems are usually unable to effectively handle data errors, omissions, or inconsistencies. For example, in the symptom data reported by patients, examination results, and imaging data, there may be problems such as data omission, information entry errors, and information inconsistencies between different departments. These problems not only lead to incomplete follow-up records of patients, but may also affect doctors' judgments on the treatment progress of patients, and even affect the formulation of subsequent treatment plans for patients. In addition, due to the lack of a systematic data review and cleaning mechanism, many hospitals still rely on the experience of medical staff to judge the validity of data in actual operations. This method relying on manual judgment is prone to subjective judgment biases, thereby affecting the accuracy and credibility of follow-up data. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides an intelligent processing service system for tumor follow-up data, which solves the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: An intelligent processing service system for tumor follow-up data, comprising a data input module, a data preprocessing module, a data quality assessment module, and a data response module;

[0006] The data input module is used to receive follow-up data from multiple channels including medical devices, patient self-reports, and experimental test reports, and format the follow-up data through a unified interface to form a standardized data set X;

[0007] The data preprocessing module performs preprocessing tasks of missing value detection and abnormal data marking on the data set X, and obtains the data vector Xclean after the completion of the preprocessing tasks;

[0008] The data quality assessment module performs quality assessment on the data vector Xclean, obtains the data missing rate Dmiss, the consistency score Dcons, and the accuracy score Dac, and integrates the data missing rate Dmiss, the consistency score Dcons, and the accuracy score Dac to obtain the quality assessment vector Q;

[0009] The data response module generates an anomaly detection score Sabn based on the obtained quality assessment vector Q, corrects the data vector Xclean based on the anomaly detection assessment Sabn to generate a response anomaly detection score XSabn, and triggers an adaptive response mechanism for data correction warning based on the response anomaly detection score XSabn.

[0010] Preferably, the data input module includes a multi-source data fusion unit and a data formatting unit;

[0011] The multi-source data fusion unit receives data of tumor patients from multiple channels including medical devices, patient self-reports, and laboratory test reports, and performs preliminary integration and marking. The marking includes marking the source channels of the data to form a preliminary data set D0;

[0012] The preliminary data set D0 includes medical device data Dd, patient report data Dp, and laboratory test data Dl;

[0013] The data formatting unit formats the obtained preliminary data set D0, unifies the timestamp formats and numerical dimensions of the medical device data Dd, the patient report data Dp, and the laboratory test data Dl, and forms the data set X;

[0014] Among them, the numerical dimensions are unified using a normalization processing method.

[0015] Preferably, the data preprocessing module includes a missing value detection unit and an abnormal data marking unit;

[0016] The missing value detection unit performs the first preprocessing task of missing value detection on the data set X, identifies missing values for each data item in the data set X. The missing values include NaN, NULL, empty strings, and 0. When a data item with a missing value is identified, mean filling, median filling, most frequent value filling, and interpolation methods are used to fill the missing values, and the data item missing Bjsj = 1 is marked;

[0017] When the total number of missing Bjsj in the data items is greater than the preset data item missing ratio threshold Bthe, delete the data items with missing Bjsj and prompt to re-enter the follow-up data.

[0018] Preferably, the abnormal data marking unit performs a second preprocessing task of abnormal data marking on the data set X after performing the first preprocessing task of missing value detection, identifies abnormal data for each data item in the data set X, marks the data item with abnormal data as A = 1 when an abnormal data item is identified, and reorganizes the data set X after performing the second preprocessing task of abnormal data marking to obtain a data group with missing Bjsj and abnormal data marking A, which is marked as the data vector Xclean;

[0019] Among them, the identification of abnormal data includes using the Z-score method and the IQR method to identify abnormal data for each data item. When the total number of abnormal data markings A is greater than the preset data item missing ratio threshold Bthe, delete the data items with abnormal data marking A and prompt to re-enter the follow-up data;

[0020] When there is a data item with abnormal data marking A = 1 in the data set X, perform abnormal marking correction on the data item with abnormal data marking A = 1, and the abnormal marking correction includes using the mean value of the data item for correction and replacement.

[0021] Preferably, the data quality assessment module includes a missing rate calculation unit, a consistency calculation unit, and an accuracy calculation unit;

[0022] The missing rate calculation unit calculates the proportion of missing values for each column feature vector in the data vector Xclean to obtain the data missing rate Dmiss;

[0023] The data missing rate Dmiss is obtained through the following calculation formula:

[0024]

[0025] In the formula, Dmiss(j) represents the missing rate of the j-th column feature in the data vector Xclean, n represents the total number of data points of the j-th column feature, Xclean(j, i) represents the i-th data in the j-th column feature of the data vector Xclean, NaN represents an invalid value, and 1{Xclean(j, i) = NaN} represents an exponential function. When the i-th data Xclean(j, i) in the j-th column feature of the data vector Xclean is an invalid value NaN, it returns 1, otherwise it returns 0.

[0026] Preferably, the consistency calculation unit measures the internal consistency of the data by measuring the vectors of each column feature in the data vector Xclean. The measurement method includes using the coefficient of variation to measure the internal consistency of the data, and obtaining a consistency score Dcons.

[0027] The consistency score Dcons is obtained through the following calculation formula:

[0028]

[0029] In the formula, Dcns(j) represents the consistency score of the j-th column feature in the data vector Xclean, and μ(j) represents the mean value of the j-th column feature.

[0030] Preferably, the accuracy calculation unit compares the vectors of each column feature in the data vector Xclean. Specifically, it compares with the historical mean value to obtain an accuracy score Dac, and integrates the data missing rate Dmiss, the consistency score Dcons, and the accuracy score Dac to obtain a quality assessment vector Q.

[0031] The accuracy score Dac is obtained through the following calculation formula:

[0032]

[0033] In the formula, μstd(j) represents the historical mean value of the j-th column feature.

[0034] Preferably, the data response module includes an anomaly detection unit and a correction response unit;

[0035] The anomaly detection unit generates an anomaly detection score Sabn based on the obtained quality assessment vector Q, and corrects the data vector Xclean based on the anomaly detection assessment Sabn. The correction is performed by comparing the anomaly detection assessment Sabn with a preset correction threshold Xthe, and triggering the correction according to the comparison result. When the correction is triggered, a response anomaly detection score XSabn is generated.

[0036] The anomaly detection score Sabn is obtained through the following calculation formula:

[0037] Sabn(j) = s1 * Dmiss(j) + s2 * Dcons(j) + s3 * Dac(j);

[0038] In the formula, Sabn(j) represents the anomaly detection score of the j-th column feature, and s1, s2, and s3 respectively represent the preset weight values of the data missing rate Dmiss, the consistency score Dcons, and the accuracy score Dac, and s1 + s2 + s3 = 1. The specific values are set by the user.

[0039] Preferably, the triggering is determined by the following comparison method:

[0040] When the abnormal detection evaluation Sabn ≥ the correction threshold Xthe, the correction mechanism is triggered to perform the correction task of the abnormal detection evaluation Sabn on the data vector Xclean;

[0041] When the abnormal detection evaluation Sabn < the correction threshold Xthe, the correction mechanism is not triggered, and the correction task of the abnormal detection evaluation Sabn on the data vector Xclean is not performed;

[0042] The response abnormal detection score XSabn is obtained by the following calculation formula:

[0043] XSabn(j,i) = Xclean(j,i) + λ(j) * Sabn(j) * ΔT(j);

[0044] In the formula, λ(j) represents the correction factor of the j-th column feature, and △T(j) represents the adjustment amount of the j-th column feature, which is specifically set by the adjustment amount when the data set X performs the first preprocessing task of missing value detection and the second preprocessing task of abnormal data marking.

[0045] Preferably, the correction response unit performs a secondary comparison of the response abnormal detection score XSabn with the correction threshold Xthe, and according to the comparison result, triggers an adaptive response mechanism for data correction warning;

[0046] The adaptive response mechanism is obtained by the following comparison method:

[0047] When the response abnormal detection score XSabn ≥ the correction threshold Xthe, the correction mechanism is triggered for the second time to prompt the abnormality of the validity of the patient's follow-up data, and the patient information is re-entered within a fixed period;

[0048] When the response abnormal detection score XSabn < the correction threshold Xthe, the correction mechanism is not triggered, it is prompted that the processing of the patient's follow-up data is completed, the data is qualified, and the patient's follow-up data is stored.

[0049] The present invention provides an intelligent processing service system for tumor follow-up data, having the following beneficial effects:

[0050] (1) When the system is running, it receives multi-channel data from medical devices, patient self-reports, experimental test reports, etc. through a unified interface, and standardizes this data to ensure a unified data format for subsequent analysis. Through the data preprocessing module, the system can automatically detect missing values and abnormal data, thus eliminating possible quality problems in the follow-up data and ensuring the reliability of the data. Subsequently, the data quality assessment module comprehensively evaluates the data quality by calculating key indicators such as the data missing rate Dmiss, the consistency score Dcons, and the accuracy score Dac, and generates a quality assessment vector Q to provide a basis for subsequent anomaly detection and correction. On this basis, the data response module automatically corrects the data based on the anomaly detection score SabnS and triggers an adaptive correction warning mechanism through the response score XSabn, and notifies the corresponding data correction measures in a timely manner according to different anomaly degrees. Compared with the traditional manual follow-up method, through the intelligent data processing and response mechanism, this system not only improves the control efficiency of data quality, but also significantly reduces the risks of human errors and omissions, ensuring the efficiency and accuracy of tumor follow-up work, thus helping to improve the follow-up effect and patient management quality.

[0051] (2) By synthesizing these three scores, the system can generate a quality assessment vector Q, providing a clear quantitative standard for the overall quality of the data, and helping the system to accurately correct and process the data. The advantage of this module is that it can timely detect and correct problems in the data through multi-dimensional quality assessment, avoiding the impact of data quality problems on the tumor follow-up effect, and providing reliable data support for subsequent treatment decisions and patient management.

[0052] (3) Through the collaborative work of the anomaly detection unit and the correction response unit, it ensures the timely identification and correction of anomalies in the follow-up data. The anomaly detection unit first generates an anomaly detection score Sabn based on the quality assessment vector Q. This score comprehensively considers the data missing rate Dmiss, the consistency score Dcons, and the accuracy score Dac, and determines whether to trigger the correction mechanism by comparing with the preset correction threshold Xthe. When the anomaly detection score exceeds the correction threshold, the system automatically starts the correction task, corrects the data vector Xclean and generates a response anomaly detection score XSabn. The correction response unit conducts a secondary comparison of the response score again to ensure the effectiveness of data correction. It can reduce unnecessary repeated data entry on the premise of ensuring data quality, thereby improving the efficiency of the patient follow-up process and the accuracy of the data, ensuring the efficient management and processing of the data, and greatly enhancing the intelligence and automation level of patient follow-up. Description of the Drawings

[0053] Figure 1 It is a schematic block diagram of an intelligent processing service system for tumor follow-up data of the present invention;

[0054] Figure 2 This is a schematic diagram of the data processing flow of an intelligent processing service system for tumor follow-up data of the present invention. Specific embodiments

[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0056] Embodiment 1

[0057] The present invention provides an intelligent processing service system for tumor follow-up data. Please refer to Figure 1 , which includes a data input module, a data preprocessing module, a data quality assessment module, and a data response module;

[0058] The data input module is used to receive follow-up data from multiple channels such as medical devices, patient self-reports, and experimental test reports, and format the follow-up data through a unified interface to form a standardized data set X;

[0059] The data preprocessing module performs preprocessing tasks of missing value detection and abnormal data marking on the data set X, and obtains a data vector Xclean after the completion of the preprocessing tasks;

[0060] The data quality assessment module performs quality assessment on the data vector Xclean, obtains a data missing rate Dmiss, a consistency score Dcons, and an accuracy score Dac, and integrates the data missing rate Dmiss, the consistency score Dcons, and the accuracy score Dac to obtain a quality assessment vector Q;

[0061] The data response module generates an anomaly detection score Sabn based on the obtained quality assessment vector Q, corrects the data vector Xclean based on the anomaly detection assessment Sabn to generate a response anomaly detection score XSabn, and triggers an adaptive response mechanism for data correction warnings based on the response anomaly detection score XSabn.

[0062] In this embodiment, multi-channel data from medical devices, patient self-reports, and laboratory test reports are received through a unified interface, and these data are standardized to ensure a unified data format for subsequent analysis. Through the data preprocessing module, the system can automatically detect missing values and abnormal data, thereby eliminating possible quality problems in the follow-up data and ensuring the reliability of the data. Subsequently, the data quality assessment module comprehensively evaluates the data quality by calculating key indicators such as the data missing rate Dmiss, the consistency score Dcons, and the accuracy score Dac, and generates a quality assessment vector Q, providing a basis for subsequent anomaly detection and correction. On this basis, the data response module automatically corrects the data based on the anomaly detection score SabnS, and triggers an adaptive correction warning mechanism through the response score XSabn, and notifies corresponding data correction measures in a timely manner according to different degrees of anomalies. Compared with the traditional manual follow-up method, through the intelligent data processing and response mechanism, this system not only improves the control efficiency of data quality, but also significantly reduces the risks of human errors and omissions, ensuring the efficiency and accuracy of tumor follow-up work, and thus contributing to improving the follow-up effect and the quality of patient management.

[0063] Embodiment 2

[0064] This embodiment is an explanatory description based on Embodiment 1. Please refer to Figure 1 , specifically: The data input module includes a multi-source data fusion unit and a data formatting unit;

[0065] The multi-source data fusion unit receives data of tumor patients from multiple channels including medical devices, patient self-reports, and laboratory test reports, and performs preliminary integration and marking. The marking includes marking the source channels of the data to form a preliminary data set D0;

[0066] The preliminary data set D0 includes medical device data Dd, patient report data Dp, and laboratory test data Dl;

[0067] The data formatting unit formats the obtained preliminary data set D0, unifies the timestamp formats and numerical dimensions of the medical device data Dd, patient report data Dp, and laboratory test data Dl, and forms a data set X;

[0068] Among them, the numerical dimensions are unified using a normalization processing method.

[0069] The data preprocessing module includes a missing value detection unit and an abnormal data marking unit;

[0070] The missing value detection unit performs the first preprocessing task of missing value detection on the data set X, identifies missing values for each data item in the data set X, where the missing values include NaN, NULL, empty strings, and 0. When a data item with a missing value is identified, mean filling, median filling, most frequent value filling, and interpolation are used to fill the missing value, and the data item is marked with missing data Bjsj = 1;

[0071] When the total number of data items with missing data Bjsj is greater than the preset data item missing ratio threshold Bthe, the data items with missing data Bjsj are deleted, and a prompt is given to re-enter the follow-up data.

[0072] The abnormal data marking unit performs the second preprocessing task of abnormal data marking on the data set X after performing the first preprocessing task of missing value detection, identifies abnormal data for each data item in the data set X, marks the data item with abnormal data A = 1 when an abnormal data item is identified, reorganizes the data set X after performing the second preprocessing task of abnormal data marking, and obtains a data group with missing data Bjsj and abnormal data marking A, which is marked as the data vector Xclean;

[0073] Among them, the identification of abnormal data includes using the Z-score method and the IQR method to identify abnormal data for each data item. When the total number of abnormal data markings A is greater than the preset data item missing ratio threshold Bthe, the data items with abnormal data marking A are deleted, and a prompt is given to re-enter the follow-up data;

[0074] When there is a data item with abnormal data marking A = 1 in the data set X, the abnormal data marking is corrected for the data item with abnormal data marking A = 1, and the abnormal marking correction includes using the mean of the data item for correction and replacement.

[0075] In this embodiment, the multi-source data fusion unit efficiently integrates multi-channel data from medical devices, patient self-reports, and laboratory test reports, and clearly marks the data sources, thereby forming a preliminary data set D0. This integration process ensures multi-dimensional coverage of the data and provides clear source information for subsequent unified formatting. The data formatting unit standardizes the timestamp format and numerical dimension of the data to ensure the consistency of data from different sources, and finally forms a data set X, enabling the data to be uniformly processed. Subsequently, the data preprocessing module further optimizes the data quality through a missing value detection unit and an abnormal data marking unit, automatically identifying and filling in missing values, and correcting them using multiple methods such as mean filling, median filling, most frequent value filling, and interpolation. The identification of abnormal data is automatically completed through the Z-score and IQR methods. The system can accurately mark the abnormal items in the data and delete or correct the abnormal data items according to the preset threshold. Finally, the data set Xclean is cleaned and corrected, providing a high-quality data basis for subsequent analysis and processing. Through these processing procedures, the system can not only automatically detect and correct missing values and abnormal data in the follow-up data, but also respond to data problems in a timely manner through the marking and correction mechanism, ensuring the high quality and high accuracy of the follow-up data, and greatly improving the reliability and management efficiency of tumor follow-up.

[0076] Embodiment 3

[0077] This embodiment is an explanatory description carried out in Embodiment 2. Please refer to Figure 1 , specifically: The data quality assessment module includes a missing rate calculation unit, a consistency calculation unit, and an accuracy calculation unit;

[0078] The missing rate calculation unit calculates the proportion of missing values for each column feature vector in the data vector Xclean to obtain the data missing rate Dmiss;

[0079] The data missing rate Dmiss is obtained through the following calculation formula:

[0080]

[0081] In the formula, Dmiss(j) represents the missing rate of the j-th column feature in the data vector Xclean, n represents the total number of data points in the j-th column feature, Xclean(j, i) represents the i-th data in the j-th column feature of the data vector Xclean, NaN represents an invalid value, and 1{Xclean(j, i) = NaN} represents an exponential function. When the i-th data Xclean(j, i) in the j-th column feature of the data vector Xclean is an invalid value NaN, it returns 1, otherwise it returns 0.

[0082] The consistency calculation unit measures the internal consistency of the data by measuring the vectors of each column feature in the data vector Xclean. The measurement method includes using the coefficient of variation to measure the internal consistency of the data, and obtaining a consistency score Dcons.

[0083] The consistency score Dcons is obtained through the following calculation formula:

[0084]

[0085] In the formula, Dcns(j) represents the consistency score of the j-th column feature in the data vector Xclean, and μ(j) represents the mean value of the j-th column feature.

[0086] The accuracy calculation unit compares the vectors of each column feature in the data vector Xclean. Specifically, it compares with the historical mean value to obtain an accuracy score Dac, and integrates the data missing rate Dmiss, the consistency score Dcons, and the accuracy score Dac to obtain a quality evaluation vector Q.

[0087] The accuracy score Dac is obtained through the following calculation formula:

[0088]

[0089] In the formula, μstd(j) represents the historical mean value of the j-th column feature.

[0090] In this embodiment, through the collaborative work of the missing rate calculation unit, the consistency calculation unit, and the accuracy calculation unit, a comprehensive analysis of each column feature in the data vector Xclean is realized. First, the missing rate calculation unit accurately calculates the missing value ratio Dmiss of each column feature based on the validity of the data, providing clear guidance for subsequent data correction. Next, the consistency calculation unit measures the internal consistency of each column feature through the coefficient of variation and calculates the consistency score Dcons. This process can reveal the volatility and reliability in the data, further improving the controllability of data quality. Finally, the accuracy calculation unit evaluates the accuracy of each column feature by comparing with the historical mean value and generates the accuracy score Dac to ensure the consistency of the data with the historical data. After comprehensively considering these three scores, the system can generate the quality evaluation vector Q, providing a clear quantitative standard for the overall quality of the data and helping the system to accurately correct and process the data. The advantage of this module is that it can detect and correct problems in the data in a timely manner through multi-dimensional quality evaluation, avoiding the impact of data quality problems on the tumor follow-up effect, and providing reliable data support for subsequent treatment decisions and patient management.

[0091] Example 4

[0092] This embodiment is explained in Embodiment 3. Please refer to Figure 1 and Figure 2 , specifically: The data response module includes an anomaly detection unit and a correction response unit;

[0093] The anomaly detection unit generates an anomaly detection score Sabn based on the obtained quality assessment vector Q, and corrects the data vector Xclean based on the anomaly detection assessment Sabn. The correction is performed by comparing the anomaly detection assessment Sabn with a preset correction threshold Xthe, and triggering the correction according to the comparison result. When the correction is triggered, a response anomaly detection score XSabn is generated;

[0094] The anomaly detection score Sabn is obtained through the following calculation formula:

[0095] Sabn(j) = s1 * Dmiss(j) + s2 * Dcons(j) + s3 * Dac(j);

[0096] In the formula, Sabn(j) represents the anomaly detection score of the j-th column feature, s1, s2, and s3 respectively represent the preset weight values of the data missing rate Dmiss, the consistency score Dcons, and the accuracy score Dac, and s1 + s2 + s3 = 1. The specific values are set by the user.

[0097] The triggering is specifically judged through the following comparison method:

[0098] When the anomaly detection assessment Sabn ≥ correction threshold Xthe, the correction mechanism is triggered, and the correction task of the data vector Xclean by the anomaly detection assessment Sabn is executed;

[0099] When the anomaly detection assessment Sabn < correction threshold Xthe, the correction mechanism is not triggered, and the correction task of the data vector Xclean by the anomaly detection assessment Sabn is not executed;

[0100] The response anomaly detection score XSabn is obtained through the following calculation formula:

[0101] XSabn(j,i) = Xclean(j,i) + λ(j) * Sabn(j) * ΔT(j);

[0102] In the formula, λ(j) represents the correction factor of the j-th column feature, and △T(j) represents the adjustment amount of the j-th column feature, which is specifically set by the adjustment amount when the data set X performs the first preprocessing task of missing value detection and the second preprocessing task of abnormal data marking.

[0103] The correction response unit makes a secondary comparison based on the response anomaly detection score XSabn and the correction threshold Xthe, and triggers an adaptive response mechanism for data correction warning according to the comparison result;

[0104] The adaptive response mechanism is obtained through the following comparison method:

[0105] When the response anomaly detection score XSabn ≥ the correction threshold Xthe, the correction mechanism is triggered again to prompt the abnormality of the validity of the patient follow-up data, and the patient information is re-entered within a fixed period;

[0106] When the response anomaly detection score XSabn < the correction threshold Xthe, the correction mechanism is not triggered, indicating that the processing of the patient follow-up data is completed, the data is qualified, and the patient follow-up data is stored;

[0107] For the specific data processing and evaluation comparison process of the triggering and the adaptive response mechanism, please refer to Figure 2 A schematic diagram of the data processing flow block diagram of an intelligent processing service system for tumor follow-up data.

[0108] In this embodiment, through the collaborative work of the anomaly detection unit and the correction response unit, the timely identification and correction of follow-up data anomalies are ensured. The anomaly detection unit first generates an anomaly detection score Sabn based on the quality assessment vector Q. This score comprehensively considers the data missing rate Dmiss, the consistency score Dcons, and the accuracy score Dac, and judges whether to trigger the correction mechanism by comparing with the preset correction threshold Xthe. When the anomaly detection score exceeds the correction threshold, the system automatically starts the correction task, corrects the data vector Xclean and generates a response anomaly detection score XSabn. The correction response unit makes a secondary comparison of the response score again to ensure the effectiveness of data correction. If the score reaches the preset threshold, the system will trigger a correction warning to prompt the patient to re-enter valid data, while if the score is lower than the threshold, the system considers the data qualified and completes the storage of the follow-up data. Through this precise adaptive response mechanism, the system can reduce unnecessary repeated data entry on the premise of ensuring data quality, thereby improving the efficiency of the patient follow-up process and the accuracy of data, ensuring the efficient management and processing of data, and greatly enhancing the intelligence and automation level of patient follow-up.

[0109] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirits of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A tumor follow-up data intelligent processing service system, characterized by: It includes a data input module, a data preprocessing module, a data quality assessment module and a data response module; The data input module is used to receive follow-up data from multiple channels such as medical equipment, patient self-reports and laboratory test reports, and format the follow-up data through a unified interface to form a standardized data set X; The data preprocessing module performs the preprocessing tasks of missing value detection and abnormal data marking on the data set X, and obtains the data vector Xclean after the preprocessing task is completed; The data quality assessment module performs quality assessment on the data vector Xclean to obtain a data missing rate Dmiss, a consistency score Dcons and an accuracy score Dac, and integrates the data missing rate Dmiss, the consistency score Dcons and the accuracy score Dac to obtain a quality assessment vector Q; The data response module generates an anomaly detection score Sabn based on the acquired quality assessment vector Q, and corrects the data vector Xclean based on the anomaly detection assessment Sabn to generate a response anomaly detection score XSabn, and triggers an adaptive response mechanism for data correction warning based on the response anomaly detection score XSabn.

2. The tumor follow-up data intelligent processing service system according to claim 1, characterized in that: The data input module includes a multi-source data fusion unit and a data formatting unit; The multi-source data fusion unit receives the data of tumor patients through multiple channels such as medical devices, patient self-reports and laboratory test reports, and performs preliminary integration and labeling, wherein the labeling includes labeling the source channels of the data, to form a preliminary data set D0; The preliminary data set D0 includes medical device data Dd, patient report data Dp and laboratory test data Dl; The data formatting unit formats the acquired preliminary data set D0, unifies the timestamp format and numerical dimension of the medical device data Dd, the patient report data Dp and the laboratory test data Dl, and forms a data set X; Among them, the numerical dimensions are unified using the normalization method.

3. The tumor follow-up data intelligent processing service system according to claim 1, characterized in that: The data preprocessing module includes a missing value detection unit and an abnormal data marking unit; The missing value detection unit performs a first preprocessing task of missing value detection on the data set X, identifies missing values ​​for each data item in the data set X, wherein the missing values ​​include NaN, NULL, an empty string, and 0. When a data item with missing values ​​is identified, the missing values ​​are filled using mean filling, median filling, most frequent value filling, and interpolation, and the data item is marked as missing Bjsj=1; When the total number of missing data items Bjsj is greater than the preset data item missing ratio threshold Bthe, the data items with missing data items Bjsj are deleted, and a prompt is given to re-enter the follow-up data.

4. The tumor follow-up data intelligent processing service system according to claim 3, characterized in that: The abnormal data marking unit performs a second abnormal data marking preprocessing task on the data set X after the first missing value detection preprocessing task, identifies abnormal data for each data item in the data set X, and marks the data item with abnormal data A=1 after the abnormal data item is identified, and reorganizes the data set X after the second abnormal data marking preprocessing task is completed to obtain a data group with missing data items Bjsj and abnormal data mark A, and marks it as a data vector Xclean; The identification of abnormal data includes using the Z-score method and the IQR method to identify abnormal data for each data item. When the total number of abnormal data markers A is greater than the preset data item missing ratio threshold Bthe, the data item with abnormal data markers A is deleted, and a prompt is given to re-enter the follow-up data. When there is a data item with abnormal data mark A=1 in the data set X, abnormal mark correction is performed on the data item with abnormal data mark A=1, and the abnormal mark correction includes using mean value correction and replacement of the data item.

5. The tumor follow-up data intelligent processing service system according to claim 1, characterized in that: The data quality assessment module includes a missing rate calculation unit, a consistency calculation unit and an accuracy calculation unit; The missing rate calculation unit calculates the ratio of missing values ​​of each column feature vector in the data vector Xclean to obtain the data missing rate Dmiss; The data missing rate Dmiss is obtained by the following calculation formula: In the formula, Dmiss(j) represents the missing rate of the j-th column feature in the data vector Xclean, n represents the total number of data points of the j-th column feature, Xclean(j, i) represents the i-th data in the j-th column feature in the data vector Xclean, NaN represents an invalid value, 1{Xclean(j, i) = NaN} represents an exponential function, and returns 1 when the i-th data Xclean(j, i) in the j-th column feature in the data vector Xclean is an invalid value NaN, otherwise it returns 0.

6. The tumor follow-up data intelligent processing service system according to claim 5, characterized in that: The consistency calculation unit measures the internal consistency of the data by measuring the vector of each column feature in the data vector Xclean, and the measurement method includes using the coefficient of variation to measure the internal consistency of the data to obtain a consistency score Dcons; The consistency score Dcons is obtained by the following calculation formula: Where Dcns(j) represents the consistency score of the j-th column feature in the data vector Xclean, and μ(j) represents the mean of the j-th column feature.

7. The tumor follow-up data intelligent processing service system according to claim 5, characterized in that: The accuracy calculation unit compares the vector of each column feature in the data vector Xclean, and specifically compares it with the historical mean to obtain the accuracy score Dac, and integrates the data missing rate Dmiss, the consistency score Dcons and the accuracy score Dac to obtain the quality assessment vector Q; The accuracy score Dac is obtained by the following calculation formula: Where μstd(j) represents the historical mean of the j-th column feature.

8. The tumor follow-up data intelligent processing service system according to claim 7, characterized in that: The data response module includes an anomaly detection unit and a correction response unit; The anomaly detection unit generates an anomaly detection score Sabn based on the acquired quality assessment vector Q, and corrects the data vector Xclean based on the anomaly detection assessment Sabn, wherein the correction is performed by comparing the anomaly detection assessment Sabn with a preset correction threshold Xthe, and triggering the correction according to the comparison result. When the correction is triggered, a response anomaly detection score XSabn is generated; The anomaly detection score Sabn is obtained by the following calculation formula: Sabn(j)=s1*Dmiss(j)+s2*Dcons(j)+s3*Dac(j); Where Sabn(j) represents the anomaly detection score of the j-th column feature, s1, s2 and s3 represent the preset weight values ​​of the data missing rate Dmiss, the consistency score Dcons and the accuracy score Dac, respectively, and s1+s2+s3=1, and the specific value is set by the user.

9. The tumor follow-up data intelligent processing service system according to claim 8, characterized in that: The trigger is specifically determined by the following comparison method: When the anomaly detection evaluation Sabn ≥ the correction threshold Xthe, the correction mechanism is triggered to execute the correction task of the anomaly detection evaluation Sabn on the data vector Xclean; When the anomaly detection evaluation Sabn is less than the correction threshold Xthe, the correction mechanism is not triggered, and the correction task of the anomaly detection evaluation Sabn on the data vector Xclean is not performed; The response anomaly detection score XSabn is obtained by the following calculation formula: XSabn(j,i)=Xclean(j,i)+λ(j)*Sabn(j)*ΔT(j); Wherein, λ(j) represents the correction factor of the j-th column feature, △T(j) represents the adjustment amount of the j-th column feature, which is specifically set by the adjustment amount when the data set X performs the first preprocessing task of missing value detection and the second preprocessing task of abnormal data marking.

10. The tumor follow-up data intelligent processing service system according to claim 9, characterized in that: The correction response unit performs a second comparison with the correction threshold Xthe based on the response anomaly detection score XSabn, and triggers an adaptive response mechanism for data correction warning according to the comparison result; The adaptive response mechanism is obtained by the following comparison method: When the response abnormality detection score XSabn ≥ the correction threshold Xthe, the correction mechanism is triggered for the second time, indicating that the validity of the patient's follow-up data is abnormal, and the patient information is re-entered within a fixed period; When the response abnormality detection score XSabn is less than the correction threshold Xthe, the correction mechanism is not triggered, indicating that the patient follow-up data processing is completed, the data is qualified, and the patient follow-up data is stored.