Joint error correction method and system for health archives of crowds in super-large area

Through the uploading of error correction requests and multi-party verification of residents' terminal devices combined with multi-scale periodic feature mapping, attention enhancement and time-aware context processing, the problems of insufficient interaction and low data accuracy in health file management are solved, and efficient and accurate filling and restoration of health file are achieved, and residents' participation and data filling accuracy are improved.

CN120407639AInactive Publication Date: 2025-08-01北京啄木鸟云健康科技有限公司 +1
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510430461.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-08-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the existing health file management methods, residents and health file are insufficient interactivity, and there is a lack of effective multi-party collaborative error correction mechanism, resulting in low file accuracy, low resident participation and satisfaction, and missing values in the health detection time series data. Traditional methods are prone to gradient disappearance and information loss when filling long-sequence data.

Method used

Through the uploading of error correction requests by residents' terminal equipment, the archive management agency conducts multi-party verification, the health file management platform is updated, and uses multi-scale periodic feature mapping, attention enhancement and time-aware context processing, combined with sliding window outlier value detection, to achieve accurate filling and restoration of health files.

Benefits of technology

It improves the accuracy of health records and residents' participation, significantly improves the filling accuracy and stability of health detection time series data, and solves the problems of information loss and gradient disappearance in long-sequence data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407639A_ABST
    Figure CN120407639A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of health archive management, in particular to a joint error correction method and system for health archives of people in a super-large area. The method comprises the following steps: a resident terminal device sends an error correction request to a health archive management platform through a feedback interface, wherein the error correction request comprises a suspected error field in a resident health archive and resident to-be-updated information of the suspected error field; the health archive management platform forwards the error correction request to a management and archive institution of the resident, and the management and archive institution of the resident carries out multi-party correctness check on resident to-be-updated information of a suspected error field; if the checking result is that the to-be-updated information of the residents is correct, the suspected error fields in the health archives of the residents are updated according to the to-be-updated information of the residents, and if the checking result is that the to-be-updated information of the residents is wrong, the health archives of the residents are not updated; the verification result and the health detection time sequence data are displayed through a health file open application; based on the error correction and feedback of the masses, the accuracy of the health archive is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of health record management, and particularly to a combined error correction method and system for health records of a large population in a large area. Background Art

[0002] With the advancement of the digital transformation of public health services, the comprehensive upgrade from traditional paper health records to an electronic health record (EHR) system has been gradually realized. Especially after the construction of a unified health service platform in a large area, by integrating the scattered archive systems of each district and county through standardized data interfaces, problems such as regional duplicate record creation have been effectively solved, ensuring the uniqueness and integrity of health records. The health record management platform mainly adopts a "one-way input-static storage" mode. Medical institutions, as the only information input parties, maintain data accuracy through manual verification. When residents find errors in their record information, they can only raise objections through traditional methods such as offline applications or phone calls. The error correction process requires multi-level manual review, resulting in a long response cycle and low processing efficiency.

[0003] Therefore, the existing health record management methods still have significant defects: the interactivity between residents and health records is seriously insufficient, lacking an effective multi-party collaborative error correction mechanism and real-time feedback channels, the accuracy of records is not high, the participation and satisfaction of residents are not strong, and the phenomenon of records being created but not used appears. Summary of the Invention

[0004] Based on this, the present invention provides a combined error correction method and system for health records of a large population in a large area to solve at least one of the above technical problems.

[0005] To achieve the above object, a combined error correction method for health records of a large population in a large area includes:

[0006] A resident terminal device sends an error correction request to a health record management platform through a feedback interface in a health record open application. The error correction request includes suspected error fields in the resident's health record and the resident's proposed updated information for the suspected error fields;

[0007] After receiving the error correction request, the health record management platform forwards the error correction request to the resident's file management institution. The resident's file management institution conducts multi-party correctness verification on the resident's proposed updated information for the suspected error fields and sends the verification result to the health record management platform;

[0008] If the verification result is that the resident's proposed updated information is correct, the health record management platform updates the suspected error fields in the resident's health record according to the resident's proposed updated information. If the verification result is that the resident's proposed updated information is incorrect, the health record management platform does not update the resident's health record;

[0009] The health record management platform displays the verification results and the health detection time series data in the resident's health record in the health record open application on the resident terminal device through the feedback interface.

[0010] In the present invention, residents can view their health records through the health record open application installed on the terminal device. Residents can upload correction requests to the health record management platform through the feedback interface of the health record open application. The health record management platform forwards the correction requests to the resident's file management agency. The file management agency jointly verifies the correctness of the information to be updated by the resident with multiple agencies and feeds back the verification results to the health record management platform. The health record management platform performs health record update processing according to the verification results. The present invention increases the interaction between residents and health records, improves the enthusiasm of residents to participate in health record management, and realizes the multi-party joint error correction of health records based on the error correction and feedback of the masses, which is beneficial to improving the accuracy of health records in a super-large area (a super-large area refers to an area of more than one city or an area of more than one province or a national area), and enhancing the resident participation rate and satisfaction.

[0011] Preferably, before the health record management platform displays the health detection time series data in the resident's health record in the health record open application on the resident terminal device through the feedback interface, the health record management platform also includes the steps of filling and restoring the health detection time series data in the resident's health record, specifically including:

[0012] Step S1: Obtain the original health detection time series data in the resident's health record; perform multi-scale periodic feature mapping on the original health detection time series data to generate multi-scale periodic time features; perform time series feature fusion according to the multi-scale periodic time features to obtain a time embedding feature vector;

[0013] Step S2: Enhance the attention of the original health detection time series data through the time embedding feature vector to obtain an enhanced variable feature matrix; identify the missing value positions in the original health detection time series data to generate missing value position mask data; calculate the missing value attention scores for the missing value position mask data through the enhanced variable feature matrix to generate a missing value attention score matrix;

[0014] Step S3: Perform time-aware context processing according to the missing value attention score matrix to generate a time-aware context vector; fill in the corresponding missing values in the original health detection time series data through the time-aware context vector to generate filled time series processed data;

[0015] Step S4: Iteratively fill and correct the filled time-series processed data to generate the final filled time-series processed data; perform time-series restoration processing based on the final filled time-series processed data to obtain the restored health detection time-series data, and store the restored health detection time-series data in the resident health record.

[0016] Through multi-scale periodic feature mapping, the method of the present invention can capture periodic patterns at different scales in the health detection time-series data and transform them into multi-scale time features. By fusing these multi-scale features, the generated time-embedded feature vector can more comprehensively reflect the dynamic changes of the health detection time-series data. The attention mechanism is introduced. Through the calculation of the missing value attention score, more attention is focused on the key information related to the missing values. This mechanism effectively solves the problem of information loss in traditional methods for long-sequence data, enabling the model to more accurately learn the correlation between the missing values and the data before and after them, thereby improving the filling accuracy. The time-aware context processing can make full use of the time dependence of the health detection time-series data. By learning the data features within the time window around the missing value, the missing value can be more accurately inferred. This context-aware ability enables the model to better understand the overall trend of the health detection time-series data and perform more reasonable missing value filling based on the time context information. The sliding window outlier detection can identify the outliers after filling and perform iterative correction, further improving the reliability and stability of the filling result. This iterative correction mechanism can effectively avoid the error accumulation caused by single filling and ensure the accuracy of the final filling result. Therefore, a joint error correction method for a large-scale population health record of the present invention captures periodic features at different time scales, automatically learns the correlation between each known data point and the missing value through the attention mechanism, and assigns different weights accordingly. The missing value attention score matrix is combined with the original health detection time-series data for missing value filling, and these outliers are iteratively filled and corrected, overcoming the limitations of traditional methods in dealing with problems such as long sequences, complex periodicity, and long-span missing values, and significantly improving the accuracy and robustness of the filling and restoration of the health detection time-series data.

[0017] Preferably, step S1 includes the following steps:

[0018] Step S11: Obtain the original health detection time-series data;

[0019] Step S12: Perform one-hot encoding on the original health detection time-series data to obtain a one-hot encoded vector of timestamps;

[0020] Step S13: Based on the one-hot encoded vector of timestamps, perform multi-scale periodic feature mapping on the original health detection time-series data to generate multi-scale periodic time features;

[0021] Step S14: Linearly scale and translate the one-hot encoded timestamp vector to generate linear time features;

[0022] Step S15: Adjust the dimensions of the multi-scale periodic time features and the linear time features, and perform feature concatenation and fusion to obtain a time embedding feature vector.

[0023] After obtaining the original health detection time series data, the present invention converts time information into a processable feature vector, which can effectively retain time information and avoid information loss caused by traditional time representation methods. Using the one-hot encoded timestamp vector for multi-scale periodic feature mapping can extract the hidden periodic patterns in the health detection time series data. By setting different time scales, periodic changes at different granularities in the data can be captured, such as daily cycles, weekly cycles, annual cycles, etc., providing more comprehensive and refined time information for subsequent missing value filling. Linearly scaling and translating the one-hot encoded timestamp vector to generate linear time features can capture the long-term trend changes in the health detection time series data. Finally, the multi-scale periodic time features and the linear time features are dimensionally adjusted and concatenated and fused to generate a time embedding feature vector, effectively combining the periodic information and trend information of the health detection time series data.

[0024] Preferably, step S2 includes the following steps:

[0025] Step S21: Perform data preprocessing on the original health detection time series data to obtain preprocessed time series data;

[0026] Step S22: Align the preprocessed time series data in the time dimension through the time embedding feature vector, and extract filling variables to obtain initial filling variable data;

[0027] Step S23: Use the time embedding feature vector to enhance the attention of the initial filling variable data to obtain an enhanced variable feature matrix;

[0028] Step S24: Identify data missing values in the preprocessed time series data, and perform time series position masking processing to generate missing value position masking data;

[0029] Step S25: Calculate the missing value attention scores for the missing value position masking data through the enhanced variable feature matrix to generate a missing value attention score matrix.

[0030] The present invention preprocesses the original health detection time series data, such as data cleaning, normalization and other operations, to improve the data quality and the model training effect. The preprocessed data is aligned in the time dimension using the time embedding feature vector, and the initial variable data for filling in the missing values is extracted. This time dimension alignment operation can effectively solve the problems of misalignment or uneven sampling existing in the health detection time series data, ensuring the accuracy of subsequent filling operations. To further improve the filling accuracy, the initial filling variable data is enhanced in attention using the time embedding feature vector. This attention enhancement mechanism can highlight the important information related to the missing values and suppress the interference of irrelevant information, thereby improving the sensitivity and accuracy of the model for filling in the missing values. The data missing values in the preprocessed data are identified, and a temporal position masking process is performed, which can clearly indicate the position information of the missing values. By calculating the missing value attention scores for the missing value position masked data through the enhanced variable feature matrix, the influence degree of different time points on the missing values can be quantitatively reflected.

[0031] Preferably, step S23 includes the following steps:

[0032] Step S231: Calculate the variable similarity of the initial filling variable data using the dynamic time warping algorithm to generate a variable correlation coefficient matrix;

[0033] Step S232: Perform threshold processing on the variable correlation coefficient matrix through a preset variable screening threshold, set the elements with similarity lower than the preset variable screening threshold to 0, and perform sparsification to obtain a sparse variable relationship matrix;

[0034] Step S233: Based on the sparse variable relationship matrix, use a preset multi-layer graph attention network to aggregate the features of the time embedding feature vector and the initial filling variable data to obtain graph attention feature data;

[0035] Step S234: Multiply the graph attention feature data and the initial filling variable data bit by bit, and perform attention mechanism weighting using the Sigmoid function to obtain attention weighted features;

[0036] Step S235: Perform regularization processing on the attention weighted features to obtain an enhanced variable feature matrix.

[0037] The present invention uses the dynamic time warping algorithm to calculate the similarity between the initial filled variable data, which can better handle the misalignment or asynchrony problems existing in the health detection time series data, so as to more accurately reflect the true correlation between variables. In order to highlight the key variables and reduce the influence of noise, the correlation coefficient matrix is thresholded using a preset variable screening threshold, the elements with lower similarity are set to zero, and sparse processing is performed, which can effectively remove redundant information, retain important variable association relationships, and improve the efficiency and accuracy of subsequent feature aggregation. Based on the sparse variable relationship matrix, a preset multi-layer graph attention network is used to perform feature aggregation on the time embedding feature vector and the initial filled variable data, which can adaptively learn the weights of different variables according to the correlation between variables, so as to better capture the interaction relationship between variables and extract more expressive feature representations. The graph attention feature data is multiplied element by element with the initial filled variable data, and the attention mechanism is weighted using the Sigmoid function, which can effectively enhance the key variable information related to the missing values and suppress the interference of irrelevant variables, thereby improving the accuracy of subsequent missing value filling.

[0038] Preferably, step S25 includes the following steps:

[0039] Step S251: Calculate the time distance from each time step to all missing value positions according to the missing value position mask data to obtain a time distance matrix;

[0040] Step S252: Perform inverse distance attenuation processing on the time distance matrix to generate missing value distance attenuation data;

[0041] Step S253: Copy three feature copies of the enhanced variable feature matrix to obtain variable feature copy data;

[0042] Step S254: Construct three independent linear transformation layers, each linear transformation layer having different weight parameters, to obtain the first linear transformation layer, the second linear transformation layer, and the third linear transformation layer respectively;

[0043] Step S255: Input the variable feature copy data into the first linear transformation layer, the second linear transformation layer, and the third linear transformation layer respectively for multi-head attention linear transformation. The first linear transformation layer outputs a query matrix, the second linear transformation layer outputs a key matrix, and the third linear transformation layer outputs a value matrix;

[0044] Step S256: Calculate the scaled dot-product attention scores for the missing value distance attenuation data using the query matrix and the key matrix to generate a missing value attention score matrix.

[0045] The present invention calculates the time distance from each time step to all missing value positions according to the missing value position mask data, and performs inverse distance attenuation processing, which can assign higher weights to the time steps closer to the missing values, so as to more accurately capture the time-dependent relationships related to the missing values. In order to learn richer feature representations, three feature copies are made of the enhanced variable feature matrix and are respectively input into three independent linear transformation layers for multi-head attention linear transformation. Each linear transformation layer has different weight parameters, which can map the original features to different subspaces, so as to capture the feature information in different aspects of the data. The scaled dot-product attention scores are calculated for the missing value distance attenuation data using the query matrix and the key matrix, which can effectively integrate the feature information from different subspaces, and calculate the influence degree of each time step on the missing values through the scaled dot-product operation.

[0046] Preferably, step S256 includes the following steps:

[0047] Perform a transpose operation on the query matrix and the key matrix, and perform scaled dot-product calculation through a preset scaling factor to obtain a scaled dot-product attention score matrix;

[0048] Perform attenuation weight processing on the missing value distance attenuation data to generate a missing value distance attenuation weight matrix;

[0049] Perform element-wise addition of the missing value distance attenuation weight matrix and the scaled dot-product attention score matrix to obtain a missing information corrected attention score matrix;

[0050] Perform non-missing value masking processing according to the missing information corrected attention score matrix to obtain a masked corrected attention score matrix;

[0051] Perform exponential function transformation on the masked corrected attention score matrix to generate a missing value attention score matrix.

[0052] The present invention performs transposition and scaled dot product calculation on the query matrix and the key matrix, making the dimension of the key matrix match that of the query matrix, so that dot product operation can be carried out. The scaling factor is used to control the magnitude of the dot product result, prevent the gradient disappearance of the subsequent softmax function, and improve the stability of numerical calculation. Through the attenuation weight processing, the degree of distance attenuation can be controlled more flexibly. For example, the attenuation rate can be adjusted according to the characteristics of the actual data, enabling the model to better adapt to different health detection time series data. The information of two dimensions, namely time and space (similarity between variables), is fused together. The element-level addition operation is simple and effective. The distance attenuation weight is directly applied to the similarity score, making the final attention score consider both the correlation between variables and the influence of time distance. Through masking processing, the attention scores corresponding to the missing value positions are set to a very small value (such as negative infinity), making the weights of these positions close to zero in the subsequent softmax calculation, thus avoiding the interference of missing values on the filling result. The attention scores are converted into a probability distribution, making the sum of the contribution degrees of all known data points to a specific missing value equal to 1. This normalization processing makes the attention weights interpretable and facilitates subsequent weighted average or other filling operations.

[0053] Preferably, step S3 includes the following steps:

[0054] Step S31: Weight the value matrix through the missing value attention score matrix to generate an attention-weighted missing value context vector;

[0055] Step S32: Perform time feature fusion on the attention-weighted missing value context vector and the time embedding feature vector to generate a time-aware context vector;

[0056] Step S33: Input the time-aware context vector into a feed-forward neural network with skip connections to predict the residual of the missing value and obtain the residual prediction value;

[0057] Step S34: Fill the corresponding missing values in the original health detection time series data according to the residual prediction value and perform inverse normalization processing to generate filled time series processing data.

[0058] The present invention uses a missing value attention score matrix to weight the value matrix, which can highlight the important time step information related to the missing values and suppress the interference of irrelevant time steps, thereby providing more accurate context information for filling in the missing values. The attention-weighted missing value context vector and the time embedding feature vector are subjected to time feature fusion, effectively combining the context information around the missing values with the overall trend information of the health detection time series data. The time-aware context vector is input into a feed-forward neural network with skip connections to predict the residuals of the missing values, which can effectively learn the complex non-linear relationships in the health detection time series data and alleviate the problem of gradient disappearance through skip connections, improving the training efficiency and prediction accuracy of the model. Filling in the corresponding missing values in the original health detection time series data according to the predicted residuals can effectively reduce the learning difficulty of the model and improve the accuracy and stability of filling.

[0059] Preferably, step S4 includes the following steps:

[0060] Step S41: Perform sliding window outlier detection on the filled time series processed data, identify and mark the outliers to obtain outlier marked data;

[0061] Step S42: Use the outlier marked data to map the outliers in the filled time series processed data and restore the original values according to the original health detection time series data to generate outlier corrected sequence data;

[0062] Step S43: Calculate the mean square error of the original health detection time series data through the outlier corrected sequence data to generate mean square difference data;

[0063] Step S44: When the mean square difference data is greater than or equal to the preset difference threshold and less than the preset maximum number of iterations, use the outlier corrected sequence data as the new time series processed data for iterative filling and correction, otherwise use the outlier corrected sequence data as the final filled time series processed data;

[0064] Step S45: Perform time series restoration processing on the final filled time series processed data to obtain the restored health detection time series data, and store the restored health detection time series data in the resident health record.

[0065] The present invention performs sliding window outlier detection on the filled health detection time series data, identifies and marks the existing outliers, and can effectively capture the local abnormal fluctuations in the health detection time series data, avoiding the errors caused by single-point judgment. Using the outlier-marked data to perform outlier mapping on the filled health detection time series data and restoring the identified outliers to the original health detection time series data can avoid the interference of outliers on the subsequent iterative filling process and improve the accuracy of the filling result. By calculating the mean square error of the original health detection time series data through the anomaly correction sequence data, if the mean square difference data is greater than or equal to the preset difference threshold and the number of iterations is less than the preset maximum number of iterations, the anomaly correction sequence data is used as the new health detection time series data for iterative filling and correction until the termination condition is met. This iterative correction mechanism can effectively reduce the filling error and improve the stability and reliability of the filling result.

[0066] Preferably, step S45 includes the following steps:

[0067] Step S451: Perform data structure analysis based on the original health detection time series data and construct a data structure restoration container to obtain the data structure restoration container;

[0068] Step S452: Extract the dimensionality of the structural data according to the data structure restoration container and perform dimensionality reshaping processing on the finally filled time series processed data to obtain the dimensionally reshaped sequence data;

[0069] Step S453: Use the data structure restoration container to restore the dimensionally reshaped sequence data to generate preliminary restored time series data;

[0070] Step S454: Perform variable loss supplementation processing on the preliminary restored time series data to obtain the restored health detection time series data, and store the restored health detection time series data in the resident health record.

[0071] The present invention performs a data structure analysis on the original health detection time series data, such as data type, dimension, hierarchical structure, etc., and constructs a data structure restoration container according to the analysis results. The data structure restoration container records these key information and provides a template for subsequent data restoration. This ensures that the restored data can be seamlessly docked with the original health detection time series data, avoiding data errors or application problems caused by inconsistent structures. Rearrange the filled data according to the dimensions of the original health detection time series data. Since the filling process involves data dimensionality reduction or dimensionality increase operations, dimensionality reshaping is required to restore the structure of the original health detection time series data. The dimensionality reshaping sequence data maintains the same dimensions and variable arrangements as the original health detection time series data, preparing for subsequent data restoration. Final inspection and supplementation of the restoration process. In some cases, for example, when there are some variables that are completely missing in the original health detection time series data and cannot be fully restored during the previous filling process. For this situation, variable loss supplementation processing is carried out for final supplementation. For example, preset values, values of adjacent variables or other methods can be used to fill these missing variables to ensure that the restored health detection time series data is also consistent with the original health detection time series data at the variable level. Preferably, the present invention also provides a joint error correction system for a large-area population health record, which executes a joint error correction method for a large-area population health record as described above. The joint error correction system for a large-area population health record includes: more than one resident terminal device, more than one file management agency, and a health record management platform. The resident terminal device is connected to the health record management platform through a feedback interface. Each resident has a corresponding file management agency. The file management agency deploys a district and county file business system, and the district and county file business system is connected to the health record management platform through the Internet. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 FIG. is a schematic structural diagram of a joint error correction system for a large-area population health record according to the present invention;

[0073] Figure 2 FIG. is a schematic flow chart of the steps of filling and restoring the health detection time series data according to the present invention;

[0074] Figure 3 is Figure 2 a detailed implementation step flow chart of step S1 in;

[0075] Figure 4 is Figure 2 a detailed implementation step flow chart of step S3 in;

[0076] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0077] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those skilled in the art within the scope of the present invention without creative work belong to the scope of protection of the present invention.

[0078] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.

[0079] It should be understood that although terms such as "first", "second", etc. may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.

[0080] To achieve the above object, please refer to Figure 1 , in a preferred embodiment, the present invention provides a joint error correction method for a super-large area population health record, including:

[0081] The resident terminal device sends an error correction request to the health record management platform through the feedback interface in the health record open application. The error correction request includes the suspected error fields in the resident health record and the resident's proposed updated information for the suspected error fields;

[0082] After receiving the error correction request, the health record management platform forwards the error correction request to the resident's file management agency. The resident's file management agency conducts multi-party correctness verification on the resident's proposed updated information for the suspected error fields and sends the verification result to the health record management platform;

[0083] If the verification result is that the resident's proposed updated information is correct, the health record management platform updates the suspected error fields in the resident health record according to the resident's proposed updated information. If the verification result is that the resident's proposed updated information is incorrect, the health record management platform does not update the resident health record;

[0084] The health record management platform displays the verification results and the health detection time series data in the resident's health record in the health record open application on the resident terminal device through the feedback interface.

[0085] In this embodiment, the resident terminal device is preferably but not limited to a self-service terminal or a mobile terminal that provides services to residents. The health record open application can be loaded on the resident terminal device in the form of a mini-program. The health record open application provides a feedback interface through which users can upload error correction requests. The suspected error fields are preferably but not limited to name, age, gender, residential address, health index detection information, and diagnosis results. The health record management platform is preferably but not limited to a cloud server, which has a database for storing all or part of the health record data, such as the front page of the record, health detection time series data, etc. Each resident has a unique file management agency, and only the file management agency has the right to apply for updating and modifying the resident's health record. The file management agency also needs to upload the medical and health records of the residents under its jurisdiction. The file management agency is generally a medical and health institution directly under the resident's household registration location or long-term residence, such as a district or county health center. The district or county file business system is deployed on the computer or server of the file management agency, and the district or county file business system is connected and communicates with the health record management platform through the Internet. The multi-party correctness verification carried out by the file management agency means that according to the different suspected error fields, it jointly verifies the correctness of the information to be updated with institutions such as the population and family basic resource database and the resident's visiting hospital (local or non-local), so as to achieve joint error correction.

[0086] Exemplarily, when a resident views that the age in his / her health record is incorrect in the health record open application, the suspected error field is age. If the age displayed in the current health record for the resident is 38 years old, while the resident's actual age is 41 years old, and the information that the resident intends to update is 41 years old, then the resident uploads an error correction request of "age, 41 years old" through the feedback interface. After the resident's file management agency obtains the error correction request, it verifies with the population and family basic resource database whether the resident is currently 41 years old. If so, it returns a verification result that the information intended to be updated by the resident is correct; if the age of the resident in the population and family basic resource database is not 41 years old, it returns a verification result that the information intended to be updated by the resident is incorrect. Further, if the age in the population and family basic resource database is not 38 years old, the verification result should also include the actual age in the population and family basic resource database, and the resident's health record should be updated according to this actual age.

[0087] Exemplarily, a resident views their medical history in the health record open application and finds that hypertension is missing from their past medical history in the health record. The suspected error field is the past medical history. The resident then uploads the error correction request "Past medical history, add hypertension, XX Hospital" through the feedback interface. After the resident's file management agency receives the error correction request, it verifies with XX Hospital whether the resident has been diagnosed with hypertension. If so, it returns that the information the resident intends to update is correct; if XX Hospital has not diagnosed the resident with hypertension, it returns that the information the resident intends to update is incorrect.

[0088] One of the main contents included in the resident health record is the record of health and health service activities. The record of health and health service activities refers to a brief record of the basic situation of relevant health and health service activities received by a resident due to diseases or health problems. Through chronological arrangement, it forms health detection time series data, which includes blood pressure health detection time series data, blood sugar health detection time series data, etc. Effectively analyzing the resident's health detection time series data is crucial for resident health prediction and disease diagnosis.

[0089] However, due to the loss of medical records, there are often missing value problems in the health detection time series data of the health and health service activities records in the health record. The existence of missing values will seriously affect the integrity and quality of the health detection time series data, and thus reduce the effect of subsequent data analysis and application. The filling and restoration of missing values in the health detection time series data aims to use the known health data information to reasonably estimate and supplement the missing health data, and restore the integrity and accuracy of the data as much as possible. However, traditional time series data filling methods usually use statistical models or recurrent neural networks for processing. However, when dealing with the filling of long sequence data, these methods are prone to problems such as gradient disappearance and information loss. Especially when the recurrent neural network processes long sequences, it needs to transmit information step by step in each time step. As the length of the health detection time series increases, the gradient will gradually decay during the backpropagation process, making it difficult to capture long-distance dependencies. This means that the network is difficult to learn the influence of the health data at earlier time points in the health detection time series on the current time point, thus reducing the filling accuracy. Especially when the span of missing values is large, the filling effect is even less satisfactory.

[0090] Therefore, in a preferred embodiment, please refer to Figures 2 to 4 Before the health record management platform provides the health detection time series data in the resident health record in the health record open application of the resident terminal device through the feedback interface, it also includes the steps of filling and restoring the health detection time series data in the resident health record by the health record management platform, specifically including:

[0091] Step S1: Obtain the original health detection time series data in the resident health record; perform multi-scale periodic feature mapping on the original health detection time series data to generate multi-scale periodic time features; perform time series feature fusion based on the multi-scale periodic time features to obtain a time embedding feature vector;

[0092] Step S2: Enhance the attention of the original health detection time series data through the time embedding feature vector to obtain an enhanced variable feature matrix; identify the missing value positions in the original health detection time series data to generate missing value position mask data; calculate the missing value attention scores for the missing value position mask data through the enhanced variable feature matrix to generate a missing value attention score matrix;

[0093] Step S3: Perform time-aware context processing based on the missing value attention score matrix to generate a time-aware context vector; fill in the corresponding missing values in the original health detection time series data through the time-aware context vector to generate filled time series processed data;

[0094] Step S4: Iteratively fill and correct the filled time series processed data to generate the final filled time series processed data; perform time series restoration processing based on the final filled time series processed data to obtain the restored health detection time series data, and store the restored health detection time series data in the resident health record.

[0095] In the embodiment of the present invention, the filling and restoration of the health detection time series data in the resident health record by the health record management platform includes the following steps:

[0096] Step S1: Obtain the original health detection time series data; perform multi-scale periodic feature mapping on the original health detection time series data to generate multi-scale periodic time features; perform time series feature fusion based on the multi-scale periodic time features to obtain a time embedding feature vector;

[0097] In the embodiments of the present invention, obtaining the original time series data to be processed includes, but is not limited to, temperature records, blood pressure records, and drug dosage adjustment records in electronic health records. The original time series data refers to an observation data sequence containing timestamp information and corresponding numerical values. For example, the temperature data recorded in the electronic health record at different time points constitutes the original time series data. Specifically, assume there is a temperature sequence containing 1000 time points, where the time points are recorded in seconds and the temperature values are in degrees Celsius. The data format can be represented as a two-dimensional array, with the first column being the timestamp (e.g., integers from 0 to 999) and the second column being the corresponding temperature values (e.g., 25.1, 25.3, 25.2,..., 26.0). For ease of subsequent processing, these data are loaded into the computer memory, and the original health detection time series data stored in a CSV file is read using the Pandas library in Python. The Pandas library converts the data into a DataFrame object for convenient access and manipulation. The time and temperature columns of the DataFrame object are extracted, and the time data is converted into a NumPy array, and the temperature data is also converted into a NumPy array. These two arrays will be used as the input for the subsequent steps. It should be noted that the quality of the original time series data directly affects the final filling and restoration effects, so it is necessary to preprocess the original health detection time series data, such as removing duplicate timestamps and handling outliers. Define multiple sine and cosine functions with different frequencies. For example, three frequencies can be defined: f1 = 1 / 24 (representing a daily cycle), f2 = 1 / 168 (representing a weekly cycle), and f3 = 1 / 8760 (representing an annual cycle). Then, for the one-hot encoded vector of each timestamp, calculate its sine and cosine values corresponding to each frequency. Use a fully connected layer to map the feature vector to a specified dimension. The fully connected layer contains a weight matrix and a bias vector. By learning the weight matrix and the bias vector, the input feature vector can be mapped to the specified dimension. Assume the input dimension required by the multi-head attention mechanism is D, then the multi-scale periodic time features need to be mapped to D dimensions respectively. Then, the multi-scale periodic time features with adjusted dimensions are concatenated. Concatenation can connect multiple feature vectors into a longer vector, thus integrating the information of multiple features.

[0098] Step S2: Enhance the attention of the original health detection time series data through the time-embedded feature vector to obtain an enhanced variable feature matrix; identify the missing value positions in the original health detection time series data to generate missing value position mask data; calculate the missing value attention scores for the missing value position mask data through the enhanced variable feature matrix to generate a missing value attention score matrix;

[0099] In the embodiments of the present invention, assuming that the lengths of the two are the same, no interpolation or truncation operations are required. After the time dimension is aligned, the preprocessed time series data and the time embedding feature vector are concatenated. The concatenation operation connects the two vectors into a longer vector, thus fusing the information of the two vectors. For example, if the dimension of the preprocessed time series data is N×1 and the dimension of the time embedding feature vector is N x D, the dimension of the concatenated vector is N×(1 + D). The concatenated vector is input into a fully connected layer, and the fully connected layer maps the vector to a specified dimension to obtain the initial filled variable data. If an element in the original time series data is a missing value, the corresponding element in the missing value position mask data is True; otherwise, it is False. For example, if the original time series data is [1, 2, NaN, 4, NaN], the missing value position mask data is [False, False, True, False, True]. The Pandas library in Python can be used to implement missing value recognition and masking processing. The Pandas library provides the isna() method to determine whether the data is a missing value. The NumPy library can be used to convert a boolean matrix into a numerical matrix. For example, True can be converted to 1 and False can be converted to 0. The enhanced variable feature matrix and the missing value position mask data are input into the fully connected layer. The fully connected layer includes a weight matrix and a bias vector. By learning the weight matrix and the bias vector, the enhanced variable feature matrix can be mapped to a specified dimension. Assuming that the output dimension of the fully connected layer is 1, each time point corresponds to a missing value attention score. The missing value attention score is multiplied by the missing value position mask data to obtain the missing value attention score. The multiplication operation sets the attention scores at non-missing value positions to 0 and only retains the attention scores at missing value positions.

[0100] Step S3: Perform time-aware context processing according to the missing value attention score matrix to generate a time-aware context vector; use the time-aware context vector to fill the corresponding missing values in the original health detection time series data to generate filled time series processed data;

[0101] In the embodiments of the present invention, for each missing value, according to its score in the missing value attention score matrix, the K missing values with the highest correlation with it are selected (for example, K = 5). Then, the timestamps corresponding to these K missing values are found, and the data values of L time points before and after these timestamps are obtained (for example, L = 3) to form a time-aware context. Next, a weighted average method is used to fill the missing values. The weights consist of two parts: one part is the score in the missing value attention score matrix, and the other part is the reciprocal of the time distance. The time distance refers to the difference between the timestamp of the current missing value and the timestamp of the context data point. The value of the context data point is multiplied by its corresponding weight and then summed to obtain the filled value of the missing value.

[0102] Step S4: Iteratively fill and correct the filled time-series processed data to generate the final filled time-series processed data; perform time-series restoration processing based on the final filled time-series processed data to obtain the restored health detection time series data, and store the restored health detection time series data in the resident health record.

[0103] In the embodiment of the present invention, a sliding window with a length of W (for example, W = 10) is used and slides on the filled time-series processed data with a step size of 1. For each window, calculate the mean and standard deviation of the data within the window. If the value of a data point within the window exceeds the mean plus or minus N times the standard deviation (for example, N = 3), then this data point is considered an outlier. For the detected outliers, treat them as missing values and repeat steps S2 and S3 for filling. Iteratively perform outlier detection and filling until no new outliers are detected or the maximum number of iterations is reached (for example, the maximum number of iterations is 5). Through iterative filling and correction, the accuracy of data filling can be further improved. Finally, the final filled time-series processed data is obtained.

[0104] Preferably, step S1 includes the following steps:

[0105] Step S11: Obtain the original health detection time series data;

[0106] Step S12: Perform one-hot encoding on the time stamps of the original health detection time series data to obtain a one-hot encoded time stamp vector;

[0107] Step S13: Based on the one-hot encoded time stamp vector, perform multi-scale periodic feature mapping on the original health detection time series data to generate multi-scale periodic time features;

[0108] Step S14: Perform linear scaling and translation on the one-hot encoded time stamp vector to generate linear time features,

[0109] Step S15: Adjust the dimensions of the multi-scale periodic time features and the linear time features, and perform feature splicing and fusion to obtain a time embedding feature vector.

[0110] As an example of the present invention, referring to Figure 3 shown, for Figure 2 the detailed implementation step flow diagram of step S1 in

[0111] Step S11: Obtain the original health detection time series data;

[0112] In an embodiment of the present invention, the original time series data to be processed is obtained. The original time series data refers to an observation data sequence containing timestamp information and corresponding numerical values. For example, the body temperature data recorded by a patient at different time points is collected, and these data constitute the original time series data. Specifically, assume that there is a temperature sequence containing 1000 time points, the time points are recorded in seconds, and the temperature values are in degrees Celsius. The data format can be represented as a two-dimensional array, where the first column is the timestamp (e.g., integers from 0 to 999), and the second column is the corresponding temperature value (e.g., 25.1, 25.3, 25.2,..., 26.0). For the convenience of subsequent processing, these data are loaded into the computer memory, and the Pandas library in Python is used to read the original health detection time series data stored in a CSV file. The Pandas library converts the data into a DataFrame object for easy access and operation. The time and temperature columns in the DataFrame object are extracted, and the time data is converted into a NumPy array, and the temperature data is also converted into a NumPy array.

[0113] Step S12: Perform one-hot encoding on the original health detection time series data to obtain a one-hot encoded vector of timestamps;

[0114] In an embodiment of the present invention, it is assumed that the data set contains N different timestamps. A zero vector with a length of N is created. For each timestamp, find its corresponding index position among all timestamps, and set the value at that index position in the zero vector to 1. For example, if the data set contains timestamps [10, 20, 30, 40] and the current timestamp is 20, then the index position of 20 in [10, 20, 30, 40] is 1, so the corresponding one-hot encoded vector is [0, 1, 0, 0]. One-hot encoding is performed on all timestamps to obtain a series of one-hot encoded vectors. These vectors can represent the uniqueness of each timestamp.

[0115] Step S13: Perform multi-scale periodic feature mapping on the original health detection time series data based on the one-hot encoded vector of timestamps to generate multi-scale periodic time features;

[0116] In the embodiments of the present invention, multiple sine and cosine functions with different frequencies are defined. For example, three frequencies can be defined: f1 = 1 / 24 (representing the daily cycle), f2 = 1 / 168 (representing the weekly cycle), and f3 = 1 / 8760 (representing the annual cycle). Then, for the one-hot encoded vector of each timestamp, calculate its sine value and cosine value corresponding to each frequency. The calculation formulas are as follows: sin(2×pi×f×t); cos(2×pi×f×t), where f is the frequency and t is the timestamp. Since only one element of the one-hot encoded vector is 1, in fact, only the sine value and cosine value of the timestamp corresponding to this element need to be calculated. For example, if the timestamp is t and the t-th element of the corresponding one-hot encoded vector is 1, then calculate sin(2×pi×f×t) and cos(2×pi×f×t). For each frequency, obtain a sine value and a cosine value, and combine these values into a vector. For example, if three frequencies are defined, then each timestamp corresponds to a vector with a length of 6. Combine the vectors corresponding to all timestamps into a matrix, and the size of this matrix is N×6, where N is the length of the time series. This matrix is the multi-scale periodic time feature.

[0117] Step S14: Linearly scale and translate the one-hot encoded vector of the timestamp to generate linear time features;

[0118] In the embodiments of the present invention, the timestamp is scaled to the interval [0, 1], and the formula is as follows: t' = (t - min(t)) / (max(t) - min(t)); where t is the timestamp, t' is the scaled timestamp, min() is the minimum value function in the mathematical formula, and max() is the maximum value function in the mathematical formula. Then, an offset can be added. For example, add 1 to all timestamps, and the formula is as follows: t” = t' + 1; Since one-hot encoding is used, it is necessary to find the position where the value in the one-hot encoding is 1, and this position represents the original timestamp. After finding this position, calculate the scaled and translated timestamp value using the above formula. For example, if the timestamp is t and the t-th element of the corresponding one-hot encoded vector is 1, then calculate t' and t”. Each timestamp corresponds to a linear time feature value, and combine these values into a vector, and the size of this vector is N×1, where N is the length of the time series. This vector is the linear time feature.

[0119] Step S15: Adjust the dimensions of the multi-scale periodic time feature and the linear time feature, and perform feature splicing and fusion to obtain a time embedding feature vector.

[0120] In the embodiments of the present invention, a fully connected layer is used to map the feature vector to a specified dimension. The fully connected layer includes a weight matrix and a bias vector. By learning the weight matrix and the bias vector, the input feature vector can be mapped to the specified dimension. Assume that the input dimension required by the multi-head attention mechanism is D. Then, the multi-scale periodic time features and the linear time features need to be mapped to the D dimension respectively. Then, the multi-scale periodic time features and the linear time features after dimension adjustment are concatenated. Concatenation can connect two feature vectors into a longer vector, thereby fusing the information of the two features. For example, if the dimension of the multi-scale periodic time features is N×D and the dimension of the linear time features is N×D, then the dimension of the time embedding feature vector after concatenation is N×2D. The concatenated feature vector is used as the final time embedding feature vector.

[0121] Preferably, step S2 includes the following steps:

[0122] Step S21: Perform data preprocessing on the original health detection time series data to obtain preprocessed time series data;

[0123] Step S22: Align the preprocessed time series data in the time dimension through the time embedding feature vector, and perform filling variable extraction to obtain initial filling variable data;

[0124] Step S23: Use the time embedding feature vector to perform attention enhancement on the initial filling variable data to obtain an enhanced variable feature matrix;

[0125] Step S24: Identify the data missing values in the preprocessed time series data, and perform time series position masking processing to generate missing value position masking data;

[0126] Step S25: Calculate the missing value attention scores for the missing value position masking data through the enhanced variable feature matrix to generate a missing value attention score matrix.

[0127] In the embodiments of the present invention, for the time point where the missing value is located, find the two nearest non-missing values before and after it, and then calculate the approximate value of the missing value through a linear function based on the time and values of these two non-missing values. For outliers, a method based on box plots is used for detection and removal. Specifically, calculate the upper quartile Q3, lower quartile Q1, and interquartile range IQR = Q3 - Q1 of the data. Then, define the upper and lower bounds of the outliers: upper bound = Q3 + 1.5 × IQR, lower bound = Q1 - 1.5 × IQR. If the value of a data point exceeds the upper and lower bounds, then this data point is considered an outlier. For the detected outliers, use the average of the adjacent data points before and after for replacement. For data standardization, the Z-score standardization method is adopted. Z-score standardization transforms the data into a standard normal distribution with a mean of 0 and a standard deviation of 1. The calculation formula is as follows: x' = (x - mean) / std; where x is the original health detection time series data, x' is the standardized data, mean is the mean of the data, and std is the standard deviation of the data. The data preprocessing can be implemented using the Pandas library and Scikit-learn library in Python. The time dimension alignment ensures the one-to-one correspondence in time between the preprocessed time series data and the time embedding feature vector. The time dimensions of the preprocessed time series data and the time embedding feature vector should be the same, that is, both have a length of N. If the lengths of the two are inconsistent, interpolation or truncation operations need to be performed. The interpolation operation refers to estimating the data values at unknown time points through the existing data points. The truncation operation refers to directly deleting the redundant data points. In this example, it is assumed that the lengths of the two are the same, and no interpolation or truncation operations are required. After the time dimension alignment, the preprocessed time series data and the time embedding feature vector are concatenated. The concatenation operation connects the two vectors into a longer vector, thereby fusing the information of the two vectors. For example, if the dimension of the preprocessed time series data is N×1 and the dimension of the time embedding feature vector is N×2D, then the dimension of the concatenated vector is N×(1 + 2D). The concatenated vector is input into a fully connected layer, and the fully connected layer maps the vector to the specified dimension to obtain the initial filled variable data. The initial filled variable data and the time embedding feature vector are input into the attention layer. The attention layer first calculates the attention weights for each feature. The calculation formula is as follows: attention_weight = softmax(query × key^T); where query is the initial filled variable data, key is the time embedding feature vector, and softmax is the softmax function. Both query and key need to undergo a linear transformation to map them to the specified dimension. The linear transformation is implemented through a fully connected layer. After calculating the attention weights, apply the attention weights to the initial filled variable data.The application formula is as follows: enhanced_feature = attention_weight × value; where value is the initial filling variable data. The feature vector after applying the attention weight contains more important information and can be used as the input for subsequent steps. Combine the enhanced variable feature vectors at all time points into a matrix. The size of this matrix is N×M, where N is the length of the time series and M is the dimension of the feature vector. This matrix is the enhanced variable feature matrix. If an element in the preprocessed time series data is a missing value, the corresponding element in the missing value position mask data is True; otherwise, it is False. For example, if the preprocessed time series data is [1, 2, NaN, 4, NaN], then the missing value position mask data is [False, False, True, False, True]. The Pandas library in Python can be used to implement missing value identification and masking processing. The Pandas library provides the isna() method to determine whether the data is a missing value. The NumPy library can be used to convert a boolean matrix into a numerical matrix. For example, True can be converted to 1 and False to 0. Input the enhanced variable feature matrix and the missing value position mask data into the fully connected layer. The fully connected layer contains a weight matrix and a bias vector. By learning the weight matrix and the bias vector, the enhanced variable feature matrix can be mapped to the specified dimension. Assume that the output dimension of the fully connected layer is 1, then each time point corresponds to a missing value attention score. Multiply the missing value attention score by the missing value position mask data to obtain the missing value attention score. The multiplication operation sets the attention scores at non-missing value positions to 0 and only retains the attention scores at missing value positions. For example, if the missing value position mask data is [0, 0, 1, 0, 1] and the missing value attention score is [0.1, 0.2, 0.3, 0.4, 0.5], then the missing value attention score is [0, 0, 0.3, 0, 0.5]. Combine the missing value attention scores at all time points into a matrix. The size of this matrix is N×1, where N is the length of the time series. This matrix is the missing value attention score matrix.

[0128] Preferably, step S23 includes the following steps:

[0129] Step S231: Use the dynamic time warping algorithm to calculate the variable similarity of the initial filling variable data and generate a variable correlation coefficient matrix;

[0130] Step S232: Perform threshold processing on the variable correlation coefficient matrix through a preset variable screening threshold, set the elements with similarity lower than the preset variable screening threshold to 0, and perform sparsification to obtain a sparse variable relationship matrix;

[0131] Step S233: Based on the sparse variable relationship matrix, use a preset multi-layer graph attention network to perform feature aggregation on the time embedding feature vector and the initial filling variable data to obtain graph attention feature data;

[0132] Step S234: Multiply the graph attention feature data and the initial filling variable data bit by bit, and use the Sigmoid function to perform attention mechanism weighting to obtain attention-weighted features;

[0133] Step S235: Perform regularization processing on the attention-weighted features to obtain an enhanced variable feature matrix.

[0134] In the embodiment of the present invention, for the initial filling variable data obtained in step S22, the dynamic time warping (DTW) algorithm is used to calculate the similarity between each variable. DTW is an algorithm for measuring the similarity of time series. Even if the time series has stretching or offset on the time axis, it can effectively calculate its similarity. The dimension of the initial filling variable data is N×M, where N is the length of the time series and M is the number of variables. Therefore, it is necessary to calculate the DTW distances between every two of the M variables to obtain an M×M variable correlation coefficient matrix. Specifically, for any two variables i and j (1≤i,j≤M), calculate the DTW distance between them. The DTW algorithm constructs an accumulated distance matrix, finds the optimal alignment path between two time series, and calculates the accumulated distance on this path. Set a variable screening threshold, such as 0.5. Then, traverse the variable correlation coefficient matrix. If the value of a certain element is less than this threshold, set this element to 0. After the threshold processing, the variable correlation coefficient matrix becomes a sparse matrix, that is, most of the elements in the matrix have a value of 0. Sparseification can further reduce the amount of calculation and prevent the graph attention network from overfitting. The threshold processing and sparseification can be implemented using the NumPy library. After the threshold processing and sparseification, a sparse variable relationship matrix is obtained. Each variable is regarded as a node in the graph, the sparse variable relationship matrix is used as the adjacency matrix of the graph, and the time embedding feature vector and the initial filling variable data are used as the feature vectors of the nodes. The multi-layer GAT is stacked by multiple GAT layers. Each GAT layer first calculates the attention weights between nodes, and then weights and sums the feature vectors of neighbor nodes to obtain a new node feature vector. The calculation method of the attention weights is as follows: attention_weight(i,j) = softmax(a^T×[W×h_i||W×h_j]), where h_i and h_j are the feature vectors of nodes i and j respectively, W is a linear transformation matrix, a is an attention vector, || represents vector concatenation, and softmax is the softmax function. The attention weights represent the degree of attention of node i to node j. In the multi-layer GAT, each GAT layer will update the feature vectors of the nodes. After being processed by the multi-layer GAT, the feature vectors of the nodes contain richer graph structure information. Take the output of the last GAT layer as the graph attention feature data. Multiply the graph attention feature data obtained in step S233 by the initial filling variable data obtained in step S22 bit by bit, and then use the Sigmoid function for attention mechanism weighting. The bit-by-bit multiplication operation multiplies the elements at the corresponding positions of the graph attention feature data and the initial filling variable data, so as to incorporate the graph structure information into the initial filling variable data. In this example, the L2 regularization method is used. L2 regularization limits the size of the model parameters by adding a penalty term to the loss function.The formula for L2 regularization is as follows: loss = original_loss + lambda × ||w||^2; where original_loss is the original loss function, lambda is the regularization coefficient, w is the model parameter, and ||w||^2 represents the L2 norm of the model parameter. Through L2 regularization, the model parameters can be made smoother, thereby improving the generalization ability of the model. Before regularization, the attention-weighted features are normalized. Common normalization methods include Z-score normalization and Min-Max normalization. In this example, Z-score normalization is selected. After normalization and regularization, an enhanced variable feature matrix is obtained.

[0135] Preferably, step S25 includes the following steps:

[0136] Step S251: Calculate the time distance from each time step to all missing value positions according to the missing value position mask data to obtain a time distance matrix;

[0137] Step S252: Perform inverse distance attenuation processing on the time distance matrix to generate missing value distance attenuation data;

[0138] Step S253: Make three feature copies of the enhanced variable feature matrix to obtain variable feature copy data;

[0139] Step S254: Construct three independent linear transformation layers, each with different weight parameters, to obtain the first linear transformation layer, the second linear transformation layer, and the third linear transformation layer respectively;

[0140] Step S255: Input the variable feature copy data into the first linear transformation layer, the second linear transformation layer, and the third linear transformation layer respectively for multi-head attention linear transformation. The first linear transformation layer outputs a query matrix, the second linear transformation layer outputs a key matrix, and the third linear transformation layer outputs a value matrix;

[0141] Step S256: Calculate the scaled dot-product attention scores for the missing value distance attenuation data using the query matrix and the key matrix to generate a missing value attention score matrix.

[0142] In the embodiments of the present invention, traverse the missing value position mask data to find the positions of all missing values. Then, for each time step i, calculate the time distance from it to each missing value position j. The calculation method of the time distance is as follows: distance(i,j) = |i - j| if missing_mask[j] == 1 else 0; where missing_mask is the missing value position mask data, and |i - j| represents the absolute distance between time step i and time step j. Through the above calculation, a time distance matrix is obtained. Use the inverse distance attenuation method to calculate the missing value distance attenuation data, so that the time steps closer to the missing value position have a greater impact on the filling of the missing value. Specifically, for each time step i, calculate its inverse distance attenuation value to all missing value positions j. Make three feature copies of the enhanced variable feature matrix obtained in step S235. The copy is to prepare different inputs for the subsequent multi-head attention mechanism, so that the model can learn features from different perspectives. The dimension of the enhanced variable feature matrix is N×M, where N is the length of the time series and M is the feature dimension. The dimension of the variable feature copy data after copy is N×M×3, where the dimension of each copy is N×M. The feature copy can be implemented using the NumPy library. For example, the numpy.tile() function or the torch.repeat() function can be used. To implement the multi-head attention mechanism, three independent linear transformation layers need to be constructed, which are used to generate the query matrix, key matrix, and value matrix respectively. Each linear transformation layer contains a weight matrix and a bias vector. The weight matrices and bias vectors of the three linear transformation layers are all different, which means they will learn different feature mappings. The input of the linear transformation layer is a copy of the enhanced variable feature matrix, and the output is the query matrix, key matrix, and value matrix. The query matrix is used to calculate the attention weights, the key matrix is used to match with the query matrix, and the value matrix is used for weighted aggregation. The weight matrix and bias vector of the linear transformation layer can be learned through training. Construct three independent linear transformation layers, each with different weight parameters. The role of the linear transformation layer is to map the input features to different spaces, so that the model can learn different feature representations. In this example, three linear transformation layers are constructed, which are used to generate the query matrix (Query), key matrix (Key), and value matrix (Value) respectively. These three linear transformation layers share the same input dimension and output dimension, but have different weight parameters, so they can map the input features to different spaces. Deep learning frameworks such as PyTorch or TensorFlow can be used to construct the linear transformation layer.Calculate the dot product of the query matrix and the key matrix to obtain the attention weights: attention_weights = softmax(queries × keys^T / sqrt(D)); where queries is the query matrix, keys is the key matrix, D is the dimension of the key matrix, and sqrt(D) is the scaling factor used to prevent the dot product from being too large, which may cause the gradient of the softmax function to vanish. The softmax function normalizes the attention weights so that their sum is 1. Then, multiply the attention weights by the missing value distance decay data to obtain the missing value attention scores: masked_attention_scores = attention_weights × normalized_decay_matrix; where normalized_decay_matrix is the missing value distance decay data. By multiplying the attention weights by the missing value distance decay data, the time steps closer to the missing value position will have a greater impact on filling the missing values. Finally, normalize the missing value attention scores so that their sum is 1.

[0143] Preferably, step S256 includes the following steps:

[0144] Perform a transpose operation on the query matrix and the key matrix, and calculate the scaled dot product through a preset scaling factor to obtain the scaled dot product attention score matrix;

[0145] Perform decay weight processing on the missing value distance decay data to generate a missing value distance decay weight matrix;

[0146] Perform an element-wise addition of the missing value distance decay weight matrix and the scaled dot product attention score matrix to obtain a missing information corrected attention score matrix;

[0147] Perform non-missing value masking on the missing information corrected attention score matrix to obtain a masked corrected attention score matrix;

[0148] Perform an exponential function transformation on the masked corrected attention score matrix to generate a missing value attention score matrix.

[0149] In the embodiment of the present invention, a transpose operation is performed on the query matrix (Query, dimension N×D) and the key matrix (Key, dimension N×D) obtained in step S255. Since a dot product operation will be performed subsequently, it is necessary to convert the dimension of the key matrix to D×N. Then, the scaled dot product calculation is performed between the transposed key matrix and the query matrix. The scaled dot product calculation is the core part of the attention mechanism, and its purpose is to measure the correlation between each query vector and all key vectors. The scaling factor is used to prevent the dot product result from being too large, resulting in problems such as vanishing gradients or exploding gradients. The formula for the scaled dot product calculation is as follows: scaled_attention = (Q@K.T) / sqrt(d_k); where Q is the query matrix, K is the key matrix, K.T is the transpose of the key matrix, d_k is the dimension of the key matrix, sqrt(d_k) is the scaling factor, and the @ symbol represents matrix multiplication. After the scaled dot product calculation, a scaled dot product attention score matrix is obtained, with a dimension of N×N. Each element in this matrix represents the correlation between corresponding time steps. The missing value distance decay weight matrix is added to the scaled dot product attention score matrix element-wise. Element-wise addition means adding the elements at the corresponding positions of the two matrices. The result after addition is the missing information corrected attention score matrix. The missing information corrected attention score matrix comprehensively considers the similarity between the query and the key and the influence of the time distance. Assuming the scaled dot product attention score matrix is S and the missing value distance decay weight matrix is W, then the calculation formula for the missing information corrected attention score matrix A is: A = S + W. Through element-wise addition, the time distance information can be directly added to the attention score, so that the missing values with a closer distance have a greater weight. The attention scores at non-missing value positions are set to a very small negative number, such as negative infinity. In this way, after normalization by the Softmax function, the attention weights at non-missing value positions will approach 0. Assuming the missing value position mask data is mask, where mask[i]=1 indicates that the i-th time step is a missing value, and mask[i]=0 indicates that the i-th time step is not a missing value. A mask matrix M of the same size as the missing information corrected attention score matrix is created. For each element M[i][j] in the mask matrix, if the j-th time step is not a missing value, then M[i][j] is set to negative infinity; otherwise, it is set to 0. Then, the missing information corrected attention score matrix is added to the mask matrix element-wise. The result after addition is the mask corrected attention score matrix. Softmax normalization is performed on the transformed result. The Softmax function can convert a vector into a probability distribution, representing the attention weight of each key to the query. Since the values at non-missing value positions in the mask corrected attention score matrix are set to negative infinity, after the exponential function transformation and Softmax normalization, the attention weights at non-missing value positions will approach 0. The finally obtained matrix is the missing value attention score matrix.The missing value attention score matrix reflects the correlation between different missing values and can be used to guide the model to consider the information of other missing values when filling in the missing values. Assuming that the masked corrected attention score matrix is A, the calculation formula for the missing value attention score matrix P is: P = softmax(exp(A)). Since the exponential function is sensitive to the value range, the masked corrected attention score matrix can be scaled to avoid the values output by the exponential function being too large or too small.

[0150] Preferably, step S3 includes the following steps:

[0151] Step S31: Weight the value matrix through the missing value attention score matrix to generate an attention-weighted missing value context vector;

[0152] Step S32: Perform time feature fusion on the attention-weighted missing value context vector and the time embedding feature vector to generate a time-aware context vector;

[0153] Step S33: Input the time-aware context vector into a feed-forward neural network with skip connections to predict the residual of the missing value and obtain the residual prediction value;

[0154] Step S34: Fill in the corresponding missing values in the original health detection time series data according to the residual prediction value and perform inverse normalization processing to generate the filled time series processing data.

[0155] As an example of the present invention, refer to Figure 4 as shown, for Figure 2 the detailed implementation step flow diagram of step S3 in

[0156] Step S31: Weight the value matrix through the missing value attention score matrix to generate an attention-weighted missing value context vector;

[0157] In the embodiments of the present invention, for each missing value position i, the i-th column of the missing value attention score matrix is extracted, and this column vector represents the attention weights of all time steps for the missing value position i. Then, this column vector is weighted and summed with the value matrix to obtain the attention-weighted missing value context vector. The formula for weighted summation is as follows: context_vector_i = sum(attention_weight_ij × value_j for j in range(N)); where attention_weight_ij represents the attention weight of time step j for the missing value position i, and value_j represents the feature vector of time step j in the value matrix. The attention-weighted missing value context vectors corresponding to all missing value positions are combined into a matrix to obtain the attention-weighted missing value context vector matrix, whose dimension is M × D, where M is the number of missing values and D is the feature dimension. If there is no missing data, then M = 0, no calculation is performed, and an empty matrix is returned. To ensure consistent output dimensions, the output dimension can be set to the weighted average over all time steps, that is, context_vector_i is calculated for all i, and finally context_vector is a vector with dimension N × D.

[0158] Step S32: Perform temporal feature fusion on the attention-weighted missing value context vector and the temporal embedding feature vector to generate a temporally aware context vector;

[0159] In the embodiments of the present invention, the concatenation method is adopted. The attention-weighted missing value context vector and the temporal embedding feature vector are concatenated in the feature dimension. The concatenated vector contains context information and temporal information and can be used as the input for subsequent steps. The formula for the concatenation operation is as follows: time_aware_context_vector = concatenate([attention_weighted_context_vector, time_embedding]); where attention_weighted_context_vector is the attention-weighted missing value context vector, time_embedding is the temporal embedding feature vector, and concatenate represents the concatenation operation. Assuming that the dimension of the attention-weighted missing value context vector is N × D and the dimension of the temporal embedding feature vector is N × T, then the dimension of the concatenated temporally aware context vector is N × (D + T).

[0160] Step S33: Input the temporally aware context vector into a feed-forward neural network with skip connections to predict the residual of the missing value and obtain the residual prediction value;

[0161] In the embodiments of the present invention, the skip connection directly connects the time-aware context vector to the output of the FFN, enabling the model to better utilize the context information of the time series. The FFN consists of multiple fully connected layers, and each fully connected layer is followed by an activation function. The activation function can introduce non-linear factors and improve the expressive ability of the model. Commonly used activation functions include ReLU, Sigmoid, Tanh, etc. In this example, the ReLU activation function is used. The output dimension of the FFN is the same as the dimension of the original time series data, i.e., N×1. For example, the time-aware context vector is input into the first fully connected layer to obtain the output of the first hidden layer. Then, the output of the first hidden layer is input into the ReLU activation function to obtain the output of the first activation layer. Next, the output of the first activation layer is input into the second fully connected layer to obtain the output of the second hidden layer. Finally, the output of the second hidden layer is skip-connected with the time-aware context vector to obtain the residual prediction value.

[0162] Step S34: Fill the corresponding missing values in the original health detection time series data according to the residual prediction value, and perform inverse normalization processing to generate the filled time series processed data.

[0163] In the embodiments of the present invention, the residual prediction value is added to the corresponding missing value position in the original time series data to achieve the filling of the missing value. The filling formula is as follows:

[0164] filled_data_normalized[missing_mask == 1] = original_data_imputed[missing_mask == 1] + residual_prediction[missing_mask == 1]; where filled_data_normalized is the filled normalized data, original_data_imputed is the data after linear interpolation of the original health detection time series data (Step S21), residual_prediction is the residual prediction value, and missing_mask is the missing value position mask. The residual prediction value (residual_prediction) obtained in Step S33 is added to the corresponding missing value position in the original time series data. It should be noted that when performing the addition operation, it is necessary to ensure that the data types of the residual prediction value and the original time series data are consistent to avoid potential type conversion errors. At the same time, the residual prediction value needs to be added to the initial filling value at the corresponding missing value position in the original time series data to obtain the updated data value. Then, inverse normalization processing is performed on the filled data to restore the data to the original scale. The inverse normalization processing requires the mean and standard deviation during the normalization processing.

[0165] Preferably, step S4 includes the following steps:

[0166] Step S41: Perform sliding window outlier detection on the filled time series processed data, identify and mark the outliers to obtain outlier marked data;

[0167] Step S42: Use the outlier marked data to perform outlier mapping on the filled time series processed data, and perform original value restoration according to the original health detection time series data to generate outlier corrected sequence data;

[0168] Step S43: Calculate the mean square error of the original health detection time series data through the outlier corrected sequence data to generate mean square difference data;

[0169] Step S44: When the mean square difference data is greater than or equal to the preset difference threshold and less than the preset maximum number of iterations, use the outlier corrected sequence data as the new time series processed data for iterative filling and correction; otherwise, use the outlier corrected sequence data as the final filled time series processed data;

[0170] Step S45: Perform time series restoration processing on the final filled time series processed data to obtain the restored health detection time series data, and store the restored health detection time series data in the resident health record.

[0171] In the embodiments of the present invention, a sliding window is used to fill the time series processed data, and the mean and standard deviation of the data within each window are calculated. If the value of a certain data point within the window deviates from the mean by more than a preset multiple (e.g., 3 times) of the standard deviation, then this data point is marked as an outlier. Other statistics such as the median and interquartile range can be selected to define outliers. To avoid misjudging normal fluctuations as outliers, the multiple of the standard deviation needs to be set reasonably. All identified outliers are marked to generate outlier marked data. The outlier marked data is a binary vector with the same size as the filled time series processed data, where 1 indicates that the corresponding position is an outlier and 0 indicates a normal value. The size of the sliding window and the multiple of the standard deviation are two important hyperparameters that need to be adjusted according to the specific data set, and methods such as grid search or Bayesian optimization can be used in the parameter optimization process. Traverse the outlier marked data. If a certain time point is marked as an outlier, then check whether this time point is a missing value in the original health detection time series data. If this time point in the original health detection time series data is not a missing value, it means that the original health detection time series data provides reliable information. Therefore, replace the value at the corresponding position in the filled time series processed data with the value at the corresponding position in the original health detection time series data. If this time point in the original health detection time series data is a missing value, it means that the original health detection time series data cannot provide information. At this time, retain the value at the corresponding position in the filled time series processed data, that is, still use the result filled by the model. After traversing the entire time series and completing all outlier mapping and original value restoration operations, the anomaly corrected sequence data can be generated. Align the anomaly corrected sequence data with the original health detection time series data. Since there are missing values in the original health detection time series data, when calculating the MSE, the influence of these missing values needs to be excluded. Only calculate the error at non-missing value positions. Then, calculate the error at each non-missing value position. The error is defined as the square of the difference between the anomaly corrected sequence data and the original health detection time series data at this position. Sum up the errors at all non-missing value positions and divide by the number of non-missing values to obtain the mean square error. Set a difference threshold and a maximum number of iterations. The difference threshold is used to determine whether the current filling effect is good enough. If the mean square difference data is less than this threshold, it is considered that the filling has been completed and no iterative correction is required. The maximum number of iterations is used to control the number of iterations to prevent the iterative process from looping infinitely. Then, determine whether the current mean square difference data is greater than or equal to the preset difference threshold and whether the current number of iterations is less than the preset maximum number of iterations. If the iterative conditions are met, then use the anomaly corrected sequence data as the new original health detection time series data and return to step S2 to re-perform missing value filling and outlier correction. This step is equivalent to using the results of the previous filling and correction as the input for the next filling and correction for iterative optimization.If the iteration condition is not satisfied, the abnormal correction sequence data is used as the final filled time series processed data, and the iteration process is ended. The data is converted into the original data format. For example, if the original health detection time series data is a CSV file, the filled data is saved as a CSV file; if the original health detection time series data is a database table, the filled data is inserted into the database table. After the above time series restoration process, the final restored health detection time series data is obtained, and the restored health detection time series data is stored in the resident health record.

[0172] Preferably, step S45 includes the following steps:

[0173] Step S451: Perform data structure analysis based on the original health detection time series data, and construct a data structure restoration container to obtain a data structure restoration container;

[0174] Step S452: Extract the structural data dimension according to the data structure restoration container, and perform dimension reshaping processing on the final filled time series processed data to obtain dimension reshaped sequence data;

[0175] Step S453: Use the data structure restoration container to restore the dimension reshaped sequence data to generate preliminary restored time series data;

[0176] Step S454: Perform variable loss supplementation processing on the preliminary restored time series data to obtain the restored health detection time series data, and store the restored health detection time series data in the resident health record.

[0177] In the embodiments of the present invention, data structure analysis is performed on the original health detection time series data. The analysis content includes: Dimension information: Determine the dimension of the data, such as one-dimensional time series, two-dimensional time series (such as multiple sensor data), or higher-dimensional time series. Data type information: Determine the data type of each dimension, such as numeric type, character type, boolean type, etc. For numeric data, it is also necessary to determine the unit and value range of the data. Timestamp format information: If the data contains timestamps, it is necessary to determine the format of the timestamps, such as year-month-day hour-minute-second, Unix timestamp, etc. It is also necessary to determine the unit of the timestamps, such as seconds, minutes, hours, etc. Categorical variable information: If the data contains categorical variables, it is necessary to determine the value range and encoding method of the categorical variables. Then, according to the results of the data structure analysis, a data structure restoration container is constructed. The data structure restoration container can be a dictionary, a class, or other data structures, which is used to store the structure information of the original health detection time series data. Extract the dimension information of the original health detection time series data from the data structure restoration container. For example, if the original health detection time series data is a two-dimensional time series, the extracted dimension information is two-dimensional. If the original health detection time series data is a three-dimensional time series, the extracted dimension information is three-dimensional. According to the extracted dimension information, dimension reshaping is performed on the finally filled time series processed data. The goal of dimension reshaping is to make the filled data have the same dimension as the original health detection time series data. The reshape function of the NumPy library can be used for dimension reshaping. According to the data type information stored in the data structure restoration container, convert the data type in the dimension reshaped sequence data to the data type of the original health detection time series data. For example, if the data type of the timestamp column in the original health detection time series data is a string, convert the data type of the timestamp column in the dimension reshaped sequence data to a string. Identify the variables lost during the data processing. The lost variables can be identified by comparing information such as column names and data types of the preliminary restored time series data and the original health detection time series data. Then, for each lost variable, according to the information in the original health detection time series data, add the variable to the preliminary restored time series data. If the lost variable is a numeric variable, the values in the corresponding column of the original health detection time series data can be directly copied to the preliminary restored time series data; if the lost variable is a character variable, character encoding conversion can be performed, and the converted values are copied to the preliminary restored time series data; if the lost variable is a timestamp variable, time format conversion is required, and the converted values are copied to the preliminary restored time series data. When supplementing the lost variables, ensure the consistency of the data types. If the addition fails due to type inconsistency, consider type conversion or redesign the data processing process to avoid information loss.

[0178] Preferably, the present invention further provides a joint error correction for a large - area population health record, which executes a joint error correction method for a large - area population health record as described above. As Figure 1 shown, the joint error correction system for a large - area population health record includes: more than one resident terminal device, more than one file - managing agency, and a health record management platform. The resident terminal device is connected to the health record management platform through a feedback interface. Each resident has a corresponding file - managing agency, and the file - managing agency deploys a district - county file service system. The district - county file service system is connected to the health record management platform through the Internet.

[0179] The present application lies in that through multi - scale analysis, it can comprehensively capture periodic features at different time scales and fuse these features into time - embedded feature vectors, providing richer and more accurate time - context information for subsequent filling and restoration. This extraction method of multi - scale periodic features avoids information deviation caused by single - scale features and enhances the model's understanding ability of complex health - detection time - series data. An attention - enhancement mechanism and a missing - value attention - score calculation are introduced. When traditional methods fill in missing values, they usually treat all known data points equally or simply use surrounding data points for interpolation, ignoring the difference in the influence of different data points on missing values. However, this method can automatically learn the correlation between each known data point and the missing value through the attention mechanism and assign different weights accordingly. This means that data points more relevant to the missing value will obtain higher weights and thus have a greater impact on the estimation of the missing value. Specifically, by processing the masked data at the missing - value position with the enhanced variable - feature matrix, a missing - value attention - score matrix is generated, and this process precisely quantifies the contribution degree of each known data point to a specific missing value. This mechanism effectively solves the problem of information loss in filling long - sequence data. Even when the span of missing values is large, it can accurately capture the dependence relationship between distant data points and missing values, avoid the influence of gradient disappearance, and significantly improve the filling accuracy. Using time - aware context processing, the missing - value attention - score matrix is combined with the original health - detection time - series data for missing - value filling. This processing method fully considers the temporal characteristics of health - detection time - series data, not only using the information of known data points but also considering the time relationship between this information and the missing value. Through the attention - score matrix, the model can "focus" on the data in the time periods most relevant to the missing value and use this data for weighted average or other more complex fusion operations to obtain a more reasonable estimation of the missing value. It overcomes the limitations of traditional methods in dealing with problems such as long sequences, complex periodicity, and long - span missing values, significantly improves the accuracy and robustness of filling and restoration of health - detection time - series data, and provides a more reliable data basis for various applications based on health - detection time - series data.

[0180] Therefore, in all respects, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Thus, all changes that fall within the meaning and scope of the equivalent elements of the application documents are intended to be embraced within the present invention.

[0181] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features invented herein.

Claims

1. A combined error correction method for health records of a large-scale population in a large area, characterized in that, Including: The resident terminal device sends an error correction request to the health record management platform through the feedback interface in the health record open application. The error correction request includes the suspected error field in the resident health record and the resident's proposed updated information for the suspected error field; After receiving the error correction request, the health record management platform forwards the error correction request to the resident's file management agency. The resident's file management agency conducts multi-party correctness verification on the resident's proposed updated information for the suspected error field and sends the verification result to the health record management platform; If the verification result is that the resident's proposed updated information is correct, the health record management platform updates the suspected error field in the resident health record according to the resident's proposed updated information. If the verification result is that the resident's proposed updated information is incorrect, the health record management platform does not update the resident health record; The health record management platform displays the verification result and the health detection time series data in the resident health record in the health record open application of the resident terminal device through the feedback interface.

2. The combined error correction method for a super-large area population health record according to claim 1, characterized in that Before the health record management platform displays the health detection time series data in the resident health record in the health record open application of the resident terminal device through the feedback interface, it also includes the steps of the health record management platform filling and restoring the health detection time series data in the resident health record, specifically including: Step S1: Obtain the original health detection time series data in the resident health record; perform multi-scale periodic feature mapping on the original health detection time series data to generate multi-scale periodic time features; perform time series feature fusion according to the multi-scale periodic time features to obtain a time embedding feature vector; Step S2: Enhance the attention of the original health detection time series data through the time embedding feature vector to obtain an enhanced variable feature matrix; identify the missing value positions in the original health detection time series data to generate missing value position mask data; calculate the missing value attention scores for the missing value position mask data through the enhanced variable feature matrix to generate a missing value attention score matrix; Step S3: Perform time-aware context processing according to the missing value attention score matrix to generate a time-aware context vector; fill in the corresponding missing values in the original health detection time series data through the time-aware context vector to generate filled time series processed data; Step S4: Perform iterative filling and correction on the filled time series processed data to generate final filled time series processed data; perform time series restoration processing according to the final filled time series processed data to obtain restored health detection time series data, and store the restored health detection time series data in the resident health record.

3. The combined error correction method for a super-large area population health record according to claim 2, characterized in that, Step S1 includes the following steps: Step S11: Obtain the original health detection time series data; Step S12: Perform one-hot encoding on the time stamps of the original health detection time series data to obtain a one-hot encoded time stamp vector; Step S13: Perform multi-scale periodic feature mapping on the original health detection time series data based on the one-hot encoded time stamp vector to generate multi-scale periodic time features; Step S14: Linearly scale and translate the one-hot encoded time stamp vector to generate linear time features; Step S15: Adjust the dimensions of the multi-scale periodic time features and the linear time features, and perform feature splicing and fusion to obtain a time embedding feature vector.

4. The combined error correction method for a super-large area population health record according to claim 2, characterized in that, Step S2 includes the following steps: Step S21: Perform data preprocessing on the original health detection time series data to obtain preprocessed time series data; Step S22: Align the preprocessed time series data in the time dimension using the time embedding feature vector, and extract filling variables to obtain initial filling variable data; Step S23: Use the time embedding feature vector to enhance the attention of the initial filling variable data to obtain an enhanced variable feature matrix; Step S24: Identify data missing values in the preprocessed time series data, and perform time series position masking processing to generate missing value position masking data; Step S25: Calculate the missing value attention scores for the missing value position masking data using the enhanced variable feature matrix to generate a missing value attention score matrix.

5. The combined error correction method for a super-large area population health record according to claim 4, characterized in that, Step S23 includes the following steps: Step S231: Calculate the variable similarity of the initial filling variable data using the dynamic time warping algorithm to generate a variable correlation coefficient matrix; Step S232: Perform threshold processing on the variable correlation coefficient matrix using a preset variable screening threshold, set elements with similarity lower than the preset variable screening threshold to 0, and perform sparsification to obtain a sparse variable relationship matrix; Step S233: Based on the sparse variable relationship matrix, use a preset multi-layer graph attention network to aggregate the time embedding feature vector and the initial filling variable data to obtain graph attention feature data; Step S234: Multiply the graph attention feature data and the initial filling variable data bit by bit, and use the Sigmoid function to perform attention mechanism weighting to obtain attention-weighted features; Step S235: Perform regularization processing on the attention-weighted features to obtain an enhanced variable feature matrix.

6. The combined error correction method for a super-large area population health record according to claim 4, characterized in that, Step S25 includes the following steps: Step S251: Calculate the time distance from each time step to all missing value positions based on the missing value position masking data to obtain a time distance matrix; Step S252: Perform inverse distance attenuation processing on the time distance matrix to generate missing value distance attenuation data; Step S253: Duplicate the enhanced variable feature matrix three times to obtain variable feature duplicate data; Step S254: Construct three independent linear transformation layers, each with different weight parameters, to obtain the first linear transformation layer, the second linear transformation layer, and the third linear transformation layer respectively; Step S255: Input the variable feature duplicate data into the first linear transformation layer, the second linear transformation layer, and the third linear transformation layer respectively for multi-head attention linear transformation. The first linear transformation layer outputs a query matrix, the second linear transformation layer outputs a key matrix, and the third linear transformation layer outputs a value matrix; Step S256: Calculate the scaled dot product attention scores for the missing value distance attenuation data using the query matrix and the key matrix to generate a missing value attention score matrix.

7. A combined error correction method for a super-large area population health record according to claim 6, characterized in that, Step S256 includes the following steps: Transpose the query matrix and the key matrix, and perform scaled dot - product calculation through a preset scaling factor to obtain a scaled dot - product attention score matrix; Perform attenuation weight processing on the missing - value distance - attenuation data to generate a missing - value distance - attenuation weight matrix; Add the missing - value distance - attenuation weight matrix and the scaled dot - product attention score matrix element - by - element to obtain a missing - information corrected attention score matrix; Perform non - missing - value masking processing according to the missing - information corrected attention score matrix to obtain a masked - corrected attention score matrix; Perform exponential function transformation on the masked - corrected attention score matrix to generate a missing - value attention score matrix.

8. A combined error correction method for a super-large area population health record according to claim 6, characterized in that Step S3 includes the following steps: Step S31: Weight the value matrix through the missing - value attention score matrix to generate an attention - weighted missing - value context vector; Step S32: Perform time - feature fusion on the attention - weighted missing - value context vector and the time - embedding feature vector to generate a time - aware context vector; Step S33: Input the time - aware context vector into a feed - forward neural network with skip connections to predict the residual of the missing value, obtaining a residual prediction value; Step S34: Fill in the corresponding missing values in the original health - detection time - series data according to the residual prediction value and perform inverse normalization processing to generate filled - time - series processed data.

9. The combined error correction method for a super-large area population health record according to claim 2, characterized in that, Step S4 includes the following steps: Step S41: Perform sliding - window outlier detection on the filled - time - series processed data, identify and mark the outliers to obtain outlier - marked data; Step S42: Map the outliers in the filled - time - series processed data using the outlier - marked data and restore the original values according to the original health - detection time - series data to generate outlier - corrected sequence data; Step S43: Calculate the mean - square error of the original health - detection time - series data through the outlier - corrected sequence data to generate mean - square difference data; Step S44: When the mean - square difference data is greater than or equal to a preset difference threshold and less than a preset maximum number of iterations, use the outlier - corrected sequence data as new time - series processed data for iterative filling and correction; otherwise, use the outlier - corrected sequence data as the final filled - time - series processed data; Step S45: Perform time - series restoration processing on the final filled - time - series processed data to obtain restored health - detection time - series data, and store the restored health - detection time - series data in the resident health record.

10. A combined error correction system for the health records of a large population in a large area, which is used to implement the combined error correction method for the health records of a large population in a large area described in any one of claims 1-9, characterized in that, The system includes: more than one resident terminal device, more than one file - management agency, and a health - record management platform. The resident terminal device is connected to the health - record management platform through a feedback interface. Each resident has a corresponding file - management agency. The file - management agency deploys a district - county file service system, and the district - county file service system is connected to the health - record management platform through the Internet.

Citation Information

Patent Citations

  • Electronic health record missing data completion method and system based on LSTM

    CN110767279A

  • Quality evaluation method and system for electronic health record

    CN113035308A

  • Health record missing value complementing method and system and storage medium

    CN114023407A

  • Resident health record opening and supervision system and method

    CN114093481A

  • Quality detection method for health archives of regional medical residents

    CN115798663A