Dynamic data modeling processing method for physical examination index track classification

CN122842970APending Publication Date: 2026-09-29SHANDONG FIRST MEDICAL UNIVERSITY FIRST AFFILIATED HOSPITAL (QIANFO MOUNTAIN HOSPITAL OF SHANDONG PROVINCE)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611091751.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-22
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

然而,不同来源的体检数据在指标名称、计量单位、采集时间、参考范围和数据完整性方面存在差异,同一受检对象的多次体检记录往往分布不均且存在缺失、冲突或低频观测,现有处理方式难以在统一时间尺度下刻画指标随时间变化的轨迹,更难区分当前数值接近但变化过程不同的受检对象类型

Benefits of technology

本发明通过匿名归集、标准名称映射、单位换算、来源冲突和数据状态标识,降低多来源体检数据不一致对建模的影响;通过统一时间网格、插值补全、统计补全和尺度归一化,使不同采集频率、不同量纲的指标共同参与轨迹分析;通过灰色预测残差、隐状态转移、动态参与量和轨迹驱动画像,强化对持续变化、过程差异和关键驱动指标的识别;通过稳定性校验和偏离程度输出,提高分型结果的可解释性、可复核性和可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842970A_ABST
    Figure CN122842970A_ABST
Patent Text Reader

Abstract

The application discloses a dynamic data modeling processing method for physical examination index track classification, relates to the technical field of medical health data processing and dynamic track modeling, and is used for solving the problem that multi-source physical examination index data is difficult to be reliably classified under time inconsistency, missing conflict and change process difference; anonymously collects, name maps, unit converts and conflict processes electronic physical examination, inspection and examination, physical measurement and questionnaire data to generate physical examination index track samples; forms dynamic track samples through time alignment, missing completion, normalization and dynamic feature extraction; then, in combination with gray prediction, hidden state modeling, dynamic participation quantity and track distance, classification is completed, and track types, key indexes and stability results are output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical and health data processing and dynamic trajectory modeling technology, and more specifically, to a dynamic data modeling and processing method for physical examination indicator trajectory classification. Background Technology

[0002] With the development of information technology in resident health management and institutional physical examinations, data such as electronic physical examination records, laboratory test records, physical measurement records, and health questionnaires are continuously collected. Existing systems can typically store single physical examination results, alert on abnormal items, or perform statistical analysis based on fixed thresholds, providing a data foundation for health record management. However, physical examination data from different sources vary in terms of indicator names, units of measurement, collection time, reference ranges, and data completeness. Multiple physical examination records for the same examinee are often unevenly distributed and contain missing, conflicting, or low-frequency observations. Existing processing methods struggle to depict the trajectory of indicator changes over time on a unified time scale, and it is even more difficult to distinguish between examinee types whose current values ​​are similar but whose change processes differ.

[0003] To address the above problems, this invention proposes a solution. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a dynamic data modeling and processing method for physical examination index trajectory classification, so as to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: In a preferred embodiment, it includes: Data from multiple sources of physical examinations are collected according to the anonymity identifier of the examinee. The original item names are matched with standard names according to the source, testing method and sample type. After converting the original index values ​​into the corresponding standard physical examination index values, physical examination index trajectory samples are constructed according to the first observation time. The physical examination indicator trajectory samples are placed into a unified time grid according to the preset alignment interval, and conflicting and missing values ​​are distinguished and processed according to the data status identifier. Then, based on the value changes of the same standard physical examination indicator name at adjacent or nearby alignment time points, dynamic trajectory samples of physical examination indicators are formed. Based on the dynamic trajectory samples of physical examination indicators, an indicator modeling sequence is established. The prediction residual is obtained through gray prediction processing, and the hidden state sequence is obtained through the hidden state transition model. Then, the observation confidence quantity is determined by the data state identifier. The dynamic participation of indicators is adjusted by the consistency of changes within consecutive aligned time points, the continuation deviation of prediction residuals, the deviation outside the reference range buffer boundary, and the correspondence between hidden state changes and indicator changes. In this way, a trajectory-driven profile is formed, and the trajectory distance is calculated based on the common standard physical examination indicator names to form the physical examination indicator trajectory classification results. The stability of the physical examination indicator trajectory classification results is verified by repeated sampling and time disturbance. The names of key standard physical examination indicators and the degree of trajectory deviation are sorted out in combination with the type center trajectory, and the physical examination indicator trajectory classification output record is generated.

[0006] In a preferred embodiment, the step of aggregating multi-source physical examination-related data according to the anonymity identifier of the examinee includes: using the anonymity identifier of the examinee as the primary key for trajectory aggregation, incorporating the observation time, original item name, original indicator value, original unit, reference range, and source fields from electronic physical examination records, test records, physical measurement records, and health questionnaire records into the same examinee's data set; using a preset indicator name mapping table, converting the original item names under different data sources, detection method identifiers, and sample type identifiers into standard physical examination indicator names, and using a preset unit conversion table, converting the corresponding original indicator values ​​into standard physical examination indicator values ​​under the target unit; and using the first observation time under the same examinee's anonymity identifier as a benchmark, scaling up various observation times to construct a standardized physical examination time.

[0007] In a preferred embodiment, after constructing the standardized physical examination time, the method further includes: for multiple candidate standard physical examination indicator values ​​under the same anonymous identifier of the examinee, the same standardized physical examination time, and the same standard physical examination indicator name, conflict screening is performed according to the detection method identifier, sample type identifier, source priority, and time proximity, and the measured values, missing values, unmapped names, unconverted units, and source conflicts are written into the data status identifier respectively; at the same time, the reference range, data source identifier, and auxiliary variable code corresponding to the standard physical examination indicator name are retained, and a physical examination indicator trajectory sample is established with the examinee's anonymous identifier, standardized physical examination time, standard physical examination indicator name, standard physical examination indicator value, reference range, data source identifier, and data status identifier as core fields.

[0008] In a preferred embodiment, the step of placing the physical examination index trajectory samples into a unified time grid according to a preset alignment interval includes: based on the anonymity identifier of the examinee, the standardized physical examination time, the name of the standard physical examination index, the value of the standard physical examination index, the reference range, the data source identifier, and the data status identifier, firstly, sorting each observation record by time according to the anonymity identifier of the examinee, and merging records and resolving source conflicts for the same standard physical examination index name under the same standardized physical examination time; then, constructing a unified time grid with the trajectory start time, trajectory end time, and preset alignment interval, mapping each standard physical examination index value to the corresponding aligned time point to form aligned physical examination index values, and writing a missing status for aligned time points with no measured values ​​or unresolved conflicts.

[0009] In a preferred embodiment, the step of processing missing values ​​according to data status identifiers and forming dynamic trajectory samples of physical examination indicators includes: around the same anonymous identifier of the examinee and the same standard physical examination indicator name, sequentially using linear interpolation of measured values ​​before and after, one-sided value continuation, and statistical completion of the population median at the missing alignment time point to generate completed physical examination indicator values, and writing them into the interpolation completion state, continuation completion state, or statistical completion state respectively; then, according to the median, first quartile, and third quartile under the same standard physical examination indicator name, the current physical examination indicator value is scaled and normalized to generate normalized physical examination indicator values; Further, the local changes and local rates of change between adjacent aligned time points are calculated within a unified time grid. Then, the normalized physical examination index values ​​are linearly fitted within a preset trend window to generate a local trend slope. Simultaneously, the reference deviation is calculated based on the lower limit, upper limit, and width of the reference range after unit conversion. Finally, the anonymous identifier of the examined subject, aligned time points, standard physical examination index names, current physical examination index values, data source identifiers, data status identifiers, normalized physical examination index values, local changes, local rates of change, local trend slopes, reference deviation, and auxiliary variable codes are written into the dynamic observation record to form a dynamic trajectory sample of the physical examination index.

[0010] In a preferred embodiment, the following are used as inputs: the anonymous identifier of the examined subject, the alignment time point, the name of the standard physical examination indicator, the data status identifier, the normalized physical examination indicator value, the local rate of change, the local trend slope, and the reference deviation degree. An indicator modeling sequence arranged according to the alignment time point is constructed based on the same anonymous identifier of the examined subject and the same standard physical examination indicator name. Different preset confidence weights are assigned to each sequence position according to the measured state, conflict resolution state, interpolation completion state, continuation completion state, statistical completion state, and missing state. For indicator modeling sequences where the number of valid sequence positions meets the preset minimum modeling quantity, the normalized physical examination indicator values ​​are extracted to form the original indicator sequence. The original indicator sequence is then subjected to an accumulation generation process. Combined with the background value and the preset confidence weights, the gray development coefficient and gray action are calculated using weighted least squares. Finally, the predicted physical examination indicator values ​​and predicted residuals for each alignment time point are calculated. Normalized physical examination index values, local rate of change, local trend slope, reference deviation degree, and prediction residuals are used to construct observation feature vectors in a fixed order. Clustering initialization is performed based on all available observation feature vectors to establish a hidden state transition model that includes initial state probability, state transition probability, and observation generation parameters. Then, through dynamic programming, the hidden state sequence corresponding to the available alignment time point is solved for each index modeling sequence. Based on the hidden state sequence, local rate of change, local trend slope, reference deviation degree, and prediction residual statistics, trajectory representation data is formed.

[0011] In a preferred embodiment, after obtaining the trajectory characterization data, for any standard physical examination indicator name and any alignment time point, an observation confidence quantity is first constructed based on the data status identifier; a change duration quantity is constructed based on the same-direction persistence of the local change rate and local trend slope within a preset duration window; a predicted deviation accumulation quantity is constructed based on the same-sign persistence and absolute value increase of the predicted residuals within a preset residual window; a reference buffer deviation quantity is constructed based on the upper and lower buffer boundaries formed by the expansion of the reference range; and a hidden state transition confirmation quantity is constructed based on the correspondence between the hidden state changes at adjacent alignment time points and the direction of the local change rate, the amplitude of the local trend slope, and the change in the absolute value of the predicted residuals. Then, using the observation confidence quantity as the basic participation condition, a participation adjustment value is generated by combining the change duration quantity, the predicted deviation accumulation quantity, the reference buffer deviation quantity, and the hidden state transition confirmation quantity. Finally, the observation confidence quantity and the participation adjustment value are used to generate the indicator dynamic participation quantity, so that different standard physical examination indicator names form a non-fixed participation degree that is adjusted according to the dynamic change cause at different alignment time points.

[0012] In a preferred embodiment, the dynamic participation of indicators is aggregated over time according to the aligned time points to identify continuous high participation intervals. The indicator trajectory participation is then generated by combining the length of the continuous high participation interval, the average dynamic participation of indicators within the interval, and the number of hidden state transitions. The main driving cause is then determined based on the duration of high-value changes, the cumulative predicted deviation, the reference buffer deviation, or the confirmed hidden state transition. The driving direction is determined based on the sign of the local trend slope within the continuous high participation interval. A trajectory-driven profile is constructed, including the standard physical examination indicator name, continuous high participation interval, indicator trajectory participation, main driving cause, and driving direction. Furthermore, between any two anonymous identifiers of the examined subjects, the participation is first determined based on their shared indicator trajectory participation. The standard physical examination indicator names are used to construct a set of commonly participating indicators. Then, level differences, persistent differences, residual differences, reference differences, and state differences are calculated from the normalized numerical level, continuous change process, prediction residual deviation, reference buffer deviation, and latent state transition performance. Based on this, indicator difference structures such as near-level heterogeneous process structure, heterogeneous near-process structure, heterogeneous heterogeneous structure, or near-level near-process structure are generated. On the basis of the indicator difference structure, process priority adjustment quantity is constructed, and the participation degree of persistent differences, residual differences, and state differences in the indicator trajectory difference quantity is increased for near-level heterogeneous process structure. Then, the indicator trajectory difference quantity is generated by combining the observation confidence quantity, indicator trajectory participation quantity, time overlap degree of continuous high participation interval, and driving direction relationship.

[0013] In a preferred embodiment, a set of dominant participating indicators is selected according to the participation amount of the indicator trajectory, and the trajectory distance between the anonymous identifiers of two subjects is synthesized based on the difference participation ratio. At the same time, when there is a near-horizontal heterogeneous process structure, a process difference identifier is written. The trajectory distance and trajectory-driven profile are used as the basis for merging trajectory clusters. The merging of trajectory clusters is controlled by the trajectory distance between clusters and the driving consistency between clusters. The near-distance heterogeneous driving relationship between trajectory clusters, the driving dispersion state of the merged trajectory clusters, the type center trajectory of the physical examination indicator trajectory type, and the pending confirmation state and pending classification state of the anonymous identifiers of the subjects are marked respectively to form the physical examination indicator trajectory classification result.

[0014] In a preferred embodiment, after obtaining the trajectory distance, the anonymous identifiers of the tested subjects are classified into physical examination indicator trajectory types through agglomerative clustering. Within each trajectory type, the type center trajectory is selected based on the average trajectory distance. Subsequently, the dynamic trajectory samples, latent state sequences, and trajectory representation data of the physical examination indicators are associated. The trajectory type identifier is recalculated by repeating the sampling standard physical examination indicator names and alignment time points, as well as the perturbation alignment time points, to generate repeating consistency scores and time perturbation consistency scores, which are then written into the stability state identifier. Next, the trajectory representation data is summarized according to the trajectory type, and the indicator contribution value of each standard physical examination indicator name is calculated. Key standard physical examination indicator names and their key indicator representation records are selected. Finally, the trajectory deviation degree is calculated based on the trajectory distance between the anonymous identifiers of the tested subjects and the type center trajectory, as well as the consistency score, and a physical examination indicator trajectory classification output record is generated.

[0015] The technical effects and advantages of the dynamic data modeling and processing method for physical examination indicator trajectory classification in this invention are as follows: This invention reduces the impact of inconsistencies in multi-source physical examination data on modeling by anonymizing data collection, standard name mapping, unit conversion, source conflict resolution, and data status identification; it enables indicators with different collection frequencies and scales to participate in trajectory analysis by unifying the time grid, interpolation completion, statistical completion, and scale normalization; it strengthens the identification of continuous changes, process differences, and key driving indicators by using gray prediction residuals, hidden state transitions, dynamic participation quantities, and trajectory-driven profiling; and it improves the interpretability, verifiability, and reliability of the classification results by using stability verification and deviation degree output. Attached Figure Description Figure 1 This diagram illustrates the unified time grid and missing data completion processing of the dynamic data modeling and processing method for physical examination indicator trajectory classification according to the present invention.

[0016] Figure 2 This is a schematic diagram of trajectory classification driven by the dynamic participation of indicators in the dynamic data modeling and processing method for trajectory classification of physical examination indicators according to the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] In this embodiment, the present invention discloses a dynamic data modeling and processing method for physical examination indicator trajectory classification, including: Step 1: Collect relevant data from multiple sources of physical examinations according to the anonymity identifier of the examinee, match the original item names with standard names according to the source, testing method and sample type, and convert the original index values ​​into the corresponding standard physical examination index values. Then, construct the physical examination index trajectory sample according to the first observation time. Step 2: Place the physical examination indicator trajectory samples into a unified time grid according to the preset alignment interval, and distinguish and process conflicting and missing values ​​according to the data status identifier. Then, based on the value changes of the same standard physical examination indicator name at adjacent or nearby alignment time points, form a dynamic trajectory sample of the physical examination indicator. Step 3: Establish an indicator modeling sequence based on the dynamic trajectory samples of physical examination indicators, obtain the prediction residuals through grey prediction processing, and obtain the hidden state sequence through the hidden state transition model; then determine the observation confidence quantity by data state identifier, and adjust the dynamic participation quantity of indicators by the consistency of changes within consecutive aligned time points, the continuation deviation of prediction residuals, the deviation outside the reference range buffer boundary, and the correspondence between hidden state changes and indicator changes, thereby forming a trajectory-driven profile, and calculate the trajectory distance based on the common standard physical examination indicator names to form the physical examination indicator trajectory classification results; Step 4: Perform repeated sampling stability verification and time disturbance stability verification on the physical examination indicator trajectory classification results, and combine the key standard physical examination indicator names and trajectory deviation degrees with the type center trajectory to generate physical examination indicator trajectory classification output records.

[0019] In step one, the relevant data of the physical examination are read with the examinee as the collection object, and the read relevant data of the physical examination are organized into a physical examination indicator trajectory sample formed by the same examinee at multiple physical examination times.

[0020] Specifically, the original identity of the examinee, examination time, examination item name, original value of the examination indicator, unit of examination indicator, reference range, data collection institution identifier, and examination batch identifier are retrieved from the electronic medical examination record. Among these, the original identity of the examinee is used to distinguish different examinees; the examination time is used to determine the position of the original value of the examination indicator in the time series; the examination item name is used to determine the category of the examination indicator corresponding to the retrieved data; the original value of the examination indicator is used to form standard examination indicator values; the unit of examination indicator is used for unit conversion; the reference range is used to identify whether the original value of the examination indicator conforms to the recording range of the corresponding item; the data collection institution identifier is used to distinguish the data generating institution; and the examination batch identifier is used to distinguish medical examination records generated in different batches within the same institution.

[0021] The following information is retrieved from the test and examination records: test item name, original value of test indicator, test time, unit of test indicator, test method identifier, sample type identifier, and test source identifier. Specifically, the test item name is used to map to the physical examination item name; the original value of the test indicator is used to generate the corresponding standard physical examination indicator value; the test time is used to determine the time position corresponding to the original value of the test indicator; the unit of test indicator is used for unit conversion; the test method identifier is used to distinguish the meaning of data generated by the same test item under different test methods; the sample type identifier is used to distinguish the meaning of the indicator under blood, urine, or other sample sources; and the test source identifier is used to record the source of the test record to which the original value of the test indicator belongs.

[0022] The physical measurement record contains the name of the physical measurement item, the original physical measurement value, the measurement time, the measurement unit, the measurement method identifier, and the measurement source identifier. Specifically, the physical measurement item name identifies physical examination indicators such as height, weight, blood pressure, and heart rate; the original physical measurement value is used to generate corresponding standard physical examination indicator values; the measurement time determines the time position corresponding to the original physical measurement value; the measurement unit is used for unit conversion; the measurement method identifier distinguishes between data generated by manual or equipment measurement; and the measurement source identifier records the source of the physical measurement record to which the original physical measurement value belongs.

[0023] If a health questionnaire record exists, retrieve the questionnaire item name, questionnaire item value, completion time, and questionnaire source identifier from the health questionnaire record. The questionnaire item name is used to identify auxiliary variable categories such as lifestyle, medical history, and family history; the questionnaire item value is used to form auxiliary variable codes; the completion time is used to determine the time position corresponding to the auxiliary variable; and the questionnaire source identifier is used to record the source of the health questionnaire record to which the questionnaire item value belongs.

[0024] After data reading is completed, the original identity of the examinee is irreversibly anonymized to obtain an anonymous identity. This irreversible anonymization process uses a salted hashing method, concatenating the original identity of the examinee with a preset salt value and performing a hash operation. The hash result is then used as the anonymous identity. The preset salt value is pre-configured by the data processing environment executing this method and is not used as an output field for the physical examination indicator trajectory sample. For records from different data sources, if the original identity of the examinee yields the same anonymous identity after the same salted hashing process, they are classified as belonging to the same examinee; if no anonymous identity can be obtained, the corresponding record is marked as an identity mismatch record and is not included in the physical examination indicator trajectory sample construction in this step.

[0025] Subsequently, using the subject's anonymity identifier as the first index and the examination time, test time, measurement time, or completion time as the second index, the retrieved electronic medical examination records, test records, physical measurement records, and health questionnaire records are aggregated. For electronic medical examination records, the examination time is used as the observation time; for test records, the test time is used; for physical measurement records, the measurement time is used; and for health questionnaire records, the completion time is used. If multiple observation times exist under the same subject's anonymity identifier, they are arranged in chronological order to form the subject's timeline record.

[0026] Next, standard name mapping is performed on the names of physical examination items, test items, and physical measurement items. Standard name mapping is performed using a preset indicator name mapping table, which includes the original item name, the source of the original item, the detection method identifier, the sample type identifier, and the standard physical examination indicator name. If an original item name, its source, detection method identifier, and sample type identifier match a mapping relationship in the preset indicator name mapping table, then the original item name is converted to the standard physical examination indicator name corresponding to that mapping relationship. If the original item names are the same but the detection method identifier or sample type identifier is different, then they are converted to different standard physical examination indicator names according to different mapping relationships. If an original item name does not match the preset indicator name mapping table, the corresponding original value of the physical examination indicator is marked as a name-unmapped record and is not written into the physical examination indicator trajectory sample.

[0027] Standard unit conversion is performed on the units of physical examination indicators, test indicators, and measurement units. The standard unit conversion is performed using a preset unit conversion table, which includes the standard physical examination indicator name, original unit, target unit, unit conversion factor, and unit conversion bias. For any original value of a physical examination indicator that has already been mapped to a standard name, the preset unit conversion table is first consulted based on its standard physical examination indicator name and original unit. Then, the standard physical examination indicator value is obtained by multiplying the original value of the physical examination indicator by the unit conversion factor and adding the unit conversion bias. The unit conversion factor and unit conversion bias are both derived from the preset unit conversion table; the target unit is the unified unit of measurement corresponding to the standard physical examination indicator name; when the original unit and target unit are consistent, the unit conversion factor is set to one, and the unit conversion bias is set to zero. If the standard physical examination indicator name and original unit cannot be matched in the preset unit conversion table, the corresponding original value of the physical examination indicator is marked as an unconverted record and is not written into the physical examination indicator trajectory sample.

[0028] Auxiliary variable coding is performed on the questionnaire item values. A pre-set questionnaire coding table is used, which includes the questionnaire item name, the original questionnaire value, and the auxiliary variable code. If the questionnaire item name and the original questionnaire value match a coding relationship in the pre-set questionnaire coding table, the original questionnaire value is converted into the corresponding auxiliary variable code; if the questionnaire item value cannot match a coding relationship, the corresponding questionnaire item value is marked as an uncoded record and is not included in the physical examination indicator trajectory sample.

[0029] Furthermore, in order to enable the preset indicator name mapping table, preset unit conversion table, and preset questionnaire code table to be directly invoked by the execution environment, in this embodiment, the preset indicator name mapping table, preset unit conversion table, and preset questionnaire code table are all pre-stored in the form of structured fields.

[0030] The preset indicator name mapping table includes the original item name, the source of the original item, the detection method identifier, the sample type identifier, the standard physical examination indicator name, and the mapping status field. For example, when the original item name is fasting blood glucose, the source of the original item is electronic physical examination record, the detection method identifier is null, and the sample type identifier is blood, the corresponding standard physical examination indicator name is fasting blood glucose; when the original item name is GLU, the source of the original item is laboratory test record, the detection method identifier is biochemical test, and the sample type identifier is serum, the corresponding standard physical examination indicator name is fasting blood glucose; when the original item name is systolic blood pressure, the source of the original item is physical measurement record, the detection method identifier is electronic blood pressure monitor, and the sample type identifier is null, the corresponding standard physical examination indicator name is systolic blood pressure. If the same original item name has different meanings under different detection method identifiers or sample type identifiers, different standard physical examination indicator names are configured for each to avoid data from different detection conditions being incorrectly merged.

[0031] The preset unit conversion table includes fields for standard physical examination indicator name, original unit, target unit, unit conversion factor, unit conversion bias, and unit conversion status. For example, when the standard physical examination indicator name is fasting blood glucose, the original unit is mg / dL, and the target unit is mmol / L, the unit conversion factor is 0.0555, and the unit conversion bias is 0; when the standard physical examination indicator name is weight, the original unit is kg, and the target unit is kg, the unit conversion factor is 1, and the unit conversion bias is 0; when the standard physical examination indicator name is height, the original unit is cm, and the target unit is m, the unit conversion factor is 0.01, and the unit conversion bias is 0. During unit conversion, the original value of the physical examination indicator is multiplied by the unit conversion factor, and then the unit conversion bias is added to obtain the standard physical examination indicator value. If the same standard physical examination indicator name has an unrecognizable original unit, the data status identifier of the corresponding record is written to the "unit not converted" state, and the original indicator value is not written as a measured state into the physical examination indicator trajectory sample.

[0032] The pre-defined questionnaire coding table includes fields for questionnaire item name, original questionnaire value, auxiliary variable code, and coding status. For example, if the questionnaire item name is "Smoking Status" and the original questionnaire value is "Never Smoker," the auxiliary variable code is 0; if the original questionnaire value is "Past Smoker," the auxiliary variable code is 1; and if the original questionnaire value is "Current Smoker," the auxiliary variable code is 2. Similarly, if the questionnaire item name is "Exercise Frequency," and the original questionnaire value is "Less than once a week," the auxiliary variable code is 0; if the original questionnaire value is "One to three times a week," the auxiliary variable code is 1; and if the original questionnaire value is "More than three times a week," the auxiliary variable code is 2. If a questionnaire item value cannot be matched with the pre-defined questionnaire coding table, the corresponding questionnaire item value will be marked as an uncoded record and will not be included in the auxiliary variable code of the physical examination indicator trajectory sample.

[0033] To ensure that data status identifiers remain consistent across the dynamic trajectory samples of physical examination indicators, indicator modeling sequences, observation confidence levels, and trajectory distance calculations, in this embodiment, data status identifiers include at least the measured status, conflict resolution status, missing status, name not mapped status, unit not converted status, source conflict status, interpolation completion status, continuation completion status, and statistical completion status. The measured status indicates that the standard physical examination index value is directly obtained from the original index value after standard name mapping and unit conversion; the conflict resolution status indicates that there are multiple candidate standard physical examination index values ​​under the same anonymous identifier of the examinee, the same standardized physical examination time, and the same standard physical examination index name, and a unique value has been obtained based on the source priority, time proximity, or dispersion of the candidate standard physical examination index values; the missing status indicates that there is no measured value at the corresponding aligned time point or the source conflict has not been resolved; the name not mapped status indicates that the original item name does not match the preset index name mapping table; the unit not converted status indicates that the original unit does not match the preset unit conversion table; the source conflict status indicates that multiple candidate standard physical examination index values ​​cannot be uniquely determined according to the preset rules; the interpolation completion status, the continuation completion status, and the statistical completion status indicate that the corresponding current physical examination index value is obtained by linear interpolation of the measured values ​​before and after, continuation of one-sided values, or statistical completion of the group median, respectively.

[0034] Then, the observation times under the same anonymous identifier for the same subject are standardized to obtain the standardized physical examination time. The standardized physical examination time is obtained by dividing the time interval between the current observation time and the first observation time of the subject by a preset time scale. The current observation time comes from the physical examination time of the electronic physical examination record, the examination time of the laboratory test record, the measurement time of the physical measurement record, or the completion time of the health questionnaire record; the first observation time is the first observation time arranged in chronological order under the same anonymous identifier for the same subject; the preset time scale is the time unit set before data processing. The standardized physical examination time is used to represent the time position of each observation record for the same subject relative to the first observation record.

[0035] Subsequently, conflict resolution is performed on standard medical examination indicator values ​​under the same anonymous identifier for the same subject, the same standardized medical examination time, and the same standard medical examination indicator name. If only one standard medical examination indicator value exists, it is directly retained. If multiple standard medical examination indicator values ​​exist, the detection method identifier and sample type identifier are compared first. If the detection method identifier and sample type identifier are the same, a standard medical examination indicator value is selected according to the preset source priority. The preset source priority is set before data processing, and the source category of electronic medical examination records, laboratory test records, physical measurement records, or health questionnaire records is used as the sorting object. If multiple standard medical examination indicator values ​​have the same source priority, the standard medical examination indicator value whose observation time is closest to the original time corresponding to the standardized medical examination time is selected. If a unique value still cannot be determined, the value of the standard medical examination indicator name under the standardized medical examination time is marked as a source conflict state, and each candidate standard medical examination indicator value and its source identifier are retained.

[0036] After conflict resolution, a data source identifier and a data status identifier are generated for each standard physical examination indicator value written into the physical examination indicator trajectory sample. The data source identifier is a combination of source category, collection institution identifier or source identifier, physical examination batch identifier or testing method identifier; the source category is taken from electronic physical examination records, laboratory test records, physical measurement records, or health questionnaire records; the collection institution identifier comes from electronic physical examination records; the laboratory source identifier and testing method identifier come from laboratory test records; the measurement source identifier and measurement method identifier come from physical measurement records; and the questionnaire source identifier comes from health questionnaire records. The data status identifier is generated based on the data processing results. Values ​​that have completed standard name mapping, unit conversion, and conflict resolution and are determined to be unique are marked as measured; values ​​lacking corresponding observations are marked as missing; values ​​that cannot be mapped to standard names are marked as unmapped; values ​​that cannot be converted to units are marked as unconverted; and values ​​that cannot be determined to be unique are marked as source conflict.

[0037] Finally, a physical examination indicator trajectory sample is generated according to the anonymous identifier of the examinee, the standardized physical examination time, the name of the standard physical examination indicator, the value of the standard physical examination indicator, the data source identifier, the data status identifier, and the auxiliary variable code. One physical examination indicator trajectory sample corresponds to one anonymous identifier of the examinee; the sample includes multiple observation records under that anonymous identifier, sorted by the standardized physical examination time; each observation record includes the standard physical examination indicator value with actual measurement status, the name of the standard physical examination indicator with missing status, the candidate standard physical examination indicator value with source conflict status, and the auxiliary variable code corresponding to that standardized physical examination time. For the same anonymous identifier of the examinee, if it has multiple observation records corresponding to different standardized physical examination times, a physical examination indicator trajectory sample is generated for that examinee; if it has only one observation record corresponding to a standardized physical examination time, that observation record is saved as a single physical examination observation record, and no physical examination indicator trajectory sample is generated.

[0038] In step two, the physical examination indicator trajectory samples obtained in step one are processed to obtain dynamic trajectory samples of physical examination indicators at the same time scale. The dynamic trajectory samples of physical examination indicators refer to data samples formed by adding aligned physical examination indicator values, supplemented physical examination indicator values, normalized physical examination indicator values, and dynamic feature records to the physical examination indicator trajectory samples.

[0039] Specifically, the following information is first retrieved from the physical examination indicator trajectory sample: the anonymous identifier of the examinee, the standardized physical examination time, the name of the standard physical examination indicator, the value of the standard physical examination indicator, the data source identifier, the data status identifier, and the auxiliary variable code. The anonymous identifier of the examinee is used to distinguish the trajectory data of different examinees; the standardized physical examination time is used to indicate the time position of each observation record for the same examinee; the name of the standard physical examination indicator is used to determine the category of the physical examination indicator being processed; the value of the standard physical examination indicator is used to form the indicator value in the time series; the data source identifier is used to record the source of the standard physical examination indicator value; the data status identifier is used to distinguish between measured status, missing status, unmapped name status, unconverted unit status, and source conflict status; and the auxiliary variable code is used to describe questionnaire-related auxiliary information associated with the corresponding observation time.

[0040] After reading the physical examination indicator trajectory samples, the observation records under the same anonymous identifier of the examinee are sorted from first to last according to the standardized physical examination time. If there are multiple observation records with the same standardized physical examination time under the same anonymous identifier of the examinee, the multiple observation records are merged according to the standard physical examination indicator name. During merging, if there is only one measured value for the same standard physical examination indicator name, the standard physical examination indicator value is retained. If there are multiple measured values ​​for the same standard physical examination indicator name, the source category, collection agency identifier, detection method identifier, or measurement method identifier is read according to the data source identifier formed in step one, and a standard physical examination indicator value is selected according to the preset source priority. If a unique standard physical examination indicator value still cannot be selected, the value of the standard physical examination indicator name under the standardized physical examination time is kept in a source conflict state.

[0041] Then, a unified time grid is constructed. The unified time grid refers to a set of standardized time points arranged at equal time intervals. The unified time grid is generated by a preset alignment interval, a trajectory start time, and a trajectory end time; the preset alignment interval is a time interval parameter set before data processing; the trajectory start time is the earliest standardized physical examination time under the same anonymous identifier of the examined object; the trajectory end time is the latest standardized physical examination time under the same anonymous identifier of the examined object. For any anonymous identifier of the examined object, its unified time grid is represented by multiple standardized time points starting from the trajectory start time and increasing sequentially according to the preset alignment interval, up to a point not exceeding the trajectory end time. Each standardized time point in the unified time grid is called an alignment time point.

[0042] Next, the standard physical examination indicator values ​​in the physical examination indicator trajectory sample are mapped to a unified time grid. For any anonymous identifier of the examinee, any standard physical examination indicator name, and any aligned time point, if there exists a standardized physical examination time with the same aligned time point, and a standard physical examination indicator value in a measured state exists under that standardized physical examination time, then that standard physical examination indicator value is used as the aligned physical examination indicator value for that aligned time point, and the corresponding data status identifier is set to the measured state. If there is no standardized physical examination time with the same aligned time point, or if there is no standard physical examination indicator value in a measured state under that aligned time point, then the value of the standard physical examination indicator name under that aligned time point is set to null, and the corresponding data status identifier is set to the missing state. The aligned physical examination indicator value refers to the standard physical examination indicator value that has fallen into the corresponding aligned time point in the unified time grid.

[0043] For standard medical examination index values ​​in a source conflict state, first read the candidate standard medical examination index values ​​and their data source identifiers retained in that source conflict state. If the standard medical examination index name, detection method identifier, and sample type identifier corresponding to each candidate standard medical examination index value are consistent, then calculate the dispersion of each candidate standard medical examination index value; the dispersion is obtained by the difference between the maximum and minimum values ​​among the candidate standard medical examination index values. If the dispersion does not exceed a preset conflict tolerance value, then take the median of each candidate standard medical examination index value as the conflict resolution value, and use this conflict resolution value as the aligned medical examination index value at that alignment time point; the preset conflict tolerance value is preset by the measurement unit and recording precision corresponding to the standard medical examination index name. If the dispersion exceeds the preset conflict tolerance value, then no conflict resolution value is generated, and the value of the standard medical examination index name at that alignment time point is kept as null, and the data status identifier is set to missing state.

[0044] After time alignment and conflict resolution, missing health check index values ​​for the missing states are filled in to obtain the filled health check index values. The filled health check index values ​​refer to alternative values ​​calculated from the standard health check index values ​​of the measured states, conflict resolution values, or statistical values ​​under the same standard health check index name. For any anonymous identifier of the examined object, any standard health check index name, and any alignment time point of the missing state, the nearest measured state or conflict resolution state alignment health check index value before the alignment time point of the missing state is first searched under the same anonymous identifier of the examined object and the same standard health check index name, and then the nearest measured state or conflict resolution state alignment health check index value after the alignment time point of the missing state is searched. If the above-mentioned alignment health check index values ​​exist on both sides, and the time interval between the alignment time points on both sides does not exceed the preset maximum filling interval, then linear interpolation is used to calculate the filled health check index values.

[0045] Linear interpolation is calculated as follows: The supplementary physical examination index value = anterior physical examination index value + (posterior physical examination index value - anterior physical examination index value) × (missing alignment time point - anterior alignment time point) / (posterior alignment time point - anterior alignment time point).

[0046] Among them, the front-side physical examination index value is the aligned physical examination index value of the same standard physical examination index name under the anonymous identifier of the examinee, in the most recent measured state or conflict resolution state before the missing alignment time point; the back-side physical examination index value is the aligned physical examination index value of the same standard physical examination index name under the anonymous identifier of the examinee, in the most recent measured state or conflict resolution state after the missing alignment time point; the missing alignment time point is the alignment time point that needs to be filled in; the front-side alignment time point and the back-side alignment time point are the alignment time points corresponding to the front-side physical examination index value and the back-side physical examination index value, respectively; the preset maximum completion interval is the maximum allowed completion time interval set for the standard physical examination index name before data processing.

[0047] If the alignment time point of the missing state only has the preceding physical examination index value and no following physical examination index value, then it is determined whether the time interval between the preceding alignment time point corresponding to the preceding physical examination index value and the missing alignment time point does not exceed the preset unilateral continuation interval. If it does not exceed the preset unilateral continuation interval, then the preceding physical examination index value is used as the supplementary physical examination index value; if it exceeds the preset unilateral continuation interval, then no unilateral continuation completion is performed. If the alignment time point of the missing state only has the following physical examination index value and no preceding physical examination index value, then unilateral continuation completion is performed using the following physical examination index value according to the same judgment method. The preset unilateral continuation interval is the range of unilateral value continuation time set for the standard physical examination index name before data processing.

[0048] If the alignment time point of the missing state cannot be completed using either linear interpolation or one-sided continuation completion, then population statistical completion is used. Population statistical completion refers to selecting aligned physical examination index values ​​with the same standard physical examination index name, the same alignment time point, and data status identifier of measured state or conflict resolution state from all physical examination index trajectory samples, calculating the median, and using this median as the completed physical examination index value. If there is no aligned physical examination index value that can be used to calculate the median under the same standard physical examination index name and the same alignment time point, then the median of all aligned physical examination index values ​​under the same standard physical examination index name in measured state or conflict resolution state is calculated, and this median is used as the completed physical examination index value. The data status identifier of the completed physical examination index value obtained using population statistical completion is set to statistical completion state; the data status identifier of the completed physical examination index value obtained using linear interpolation or one-sided continuation completion is set to interpolation completion state or continuation completion state, respectively, such as... Figure 1 As shown.

[0049] After completing the missing data completion, the aligned physical examination index values ​​are normalized to obtain normalized physical examination index values. Scale normalization refers to converting physical examination index values ​​with different dimensions and numerical ranges under different standard physical examination index names into comparable numerical scales. For any standard physical examination index name, firstly, the physical examination index values ​​under that standard physical examination index name with data status identified as measured, conflict resolution, interpolation completion, or continuation completion are read from the entire physical examination index trajectory sample, and the median, first quartile, and third quartile are calculated; the median, first quartile, and third quartile are all obtained by sorting the existing physical examination index values ​​under that standard physical examination index name. Then, the normalized physical examination index values ​​are calculated as follows: Normalized physical examination index value = (current physical examination index value - index median) / (third quartile - first quartile + stability constant).

[0050] The current physical examination indicator value is the value after time alignment and missing data completion; the median is the median of the data participating in the calculation under the same standard physical examination indicator name; the third quartile and the first quartile are the third quartile and the first quartile of the data participating in the calculation under the same standard physical examination indicator name, respectively; the stability constant is a preset value greater than zero, used to keep the denominator non-zero when the third quartile and the first quartile are equal. If the number of physical examination indicator values ​​that can be used for calculation under a certain standard physical examination indicator name is insufficient to obtain the first quartile and the third quartile, then the preset scale parameter corresponding to the standard physical examination indicator name is used for normalization; the preset scale parameter includes a preset central value and a preset discrete value, the preset central value is used to replace the indicator median, and the preset discrete value is used to replace the difference between the third quartile and the first quartile.

[0051] After obtaining the normalized physical examination index values, dynamic feature records are generated. These dynamic feature records describe the changes in the same anonymous identifier of the same subject and the same standard physical examination index name between adjacent alignment time points. For any anonymous identifier of the subject, any standard physical examination index name, and any non-first alignment time point, the local change amount and local change rate are calculated. The local change amount is the normalized physical examination index value at the current alignment time point minus the normalized physical examination index value at the previous alignment time point; the local change rate is the local change amount divided by the time interval between the current alignment time point and the previous alignment time point. Both the current and previous alignment time points are from a unified time grid, the normalized physical examination index values ​​are obtained through scale normalization, and the time interval is obtained by subtracting adjacent alignment time points in the unified time grid.

[0052] For any anonymous identifier of the examined subject, any standard physical examination indicator name, and any alignment time point, the local trend slope is further calculated. The local trend slope refers to the linear direction and magnitude of the change in the normalized physical examination indicator value with respect to the alignment time point within a preset trend window near the current alignment time point. The preset trend window is the range of alignment time points set before data processing. During calculation, multiple alignment time points and their normalized physical examination indicator values ​​falling within the preset trend window for the same anonymous identifier of the examined subject and the same standard physical examination indicator name are first selected. Then, the slope is obtained by least-squares linear fitting, and this slope is used as the local trend slope. If the number of normalized physical examination indicator values ​​that can participate in the fitting within the preset trend window is insufficient, the local trend slope is set to a null value, and the corresponding dynamic feature state identifier is set to a trend-incalculable state.

[0053] Simultaneously, the reference deviation degree is calculated based on the reference range. The reference deviation degree refers to the magnitude of the deviation of the current physical examination indicator value from the reference range of the corresponding standard physical examination indicator name. The reference range is derived from the reference range field in the electronic physical examination record or test record read in step one, and converted to the same target unit as the standard physical examination indicator value during unit conversion. If the current physical examination indicator value is between the lower limit and the upper limit of the reference range, the reference deviation degree is zero; if the current physical examination indicator value is less than the lower limit of the reference range, the reference deviation degree is the difference between the lower limit and the current physical examination indicator value divided by the reference range width; if the current physical examination indicator value is greater than the upper limit of the reference range, the reference deviation degree is the difference between the current physical examination indicator value and the upper limit of the reference range divided by the reference range width. The reference range width is the upper limit of the reference range minus the lower limit of the reference range. If no available reference range exists for the corresponding standard physical examination indicator name, the reference deviation degree is not calculated, and the corresponding dynamic feature status is set to a reference range missing state.

[0054] Finally, the anonymity identifier of the examinee, the alignment time point, the name of the standard physical examination indicator, the current physical examination indicator value, the data source identifier, the data status identifier, the normalized physical examination indicator value, the local change amount, the local change rate, the local trend slope, the reference deviation degree, and the auxiliary variable code are written into the same dynamic observation record. All dynamic observation records under the same anonymity identifier of the examinee are arranged from first to last according to the alignment time point to form a dynamic trajectory sample of the physical examination indicators of that examinee.

[0055] In step three, dynamic data modeling is performed on the dynamic trajectory samples of physical examination indicators obtained in step two to obtain the trajectory type of physical examination indicators corresponding to each anonymous identifier of the examinee. The trajectory type of physical examination indicators refers to the type identifier obtained by classifying the overall change pattern formed by the changes of multiple standard physical examination indicator names under the same anonymous identifier of the examinee as the alignment time point.

[0056] Specifically, the following information is first retrieved from the dynamic trajectory sample of the physical examination indicators: the anonymous identifier of the examinee, the alignment time point, the standard physical examination indicator name, the current physical examination indicator value, the data source identifier, the data status identifier, the normalized physical examination indicator value, the local change amount, the local change rate, the local trend slope, the reference deviation degree, and the auxiliary variable code. The alignment time point is the time position in the unified time grid in step two; the normalized physical examination indicator value is the value obtained after scaling the current physical examination indicator value in step two; the local change amount, the local change rate, and the local trend slope are used to represent the direction and magnitude of change of the same standard physical examination indicator name between adjacent or nearby alignment time points; the reference deviation degree is used to represent the magnitude of the deviation of the current physical examination indicator value from the corresponding reference range; and the data status identifier is used to distinguish between the measured state, the conflict resolution state, the interpolation completion state, the continuation completion state, the statistical completion state, and the missing state.

[0057] After reading the dynamic trajectory samples of physical examination indicators, the dynamic observation records under the same anonymous identifier of the examinee and the same standard physical examination indicator name are arranged from first to last according to the alignment time point to form an indicator modeling sequence. The indicator modeling sequence refers to the data sequence arranged by time corresponding to a certain standard physical examination indicator name under the same anonymous identifier of the examinee. Each sequence position in the indicator modeling sequence corresponds to an alignment time point and includes the normalized physical examination indicator value, local change, local change rate, local trend slope, reference deviation degree, and data status identifier at that alignment time point. If the data status identifier of a certain sequence position is missing, then that sequence position will not participate in the calculation of gray prediction parameters and hidden state observation distribution; if the data status identifier of a certain sequence position is statistically complete, then when that sequence position participates in the calculation, a preset confidence weight lower than that of the measured state is used; the preset confidence weight is set according to the data status identifier before data processing, with the measured state and conflict resolution state corresponding to higher confidence weight, the interpolation completion state and continuation completion state corresponding to medium confidence weight, and the statistical completion state corresponding to lower confidence weight.

[0058] Subsequently, grey prediction processing is performed on each indicator modeling sequence to obtain predicted physical examination indicator values ​​and prediction residuals. Grey prediction processing refers to using an accumulated sequence to reduce local fluctuations and estimate the short-term trend of the standard physical examination indicator name when the number of observation points in the indicator modeling sequence is limited and the time change trend needs to be estimated. The predicted physical examination indicator value refers to the estimated value of the next aligned time point or the current aligned time point calculated by grey prediction processing based on existing normalized physical examination indicator values; the prediction residual refers to the difference between the actual normalized physical examination indicator value and the corresponding predicted physical examination indicator value.

[0059] In grey prediction processing, normalized health check index values, categorized as measured state, conflict resolution state, interpolation completion state, or continuation completion state, are first extracted from the index modeling sequence. These values ​​are then arranged in chronological order according to the aligned time points to form the original index sequence. If the number of valid sequence positions in the original index sequence reaches the preset minimum modeling quantity, a cumulative generation process is performed on the original index sequence. This cumulative generation process involves sequentially summing the normalized health check index values ​​from the starting sequence position to the current sequence position in the original index sequence to obtain the cumulative index sequence. For two adjacent sequence positions in the cumulative index sequence, their average value is taken as the background value; this background value represents the cumulative change level between two adjacent aligned time points.

[0060] Then, based on the original indicator sequence, background values, and corresponding preset confidence weights, weighted least squares processing is used to calculate the gray development coefficient and gray action quantity. The gray development coefficient represents the overall rate of change of the standard physical examination indicator name under the anonymous identifier of the examinee; the gray action quantity represents the basic intensity of change of the standard physical examination indicator name under the anonymous identifier of the examinee; the weighted least squares processing refers to using the preset confidence weights corresponding to different data state identifiers as weights during the fitting process, ensuring that the influence of the normalized physical examination indicator value of the measured state on the parameter calculation is greater than the normalized physical examination indicator value obtained after completion. After obtaining the gray development coefficient and gray action quantity, the predicted physical examination indicator value for each aligned time point is calculated according to the gray prediction equation, and the predicted physical examination indicator value is subtracted from the actual normalized physical examination indicator value corresponding to each aligned time point with an actual normalized physical examination indicator value to obtain the prediction residual.

[0061] If the number of valid sequence positions in a modeling sequence for a certain indicator does not reach the preset minimum modeling number, then gray prediction processing will not be performed, the predicted physical examination index value of the modeling sequence will be set to null, and the prediction residual will also be set to null. The preset minimum modeling number is a sequence length condition set before data processing, and its value is determined based on the minimum number of valid observations required for gray prediction processing.

[0062] Next, an observation feature vector is generated based on the normalized physical examination index value, local rate of change, local trend slope, reference deviation, and prediction residual. The observation feature vector refers to a data combination representing an anonymous identifier for an examined object, a standard physical examination index name, and dynamic change characteristics at an aligned time point. For any dynamic observation record, the normalized physical examination index value, local rate of change, local trend slope, reference deviation, and prediction residual corresponding to that dynamic observation record are arranged in a fixed order to form the observation feature vector; where the normalized physical examination index value represents the current level, the local rate of change represents the rate of change between adjacent aligned time points, the local trend slope represents the overall direction of change within a preset trend window, the reference deviation represents the degree of deviation relative to a reference range, and the prediction residual represents the degree of deviation between the actual change and the gray predicted change. If a dynamic observation record does not have a usable prediction residual, a preset residual null value identifier is written into the observation feature vector, and the dimension corresponding to the prediction residual is ignored when calculating the distance between the observation feature vectors.

[0063] Then, latent state transition modeling is performed. The latent state refers to the dynamic change state of the physical examination indicators that is not directly given by the original physical examination records but inferred from the observed feature vectors. Latent state transition modeling refers to calculating the temporal relationship of the latent states based on the observed feature vectors of adjacent aligned time points in the same indicator modeling sequence. The number of latent states is set before data processing, or selected using a silhouette coefficient within a preset candidate range. The silhouette coefficient is used to measure the proximity of observed feature vectors within the same latent state and the distinguishability of observed feature vectors between different latent states.

[0064] When modeling hidden state transitions, the anonymous identifiers of all examined subjects, the names of all standard physical examination indicators, and the available observation feature vectors at all aligned time points are first aggregated into an observation feature set. Clustering initialization is then performed on the observation feature set to obtain multiple initial hidden state centers; each initial hidden state center refers to the initial representative position of each hidden state in the observation feature space. The clustering initialization process employs a distance-based iterative partitioning method. Several initial centers are first selected from the observation feature set, and then the observation feature vectors are assigned to corresponding centers according to the distance between the observation feature vectors and the initial centers. The center positions are repeatedly updated until the change in center position is less than a preset convergence difference or a preset number of iterations is reached. The preset convergence difference and the preset number of iterations are calculation stopping conditions set before data processing.

[0065] After obtaining the initial hidden state center, the initial state probability, state transition probability, and observation generation parameters are calculated based on the sequential relationship between adjacent aligned time points in the same indicator modeling sequence. The initial state probability refers to the probability that the indicator modeling sequence belongs to a certain hidden state at the first available aligned time point; the state transition probability refers to the probability of transitioning from one hidden state to another between two adjacent available aligned time points; and the observation generation parameters refer to the distribution parameters of the observed feature vectors in a certain hidden state. The observation generation parameters include the mean and variance of each dimension of the observed feature vectors in that hidden state; the mean is obtained by averaging the observed feature vectors assigned to that hidden state; and the variance is obtained by averaging the squared deviations between the observed feature vectors assigned to that hidden state and their corresponding mean. When calculating the above probabilities and parameters, the corresponding preset confidence weights are read according to the data state identifier of the dynamic observation records, and these preset confidence weights are written into the probability statistics and mean / variance calculation process.

[0066] As an executable example, for the same subject with the anonymous identifier A001 and the same standard physical examination indicator name (fasting blood glucose), if the normalized physical examination indicator values ​​at alignment time points 0, 1, 2, and 3 are -0.20, 0.05, 0.32, and 0.65 respectively, and the corresponding data status identifiers are measured state, measured state, interpolation completion state, and measured state respectively, then a preset confidence weight is configured for each sequence position according to the data status identifier. The preset confidence weight corresponding to the measured state can be 1, and the preset confidence weight corresponding to the interpolation completion state can be 0.7. Subsequently, the original indicator sequence formed by the normalized physical examination indicator values ​​is accumulated to obtain the accumulated indicator sequence; then, a background value is generated based on the average of adjacent accumulated values, and the gray development coefficient and gray action are calculated using weighted least squares in combination with the preset confidence weights. The predicted physical examination indicator value corresponding to each alignment time point is calculated based on the gray development coefficient and gray action, and the predicted physical examination indicator value is subtracted from the actual normalized physical examination indicator value to obtain the prediction residual.

[0067] For example, if the normalized health index value at aligned time point 2 is 0.32, and the corresponding predicted health index value is 0.24, then the prediction residual for that aligned time point is 0.08; if the normalized health index value at aligned time point 3 is 0.65, and the corresponding predicted health index value is 0.41, then the prediction residual for that aligned time point is 0.24. Thus, the prediction residual can represent the degree of deviation of the actual change from the gray predicted change, and serves as input for the cumulative prediction deviation and trajectory representation data.

[0068] When constructing the observation feature vector, the vector is generated in a fixed order: normalized health index value, local rate of change, local trend slope, reference deviation, and prediction residual. For example, at a certain alignment time point, if the normalized health index value is 0.65, the local rate of change is 0.33, the local trend slope is 0.28, the reference deviation is 0.12, and the prediction residual is 0.24, then the corresponding observation feature vector is (0.65, 0.33, 0.28, 0.12, 0.24). If there is no available prediction residual at a certain alignment time point, a preset residual null value identifier is written into the observation feature vector, and the dimension corresponding to the prediction residual is ignored when calculating the distance between observation feature vectors or the observation generation probability.

[0069] When establishing a hidden state transition model, the hidden states can be categorized as stable, rising, falling, and fluctuating states. A stable state represents a state where the normalized physical examination index value changes little, the local rate of change is close to zero, and the prediction residual is small. A rising state represents a state where the overall local rate of change and local trend slope are greater than zero. A falling state represents a state where the overall local rate of change and local trend slope are less than zero. A fluctuating state represents a state where the local rate of change or the direction of the prediction residual changes frequently. The number of hidden states can also be selected using silhouette coefficients within a preset candidate range. After obtaining the initial hidden state centers through cluster initialization, the initial state probability, state transition probability, and observation generation parameters are statistically analyzed based on the state allocation results of adjacent aligned time points in the same index modeling sequence.

[0070] For example, if the observed feature vectors of the same indicator modeling sequence at alignment time points 0, 1, 2, and 3 successively approach a stable state, a stable state, an ascending state, and an ascending state, then the hidden state sequence of this indicator modeling sequence can be represented as a stable state, a stable state, an ascending state, and an ascending state. A hidden state change from a stable state to an ascending state occurs between adjacent alignment time points 1 and 2.

[0071] Subsequently, the hidden state sequence is solved for each indicator modeling sequence. The hidden state sequence refers to the arrangement of hidden states corresponding one-to-one with the alignment time points of the indicator modeling sequence. Dynamic programming is used to solve the hidden state sequence. During dynamic programming, at the first available alignment time point, the cumulative probability of each hidden state is calculated based on the initial state probability and the observation generation probability corresponding to the observation feature vector at that alignment time point. At each available alignment time point, the cumulative probability of the current hidden state is calculated based on the cumulative probability of the previous available alignment time point, the state transition probability, and the observation generation probability of the current alignment time point. After completing the calculation for all available alignment time points, the hidden state sequence corresponding to the indicator modeling sequence is obtained by backtracking along the path with the highest cumulative probability. For sequence positions with missing states, no hidden state is generated; for sequence positions with statistically completed states, the influence of the sequence position on the path probability is reduced according to the corresponding preset confidence weight.

[0072] After obtaining the hidden state sequence, trajectory representation data is generated for each anonymous identifier of the tested subject. The trajectory representation data refers to the dataset used to represent the dynamic change pattern of the physical examination indicators for that anonymous identifier of the tested subject. For the same anonymous identifier of the tested subject and the same standard physical examination indicator name, the following are first statistically analyzed: the proportion of hidden state retention, the number of hidden state transitions, the average local rate of change, the average local trend slope, the average reference deviation, the mean of the prediction residuals, and the fluctuation value of the prediction residuals corresponding to that standard physical examination indicator name. The hidden state retention ratio refers to the proportion of the number of times each hidden state appears under the standard physical examination indicator name to the total number of available hidden states under the standard physical examination indicator name; the number of hidden state transitions refers to the number of times the hidden state changes at adjacent available aligned time points under the standard physical examination indicator name; the average local change rate is the weighted average of the available local change rates under the standard physical examination indicator name; the average local trend slope is the weighted average of the available local trend slopes under the standard physical examination indicator name; the average reference deviation is the weighted average of the available reference deviation under the standard physical examination indicator name; the mean of the prediction residuals is the weighted average of the available prediction residuals under the standard physical examination indicator name; the prediction residual fluctuation is the weighted average of the squared deviations of the available prediction residuals under the standard physical examination indicator name relative to the mean of the prediction residuals. The weights in the above weighted averages are all taken from the preset confidence weights of the corresponding data state identifiers.

[0073] It should be noted that during the classification of physical examination indicator trajectories, multiple standard physical examination indicator names corresponding to the same anonymous identifier of the examinee usually have different collection frequencies, different data status identifiers, and different temporal changes. Although the current physical examination indicator value of some standard physical examination indicator names is close to the reference range, their local rate of change, local trend slope, or prediction residuals change continuously at multiple aligned time points; although there is a lot of measured data for some standard physical examination indicator names, their latent state sequences remain stable for a long time, and their ability to distinguish trajectory differences is weak; although the number of observations for some standard physical examination indicator names is small, their latent state sequences undergo continuous shifts, and their prediction residuals show continuous deviations.

[0074] In the above situations, if the physical examination index trajectory is classified solely based on the numerical differences between normalized physical examination index values, it is easy to classify the anonymous identifiers of examinees with similar current numerical levels but different dynamic change processes into the same physical examination index trajectory type; if the importance of the standard physical examination index name is determined solely based on the number of effective observations, it is easy to weaken the importance of the standard physical examination index name that is collected in low frequency but has obvious dynamic changes; if the judgment is made solely based on the reference deviation of a single alignment time point, it is easy to mistake random fluctuations for stable trajectory differences.

[0075] Therefore, in this embodiment, after obtaining the trajectory representation data, a dynamic participation value is generated for each standard physical examination indicator name. The dynamic participation value refers to the degree to which the same standard physical examination indicator name participates in the trajectory classification of the physical examination indicator at different alignment time points. The dynamic participation value is not a fixed value, but is determined jointly based on the data state identifier, local rate of change, local trend slope, reference deviation degree, prediction residual, and latent state sequence of the standard physical examination indicator name at the corresponding alignment time point.

[0076] Specifically, for any standard physical examination indicator name and any aligned time point, the observation confidence level is first calculated. The observation confidence level indicates the degree to which the value of the standard physical examination indicator name at that aligned time point can be used in the calculation. The observation confidence level is determined by the data status identifier; if the data status identifier is in the measured state or conflict resolution state, the observation confidence level is a first preset confidence value; if the data status identifier is in the interpolation completion state or continuation completion state, the observation confidence level is a second preset confidence value; if the data status identifier is in the statistical completion state, the observation confidence level is a third preset confidence value; if the data status identifier is in the missing state, the observation confidence level is zero. The first, second, and third preset confidence values ​​are all parameters set before data processing, and the first preset confidence value is greater than the second preset confidence value, and the second preset confidence value is greater than the third preset confidence value. The input of the observation confidence level comes from the data status identifier formed in step two.

[0077] Then, the duration of change is calculated. The duration of change indicates whether the changes of the same standard physical examination indicator name at multiple consecutive aligned time points have a consistent direction. When calculating the duration of change, the local rate of change and the local trend slope within a preset duration window are first read, centered on the current aligned time point; the preset duration window is the range of consecutive aligned time points set before data processing. If multiple local rates of change within the preset duration window have the same direction of change, and multiple local trend slopes have the same direction of change, the duration of change increases with the increase in the number of consistent directions; if the local rates of change and local trend slopes frequently change direction within the preset duration window, the duration of change decreases. Both the local rates of change and local trend slopes are derived from the dynamic feature records generated in step two, and the direction of change is obtained by determining whether the corresponding value is greater than zero, less than zero, or equal to zero.

[0078] Next, the cumulative prediction deviation is calculated. The cumulative prediction deviation represents the degree of accumulated deviation between the actual changes of the same standard physical examination indicator name at multiple consecutive aligned time points and the predicted changes obtained from gray prediction processing. When calculating the cumulative prediction deviation, the predicted residuals within the current aligned time point and a preset residual window prior to it are read; the preset residual window is a range of consecutive aligned time points set before data processing; the predicted residuals come from gray prediction processing. For aligned time points with available predicted residuals within the preset residual window, if the predicted residuals continuously maintain the same sign and the absolute value of the predicted residuals gradually increases, the cumulative prediction deviation is increased; if the signs of the predicted residuals alternate, or the absolute value of the predicted residuals does not show a continuous increase, the cumulative prediction deviation is decreased. The sign of the predicted residuals is obtained by determining whether the predicted residuals are greater than zero, less than zero, or equal to zero, and the absolute value of the predicted residuals is obtained by removing the sign from the predicted residuals.

[0079] Subsequently, the reference buffer deviation is calculated. The reference buffer deviation indicates whether the current physical examination indicator value has deviated from the allowable buffer zone relative to the reference range. When calculating the reference buffer deviation, the lower limit, upper limit, and width of the reference range for the corresponding standard physical examination indicator name are read; the lower and upper limits are obtained from the reference range field read in step one and converted to the correct unit; the width is the difference between the upper and lower limits. The range corresponding to a preset buffer ratio is expanded outward from the lower and upper limits respectively to obtain the lower and upper buffer boundaries; the preset buffer ratio is a proportional parameter set for the standard physical examination indicator name before data processing. If the current physical examination indicator value is between the lower and upper buffer boundaries, the reference buffer deviation is set to a lower value; if the current physical examination indicator value is lower than the lower buffer boundary or higher than the upper buffer boundary, the reference buffer deviation increases with the deviation distance. The current physical examination indicator value comes from the dynamic observation record in step two.

[0080] Next, the hidden state transition confirmation quantity is calculated. This quantity indicates whether the state changes in the hidden state sequence are consistent with the dynamic changes of the physical examination indicators. When calculating the hidden state transition confirmation quantity, the hidden state, local rate of change, local trend slope, and prediction residual corresponding to the current aligned time point and its adjacent aligned time points under the same standard physical examination indicator name are read. If a hidden state change occurs between adjacent aligned time points, and the preceding and following aligned time points corresponding to the hidden state change simultaneously exhibit changes in the direction of the local rate of change, an increase in the magnitude of the local trend slope, or an increase in the absolute value of the prediction residual, the hidden state transition confirmation quantity is increased. If a hidden state change occurs between adjacent aligned time points, but the corresponding local rate of change, local trend slope, and prediction residual do not change accordingly, the hidden state transition confirmation quantity is decreased. The hidden state comes from the hidden state sequence solution result in step three, while the local rate of change, local trend slope, and prediction residual come from the dynamic feature record in step two and the gray prediction processing result in step three, respectively.

[0081] After obtaining the observed confidence level, duration of change, cumulative predicted deviation, reference buffer deviation, and latent state transition confirmation, the dynamic participation value of the indicator is generated. During generation, the observed confidence level is first used as the basic participation condition; if the observed confidence level is zero, the dynamic participation value of the standard physical examination indicator name at the alignment time point is zero. If the observed confidence level is greater than zero, the participation adjustment value is determined based on the duration of change, cumulative predicted deviation, reference buffer deviation, and latent state transition confirmation. The participation adjustment value increases with the increase of the duration of change, cumulative predicted deviation, and latent state transition confirmation; when the reference buffer deviation is generated only by a single alignment time point and the duration of change is lower than a preset duration threshold, the participation adjustment value is decreased; when the reference buffer deviation and cumulative predicted deviation increase simultaneously at consecutive alignment time points, the participation adjustment value is increased. The preset duration threshold is a parameter set before data processing to determine whether a continuous change is valid. Multiplying the observed confidence level by the participation adjustment value yields the dynamic participation value of the standard physical examination indicator name at the alignment time point.

[0082] To ensure that the observation confidence quantity, change duration quantity, prediction deviation accumulation quantity, reference buffer deviation quantity, hidden state transition confirmation quantity, participation adjustment value, and indicator dynamic participation quantity can be specifically implemented, this embodiment further provides the following value selection method.

[0083] The observation confidence level is determined by the data status identifier. When the data status identifier is in the measured state or conflict resolution state, the observation confidence level is the first preset confidence value; when the data status identifier is in the interpolation completion state or continuation completion state, the observation confidence level is the second preset confidence value; when the data status identifier is in the statistical completion state, the observation confidence level is the third preset confidence value; when the data status identifier is in the missing state, name unmapped state, unit unconverted state, or source conflict state, the observation confidence level is zero. As an executable example, the first preset confidence value is 1, the second preset confidence value is 0.7, the third preset confidence value is 0.4, and the observation confidence level corresponding to the missing state, name unmapped state, unit unconverted state, or source conflict state is 0. The above values ​​are used to represent the confidence level of the trajectory classification corresponding to different data status identifier values, and the first preset confidence value is greater than the second preset confidence value, and the second preset confidence value is greater than the third preset confidence value.

[0084] The duration of change is used to characterize whether the local rate of change and the local trend slope of the same standard physical examination indicator name remain in the same direction within a preset duration window. In this embodiment, at least one alignment time point is selected before and after the current alignment time point to form a preset duration window. The number of local rates of change in the same direction and the number of local trend slopes in the same direction within the preset duration window are counted, and both are normalized to between 0 and 1 to obtain the duration of change. If the local rates of change within the preset duration window are all greater than zero and the local trend slopes are all greater than zero, or if the local rates of change are all less than zero and the local trend slopes are all less than zero, then the duration of change takes a higher value; if the local rates of change and local trend slopes change frequently, then the duration of change takes a lower value. For example, for three consecutive alignment time points, when the local change rates are 0.10, 0.12, and 0.15, and the local trend slopes are 0.08, 0.09, and 0.11, the duration of change can be taken as 1; if the local change rates are 0.10, -0.08, and 0.05, and the local trend slopes are 0.03, -0.04, and 0.02, the duration of change can be taken as 0.33.

[0085] The cumulative predicted deviation is used to characterize the continuity of the same sign and the increasing absolute value of the predicted residuals for the same standard physical examination indicator name within a preset residual window. In this embodiment, it is first determined whether the signs of the predicted residuals within the preset residual window are consistent, and then it is determined whether the absolute value of the predicted residuals increases with the alignment time point. If the predicted residuals maintain the same sign within the preset residual window and the absolute value of the predicted residuals gradually increases, then the cumulative predicted deviation takes a higher value; if the signs of the predicted residuals change alternately or the absolute value of the predicted residuals does not continuously increase, then the cumulative predicted deviation takes a lower value. For example, when the predicted residuals are 0.05, 0.09, and 0.14 respectively, the cumulative predicted deviation can be 1; when the predicted residuals are 0.05, -0.03, and 0.04 respectively, the cumulative predicted deviation can be 0.33. Alignment time points with null predicted residuals are not included in the calculation of the cumulative predicted deviation.

[0086] The reference buffer deviation is used to characterize the degree of deviation of the current physical examination index value from the upper and lower buffer boundaries formed by the expansion of the reference range. In this embodiment, the lower and upper buffer boundaries are used as the allowable buffer areas. If the current physical examination index value is between the lower and upper buffer boundaries, the reference buffer deviation is 0. If the current physical examination index value is lower than the lower buffer boundary, the reference buffer deviation is obtained by dividing the difference between the lower buffer boundary and the current physical examination index value by the width of the reference range. If the current physical examination index value is higher than the upper buffer boundary, the reference buffer deviation is obtained by dividing the difference between the current physical examination index value and the upper buffer boundary by the width of the reference range. When the reference buffer deviation is greater than 1, it can be truncated to 1. As an executable example, the lower limit of the reference range for a certain standard physical examination indicator is 3.9, the upper limit is 6.1, the width is 2.2, and the preset buffer ratio is 0.1. Then the lower buffer boundary is 3.68 and the upper buffer boundary is 6.32. If the current physical examination indicator value is 6.76, then the reference buffer deviation is (6.76-6.32) / 2.2.

[0087] The latent state transition confirmation value is used to characterize whether a change in the latent state is supported by changes in the local rate of change, the local trend slope, and the predicted residual. In this embodiment, if no latent state change occurs between adjacent alignment time points, the latent state transition confirmation value is 0; if a latent state change occurs between adjacent alignment time points, and at least two of the following three conditions are met: a change in the direction of the local rate of change, an increase in the magnitude of the local trend slope, and an increase in the absolute value of the predicted residual, the latent state transition confirmation value is a higher value; if only one condition is met, the latent state transition confirmation value is an intermediate value; if none of the three conditions are met, the latent state transition confirmation value is a lower value. As an executable example, the higher value is 1, the intermediate value is 0.5, and the lower value is 0.2.

[0088] The adjustment value is generated jointly by the duration of change, the cumulative predicted deviation, the reference buffer deviation, and the latent state transition confirmation. As an executable example, the adjustment value equals the weighted sum of the duration of change, the cumulative predicted deviation, the reference buffer deviation, and the latent state transition confirmation, where the weights for the duration of change, the cumulative predicted deviation, the reference buffer deviation, and the latent state transition confirmation are all 0.3 and 0.2 respectively. If the reference buffer deviation is generated only by a single alignment time point and the duration of change is below a preset duration threshold, the adjustment value is multiplied by 0.5; if the reference buffer deviation and the cumulative predicted deviation increase simultaneously at consecutive alignment time points, the adjustment value is multiplied by 1.2, and any result exceeding 1 is truncated to 1. The preset duration threshold can be 0.5.

[0089] The dynamic participation of an indicator is obtained by multiplying the observation confidence level by the participation adjustment value. If the observation confidence level is 0, the dynamic participation of the indicator is 0; if the observation confidence level is greater than 0, the dynamic participation of the indicator increases with the increase of the participation adjustment value. Therefore, data in the measured state or conflict resolution state have a higher degree of participation in trajectory classification when there are continuous changes, continued deviations in prediction residuals, deviations outside the reference range buffer boundary, or confirmation of hidden state transitions; data in the statistical completion state, even if there are changes, will have a lower degree of participation due to the lower observation confidence level.

[0090] Subsequently, the dynamic participation of the same standard physical examination indicator name at all aligned time points is aggregated over time to obtain the indicator trajectory participation. The indicator trajectory participation represents the degree to which the standard physical examination indicator name participates in the trajectory categorization within the entire dynamic trajectory sample of physical examination indicators. During time aggregation, the dynamic participation of the standard physical examination indicator name is first arranged in order of alignment time points; then, consecutive high participation intervals are identified. A consecutive high participation interval refers to a time interval where the dynamic participation of the indicator at multiple consecutive aligned time points is not lower than a preset participation threshold; the preset participation threshold is a participation degree judgment parameter set before data processing. For standard physical examination indicator names with consecutive high participation intervals, the indicator trajectory participation is determined according to the length of the consecutive high participation interval, the average value of the indicator dynamic participation within the interval, and the number of hidden state transitions within the interval; for standard physical examination indicator names without consecutive high participation intervals, the indicator trajectory participation is determined according to the median of the indicator dynamic participation at all aligned time points. If the observation confidence of a certain standard physical examination indicator name is zero at all aligned time points, its indicator trajectory participation is zero.

[0091] Then, a trajectory-driven profile is generated for each anonymous identifier of the examined subject. The trajectory-driven profile refers to a data set describing the dynamic trajectory samples of the physical examination indicators of the anonymous identifier of the examined subject, which standard physical examination indicator names, which aligned time points, and which dynamic change causes jointly drive the formation. The trajectory-driven profile includes the standard physical examination indicator name, continuous high-involvement interval, indicator trajectory participation amount, main driving cause, and driving direction. The main driving cause is determined based on the higher value among the change duration, prediction deviation accumulation, reference buffer deviation, and hidden state transition confirmation amount; if the change duration is high, the main driving cause is recorded as continuous change driving; if the prediction deviation accumulation is high, the main driving cause is recorded as prediction deviation driving; if the reference buffer deviation is high, the main driving cause is recorded as reference deviation driving; if the hidden state transition confirmation amount is high, the main driving cause is recorded as state transition driving. The driving direction is determined by the sign of the local trend slope within a continuous high-participation interval; if the overall local trend slope is greater than zero, the driving direction is recorded as an upward direction; if the overall local trend slope is less than zero, the driving direction is recorded as a downward direction; if the local trend slope alternates between positive and negative and there is no dominant sign, the driving direction is recorded as a fluctuating direction, such as... Figure 2 As shown.

[0092] Next, the dynamic trajectory difference between any two anonymous identifiers of the tested subjects is calculated. The dynamic trajectory difference refers to the differences in numerical changes, continuous changes, prediction deviations, reference deviations, and latent state transitions of the dynamic trajectory samples of the physical examination indicators corresponding to the anonymous identifiers of the two tested subjects. To ensure that level differences, continuous differences, residual differences, reference differences, state differences, process priority adjustment amounts, indicator trajectory differences, and trajectory distances can be specifically implemented, this embodiment further provides the following calculation method.

[0093] For any standard health checkup indicator name in the shared indicator set, first determine the set of aligned time points where the anonymized identifiers of the two subjects share available dynamic observation records under that standard health checkup indicator name. If the number of aligned time points with available dynamic observation records is less than the preset minimum comparison number, then that standard health checkup indicator name will not participate in the calculation of the difference in indicator trajectories between the two subjects. The preset minimum comparison number is a quantity parameter set before data processing, for example, 2.

[0094] The level difference is determined by the difference in normalized health indicator values ​​between the anonymized identifiers of two subjects within a common alignment time point set. As an executable example, the absolute difference between the two normalized health indicator values ​​at each common alignment time point is first calculated, then the average of all absolute differences is calculated, and this average is used as the level difference. If this average is greater than 1, the level difference can be truncated to 1.

[0095] The persistence difference is determined by the differences between the duration of change, the local rate of change, and the local trend slope. As an executable example, the persistence difference equals the average of the differences in the duration of change, the differences in the direction of the local rate of change, and the differences in the direction of the local trend slope; where the difference in the duration of change is the absolute difference in the duration of change corresponding to the anonymized identifiers of the two tested objects, the difference in the direction of the local rate of change is 1 when the directions of the local rates of change are opposite and 0 when the directions are the same, and the difference in the direction of the local trend slope is 1 when the directions of the local trend slopes are opposite and 0 when the directions are the same.

[0096] The residual variance is determined by the differences between the cumulative predicted deviation, the mean predicted residual, and the fluctuation of the predicted residual. As an executable example, the residual variance equals the weighted average of the differences in the cumulative predicted deviation, the mean predicted residual, and the fluctuation of the predicted residual; where the cumulative predicted deviation difference is the absolute difference between the cumulative predicted deviations corresponding to the anonymized identifiers of the two tested objects, the mean predicted residual difference is the normalized result of the absolute difference between the means of the two predicted residuals, and the fluctuation of the predicted residual difference is the normalized result of the absolute difference between the fluctuations of the two predicted residuals. If the predicted residual at a certain alignment time point is null, then that alignment time point is not included in the residual variance calculation.

[0097] The reference difference is determined by the difference between the reference buffer deviation and the average reference deviation. As an executable example, the reference difference equals the average of the absolute differences between the reference buffer deviations corresponding to the anonymized identifiers of two subjects and the absolute differences between the average reference deviations. If no available reference range exists for a given standard health indicator name, the reference difference is not included in the calculation of the indicator trajectory difference for that standard health indicator name.

[0098] The state difference is determined by the differences between the hidden state transition confirmation quantity, the hidden state dwell ratio, and the number of hidden state transitions. As an executable example, the state difference is equal to the weighted average of the differences in hidden state transition confirmation quantity, hidden state dwell ratio, and number of hidden state transitions; where the hidden state dwell ratio difference is the distance between the hidden state dwell ratio vectors corresponding to the anonymous identifiers of the two inspected objects, and the hidden state transition number difference is the normalized result of the difference between the two hidden state transition numbers.

[0099] After obtaining the level difference, persistence difference, residual difference, reference difference, and state difference, an indicator difference structure is generated based on preset level difference thresholds, preset process difference thresholds, preset residual difference thresholds, and preset state difference thresholds. As an executable example, if the level difference is less than the preset level difference threshold, and at least one of the persistence difference, residual difference, or state difference is greater than its corresponding threshold, the indicator difference structure is recorded as a near-level, different-process structure; if the level difference is not less than the preset level difference threshold, and the persistence difference, residual difference, and state difference are all less than their corresponding thresholds, the indicator difference structure is recorded as a different-level, near-process structure; if the level difference is not less than the preset level difference threshold, and at least one of the persistence difference, residual difference, or state difference is greater than its corresponding threshold, the indicator difference structure is recorded as a different-level, different-process structure; if the level difference is less than the preset level difference threshold, and the persistence difference, residual difference, and state difference are all less than their corresponding thresholds, the indicator difference structure is recorded as a near-level, near-process structure.

[0100] The process priority adjustment amount is determined based on the indicator difference structure. As an executable example, when the indicator difference structure is a near-level, different-process structure, the process priority adjustment amount is 0.5; when the indicator difference structure is a different-level, near-process structure, the process priority adjustment amount is 0.1; when the indicator difference structure is a different-level, different-process structure, the process priority adjustment amount is the maximum value among the persistence difference, residual difference, and state difference multiplied by 0.3; when the indicator difference structure is a near-level, near-process structure, the process priority adjustment amount is 0.

[0101] The indicator trajectory variance is determined by a combination of level variance, persistence variance, residual variance, reference variance, state variance, process priority adjustment, observation confidence level, indicator trajectory participation, temporal overlap of consecutive high-participation intervals, and driving direction relationship. As an executable example, the basic variance is first obtained by weighting the level variance, persistence variance, residual variance, reference variance, and state variance; where the weights are: level variance 0.25, persistence variance 0.25, residual variance 0.2, reference variance 0.1, and state variance 0.2. If the indicator variance structure is a near-level, different-process structure, the weights of persistence variance, residual variance, and state variance are increased, while the total weight remains unchanged by decreasing the weight of level variance. Then, the basic variance is added to the process priority adjustment to obtain the process-adjusted variance; this is then multiplied by the average of the observation confidence levels and the average of the indicator trajectory participation under the same standard health check indicator name for the anonymized identifiers of the two subjects to obtain the indicator trajectory variance. If the driving directions of the two subjects' anonymous identifiers under the name of the standard physical examination indicator are opposite, the indicator trajectory difference is increased; if the time overlap of the consecutive high participation intervals of the two subjects' anonymous identifiers is low, the indicator trajectory difference is decreased or increased according to the time overlap.

[0102] After obtaining the indicator trajectory difference of each standard physical examination indicator name in the common participation indicator set, the standard physical examination indicator names whose cumulative indicator trajectory participation reaches a preset coverage ratio are selected as the dominant participation indicator set, sorted from high to low according to the indicator trajectory participation. The preset coverage ratio can be 0.8. For each standard physical examination indicator name in the dominant participation indicator set, the difference participation ratio is determined based on its indicator trajectory participation and observation confidence. As an executable example, the difference participation ratio is equal to the product of the indicator trajectory participation and observation confidence corresponding to the standard physical examination indicator name, divided by the sum of the products of the indicator trajectory participation and observation confidence corresponding to all standard physical examination indicator names in the dominant participation indicator set. Then, the indicator trajectory difference of each standard physical examination indicator name is multiplied by the corresponding difference participation ratio and summed to obtain the trajectory distance between the anonymity identifiers of the two examined objects. If there is at least one standard physical examination indicator name in the dominant participation indicator set whose indicator difference structure is a near-horizontal heterogeneous process structure, then the process difference identifier is written into the trajectory distance.

[0103] For any two anonymous test subjects, first determine the common standard medical examination indicator names that both have a greater than zero indicator trajectory participation value, thus obtaining a common participation indicator set. The common participation indicator set refers to the set of standard medical examination indicator names for which both anonymous test subjects have available dynamic participation results. If the common participation indicator set is empty, the dynamic trajectory difference between the two is not calculated, and their distance status is marked as incomparable.

[0104] For any standard health checkup indicator name in the shared indicator set, the level difference, persistence difference, residual difference, reference difference, and state difference are calculated respectively. The level difference is obtained by the difference between the normalized health checkup indicator values ​​of the two anonymous identifiers under that standard health checkup indicator name; the persistence difference is obtained by the difference between the duration of change, local rate of change, and local trend slope of the two anonymous identifiers under that standard health checkup indicator name; the residual difference is obtained by the difference between the cumulative predicted deviation, mean predicted residual, and predicted residual fluctuation value of the two anonymous identifiers under that standard health checkup indicator name; the reference difference is obtained by the difference between the reference buffer deviation and average reference deviation of the two anonymous identifiers under that standard health checkup indicator name; and the state difference is obtained by the difference between the number of confirmed latent state transitions, the proportion of latent state retention, and the number of latent state transitions of the two anonymous identifiers under that standard health checkup indicator name. The above differences are all calculated within the range of aligned time points where both have available dynamic observation records; if the data status of any anonymous identifier of an examined object is missing at a certain aligned time point, then that aligned time point will not participate in the difference calculation under the name of that standard physical examination indicator.

[0105] After obtaining the level difference, persistence difference, residual difference, reference difference, and state difference, an indicator difference structure is generated. This indicator difference structure represents the source of difference between two trajectories under the same standard medical examination indicator name. If the level difference is small but the persistence difference, residual difference, or state difference is large, the indicator difference structure corresponding to that standard medical examination indicator name is recorded as a near-level, different-process structure; this near-level, different-process structure indicates that the current numerical levels of the anonymous identifiers of two subjects are similar, but their dynamic change processes are different. If the level difference is large but the persistence difference, residual difference, and state difference are all small, it is recorded as a different-level, near-process structure; this different-level, near-process structure indicates that the numerical levels of the anonymous identifiers of two subjects are different, but their change processes are similar. If the level difference, persistence difference, residual difference, and state difference are all large, it is recorded as a different-level, different-process structure. If all the above differences are small, it is recorded as a near-level, near-process structure. The thresholds used to determine the magnitude of the difference are set separately for each type of difference before data processing.

[0106] Subsequently, a process priority adjustment amount is generated based on the indicator difference structure. This process priority adjustment amount is used to enhance the impact of dynamic changes on trajectory differences when the anonymized identifier values ​​of two tested objects are similar but their change processes differ. If the indicator difference structure is a near-level, different-process structure, the process priority adjustment amount takes a higher value; if the indicator difference structure is a different-level, near-process structure, the process priority adjustment amount takes a lower value; if the indicator difference structure is a different-level, different-process structure, the process priority adjustment amount is determined based on the largest of the persistence difference, residual difference, and state difference; if the indicator difference structure is a near-level, near-process structure, the process priority adjustment amount is zero. The value range of the process priority adjustment amount and the adjustment rules corresponding to various indicator difference structures are set before data processing.

[0107] Next, the difference in indicator trajectories between the two anonymous identifiers of the tested subjects under the name of the standard medical examination indicator is generated. The difference in indicator trajectory is determined by level difference, persistence difference, residual difference, reference difference, state difference, and process priority adjustment. When the process priority adjustment increases, the impact of persistence difference, residual difference, and state difference on the difference in indicator trajectory increases; when the observation confidence level is low or the indicator trajectory participation level is low, the difference in indicator trajectory corresponding to the name of the standard medical examination indicator decreases; when both anonymous identifiers of the tested subjects have consecutive high participation intervals under the name of the standard medical examination indicator, and the two consecutive high participation intervals occur at similar alignment time points but with opposite driving directions, the difference in indicator trajectory is increased; when the two consecutive high participation intervals occur at significantly different alignment time points, the difference in indicator trajectory is adjusted according to the degree of temporal overlap between the two consecutive high participation intervals. The degree of temporal overlap is obtained by dividing the number of alignment time points shared by the two consecutive high participation intervals by the number of alignment time points after merging the two consecutive high participation intervals.

[0108] After obtaining the trajectory difference of each standard medical examination indicator name in the common participation indicator set, the trajectory distance between the anonymous identifiers of two subjects is generated. When generating the trajectory distance, instead of directly treating all standard medical examination indicator names equally, the standard medical examination indicator names in the common participation indicator set are first arranged in descending order of their trajectory participation. Then, the standard medical examination indicator names that rank highest and whose cumulative trajectory participation reaches a preset coverage ratio are selected to form the dominant participation indicator set. The preset coverage ratio is a ratio parameter set before data processing; the dominant participation indicator set refers to the set of standard medical examination indicator names that plays a major role in the trajectory difference formation between the anonymous identifiers of the two subjects. For the standard medical examination indicator names in the dominant participation indicator set, the difference participation ratio is determined according to their trajectory participation and observation confidence level; then, the corresponding trajectory difference is synthesized based on the difference participation ratio to obtain the trajectory distance. If at least one standard medical examination indicator name in the dominant participation indicator set has a near-horizontal heterogeneous process structure for its indicator difference, a process difference identifier is written into the trajectory distance; the process difference identifier indicates that the trajectory distance is mainly formed by dynamically changing process differences.

[0109] Then, trajectory classification is performed based on trajectory distance and trajectory-driven profiles. During trajectory classification, each anonymous identifier of the tested subject is initially treated as an initial trajectory cluster. A trajectory cluster refers to a group of anonymous identifiers of the tested subject that are close to each other in terms of trajectory distance and trajectory-driven profile. Before merging trajectory clusters, the inter-cluster trajectory distance and inter-cluster driving consistency between two trajectory clusters are calculated. The inter-cluster trajectory distance is determined by the median of the trajectory distances between any two anonymous identifiers of the tested subject within the two trajectory clusters; if the distance between two anonymous identifiers of the tested subject is not comparable, then that pair of anonymous identifiers is not included in the inter-cluster trajectory distance calculation. The inter-cluster driving consistency is used to indicate whether the trajectory differences between the anonymous identifiers of the tested subject in two trajectory clusters are caused by similar standard medical examination indicator names, similar continuous high-involvement intervals, and similar driving directions; the inter-cluster driving consistency is jointly determined by the degree of overlap of the dominant standard medical examination indicator names, the degree of overlap of continuous high-involvement intervals, and the degree of consistency of driving directions in the trajectory-driven profiles of the two trajectory clusters.

[0110] If the inter-cluster trajectory distance between two trajectory clusters is less than a preset merging distance threshold, and the inter-cluster driving consistency is not lower than a preset driving consistency threshold, then the two trajectory clusters are merged. The preset merging distance threshold limits the maximum trajectory difference between mergeable trajectory clusters, and the preset driving consistency threshold limits the dynamic driving similarity between mergeable trajectory clusters; both are parameters set before data processing. If the inter-cluster trajectory distance between two trajectory clusters is small, but the inter-cluster driving consistency is lower than the preset driving consistency threshold, then the two trajectory clusters are not merged, and the relationship between the two trajectory clusters is marked as a close-range, different-driving relationship. The close-range, different-driving relationship indicates that although the overall trajectory distance between the two trajectory clusters is close, the names of the standard physical examination indicators that form the trajectory difference, the continuous high-involvement intervals, or the driving direction are different.

[0111] During trajectory cluster merging, the trajectory driving profile of the merged trajectory cluster is recalculated after each merging. During recalculation, the frequency of each standard health indicator name appearing as the dominant standard health indicator name in the merged trajectory cluster is first counted. Then, the continuous high-involvement intervals and driving directions corresponding to each standard health indicator name are counted. Standard health indicator names with high frequency of occurrence, high overlap of continuous high-involvement intervals, and consistent driving directions are determined as the dominant standard health indicator names of the merged trajectory cluster. If no standard health indicator name meets the above conditions in the merged trajectory cluster, the trajectory cluster is marked as having a dispersed driving state. Trajectory clusters marked as having dispersed driving states need to simultaneously meet a smaller preset merging distance threshold and a higher preset driving consistency threshold during merging; both the smaller preset merging distance threshold and the higher preset driving consistency threshold are set before data processing.

[0112] When no two trajectory clusters meet the merging conditions, or when the number of trajectory clusters reaches the preset number of trajectory types, trajectory classification processing stops. After stopping, each trajectory cluster is determined as a physical examination indicator trajectory type, and a trajectory type identifier is generated for each physical examination indicator trajectory type. For each physical examination indicator trajectory type, the type center trajectory for that physical examination indicator trajectory type is calculated. The type center trajectory refers to the dynamic trajectory sample of the physical examination indicator corresponding to the anonymous identifier of the examined object within the physical examination indicator trajectory type that simultaneously meets the following conditions: small average trajectory distance, trajectory driving profile consistent with the name of the dominant standard physical examination indicator of the physical examination indicator trajectory type, and stable with a continuous high participation interval. If multiple anonymous identifiers of examined objects meet the above conditions, the anonymous identifier of the examined object with the largest amount of available trajectory representation data is selected as the type center trajectory.

[0113] For anonymous subject identifiers that fail to be included in any physical examination indicator trajectory type, their trajectory-driven profiles and trajectory distances to the central trajectories of each type are read. If the trajectory distance between the anonymous subject identifier and a certain type's central trajectory is not greater than a preset proximity threshold, and both trajectory-driven profiles contain the same dominant standard physical examination indicator name, overlapping continuous high-participation intervals, and consistent driving directions, then the physical examination indicator trajectory type is written into the candidate trajectory type identifier, and the trajectory type identifier of the anonymous subject identifier is set to a pending confirmation state. If no physical examination indicator trajectory type meets the above conditions, then the trajectory type identifier of the anonymous subject identifier is set to a pending classification state, and its non-attribution reason is recorded as incomparable distance, inconsistent driving profiles, or insufficient available dynamic observation records.

[0114] Finally, the anonymity identifier of the examined subject, trajectory type identifier, candidate trajectory type identifier, trajectory-driven profile, indicator dynamic participation, indicator trajectory participation, indicator difference structure, process difference identifier, trajectory distance, distance status, and type center trajectory identifier are written into the physical examination indicator trajectory classification results. For the anonymity identifier of the examined subject in the pending classification or pending confirmation status, the corresponding reason for non-attribution or candidate attribution reason is also written.

[0115] After obtaining the trajectory distance, the anonymous identifiers of the inspected objects are categorized by trajectory. Trajectory categorization employs agglomerative clustering based on trajectory distance. This agglomerative clustering process involves first treating each anonymous identifier of the inspected object as an initial trajectory cluster, then repeatedly merging the two trajectory clusters with the smallest trajectory distance until a preset number of trajectory types is reached or the minimum inter-cluster distance is greater than a preset merging distance threshold. A trajectory cluster refers to a group of anonymous identifiers of the inspected object that are close to each other in trajectory distance; the preset number of trajectory types is a type quantity condition set before data processing; the preset merging distance threshold is the maximum allowed merging distance set before data processing. The inter-cluster distance between two trajectory clusters is calculated using the average of the pairwise trajectory distances of the anonymous identifiers of the inspected objects within the cluster; if the distance between two anonymous identifiers of the inspected objects is not comparable, then that pair of anonymous identifiers is not included in the calculation of the average inter-cluster distance.

[0116] After the agglomerative clustering process stops, each trajectory cluster is identified as a physical examination indicator trajectory type, and a trajectory type identifier is generated for each type. These identifiers are generated sequentially according to the trajectory clusters after clustering and are used to distinguish different physical examination indicator trajectory types. For each physical examination indicator trajectory type, the type center trajectory is calculated. The type center trajectory refers to the dynamic trajectory sample of the physical examination indicator corresponding to the anonymous identifier of the examined object with the smallest average trajectory distance to other anonymous identifiers within that trajectory type. If multiple anonymous identifiers have the same average trajectory distance, the anonymous identifier with the largest amount of available trajectory representation data is selected as the type center trajectory.

[0117] Finally, the anonymous identifier of the tested subject, the trajectory type identifier, the hidden state sequence corresponding to the anonymous identifier, the trajectory representation data, the trajectory distance state, and the type center trajectory identifier are written into the trajectory typing results of the physical examination indicators. For the anonymous identifier of the tested subject that is marked as non-comparable and fails to be merged into any trajectory cluster, its trajectory type identifier is set to the pending typing state, and its dynamic trajectory samples of physical examination indicators, hidden state sequences, and reasons for non-comparable states are retained.

[0118] In step four, the physical examination indicator trajectory classification results obtained in step three are verified and organized to generate a physical examination indicator trajectory classification output record corresponding to each anonymous identifier of the examinee. The physical examination indicator trajectory classification output record refers to a data record composed of trajectory type identifier, stability status identifier, type characterization data, key standard physical examination indicator name, and trajectory deviation degree. The stability status identifier indicates whether the trajectory type identifier corresponding to the anonymous identifier of the examinee remains consistent under repeated calculation conditions; the type characterization data indicates the common dynamic change characteristics within the same physical examination indicator trajectory type; the key standard physical examination indicator name indicates the name of the standard physical examination indicator that contributes significantly to the formation of the physical examination indicator trajectory type; and the trajectory deviation degree indicates the magnitude of the difference between the dynamic trajectory sample of a certain anonymous identifier of the physical examination indicator and the type center trajectory of its corresponding physical examination indicator trajectory type.

[0119] Specifically, the physical examination indicator trajectory classification results obtained in step three are first read. These results include the anonymous identifier of the examinee, trajectory type identifier, hidden state sequence, trajectory representation data, trajectory distance status, and type center trajectory identifier. The trajectory type identifier is used to distinguish different physical examination indicator trajectory types; the hidden state sequence represents the arrangement of hidden states of the indicator modeling sequence under the same anonymous identifier of the examinee at each aligned time point; the trajectory representation data represents the dynamic change pattern of the physical examination indicators of the anonymous identifier of the examinee; the trajectory distance status represents whether the trajectory distances between different anonymous identifiers of the examinee are comparable; and the type center trajectory identifier represents the anonymous identifier of the examinee that serves as the type center trajectory within the corresponding physical examination indicator trajectory type.

[0120] Simultaneously, the dynamic trajectory samples of the physical examination indicators obtained in step two are read. These dynamic trajectory samples include the anonymous identifier of the examinee, alignment time point, standard physical examination indicator name, current physical examination indicator value, data source identifier, data status identifier, normalized physical examination indicator value, local change amount, local change rate, local trend slope, reference deviation degree, and auxiliary variable code. When reading the dynamic trajectory samples, the anonymous identifier of the examinee and the standard physical examination indicator name are used as indexes to map the dynamic trajectory samples to the trajectory classification results, ensuring that each trajectory type identifier corresponds to the original dynamic observation record, hidden state sequence, and trajectory representation data that formed that trajectory type identifier.

[0121] Then, the stability of the trajectory classification results of the physical examination indicators is verified by repeated sampling. This repeated sampling stability verification involves, while retaining the original set of anonymous identifiers for the examined subjects, repeatedly extracting a portion of the standard physical examination indicator names and a portion of the aligned time points to re-execute the trajectory distance calculation and trajectory classification processing, and comparing whether the re-obtained trajectory type identifiers are consistent with those obtained in step three. The partial standard physical examination indicator names are extracted from the standard physical examination indicator names with indicator weights in step three; the partial aligned time points are extracted from the unified time grid formed in step two; the extraction ratio, the number of extractions, and the random seed are all repeated sampling parameters set before data processing, and the random seed is used to ensure that the same input data yields the same extraction results under the same repeated sampling parameters.

[0122] In each repeated sampling, firstly, according to the extracted standard physical examination indicator names and aligned time points, corresponding dynamic observation records are selected from the dynamic trajectory samples of the physical examination indicators; then, following the trajectory representation data generation method in step three, the hidden state dwell ratio, hidden state transition times, average local change rate, average local trend slope, average reference deviation, mean predicted residual, and predicted residual fluctuation value corresponding to the selected dynamic observation records are recalculated; subsequently, the trajectory distance between any two anonymous identifiers of the examined objects is recalculated according to the trajectory distance calculation method in step three; finally, the agglomerative clustering processing in step three is used to regenerate the repeated sampling trajectory type identifier. The repeated sampling trajectory type identifier refers to the trajectory type identifier re-obtained under a certain repeated sampling condition.

[0123] After obtaining multiple repeated sampling trajectory type identifiers, the repeated sampling consistency score of the anonymous identifier of the examinee is calculated. The repeated sampling consistency score refers to the proportion of the number of times the same anonymous identifier of the examinee is still classified into the same physical examination indicator trajectory type as in step three during multiple repeated sampling trajectory classification processes, out of the total number of repeated samplings. If, in a certain repeated sampling, the anonymous identifier of the examinee does not generate a repeated sampling trajectory type identifier because the trajectory distance status is incomparable, then this repeated sampling is not included in the denominator, and the result is recorded as a sampling incomparable state. The higher the repeated sampling consistency score, the less sensitive the trajectory type identifier corresponding to the anonymous identifier of the examinee is to changes in the standard physical examination indicator name and alignment time point.

[0124] Next, the time perturbation stability of the physical examination index trajectory classification results is verified. This time perturbation stability verification involves applying a preset time perturbation to the alignment time points without changing the source of the standard physical examination index values, and recalculating the local change rate, local trend slope, prediction residual, and trajectory representation data in the dynamic trajectory samples of the physical examination indicators. The trajectory type identifier obtained after the perturbation is then compared with the trajectory type identifier obtained in step three to see if they are consistent. The preset time perturbation is a time offset range set before data processing, and this time offset range is less than the preset alignment interval. The perturbation value for each alignment time point is generated by the preset time perturbation range and a random seed.

[0125] When performing time-perturbation stability verification, firstly, each aligned time point under the anonymous identifier of the same tested object is added with the corresponding perturbation value to obtain the perturbation aligned time point; then, the time interval between adjacent time positions is recalculated according to the perturbation aligned time point; subsequently, the local change rate and local trend slope are recalculated based on the original normalized physical examination index values; for the index modeling sequence that has already completed gray prediction processing in step three, the predicted physical examination index values ​​and prediction residuals are recalculated using the perturbation aligned time point; then, according to the trajectory representation data generation method and agglomerative clustering processing method in step three, the time-perturbation trajectory type identifier is obtained. The time-perturbation trajectory type identifier refers to the trajectory type identifier obtained again after the aligned time point is perturbed.

[0126] After obtaining multiple time-perturbation trajectory type identifiers, the time-perturbation consistency score of the anonymous identifier of the examined object is calculated. The time-perturbation consistency score refers to the proportion of the number of times the same anonymous identifier of the examined object is still classified into the same physical examination indicator trajectory type as in step three across multiple time-perturbation trajectory classification processes, relative to the total number of time perturbations. If a time perturbation process causes the trajectory distance state of the anonymous identifier of the examined object to become incomparable, then that time perturbation process is not included in the denominator, and the result is recorded as a perturbation incomparable state.

[0127] Then, a stability status identifier is generated based on the repeat sampling consistency score and the time perturbation consistency score. The stability status identifier includes a stable classification status, an unstable classification status, and a classification status pending review. If both the repeat sampling consistency score and the time perturbation consistency score of the same anonymous identifier are not lower than a preset stability threshold, the stability status identifier corresponding to that anonymous identifier is set to a stable classification status. If the repeat sampling consistency score or the time perturbation consistency score is lower than the preset stability threshold, and the anonymous identifier still has a comparable trajectory distance, its stability status identifier is set to an unstable classification status. If the repeat sampling consistency score or the time perturbation consistency score cannot be calculated, or the trajectory distance is not comparable, its stability status identifier is set to a classification status pending review. The preset stability threshold is a proportional threshold set before data processing, used to classify the stability of trajectory type identifiers under repeated calculation conditions.

[0128] Subsequently, type characterization data is generated for each physical examination indicator trajectory type. When generating type characterization data, the anonymous identifiers of all subjects belonging to the same trajectory type and their trajectory characterization data are first read; then, according to the standard physical examination indicator name, the type average local change rate, type average local trend slope, type average reference deviation, type prediction residual mean, type prediction residual fluctuation value, type latent state retention ratio, and type latent state transition number are calculated for each physical examination indicator trajectory type. The type average local change rate is the weighted average of the average local change rates of the corresponding standard physical examination indicator names within the same physical examination indicator trajectory type; the type average local trend slope is the weighted average of the average local trend slopes of the corresponding standard physical examination indicator names within the same physical examination indicator trajectory type; the type average reference deviation is the weighted average of the average reference deviation of the corresponding standard physical examination indicator names within the same physical examination indicator trajectory type; the type prediction residual mean is the weighted average of the prediction residual mean values ​​of the corresponding standard physical examination indicator names within the same physical examination indicator trajectory type; the type prediction residual fluctuation is the weighted average of the prediction residual fluctuation values ​​of the corresponding standard physical examination indicator names within the same physical examination indicator trajectory type; the type hidden state retention ratio is the weighted average of the hidden state retention ratios of each corresponding standard physical examination indicator name within the same physical examination indicator trajectory type; the type hidden state transition number is the weighted average of the hidden state transition numbers of the corresponding standard physical examination indicator name within the same physical examination indicator trajectory type. The weights in the above weighted averages are jointly determined by the indicator weights obtained in step three and the preset reliable weights corresponding to the data state identifiers of the corresponding dynamic observation records.

[0129] Next, the contribution value of the standard physical examination indicator name to the physical examination indicator trajectory type is calculated. The contribution value refers to the degree of difference generated by a certain standard physical examination indicator name when distinguishing different physical examination indicator trajectory types. For any standard physical examination indicator name, the intra-type difference value within the same physical examination indicator trajectory type is first calculated, and then the inter-type difference value between different physical examination indicator trajectory types is calculated. The intra-type difference value is calculated from the differences in the average local change rate, average local trend slope, average reference deviation, mean predicted residual, predicted residual fluctuation value, and latent state retention ratio of each anonymous identifier of each examinee within the same physical examination indicator trajectory type relative to the corresponding type representation data; the inter-type difference value is calculated from the differences in the corresponding type representation data between different physical examination indicator trajectory types. The contribution value of the standard physical examination indicator name is obtained by dividing the inter-type difference value by the sum of the intra-type difference value and the stability constant. The stability constant is a preset value greater than zero, used to maintain the availability of calculation results when the intra-type difference value is zero. The larger the contribution value of the indicator, the stronger the distinguishing effect of the standard physical examination indicator name among different physical examination indicator trajectory types.

[0130] After obtaining the indicator contribution values, key standard physical examination indicator names are determined for each physical examination indicator trajectory type. The key standard physical examination indicator name refers to the name of a standard physical examination indicator whose indicator contribution value reaches a preset contribution threshold in that physical examination indicator trajectory type, or whose indicator contribution value, sorted from largest to smallest, falls within a preset selection range. The preset contribution threshold and preset selection range are selection conditions set before data processing. For each key standard physical examination indicator name, its type average local change rate, type average local trend slope, type average reference deviation, type prediction residual mean, and type latent state retention ratio are read, and the above data are written into the key indicator characterization record for that physical examination indicator trajectory type. The key indicator characterization record is used to describe the dynamic change characteristics of the key standard physical examination indicator name in the corresponding physical examination indicator trajectory type.

[0131] Then, the trajectory deviation degree of each anonymous subject identifier is calculated. The trajectory deviation degree is determined based on the trajectory representation data corresponding to the anonymous subject identifier and the trajectory representation data corresponding to the type center trajectory of the corresponding physical examination indicator trajectory type. If the anonymous subject identifier belongs to a stable or unstable stratification state, and there is a comparable trajectory distance between it and the anonymous subject identifier corresponding to the type center trajectory identifier, then this trajectory distance is used as the basic deviation value; then, the repeated sampling consistency score and time perturbation consistency score of the anonymous subject identifier within the corresponding physical examination indicator trajectory type are read, and the average of the two is used as the stability correction value; finally, the basic deviation value is divided by the sum of the stability correction value and the stability constant to obtain the trajectory deviation degree. If the trajectory distance between the anonymous subject identifier and the anonymous subject identifier corresponding to the type center trajectory identifier is not comparable, then the trajectory deviation degree is not calculated, and the trajectory deviation degree state is set to an incalculable state.

[0132] Subsequently, the anonymous identifiers of the subjects in the genotyping state undergo neighbor trajectory assignment processing. This neighbor trajectory assignment processing refers to determining whether candidate trajectory type identifiers can be generated based on the trajectory distance between the anonymous identifier of the subject in the genotyping state and the type center trajectory of each physical examination indicator trajectory type when the anonymous identifier has some usable trajectory representation data. During neighbor trajectory assignment processing, the trajectory distance between the anonymous identifier of the subject in the genotyping state and the anonymous identifier of the subject corresponding to each type center trajectory identifier is first calculated. If there are physical examination indicator trajectory types with comparable trajectory distances, the physical examination indicator trajectory type with the smallest trajectory distance is selected as the candidate physical examination indicator trajectory type. If the minimum trajectory distance is not greater than a preset neighbor assignment threshold, the trajectory type identifier corresponding to the candidate physical examination indicator trajectory type is written into the candidate trajectory type identifier, and the stability state identifier is set to the genotyping state pending review. If the minimum trajectory distance is greater than the preset neighbor assignment threshold, no candidate trajectory type identifier is generated. The preset neighbor assignment threshold is the maximum allowable neighbor distance set before data processing.

[0133] Finally, the physical examination indicator trajectory typing output record is generated. For each subject's anonymous identifier, the following information is written: subject anonymity identifier, trajectory type identifier, candidate trajectory type identifier, stability status identifier, repeated sampling consistency score, time perturbation consistency score, trajectory deviation degree, trajectory deviation degree status, type center trajectory identifier, key standard physical examination indicator name, key indicator representation record, hidden state sequence, trajectory representation data, data source identifier, and data status identifier. As an executable output example, the physical examination indicator trajectory typing output record can be saved in the form of fields. For a specific subject with an anonymous identifier A001, its trajectory type is identified as T1, the candidate trajectory type is identified as null, the stability status is identified as stable subtype status, the repeatability score is 0.92, the temporal perturbation consistency score is 0.88, the type center trajectory is identified as A015, the trajectory deviation is 0.18, and the trajectory deviation status is calculable. The key standard physical examination indicators include fasting blood glucose, systolic blood pressure, and body mass index. The key indicator characterization records include the continuous high participation interval corresponding to fasting blood glucose, the participation amount of the indicator trajectory, the main driving cause, and the driving direction. The hidden state sequence includes the hidden state corresponding to each aligned time point. The trajectory characterization data includes the hidden state dwell ratio, the number of hidden state transitions, the average local change rate, the average local trend slope, the average reference deviation, the mean of the prediction residual, and the prediction residual fluctuation value. The data source identifier and data status identifier record the data source and processing status corresponding to each dynamic observation record, respectively.

[0134] For an anonymous identifier A002 of a certain subject, if the trajectory distance between it and a certain type of center trajectory is not greater than the preset neighboring attribution threshold, but the consistency score of repeated sampling or the consistency score of time disturbance is lower than the preset stability threshold, then the trajectory type identifier is written to the pending confirmation status, the candidate trajectory type identifier is written to the corresponding physical examination indicator trajectory type, the stability status identifier is written to the pending review classification status, and the candidate attribution reason is written as the trajectory distance is close, the trajectory-driven profile is partially consistent, or the continuous high participation interval overlaps.

[0135] For a certain anonymous identifier A003 of the tested object, if its common participating indicator set is empty, or the trajectory distance status with each type of center trajectory is incomparable, then the trajectory type identifier is written to the pending classification status, the candidate trajectory type identifier is empty, the trajectory deviation status is written to the incalculable status, and the reason for the incomparable status is written as insufficient number of available standard physical examination indicator names, insufficient number of available alignment time points, or inconsistent driving profile.

[0136] For anonymous subjects whose trajectory type is marked as "pending classification" and for which no candidate trajectory type identifiers have been generated, the following information is recorded: pending classification status, reason for incomparability, number of available standard medical examination indicator names, and number of available alignment time points. Output records are generated one by one according to the anonymous identifiers of the subjects, and type representation data under the same medical examination indicator trajectory type are saved separately according to the trajectory type identifier.

[0137] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0138] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.

[0139] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and inventive constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0140] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0141] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0142] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A dynamic data modeling processing method for physical examination index trajectory classification, characterized in that, include: Data from multiple sources of physical examinations are collected according to the anonymity identifier of the examinee. The original item names are matched with standard names according to the source, testing method and sample type. After converting the original index values ​​into the corresponding standard physical examination index values, physical examination index trajectory samples are constructed according to the first observation time. The physical examination indicator trajectory samples are placed into a unified time grid according to the preset alignment interval, and conflicting and missing values ​​are distinguished and processed according to the data status identifier. Then, based on the value changes of the same standard physical examination indicator name at adjacent or nearby alignment time points, dynamic trajectory samples of physical examination indicators are formed. Based on the dynamic trajectory samples of physical examination indicators, an indicator modeling sequence is established. The prediction residual is obtained through gray prediction processing, and the hidden state sequence is obtained through the hidden state transition model. Then, the observation confidence quantity is determined by the data state identifier. The dynamic participation of indicators is adjusted by the consistency of changes within consecutive aligned time points, the continuation deviation of prediction residuals, the deviation outside the reference range buffer boundary, and the correspondence between hidden state changes and indicator changes. In this way, a trajectory-driven profile is formed, and the trajectory distance is calculated based on the common standard physical examination indicator names to form the physical examination indicator trajectory classification results. The stability of the physical examination indicator trajectory classification results is verified by repeated sampling and time disturbance. The names of key standard physical examination indicators and the degree of trajectory deviation are sorted out in combination with the type center trajectory, and the physical examination indicator trajectory classification output record is generated.

2. The dynamic data modeling and processing method for physical examination indicator trajectory classification according to claim 1, characterized in that: The process of aggregating multi-source physical examination data according to the anonymity identifier of the examinee includes: using the anonymity identifier of the examinee as the primary key for trajectory aggregation, incorporating the observation time, original item name, original indicator value, original unit, reference range, and source fields from electronic physical examination records, test records, physical measurement records, and health questionnaire records into the same examinee's data set; using a preset indicator name mapping table, converting the original item names under different data sources, detection method identifiers, and sample type identifiers into standard physical examination indicator names, and using a preset unit conversion table, converting the corresponding original indicator values ​​into standard physical examination indicator values ​​under the target unit; and using the first observation time under the same examinee's anonymity identifier as a benchmark, scaling up various observation times to construct a standardized physical examination time.

3. The dynamic data modeling and processing method for physical examination indicator trajectory classification according to claim 2, characterized in that: After constructing the standardized physical examination time, the process further includes: for multiple candidate standard physical examination indicator values ​​under the same anonymous identifier of the examinee, the same standardized physical examination time, and the same standard physical examination indicator name, conflict screening is performed according to the detection method identifier, sample type identifier, source priority, and time proximity. Measured values, missing values, unmapped names, unconverted units, and source conflicts are written into the data status identifier. Simultaneously, the reference range, data source identifier, and auxiliary variable code corresponding to the standard physical examination indicator name are retained, establishing a physical examination indicator trajectory sample with the examinee's anonymous identifier, standardized physical examination time, standard physical examination indicator name, standard physical examination indicator value, reference range, data source identifier, and data status identifier as core fields.

4. The dynamic data modeling and processing method for physical examination indicator trajectory classification according to claim 1, characterized in that: The step of placing the physical examination indicator trajectory samples into a unified time grid according to a preset alignment interval includes: based on the anonymous identifier of the examinee, the standardized physical examination time, the name of the standard physical examination indicator, the value of the standard physical examination indicator, the reference range, the data source identifier, and the data status identifier, firstly, sorting each observation record by time according to the anonymous identifier of the examinee, and merging records and resolving source conflicts for the same standard physical examination indicator name under the same standardized physical examination time; then, constructing a unified time grid with the trajectory start time, trajectory end time, and preset alignment interval, mapping each standard physical examination indicator value to the corresponding aligned time point to form aligned physical examination indicator values, and writing a missing status for aligned time points with no measured values ​​or unresolved conflicts.

5. The dynamic data modeling and processing method for physical examination indicator trajectory classification according to claim 4, characterized in that: The process of processing missing values ​​according to data status identifiers and forming dynamic trajectory samples of physical examination indicators includes: around the same anonymous identifier of the examinee and the same standard physical examination indicator name, at the missing alignment time point, sequentially using linear interpolation of measured values ​​before and after, one-sided value continuation, and statistical completion of the population median to generate completed physical examination indicator values, and writing them into the interpolation completion state, continuation completion state, or statistical completion state respectively; then, according to the median, first quartile, and third quartile under the same standard physical examination indicator name, the current physical examination indicator value is scaled and normalized to generate normalized physical examination indicator values; Further, the local changes and local rates of change between adjacent aligned time points are calculated within a unified time grid. Then, the normalized physical examination index values ​​are linearly fitted within a preset trend window to generate a local trend slope. Simultaneously, the reference deviation is calculated based on the lower limit, upper limit, and width of the reference range after unit conversion. Finally, the anonymous identifier of the examined subject, aligned time points, standard physical examination index names, current physical examination index values, data source identifiers, data status identifiers, normalized physical examination index values, local changes, local rates of change, local trend slopes, reference deviation, and auxiliary variable codes are written into the dynamic observation record to form a dynamic trajectory sample of the physical examination index.

6. The dynamic data modeling and processing method for physical examination indicator trajectory classification according to claim 1, characterized in that: Using the anonymous identifier of the examinee, the alignment time point, the name of the standard physical examination indicator, the data status identifier, the normalized physical examination indicator value, the local rate of change, the local trend slope, and the reference deviation as inputs, an indicator modeling sequence arranged according to the alignment time point is constructed according to the same anonymous identifier of the examinee and the same standard physical examination indicator name. Different preset confidence weights are assigned to each sequence position based on the measured state, conflict resolution state, interpolation completion state, continuation completion state, statistical completion state, and missing state. For the indicator modeling sequence whose number of effective sequence positions meets the preset minimum modeling number, the normalized physical examination indicator value is extracted to form the original indicator sequence. The original indicator sequence is accumulated and generated once. Combined with the background value and the preset confidence weight, the gray development coefficient and gray action are calculated by weighted least squares, and then the predicted physical examination indicator value and prediction residual at each alignment time point are calculated. Normalized physical examination index values, local rate of change, local trend slope, reference deviation degree, and prediction residuals are used to construct observation feature vectors in a fixed order. Clustering initialization is performed based on all available observation feature vectors to establish a hidden state transition model that includes initial state probability, state transition probability, and observation generation parameters. Then, through dynamic programming, the hidden state sequence corresponding to the available alignment time point is solved for each index modeling sequence. Based on the hidden state sequence, local rate of change, local trend slope, reference deviation degree, and prediction residual statistics, trajectory representation data is formed.

7. The dynamic data modeling and processing method for physical examination indicator trajectory classification according to claim 6, characterized in that: After obtaining the trajectory representation data, for any standard physical examination indicator name and any alignment time point, an observation confidence quantity is constructed based on the data status identifier. A change duration quantity is constructed based on the same-direction persistence of the local change rate and local trend slope within the preset duration window. A predicted deviation accumulation quantity is constructed based on the same-sign persistence and absolute value increase of the predicted residuals within the preset residual window. A reference buffer deviation quantity is constructed based on the upper and lower buffer boundaries formed by the expansion of the reference range. A hidden state transition confirmation quantity is constructed based on the correspondence between the hidden state changes at adjacent alignment time points and the direction of the local change rate, the amplitude of the local trend slope, and the change in the absolute value of the predicted residuals. Then, using the observation confidence quantity as the basic participation condition, a participation adjustment value is generated by combining the change duration quantity, the predicted deviation accumulation quantity, the reference buffer deviation quantity, and the hidden state transition confirmation quantity. Finally, the observation confidence quantity and the participation adjustment value are used to generate the indicator dynamic participation quantity, so that different standard physical examination indicator names form a non-fixed participation degree that is adjusted according to the dynamic change cause at different alignment time points.

8. The dynamic data modeling and processing method for physical examination indicator trajectory classification according to claim 1, characterized in that: The dynamic participation of indicators is aggregated over time according to the aligned time point to identify continuous high participation intervals. The indicator trajectory participation is generated by combining the length of the continuous high participation interval, the average value of the indicator dynamic participation within the interval, and the number of hidden state transitions. The main driving cause is determined based on the duration of high value changes, the cumulative amount of predicted deviation, the amount of reference buffer deviation, or the amount of hidden state transition confirmation. The driving direction is determined based on the sign of the local trend slope within the continuous high participation interval. A trajectory-driven profile is constructed, which includes the standard physical examination indicator name, the continuous high participation interval, the indicator trajectory participation, the main driving cause, and the driving direction. Furthermore, between any two anonymous identifiers of the examinees, a set of common participating indicators is first constructed based on the standard physical examination indicator names that share the indicator trajectory participation. Then, level differences, duration differences, residual differences, reference differences, and state differences are calculated from the normalized numerical level, continuous change process, predicted residual deviation, reference buffer deviation, and hidden state transition performance. Based on this, indicator difference structures such as near-level heterogeneous process structure, heterogeneous near-level structure, heterogeneous heterogeneous process structure, or near-level near-process structure are generated. Based on the index difference structure, a process priority adjustment quantity is constructed, and the participation of continuous difference, residual difference and state difference in the index trajectory difference quantity is improved for near-level heterogeneous process structure. Then, the index trajectory difference quantity is generated by combining the observation confidence quantity, index trajectory participation quantity, time overlap degree of continuous high participation interval and driving direction relationship.

9. The dynamic data modeling and processing method for physical examination indicator trajectory classification according to claim 8, characterized in that: The dominant participating indicator set is selected according to the participation of the indicator trajectory, and the trajectory distance between the anonymous identifiers of two subjects is synthesized based on the difference participation ratio. At the same time, when there is a near-horizontal heterogeneous process structure, the process difference identifier is written. The trajectory distance and trajectory-driven profile are used as the basis for merging trajectory clusters. The merging of trajectory clusters is controlled by the trajectory distance between clusters and the driving consistency between clusters. The near-distance heterogeneous driving relationship between trajectory clusters, the driving dispersion state of the merged trajectory clusters, the type center trajectory of the physical examination indicator trajectory type, and the pending confirmation state and pending classification state of the anonymous identifiers of the subjects are marked to form the physical examination indicator trajectory classification result.

10. The dynamic data modeling and processing method for physical examination indicator trajectory classification according to claim 1, characterized in that: After obtaining the trajectory distance, the anonymous identifiers of the tested subjects are classified into physical examination indicator trajectory types through agglomerative clustering. Within each trajectory type, the type center trajectory is selected based on the average trajectory distance. Subsequently, the dynamic trajectory samples, latent state sequences, and trajectory representation data of the physical examination indicators are associated. The trajectory type identifier is recalculated by repeating the sampling of standard physical examination indicator names and the alignment time points and perturbation alignment time points, generating the repeating consistency score and the time perturbation consistency score, which are then written into the stability state identifier. Next, the trajectory representation data is summarized according to the trajectory type, and the indicator contribution value of each standard physical examination indicator name is calculated. The key standard physical examination indicator names and their key indicator representation records are selected. Finally, the trajectory deviation is calculated based on the trajectory distance between the anonymous identifiers of the tested subjects and the type center trajectory, as well as the consistency score, and the physical examination indicator trajectory classification output record is generated.