Data tag system automatic construction system
Patent Information
- Application Number
- CN202611073163.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-08-18
AI Technical Summary
[0003]在提取连续分布的时序血压特征时,若采用固定时间窗口内的均值或极值进行统计,虽然能够反映该时段内的整体血压水平,但在面对相邻采样节点间因时间间隔差异而产生的瞬时变化时,这种统计方式难以兼顾血压数值变化量随时间间隔的动态响应,随着对血管压力突变趋势分析要求的提高,这种处理方式在定位血压波动趋势中的方向反转节点及状态突变时序位置时,往往存在颗粒度较粗的情况;在处理跨段落分散分布的生化指标数值与影像描述病灶体积时,通常会对各属性维度分别进行阈值判定并进行叠加汇总,以评估器官的异常状态,然而,人体生理状态是一个复杂的协同网络,这种独立判定后叠加的方式,在融合血压波动的时序累积效应与各器官部位的生化异常频次、病灶体积时,难以充分体现多源数据间的跨维度联合核算关系
[0015] By locating segments of physical examination reports to extract time-series indicators and analyze semantic relationships, the problem of scattered information in multi-source heterogeneous texts is overcome, and standardized conversion of structured indicator sequences and entity relationship graphs is achieved. By calculating the rate of change of blood pressure intervals point by point and screening reversal nodes, the problem of fixed windows being unable to capture time interval differences is overcome, and accurate identification of turning points in blood pressure fluctuations is achieved. By calculating the frequency of biochemical abnormalities and combining lesion volume with the effect of blood pressure fluctuations for cross-dimensional fusion, the problem of simple superposition after independent determination of each attribute is overcome, and unified quantification of multi-source risk information is achieved. By decoupling attributes and removing redundant dimensions and combining age and tumor markers into a tree-like risk grading network, the problem of high-dimensional features interfering with grading accuracy is overcome, and automated generation of health rating labels is achieved. By comparing underwriting stratification thresholds and associating with insurance liability configuration, the problem of separation between health rating and insurance configuration is overcome, and end-to-end automated construction from physical examination text to underwriting labels and insurance plans is achieved.
Smart Images

Figure CN122594472A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to an automated data tagging system construction system. Background Technology
[0002] Physical examination reports typically contain heterogeneous text from multiple sources, including biochemistry sections from the laboratory and descriptive sections from the radiology department. Each section contains structured elements such as time-series blood pressure indicators, biochemical values, lesion volumes described in imaging, and tumor marker values. These elements form a semantic network through organ location attribution, modification constraints, and hierarchical relationships. Extracting these structured elements from the physical examination report and organizing them according to sampling timestamps to form an initial sequence of medical indicators serves as the input basis for subsequent health status analysis and label generation.
[0003] When extracting time-series blood pressure features from continuous distributions, using the mean or extreme values within a fixed time window for statistical analysis can reflect the overall blood pressure level within that period. However, when faced with instantaneous changes caused by time interval differences between adjacent sampling nodes, this statistical method struggles to account for the dynamic response of blood pressure value changes over time. As the requirements for analyzing vascular pressure abrupt change trends increase, this approach often exhibits coarse granularity when locating reversal points and temporal locations of state changes in blood pressure fluctuation trends. When processing biochemical index values and lesion volumes described by images that are distributed across segments, threshold judgments are usually applied to each attribute dimension and then superimposed to assess the abnormal state of organs. However, the human physiological state is a complex collaborative network. This method of independent judgment followed by superposition fails to fully reflect the cross-dimensional joint accounting relationship between multi-source data when integrating the temporal cumulative effect of blood pressure fluctuations with the frequency of biochemical abnormalities and lesion volumes of various organ sites. Summary of the Invention
[0004] This invention provides an automated data tagging system construction system. By combining time-series blood pressure interval change tracking with cross-dimensional organ load joint calculation, it realizes end-to-end automated construction from the original text of the physical examination report to a hierarchical data tagging system.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0006] The first aspect is the automated construction system for the data tagging system, which includes:
[0007] The acquisition module is used to acquire physical examination reports and locate biochemical and imaging segments, extract time-series blood pressure, biochemical values, lesion volume, age and tumor markers, and generate an initial medical indicator sequence containing sampling nodes by sorting by timestamp;
[0008] The parsing module is used to parse semantic associations and identify organ attributes and suspicious occupants based on the initial medical indicator sequence, and generate a structured entity relationship graph.
[0009] The calculation module is used to extract time-series blood pressure based on the structured entity relationship graph, measure the rate of change of blood pressure between adjacent sampling nodes over time intervals, and generate a feature vector of index fluctuation trend.
[0010] The extraction module is used to extract rate differences and direction identifiers based on the indicator fluctuation trend feature vector, filter sampling nodes with rising, falling and reversing characteristics and extract time sequence position identifiers to generate local fluctuation difference sequences.
[0011] The comparison module is used to compare the frequency of biochemical abnormalities with the medical reference range based on the local fluctuation difference sequence and the structured entity relationship map, summarize and calculate the total abnormal load and lesion volume of each organ across the entire domain, perform feature splicing, and generate a multidimensional risk accumulation matrix.
[0012] The transmission module is used to perform attribute decoupling operations to remove redundant dimensions based on the multidimensional risk accumulation matrix and generate a core health status feature set. Based on the core health status feature set, combined with the continuously transmitted age and tumor markers, the module is input into a tree-like risk grading network to extract the risk level corresponding to the maximum probability value and generate an initial health rating label.
[0013] The segmentation module is used to divide the first and second underwriting tiers based on the initial health rating labels and the underwriting tier thresholds, associate the corresponding set of coverage liability configuration parameters, and output the data label system.
[0014] The above-described solution of the present invention has at least the following beneficial effects:
[0015] By locating segments of physical examination reports to extract time-series indicators and analyze semantic relationships, the problem of scattered information in multi-source heterogeneous texts is overcome, and standardized conversion of structured indicator sequences and entity relationship graphs is achieved. By calculating the rate of change of blood pressure intervals point by point and screening reversal nodes, the problem of fixed windows being unable to capture time interval differences is overcome, and accurate identification of turning points in blood pressure fluctuations is achieved. By calculating the frequency of biochemical abnormalities and combining lesion volume with the effect of blood pressure fluctuations for cross-dimensional fusion, the problem of simple superposition after independent determination of each attribute is overcome, and unified quantification of multi-source risk information is achieved. By decoupling attributes and removing redundant dimensions and combining age and tumor markers into a tree-like risk grading network, the problem of high-dimensional features interfering with grading accuracy is overcome, and automated generation of health rating labels is achieved. By comparing underwriting stratification thresholds and associating with insurance liability configuration, the problem of separation between health rating and insurance configuration is overcome, and end-to-end automated construction from physical examination text to underwriting labels and insurance plans is achieved. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of an automated data tagging system construction system provided by an embodiment of the present invention.
[0017] Figure 2 This is a flowchart illustrating the process of dividing the first and second underwriting tiers based on the initial health rating label, comparing the underwriting tiering thresholds, associating the corresponding set of coverage responsibility configuration parameters, and outputting a data label system, as provided in an embodiment of the present invention.
[0018] Figure 3 This is a schematic diagram of the process provided by an embodiment of the present invention, which performs attribute decoupling operation to remove redundant dimensions based on a multidimensional risk accumulation matrix and generates a core health status feature set. Detailed Implementation
[0019] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0020] like Figure 1 As shown, an embodiment of the present invention proposes an automated data tagging system construction system, comprising:
[0021] The acquisition module is used to acquire physical examination reports and locate biochemical and imaging segments, extract time-series blood pressure, biochemical values, lesion volume, age and tumor markers, and generate an initial medical indicator sequence containing sampling nodes by sorting by timestamp;
[0022] The parsing module is used to parse semantic associations and identify organ attributes and suspicious occupants based on the initial medical indicator sequence, and generate a structured entity relationship graph.
[0023] The calculation module is used to extract time-series blood pressure based on the structured entity relationship graph, measure the rate of change of blood pressure between adjacent sampling nodes over time intervals, and generate a feature vector of index fluctuation trend.
[0024] The extraction module is used to extract rate differences and direction identifiers based on the indicator fluctuation trend feature vector, filter sampling nodes with rising, falling and reversing characteristics and extract time sequence position identifiers to generate local fluctuation difference sequences.
[0025] The comparison module is used to compare the frequency of biochemical abnormalities with the medical reference range based on the local fluctuation difference sequence and the structured entity relationship map, summarize and calculate the total abnormal load and lesion volume of each organ across the entire domain, perform feature splicing, and generate a multidimensional risk accumulation matrix.
[0026] The transmission module is used to perform attribute decoupling operations to remove redundant dimensions based on the multidimensional risk accumulation matrix and generate a core health status feature set. Based on the core health status feature set, combined with the continuously transmitted age and tumor markers, the module is input into a tree-like risk grading network to extract the risk level corresponding to the maximum probability value and generate an initial health rating label.
[0027] The segmentation module is used to divide the first and second underwriting tiers based on the initial health rating labels and the underwriting tier thresholds, associate the corresponding set of coverage liability configuration parameters, and output the data label system.
[0028] In this embodiment of the invention, by locating paragraphs in the physical examination report to extract time-series indicators and analyze semantic relationships, the problem of scattered multi-source heterogeneous text information is overcome, and the standardized conversion of structured indicator sequences and entity relationship graphs is achieved. By calculating the rate of change of blood pressure intervals point by point and screening reversal nodes, the problem of fixed windows being unable to capture time interval differences is overcome, and the accurate identification of blood pressure fluctuation turning points is achieved. By calculating the frequency of biochemical abnormalities and combining lesion volume with the effect of blood pressure fluctuations for cross-dimensional fusion, the problem of simple superposition after independent determination of each attribute is overcome, and the unified quantification of multi-source risk information is achieved. By decoupling attributes and stripping redundant dimensions and combining age and tumor markers into a tree-like risk grading network, the problem of high-dimensional features interfering with grading accuracy is overcome, and the automatic generation of health rating labels is achieved. By comparing underwriting stratification thresholds and associating with insurance liability configuration, the problem of the separation between health rating and insurance configuration is overcome, and the end-to-end automated construction from physical examination text to underwriting labels and insurance plans is achieved.
[0029] In a preferred embodiment of the present invention, a physical examination report is obtained and biochemical and imaging segments are located. Time-series blood pressure, biochemical values, lesion volume, age, and tumor markers are extracted. An initial medical indicator sequence containing sampling nodes is generated by sorting the data by timestamps, including:
[0030] Obtain the original text of the physical examination report and locate the biochemistry section from the laboratory and the description section from the radiology department, generating a multi-source text paragraph set, specifically including:
[0031] Scan the full text of the original medical examination report, using blank lines, line indentation, and chapter title lines as paragraph separators to divide the entire text into several independent paragraph units, and arrange them in the original order to obtain the original paragraph sequence.
[0032] Extract the first-line text content of each paragraph unit in the original paragraph sequence, perform keyword matching judgment. If the first-line text contains both inspection attribution keywords and biochemical item keywords, mark this paragraph unit as a candidate paragraph for biochemical inspection in the laboratory. Among them, inspection attribution keywords include laboratory, laboratory testing, and laboratory. Biochemical item keywords include biochemistry, blood routine, liver function, kidney function, and blood lipid. If the first-line text contains both imaging attribution keywords and description keywords, mark this paragraph unit as a candidate paragraph for imaging description. Among them, imaging attribution keywords include radiology department, ultrasound, radiology, CT, and magnetic resonance imaging; description keywords include findings, description, diagnosis opinion, and imaging findings.
[0033] Extract the main text content of all candidate paragraphs for biochemical inspection in the laboratory, and splice them line by line into a continuous biochemical text block, which is classified into the biochemical text subset. Extract the main text content of all candidate paragraphs for imaging description, and splice them line by line into a continuous imaging text block, which is classified into the imaging text subset. Combine the biochemical text subset and the imaging text subset into a multi-source text paragraph set.
[0034] According to the multi-source text paragraph set, extract the blood pressure measurement values and the bound date and time strings, and generate a time-series blood pressure index group in chronological order, specifically including:
[0035] In the biochemical text subset of the multi-source text paragraph set, scan each line to find the numerical field lines prefixed with systolic blood pressure, diastolic blood pressure, or blood pressure. Extract the numerical string after the prefix as the blood pressure measurement value; at the same time, extract the date and time string bound within or adjacent to this numerical field line, and parse it into a sampling timestamp. For the case where only systolic blood pressure or only diastolic blood pressure appears in the same numerical field line, set the missing dimension value as an empty mark, and still save the extracted dimension values with this sampling timestamp as the key. In subsequent steps, the missing dimension at this sampling node will not be included in the calculation. For the case where there are multiple blood pressure records corresponding to the same sampling timestamp, preferentially select the record marked as the right upper arm or without a marked measurement site. If multiple records are not marked with a measurement site, select the one with the largest value. The common formats of the date and time string include YYYY-MM-DD HH:MM, YYYY year MM month DD day, and DD / MM / YYYY. When parsing, try to match in this order in turn. After successful matching, uniformly convert it into the cumulative number of hours since 0:00 on January 1, 1 AD as the sampling timestamp. Bind the systolic blood pressure value and the diastolic blood pressure value corresponding to the same sampling timestamp into a blood pressure record, and obtain the time-series blood pressure index group in chronological order.
[0036] According to the multi-source text paragraph set, identify the names of biochemical items and the corresponding detection values, and generate a biochemical index value group, specifically including:
[0037] Based on the multi-source text segment set, the biochemical text subset is further scanned, and line-by-line identification of numerical fields prefixed with preset biochemical item names is performed. These preset biochemical item names include serum total protein, albumin, alanine aminotransferase, serum creatinine, fasting blood glucose, and total cholesterol. Each item name, along with the following numerical value string and unit string, is extracted and categorized by item name. For numerical values containing the prefix "<", it indicates a value below the detection limit, and the value is recorded as half of the detection limit. Values containing the prefix ">" indicate values above the detection limit. The upper limit is recorded as the upper limit value of the test. In cases where the unit string is inconsistent with the preset standard unit, it is uniformly converted according to the following conversion relationship: serum creatinine 1 mg / dL = 88.4 μmol / L, fasting blood glucose 1 mg / dL = 0.0555 mmol / L. For other items, if the unit is inconsistent with the preset standard unit, the original value is recorded and a unit mark is added. In cases where there are multiple test records for the same biochemical item at the same sampling time stamp, the arithmetic mean of the test values is taken as the representative value of the item at that time stamp to obtain the biochemical indicator value group.
[0038] Based on a multi-source text segment set, age and tumor marker values are extracted to generate user age and tumor marker value sets, specifically including:
[0039] Based on a multi-source text segment set, scan the numerical fields in the biochemical text subset that are prefixed with age or years, and extract the integer value immediately following them as the user's age value. For cases where age information is given in the format YYYY year MM month DD day, extract the birth date string and subtract the birth year from the sampling timestamp year of the medical examination report to obtain the user's age value. If the medical examination report contains both a direct age value and a birth date, the age calculated from the birth date is given priority. If the age values in multiple medical examination reports for the same user are inconsistent, the age is calculated based on the sampling timestamp and birth date of each report. The values are verified and corrected. The rows of numerical fields in the biochemical text subset are scanned and identified by prefixes such as alpha-fetoprotein, AFP, carcinoembryonic antigen, and CEA. The item name and the detection value are extracted. If the unit string is IU / mL, it is uniformly converted to ng / mL according to the conversion relationship of AFP: 1IU / mL≈1.21ng / mL and CEA: 1IU / mL≈1.0ng / mL before storage. For the case where there are multiple detection records for the same tumor marker item at the same sampling time stamp, the arithmetic mean of the detection values is taken as the representative value of the time stamp. The values are classified and stored according to the item to obtain the tumor marker value group.
[0040] Based on a multi-source text segment set, the diameter values and associated organ names are extracted from the lesion description segments. The lesion volume is calculated and stored in pairs to generate an image lesion volume group, specifically including:
[0041] In the image text subset of the multi-source text segment set, the lesion description segments are scanned sentence by sentence and the number of radial values contained therein is counted to generate radial dimension determination results. Based on the radial dimension determination results, when the number of radial values is at least two, the largest radial value is extracted as the major axis parameter. The minimum diameter value is extracted as the minor axis parameter. Calculate the major axis parameters With minor axis parameters The arithmetic mean is used as the parameter of the vertical axis. Based on major axis parameters minor axis parameters and vertical axis parameters Construct a basic three-dimensional geometric envelope, establish a three-dimensional rectangular coordinate system with the lesion center as the origin, and define the surface point set of the basic three-dimensional geometric envelope. Satisfies the fundamental ellipsoid equation: ;
[0042] in, Indicates the major axis parameter. Indicates the minor axis parameter. Indicates the vertical axis parameter. This represents the coordinates of the lesion's basic ellipsoid on the horizontal plane (horizontal axis). This represents the coordinates of the lesion's basic ellipsoid on the horizontal plane (vertical axis). This represents the coordinates of the lesion's basic ellipsoid on the vertical axis. Morphological modification phrases adjacent to the lesion description fragments are extracted from the image text subset. These morphological modification phrases are then mapped to surface morphology compensation factors. The specific mapping and calculation rules are as follows: If the morphological modification phrase contains lobes, the number of lobes is extracted. And lobule depth descriptors, mapping lobule depth descriptors to lobule depth coefficients. For example, a shallow foliation mapping is 0.1, and a deep foliation mapping is 0.3, generating a result containing... and The first morphological compensation factor; if the morphological modification phrase contains spiculation characteristics, then extract the spiculation density descriptor and map the spiculation density descriptor to the number of spiculations per unit area. and burr length coefficient A second-morphological compensation factor is generated. If the morphological modifier phrase contains unclear boundaries, descriptive terms for the degree of boundary ambiguity are extracted, mapping slightly ambiguous to 0.05, unclear boundaries to 0.10, and unclear boundaries to 0.15, thus generating a boundary ambiguity coefficient. The corresponding compensation method is as follows:
[0043] The surface determination condition of the basic three-dimensional geometric envelope is relaxed from an equality constraint to an inequality constraint. This creates a fuzzy transition layer radially extending outward from the outer boundary of the envelope, with voxels within this transition layer directly included in the effective voxel statistics. If the morphological modification phrase includes cystic changes, the descriptive term for the proportion of cystic changes is extracted, mapping small cystic changes to 0.05, partial cystic changes to 0.15, and widespread cystic changes to 0.30, generating a cystic change concavity volume coefficient. The corresponding compensation method is as follows:
[0044] Randomly generated within the envelope, accounting for a percentage of the total volume. For the spherical cavity region, voxels within this cavity region are directly deducted from the subsequent effective voxel statistics. Based on the surface morphology compensation factor, local spatial deformation compensation is performed on the surface of the basic three-dimensional geometric envelope to generate a compensated three-dimensional lesion envelope. The specific compensation process is as follows: for the surface point set of the basic three-dimensional geometric envelope... For any surface point in the array, calculate the direction of its normal vector. The polar angle of that point on the equatorial plane If a first-morphological compensation factor exists, a periodic radial perturbation is applied along the normal vector direction, and the lobed perturbation displacement is... If a second-morphological compensation factor exists, a high-frequency random radial perturbation is applied along the normal vector direction, with the spur perturbation displacement as follows: ;
[0045] in, This refers to the displacement caused by the burr disturbance, which is the distance and direction of the surface point's movement along the normal vector caused solely by the burr effect. This is the burr length coefficient (relative proportion), which controls the extension length of the burr protrusion relative to the long semi-axis. The larger the value, the longer and sharper the burr. To determine the burr density The expected value is a Poisson distribution random function used to control the frequency and spatial location of burr protrusions per unit area; the surface points are updated to... The process involves traversing all surface points to complete local spatial deformation compensation, generating a closed, compensated 3D lesion envelope. Spatial voxel accumulation calculation is then performed on the compensated 3D lesion envelope to generate the first lesion volume. The specific calculation process is as follows: The spatial region containing the compensated 3D lesion envelope is divided into a cubic voxel mesh with a side length of 0.1 mm; the ray casting method is used to determine whether the center point coordinates of each voxel are located within the compensated surface point set. The interior of the closed surface formed; if there is a boundary ambiguity coefficient Then the distance from the closed curved surface will be extended outward in the normal direction. The spatial region is defined as a boundary fuzzy transition layer, and voxels located within this transition layer are directly included in the effective voxel statistics; if a cystic concave volume coefficient exists... Then, multiple spherical cavity regions are randomly generated inside the closed surface, such that their total volume accounts for a certain percentage. The voxels located within these cavity regions are directly deducted from the effective voxel count; the total number of effective voxels is then counted. According to the formula The volume of the first lesion is calculated. Based on the results of the diameter dimension determination, when the diameter value is a single value, that single diameter value is extracted as the diameter parameter. Based on diameter parameter Construct a basic spherical geometric envelope whose surface point set satisfies the sphere equation. Using the same spatial voxel accumulation calculation method as described above, the number of effective voxels inside the basic spherical geometric envelope is counted and the volume of the second lesion is calculated.
[0046] Based on the volume of the first or second lesion, and combined with the name of the attributing organ extracted backward from the sentence containing the lesion description fragment, a pairing and binding is performed to generate an image lesion volume group.
[0047] Based on the time-series blood pressure index group, biochemical index value group, and imaging lesion volume group, sampling nodes are generated by aggregating various data using date and time strings as the primary key, and then sorted in ascending order of time to generate an initial medical index sequence, specifically including:
[0048] Based on the time-series blood pressure index group, biochemical index value group, user age value, tumor marker value group, and imaging lesion volume group, using the sampling timestamps in the time-series blood pressure index group as the primary key, for each baseline timestamp, records with precisely matching timestamps are retrieved in the biochemical index value group, user age value group, and tumor marker value group. For the imaging lesion volume group, since its sampling timestamp may not be on the same day as the biochemical test timestamp, a time alignment tolerance of ±7 calendar days is set, that is, imaging examination records within 7 days before and after the baseline timestamp are retrieved. If multiple records are found, the one with the closest timestamp to the baseline timestamp is selected; if no record is found... Upon arrival, the image lesion volume field of the sampling node is set to an empty marker. In cases where some indicators are missing at a specific sampling node, the corresponding indicator field of the sampling node is filled with an empty marker. When the calculation of this indicator is involved in subsequent steps, the sampling node automatically skips the calculation of this dimension, aggregates all data into a single sampling node, and arranges each sampling node in ascending order according to the sampling timestamp. The time interval between adjacent nodes is not uniformized, and the non-uniform interval characteristics of the original sampling are preserved. The calculation of the interval change rate in subsequent steps has been adapted to the non-uniform sampling situation by dividing point by the actual time interval, generating the initial medical indicator sequence.
[0049] In a preferred embodiment of the present invention, based on an initial medical indicator sequence, semantic associations are parsed and organ attributes and suspected occupants are identified to generate a structured entity relationship graph, including:
[0050] Organ site names are extracted from the image lesion volume group of the initial medical indicator sequence to create entity nodes. Then, lesion modification phrases adjacent to each organ site are extracted from the image text subset of the multi-source text paragraph set and bound to the corresponding entity nodes, generating an entity-modification attribute mapping. Specifically, this includes:
[0051] The initial medical indicator sequence is traversed, and the image lesion volume groups associated with each sampling node are extracted. All organ site names that appear are extracted, and an independent entity node is created for each organ site name. If the same organ site name corresponds to multiple lesion records with different sampling timestamps, only one entity node is created. The lesion volume field of this node contains multiple records for each timestamp in list form, resulting in the initial entity node set. Lesion modification phrases adjacent to each organ site name are extracted from the image text subset. The adjacency determination range is:
[0052] Starting with the organ / part name, scan one complete sentence separated by a Chinese period before and after it. Dimensional markers or morphological descriptive words appearing within this range are considered adjacent. Modifying phrases are divided into two categories: size modifiers and morphological modifiers.
[0053] Size modifications include descriptions of size variations such as approximately …cm, …×…mm, increase, and decrease; morphological modifications include descriptions of morphological features such as hypoechoic, hyperechoic, indistinct borders, spiculation, lobulation, calcification, and cystic changes. The extracted modification phrases are used as attribute tags and attached to the corresponding organ entity nodes to establish a mapping between entities and modification attributes.
[0054] The values of biochemical indicators in the initial medical indicator sequence are compared with preset medical reference intervals to determine abnormal indications. The associated organ locations are recorded according to the indicator-organ attribution relationship, generating an entity-abnormal association mapping. Entity nodes that simultaneously contain both modified records and abnormal indication records are marked as carrying suspicious placeholder attributes, generating a basic set of entity nodes, specifically including:
[0055] The system iterates through the biochemical indicator values in the initial medical indicator sequence, comparing each biochemical item's value with its corresponding preset medical reference interval. Before comparison, the user's gender field is extracted from the original medical examination report. If extraction is successful, alanine aminotransferase (ALT) and serum creatinine are compared using the corresponding gender-differentiated reference intervals. If the gender field is missing, ALT is uniformly compared to the male reference interval of 9-50 U / L, and serum creatinine is uniformly compared to the male reference interval of 54-106 μmol / L. The preset medical reference intervals are:
[0056] The following biochemical indicators are considered abnormal: serum total protein 60-80 g / L, albumin 35-55 g / L, alanine aminotransferase 9-50 U / L for men and 7-40 U / L for women, serum creatinine 54-106 μmol / L for men and 44-97 μmol / L for women, fasting blood glucose 3.9-6.1 mmol / L, and total cholesterol 2.8-5.2 mmol / L. Values below the lower limit or above the upper limit of the reference range are considered abnormal. Record the name of the abnormal indicator, the direction of the exceedance, and the associated organ location. (Alanine aminotransferase is not mentioned in the original text.) The liver is associated with serum creatinine, the kidney with fasting blood glucose and total cholesterol with the cardiovascular system. This information is bound to an abnormal pointing record and included in the entity-abnormal association mapping table. For each organ part entity node in the initial entity node set, a cross-query is performed to see if it meets two conditions simultaneously: there is at least one modification record in the entity-modification mapping and at least one abnormal pointing record in the entity-abnormal association mapping table. Entity nodes that meet both conditions are marked as carrying suspicious placeholder attributes. After completing the full annotation, the basic entity node set is output.
[0057] Based on each entity node in the basic entity node set, the corresponding biochemical indicator values, lesion volume, blood pressure indicators, and tumor marker values from the initial medical indicator sequence are written in on a time-stamp basis to generate an entity node set with quantitative attributes, specifically including:
[0058] For each entity node in the basic entity node set, a node attribute structure is created. This attribute structure contains a unique entity identifier, an organ / part name, and a mapping table with sampling timestamps as keys and indicator data blocks as values. Each indicator data block contains six sub-fields: systolic blood pressure value, diastolic blood pressure value, biochemical indicator sub-table, lesion volume value, tumor marker sub-table, and suspicious lesion marker. The associated data for the organ / part corresponding to the entity node at each sampling timestamp is extracted from the initial medical indicator sequence. This includes biochemical indicator values, lesion volume values, blood pressure values, and tumor marker values. These are then written to the indicator data block corresponding to the timestamp key in the attribute structure, with the specific writing rules as follows:
[0059] Using each sampling timestamp in the initial medical indicator sequence as the key name of the mapping table in the attribute structure, the indicator data related to the organ part of this entity node under that timestamp are filled into the corresponding sub-field. If there are biochemical items belonging to this organ part in the biochemical indicator value group of that timestamp, they are filled into the BIOCHEM sub-table; if there are lesion records belonging to this organ part in the image lesion volume group, the calculated volume is... Fill in the VOLUME subfield; blood pressure indicators and tumor marker indicators are filled in with general data based on the full timestamp, without using organ location as a filtering condition. If there is no corresponding data for a subfield under a certain timestamp, write a null value marker to the subfield to maintain the integrity of the mapping table structure, so that each entity node carries complete quantitative attributes and time series information, and obtain a set of entity nodes with quantitative attributes.
[0060] According to the organ hierarchy classification rules, entity nodes with quantified attributes are assigned to their parent category nodes and hierarchical classification edges are established. Furthermore, for two entity nodes within the same organ region carrying suspicious occupancy attributes and those indicating biochemical abnormalities, cross-index association edges are established, specifically including:
[0061] Based on the hierarchical classification of organ systems in systemic anatomy, entity nodes with quantified attributes are assigned to their corresponding system category nodes. First, four system category nodes are created:
[0062] Cardiovascular system category nodes, liver system category nodes, kidney system category nodes, and tumor marker system category nodes. System category nodes themselves do not carry quantitative attributes; they only serve as organizational nodes in a hierarchical structure. The specific attribution rules are as follows:
[0063] The cardiac entity node and its co-occurring blood pressure record vascular associated entity nodes are classified into the cardiovascular system category node; the liver and gallbladder entity nodes are classified into the liver system category node; the left and right kidney entity nodes are classified into the kidney system category node; the entity nodes representing the source organs of alpha-fetoprotein and carcinoembryonic antigen values are classified into the tumor marker system category node. A hierarchical relationship is established from each system category node to its subordinate organ / part entity nodes, with the edges uniformly directed from the higher-level category node to the lower-level entity node. The edge type is denoted as BELONGSTO. Entity nodes with quantified attributes are traversed. For any two entity nodes, an association determination is performed. If the two entity nodes are associated through the same organ part name, and one node carries a suspicious placeholder attribute while the other node has a corresponding biochemical anomaly pointing record in the entity-anomaly association mapping table, then a cross-index association relationship edge is created between the two entity nodes. The edge direction is uniformly from the entity node carrying the suspicious placeholder attribute to the entity node with the biochemical anomaly pointing record. The edge type is identified as CROSSINDICATOR, and the edge attribute records the common organ part name of the two entity nodes and the biochemical anomaly item name on which the association determination is based.
[0064] A structured entity relationship graph is constructed using all entity nodes as graph nodes and hierarchical affiliation edges and cross-index association edges as graph edges. Specifically, this includes:
[0065] All nodes in the set of entity nodes with quantified attributes are treated as graph nodes. The system category nodes are also included in the graph node set. The node type identifier of the system category nodes is denoted as SYSTEM, and the node type identifier of the organ / part entity nodes is denoted as ORGAN. All edges of hierarchical affiliation edges and cross-index association edges are treated as graph edges. Each graph edge is assigned a globally unique edge number. The edge attributes store the edge type identifier, starting node identifier, ending node identifier, and association determination criteria. A structured entity relationship graph is constructed. The graph is organized and stored in the form of an adjacency list. Each node corresponds to an adjacency entry, which lists the edge numbers of all directed edges originating from that node and the ending node identifier. Each node in the graph carries the index value and sampling timestamp in the attribute field. Each edge distinguishes hierarchical affiliation and cross-index association by the edge type identifier. The graph also constructs a reverse index table with the organ / part name as the index key, which supports direct location of the corresponding entity node and all its associated edges by the organ / part name. This forms a complete index structure that stores quantified data with node attributes and organizes the structural relationships between entities and cross-index associations with connecting edges.
[0066] In a preferred embodiment of the present invention, based on a structured entity relationship graph, time-series blood pressure is extracted, and the rate of change of blood pressure between adjacent sampling nodes over time is calculated point by point to generate an indicator fluctuation trend feature vector, including:
[0067] Based on the structured entity relationship graph, time-series blood pressure indicators bound to timestamps are extracted and arranged in ascending order of sampling time to generate a set of blood pressure time-series nodes, specifically including:
[0068] In the structured entity relationship graph, starting from the cardiovascular system category node, we traverse downwards along the hierarchical belonging relationship edge to locate all entity nodes carrying blood pressure index attribute fields, and read the sampling timestamps and their corresponding systolic and diastolic blood pressure values stored in the attribute structure of each entity node one by one.
[0069] Each read record is instantiated as a blood pressure time-series node. Each node contains three data items: sampling timestamp, systolic blood pressure value, and diastolic blood pressure value. After traversing all entity nodes and extracting records, all blood pressure time-series nodes are sorted in ascending order according to the time of sampling timestamp, with the earliest timestamp at the top and the latest timestamp at the bottom. After sorting, a set of blood pressure time-series nodes is obtained. Adjacent nodes naturally form a sequential relationship. The above process unifies the blood pressure data that are scattered in different entity nodes in the graph into a linear sequence organized along the time axis, making the continuous time-series characteristics of blood pressure indicators explicit from the graph structure, and turning it into an ordered dataset that can be directly used for point-by-point numerical calculations.
[0070] Based on the blood pressure time series node set, the blood pressure value difference and time interval length between adjacent sampling nodes are extracted to generate a node difference parameter set, specifically including:
[0071] Starting from the beginning of the blood pressure time series node set, take two adjacent blood pressure time series nodes to form a node pair. The node that appears earlier is called the preceding node, and the node that appears later is called the following node. For each node pair, extract the following three difference parameters:
[0072] Subtracting the systolic blood pressure value of the preceding node from the systolic blood pressure value of the subsequent node yields the net change in systolic blood pressure between two adjacent sampling times. A positive value indicates an increase in systolic blood pressure, while a negative value indicates a decrease. The absolute value reflects the absolute magnitude of the systolic blood pressure variation within that interval. Subtracting the diastolic blood pressure value of the preceding node from the diastolic blood pressure value of the subsequent node has a similar meaning to the systolic blood pressure difference, reflecting the net change in diastolic blood pressure between two sampling times. Subtracting the sampling timestamp of the preceding node from the sampling timestamp of the subsequent node yields the actual interval between two adjacent sampling times. For timestamps in date and time format, both timestamps are first converted to the cumulative number of hours calculated from the reference time before subtraction.
[0073] The above three difference parameters are encapsulated into a group as the difference parameter group for the node pair. Starting from the first node pair, the process is slid forward pair by pair until the last pair of adjacent nodes is processed. This results in a complete sequence consisting of difference parameter groups equal in number to the adjacent node pairs. After this step, the absolute values in the original blood pressure time series node set have been transformed into relative changes between adjacent nodes and the corresponding time span.
[0074] Based on the node difference parameter set, the rate of change of blood pressure values between adjacent sampling nodes with the length of the time interval is calculated point by point to generate an interval change rate sequence, which specifically includes:
[0075] For each set of difference parameters obtained in the previous step, the change in blood pressure is divided by the corresponding time interval for both systolic and diastolic blood pressure dimensions to calculate the intensity of change per unit time. For the first... Take the differential parameters, extract the systolic blood pressure difference, diastolic blood pressure difference, and time interval length, and calculate the rate of change of systolic blood pressure interval and the rate of change of diastolic blood pressure interval: ; ;
[0076] In the formula, For the first The rate of change of systolic blood pressure within each adjacent sampling interval; This represents the rate of change of the corresponding diastolic blood pressure range; This refers to the systolic pressure difference value extracted earlier for this interval; This represents the diastolic pressure difference within that range. Given the time interval length extracted earlier, the interval change rate calculated in this step is based on the ratio of the blood pressure difference between two discrete physical examination sampling points to the actual number of calendar days. It reflects the long-term linear trend of blood pressure levels within this period. This indicator is a statistical characteristic describing long-term changes and does not characterize physiological instantaneous fluctuations within the cardiac cycle, nor is it used to predict acute hemodynamic events. Within this adjacent sampling interval, it represents the average rate of change of blood pressure values over time. Positive values indicate an upward trend in blood pressure, while negative values indicate a downward trend. The larger the absolute value, the greater the cumulative change in blood pressure within that interval. The calculated rates of change for each pair of systolic and diastolic blood pressure intervals are arranged sequentially according to the original adjacent node pairs. Each record synchronously carries the sampling timestamp of the corresponding preceding node, generating an interval change rate sequence.
[0077] Based on the interval change rate sequence, the direction identifier and magnitude of the rate change at each sampling node are extracted to generate a feature vector of the index fluctuation trend, specifically including:
[0078] The algorithm iterates through each record in the interval change rate sequence, determining the direction and extracting the amplitude of the change trend in each sampling interval. For the systolic blood pressure interval change rate in each record, the direction of rise or fall is determined based on its numerical sign: a rate value greater than zero indicates an upward trend; a rate value less than zero indicates a downward trend; and a rate value equal to zero indicates a stable trend. The absolute value of this rate is taken as the systolic blood pressure rate amplitude, and the magnitude of the amplitude reflects the severity of the blood pressure change in that interval. The same direction determination and amplitude extraction operation is performed on the diastolic blood pressure interval change rate. For sampling intervals where the systolic and diastolic blood pressure direction determination results are consistent, the overall direction identifier for that interval is taken from the consistent direction. For intervals where the two directions are inconsistent, the overall direction identifier is taken from the direction corresponding to the side with the larger rate amplitude, and a mark indicating that there is a divergence between the systolic and diastolic blood pressure trends in that interval is added.
[0079] Each record's interval start timestamp, comprehensive direction identifier, systolic blood pressure rate amplitude, and diastolic blood pressure rate amplitude are encapsulated into a feature dimension unit. All feature dimension units are arranged in chronological order of their timestamps to generate an indicator fluctuation trend feature vector. The structure of the indicator fluctuation trend feature vector is as follows:
[0080] Using the sampling timestamp as the horizontal axis index, each timestamp position corresponds to a set of feature values, including the direction of blood pressure change in the adjacent intervals at that position, as well as the intensity of change in systolic and diastolic blood pressure.
[0081] In this embodiment of the invention, by extracting time-series blood pressure data bound with timestamps from the structured entity relationship graph and sorting it to generate a set of blood pressure time-series nodes, the technical problems of traditional fixed time window mean or extreme value descriptions being unable to capture the differences in blood pressure changes between adjacent sampling nodes with the sampling time interval, and the inability to effectively locate the time-series positions of reversal nodes and state changes in the blood pressure fluctuation trend are overcome. This achieves the technical effects of point-by-point quantitative characterization of the rate of change of blood pressure intervals, simultaneous characterization of the dual-dimensional fluctuation features of systolic and diastolic blood pressure, and direct location and retrieval of rising and falling reversal nodes in subsequent steps.
[0082] In a preferred embodiment of the present invention, based on the indicator fluctuation trend feature vector, the rate difference and direction identifier are extracted, sampling nodes with rising and falling reversals are screened and time-series position identifiers are extracted, and a local fluctuation difference sequence is generated, including:
[0083] Based on the indicator fluctuation trend feature vector, the interval change rate value and rate change direction identifier corresponding to each continuous sampling node are extracted to generate a node rate feature set, which specifically includes:
[0084] The indicator fluctuation trend feature vector has been constructed as described above. Internally, it uses the sampling timestamp as the horizontal axis index. Each timestamp position is bound to a set of feature values, including the comprehensive direction identifier of that interval, the systolic blood pressure rate amplitude, and the diastolic blood pressure rate amplitude. The comprehensive direction identifier takes one of three states: rising, falling, or stable. Iterates through each position from front to back along the time axis of the indicator fluctuation trend feature vector, and extracts the comprehensive direction identifier and systolic blood pressure rate amplitude bound to that position. If the direction identifier is rising, the systolic blood pressure rate amplitude is given a positive sign as the signed systolic blood pressure rate value for that position; if the direction identifier is falling, it is given a negative sign; if it is stable, it is taken as zero. The diastolic blood pressure rate amplitude is processed according to the same rule to obtain the signed diastolic blood pressure rate value. The positive or negative sign of the signed rate value directly encodes the direction of blood pressure rise and fall in that interval, and the absolute value is equal to the original rate amplitude. After processing each timestamp, a rate feature record is obtained, containing three fields: the interval start timestamp, the signed systolic blood pressure rate value, and the signed diastolic blood pressure rate value. All rate feature records for all timestamp positions are arranged in ascending order of timestamp to generate a node rate feature set. Adjacent records in this set are arranged sequentially on the time axis, and each record fully describes the direction and intensity of blood pressure changes in both the systolic and diastolic phases within the corresponding sampling interval.
[0085] Based on the node rate feature set, the rate change between adjacent sampling nodes is calculated point by point, and the direction of rise and fall is identified, generating a node rate change parameter set, specifically including:
[0086] Starting from the beginning of the node rate feature set, two adjacent rate feature records are sequentially selected to form an adjacent feature pair. The preceding record is called the previous feature, and the following record is called the subsequent feature. For each adjacent feature pair, the rate jump amplitude between the two intervals is first calculated. Let the signed systolic pressure rate value of the previous feature be... The signed systolic blood pressure rate value of the post-position characteristic Calculate the difference between the two. , A value greater than zero indicates that the systolic blood pressure fluctuation intensity of the later interval is greater than that of the previous interval, while a value less than zero indicates that it is less. The absolute value reflects the severity of the fluctuation intensity jump between the two intervals. The same difference operation is performed on the signed diastolic blood pressure rate value to obtain the rate jump amplitude of the diastolic blood pressure dimension.
[0087] like >0 and <0 indicates that blood pressure was rising in the previous interval and falling in the adjacent next interval; the change in the direction of rise and fall of this adjacent feature pair is marked as falling. <0 and A value >0 indicates that blood pressure has changed from decreasing to increasing, and is recorded as an increase; if and Same sign, denoted as same direction, is used to compare the signs of signed rate values in the diastolic blood pressure dimension according to the same rules and mark the direction of change. The above calculation results are encapsulated into a set of rate change parameters: the interval start time stamp of the preceding feature, the interval start time stamp of the following feature, and the amplitude of the systolic blood pressure rate jump. The parameters for the change in diastolic blood pressure rate, the change in systolic blood pressure direction, and the change in diastolic blood pressure direction are processed one by one from the first adjacent feature pair until the last pair is processed, thus generating a set of node rate change parameters.
[0088] Based on the reversal of the rising and falling directions in adjacent intervals of the node rate change parameter set, sampling nodes where the blood pressure fluctuation trend reverses are selected, generating an initial state inflection node set, specifically including:
[0089] Iterate through each set of parameters in the node rate change parameter set, and examine the values of the systolic blood pressure rise / fall direction change indicators and diastolic blood pressure rise / fall direction change indicators one by one. The core criterion for screening is whether the direction change indicator is equal to rising or falling. The screening process is classified and handled according to the following three cases:
[0090] The change in the direction of systolic blood pressure rise or fall is indicated as an increase or decrease, while diastolic blood pressure is indicated as the same direction. The time sequence position corresponding to this set of parameters is marked as a systolic-related inflection point, and the reversal type is recorded as a unidirectional reversal. The specific direction of the reversal is recorded simultaneously. This point indicates that only the blood pressure fluctuation trend of the systolic phase changes direction at this point, while the diastolic phase maintains its original trend. The change in the direction of diastolic blood pressure rise or fall is indicated as an increase or decrease, while systolic blood pressure is indicated as the same direction. The time sequence position corresponding to this set of parameters is marked as a diastolic-related inflection point, and the processing method is symmetrical to case one. The change in the direction of systolic and diastolic blood pressure rise or fall simultaneously is indicated as an increase or decrease, that is, the systolic and diastolic phases reverse their trends synchronously at the same sampling position. This position is marked as a bidirectional reversal inflection point, and the reversal type is recorded as a bidirectional reversal. The reversal direction of each phase is recorded simultaneously. A bidirectional reversal node means that the regulatory trends of systolic and diastolic blood pressure turn at the same point. For each marked turning point, the interval start time stamp of the preceding feature in the corresponding parameter group is extracted as the temporal position identifier of the turning point. At the same time, the systolic blood pressure rate jump amplitude Δr, the diastolic blood pressure rate jump amplitude, and the reversal type at this position are extracted.
[0091] All the selected inflection points are collected to generate an initial state inflection point set. Each element in the set contains four types of information: time-series location timestamp, systolic blood pressure rate jump magnitude, diastolic blood pressure rate jump magnitude, and inversion type.
[0092] Based on the initial state inflection point set, the sampling timestamp and rate change corresponding to each inflection point in the blood pressure time series node set are extracted to generate a local fluctuation difference sequence, specifically including:
[0093] Using the time-series timestamp of each transition node in the initial state transition node set as the query key, the blood pressure time-series node set is retrieved. Each node in the blood pressure time-series node set uses the sampling timestamp as the primary key and carries the original measured systolic and diastolic blood pressure values at that moment. The node whose timestamp precisely matches the query key is located in the blood pressure time-series node set, and the original systolic and diastolic blood pressure values from that node are extracted and attached to the corresponding transition node. After attachment, each transition node carries two sets of information: the rate jump amplitude and the original blood pressure values. Along the time axis, two temporally adjacent transition nodes in the transition node set with attached original blood pressure values form a transition interval pair. The one with the earlier time sequence is the starting transition node, and the one with the later time sequence is the ending transition node. For each transition interval pair, the following information is extracted sequentially:
[0094] The time-series timestamp of the starting inflection point serves as the interval start timetamp; the time-series timestamp of the ending inflection point serves as the interval end timetamp; the amplitude of the systolic and diastolic blood pressure rate jumps at the starting inflection point serves as the starting boundary rate characteristic; the amplitude of the systolic and diastolic blood pressure rate jumps at the ending inflection point serves as the ending boundary rate characteristic; the original systolic and diastolic blood pressure values at the starting inflection point serve as the starting boundary blood pressure benchmark; the original systolic and diastolic blood pressure values at the ending inflection point serve as the ending boundary blood pressure benchmark; and the reversal type identifiers for the starting and ending inflection points are also provided. All of the above information is integrated into a local fluctuation difference unit, which fully describes the fluctuation segment between one trend inflection point and the next trend inflection point.
[0095] The start and end times of the interval indicate the time span of the segment; the rate jump amplitude at the start and end boundaries indicates the intensity of fluctuations at the beginning and end of the segment; the original blood pressure values at the start and end boundaries indicate the absolute blood pressure levels at the beginning and end of the segment; the reversal type indicates whether the turning point is unidirectional or bidirectional; all turning point intervals are arranged sequentially in chronological order to generate a local fluctuation difference sequence. This sequence uses adjacent trend turning points as boundaries to divide the entire time-series blood pressure record into several independent fluctuation segments. The direction of blood pressure change within each segment remains consistent, with no direction reversal; the trend direction changes between segments at the boundaries.
[0096] In this embodiment of the invention, a set of node rate features is generated by extracting the signed rate value and direction identifier from the indicator fluctuation trend feature vector. This overcomes the technical problems of traditional mean statistics, which cannot capture the rate of change of blood pressure intervals point by point, cannot accurately locate the trend reversal node, and cannot quantify the amplitude of the fluctuation intensity jump before and after the reversal. It achieves the technical effects of accurately screening the rise and fall reversal nodes in the blood pressure fluctuation trend, automatically distinguishing between unidirectional and bidirectional reversal types, and cutting the continuous time series into fluctuation segments with the trend turning point as the boundary.
[0097] In a preferred embodiment of the present invention, based on the local fluctuation difference sequence and the structured entity relationship map, the frequency of biochemical abnormalities is calculated by comparing with the medical reference range, the total abnormal load and lesion volume of each organ are calculated by summarizing the entire domain, feature splicing is performed, and a multidimensional risk accumulation matrix is generated, including:
[0098] Based on the structured entity relationship graph, the values of each biochemical indicator are extracted and compared with preset medical reference intervals to calculate the frequency of indicators exceeding the normal range, generating the frequency of abnormal biochemical indicators, specifically including:
[0099] In the structured entity relationship graph, starting from the liver system category node and the kidney system category node, the system traverses downwards along the hierarchical relationship edge to locate the entity node carrying biochemical indicator attribute fields. It then reads the biochemical indicator values stored in the attribute structure of each entity node under each sampling timestamp, categorizes and aggregates them according to the indicator item name, and pre-sets the medical reference intervals as follows:
[0100] Serum total protein 60–80 g / L, albumin 35–55 g / L, alanine aminotransferase 9–50 U / L for men and 7–40 U / L for women, serum creatinine 54–106 μmol / L for men and 44–97 μmol / L for women, fasting blood glucose 3.9–6.1 mmol / L, and total cholesterol 2.8–5.2 mmol / L.
[0101] For each indicator item, all timestamp detection values are compared one by one with the upper and lower limits of the corresponding reference interval. Values below the lower limit are recorded as low-value anomalies, and values above the upper limit are recorded as high-value anomalies. Values falling within the interval are not counted as anomalies. During comparison, the corresponding gender-differentiated reference interval is selected based on the user's gender. After comparing all timestamps, statistics are compiled separately for each organ site. Alanine aminotransferase, serum total protein, and albumin are categorized under liver site statistics; serum creatinine is categorized under kidney site statistics; and fasting blood glucose and total cholesterol are categorized under cardiovascular system statistics. For each organ site, the total number of abnormal occurrences of each indicator item across all sampling timestamps is summarized to obtain the biochemical indicator abnormality frequency for that organ site. The names of each organ site and their corresponding biochemical indicator abnormality frequencies are paired and stored to generate a biochemical indicator abnormality frequency table. This table provides a scalar indicator for each organ site, representing the total number of times the biochemical detection values associated with that organ deviate from the normal range throughout the entire sampling period. The larger the value, the denser the accumulated abnormal signals at the biochemical level for that organ.
[0102] Based on the frequency of abnormal biochemical indicators and the volume of lesions described in imaging, the total abnormal burden of each organ site is calculated and aggregated across the entire domain, generating a global organ burden feature vector, which specifically includes:
[0103] The frequency table of abnormal biochemical indicators and the volume group of imaging lesions were merged across sources. Using organ site name as the association key, the total abnormality burden of each organ site was calculated. This total was a weighted composite of the contributions from biochemical abnormalities and structural abnormalities. abnormal frequency of biochemical indicators Divide by the total number of biochemical tests performed on that organ site across all sampling cycles. ,Right now: ;
[0104] in, This is a biochemical anomaly density index, with values ranging from... Dimensionless values of intervals; organ parts Image description of lesion volume Divide by the mean reference volume of that organ site in the normal population. ,Right now: ;
[0105] The liver reference volume is taken as a relative lesion volume ratio. Kidney reference volume thyroid reference volume is taken Let the organ location be... The total abnormal load on that organ site Calculate using the following formula: In the formula, This is the biochemical anomaly density weighting coefficient, with a value of 0.3; The relative lesion volume ratio weighting coefficient is set to 0.7, and ,exist , Under the constraint, traverse all [variables] with a step size of 0.05. The optimization objective is to maximize the F1-score of the underwriting hierarchical classification macro-average in the subsequent GBDT model, which is obtained by combining the global organ burden feature vectors generated by each combination and performing five-fold cross-validation. , For the optimal combination, during the grid search process, each The average F1-score of the combined five-fold cross-validation macros is as follows: It is 0.801. It is 0.807. It is 0.812. It is 0.809. It is 0.803, based on which it is determined , For optimal combination, in practical applications, users can re-execute the above grid search using their own training data to adapt it to specific population characteristics, or directly adopt the optimized values provided by this method. , As a general default configuration, this applies to organ sites for which no corresponding lesion is recorded in the imaging lesion volume group. If the value is zero, the abnormal burden of the organ is contributed solely by the frequency of biochemical abnormalities. For organ sites without corresponding records in the biochemical index abnormality frequency table, Take zero, and set all organ parts to zero. Arranged by organ location name, a global organ load feature vector is generated. Each dimension of the vector corresponds to an organ location, and the dimension value is the total abnormal load of that organ. This vector, from a global perspective, quantifies abnormal information across two independent dimensions—biochemical testing and imaging examination—into a unified comprehensive risk load for each organ.
[0106] Based on the local fluctuation difference sequence, the coefficient of variation of blood pressure values within each fluctuation segment is calculated segment by segment, and the cumulative mean is calculated to generate a cumulative blood pressure variation sequence, specifically including:
[0107] The local fluctuation difference sequence, as previously generated, divides the time-series blood pressure record into several independent fluctuation segments, with adjacent trend inflection points serving as boundaries. Within each fluctuation segment, the direction of blood pressure change remains consistent. Each fluctuation segment in the local fluctuation difference sequence is traversed, and the systolic and diastolic blood pressure values of all sampling points within that segment are extracted. For any given fluctuation segment, the standard deviation and mean of systolic blood pressure, as well as the standard deviation and mean of diastolic blood pressure, are calculated using the standard deviation and mean calculation method. When the number of sampling points within a segment is greater than or equal to 2, the coefficient of variation of systolic blood pressure is calculated using the following formula. With diastolic blood pressure coefficient of variation : ;
[0108] in, This represents the standard deviation of systolic blood pressure within this fluctuation range. This represents the average systolic blood pressure within this fluctuation range. The standard deviation of diastolic pressure within this fluctuation range. This represents the mean diastolic blood pressure within this fluctuation range. The average of these two values is taken as the combined blood pressure variability coefficient for this fluctuation range. : ;
[0109] When the number of sampling points within a fluctuation segment is only 1, the standard deviation cannot be calculated, and the composite blood pressure variability coefficient for that segment is also unavailable. The value is set to 0. The coefficient of variation (BPV) is a clinically recognized quantitative indicator of blood pressure variability. A higher BPV value indicates greater blood pressure fluctuations within that time period and is positively correlated with the risk of cardiovascular events. The combined BPV of each fluctuation segment is then used to calculate the overall blood pressure variability. Arrange the fluctuation segments in chronological order along the time axis to form a blood pressure fluctuation characteristic sequence. ,in The total number of fluctuation segments, for any sampling time section. Extract the coefficient of variation of all fluctuation segments that have occurred before this cross section. ,in As of The total number of fluctuation segments that have occurred at any given time is used to calculate the cumulative average value of blood pressure variability at that cross-section. : ;
[0110] in, For the first The overall blood pressure variability coefficient for each fluctuation segment, For the sampling time section The total number of fluctuating segments that have appeared. The dimensionless value of indicates that the larger the value, the greater the overall fluctuation of blood pressure since the start of monitoring. The values at each sampling time point are... Arranged in timestamp order, a cumulative blood pressure variation sequence is generated.
[0111] Based on the global organ load feature vector and the cumulative blood pressure variability sequence, a multidimensional attribute weighted concatenation operation is performed to reconstruct the physiological state dimension, resulting in a multidimensional risk accumulation matrix, which specifically includes:
[0112] The global organ load feature vector describes the comprehensive abnormal load of various organ sites from a static organ dimension, while the cumulative blood pressure variability sequence describes the cumulative process of blood pressure fluctuations over time from a dynamic time dimension. These two datasets capture the static structural and dynamic functional aspects of health risk, respectively. Dimensional reorganization is required within a unified framework. The cumulative blood pressure variability sequence is time-aligned with the global organ load feature vector, using the sampling timestamps of each organ site in the global organ load feature vector as reference points. The cumulative blood pressure variability value with the closest timestamp is then found within the cumulative blood pressure variability sequence. (t), this value is used as a supplementary dimension of vascular system load for this time segment. The organ load and vascular load for each time segment are weighted and concatenated to construct the risk accumulation vector for that segment. Let the total abnormal load of each organ in the global organ load feature vector of this segment be . ( (Total number of organs / sites), the cumulative blood pressure variation corresponding to this cross-section is Then the risk accumulation vector of this section for: ;
[0113] In the formula, For the first The total abnormal load of each organ site, dimensionless; For the first The physiological weight coefficients for each organ site are dimensionless and are configured according to the following rules: liver 1.2, kidney 1.0, cardiovascular system 1.5, and other organs 0.8. The weight differences reflect the different contribution strengths of different organs to overall health risks. The cross-dimensional mapping coefficient of blood pressure variability is dimensionless and has a value of 1.0, which maps vascular load from its original dimension to a risk space comparable to organ load. For this time section, the multidimensional risk accumulation vector is... The risk accumulation matrix is formed by stacking the risk accumulation vectors of all time sections in ascending order of time into a matrix. The rows of the matrix represent the time sections, and the columns represent the organ load dimension and the vascular load dimension.
[0114] like Figure 3 As shown, in a preferred embodiment of the present invention, based on the multidimensional risk accumulation matrix, an attribute decoupling operation is performed to remove redundant dimensions and generate a core health status feature set, including:
[0115] Based on the multidimensional risk accumulation matrix, the central reference value of each physiological risk dimension is calculated column by column across all sampling time sections to generate a baseline zero-return fluctuation matrix. Based on the baseline zero-return fluctuation matrix, the degree of co-variance between each pair of physiological risk dimensions is calculated to generate a dimensional co-variance distribution matrix, specifically including:
[0116] Based on the multidimensional risk accumulation matrix, whose rows represent the sampling time sections, a total of The columns are categorized by physiological risk dimensions, totaling [number]. Dimension, among which For the total number of organ sites, the first The dimension is the cumulative dimension of blood pressure variability, and the matrix is denoted as... ,for OK Column matrix, where Before performing centralized processing, first verify the number of sampling time segments. ,like If so, skip this step and directly output each row of the original multidimensional risk accumulation matrix as the core health status feature set. At that time, continue to perform the following processing on the matrix. Centralized processing is performed, and each physiological risk dimension is calculated column by column. The mean value at the nth time cross section, let the nth time cross section be the mean value. The mean of the column is ,Will The Middle Subtract each element of the column The centered matrix is obtained. , This is the baseline zero-range fluctuation matrix, where the mean of each column is zero. The relative fluctuation characteristics between columns are highlighted, and the baseline offset caused by differences in dimensions or value ranges in the original dimensions is eliminated. Based on the baseline zero-range fluctuation matrix... Constructing a dimensional collaborative distribution matrix , for The square array, its first Line number The column element is the first Wei and Di The degree of coordinated fluctuation between dimensions: ;
[0117] In the formula, The baseline zero-return fluctuation matrix is the first Line number Column elements; The baseline zero-return fluctuation matrix is the first Line number Column elements; For the first Wei and Di The degree of coordinated fluctuation between dimensions; when hour, For the first The independent variance of a dimension reflects the amplitude of the fluctuation of that dimension itself; when hour, This reflects the degree of coordination between fluctuations in two dimensions; positive values indicate fluctuations in the same direction, and negative values indicate fluctuations in opposite directions. (Dimensional Coordination Distribution Matrix) Fully depicted The fluctuation coupling structure between each pair of physiological risk dimensions is such that the main diagonal elements of the matrix constitute the independent fluctuation variance of each dimension. The closer the absolute value of the off-diagonal elements is to the square root of the product of the main diagonal elements, the stronger the linear correlation between the corresponding two dimensions.
[0118] Based on the dimensional cooperative distribution matrix, the independent fluctuation directions and their respective proportions of independent information are calculated. These are then arranged in descending order of their independent information proportions to generate a sequence of independent fluctuation directions, specifically including:
[0119] For dimensionally co-distributed matrix Perform eigenvalue decomposition to obtain the solution. There are eigenvalues, denoted as . and the corresponding Unit eigenvectors Each eigenvector is pairwise orthogonal and has a magnitude of 1, representing the original vector. An independent fluctuation direction in 3D space, each eigenvalue Equals the original data in its corresponding feature vector The projection variance along a direction; the larger the eigenvalue, the wider the original data is distributed along that eigenvector direction, meaning the greater the amount of independent information carried by that independent fluctuation direction. The proportion of independent information in each independent fluctuation direction Calculate using the following formula: ;
[0120] In the formula, For the first The eigenvalues in each direction, the magnitude of which is directly equal to the original data projected onto the image. The degree of dispersion (variance) after direction. For all The sum of all eigenvalues represents the total amount of information (total variance) contained in the original data. The value of is between 0 and 1, all One direction The sum equals 1, eigenvector Each component represents the original The contribution weights of each physiological risk dimension to the independent fluctuation direction are calculated, with larger absolute values indicating a stronger correlation between the original dimension and the direction. The independent fluctuation directions are then arranged from largest to smallest proportion of their corresponding independent information content. Corresponding to the maximum information content ratio, The independent fluctuation direction sequence is obtained by arranging the corresponding minimum information content proportions.
[0121] Based on the independent fluctuation direction sequence, the proportion of independent information in each direction is accumulated item by item, starting from the first item, until the accumulated value reaches a preset energy retention threshold. The independent fluctuation directions within the accumulated range are then extracted as the core feature direction set, specifically including:
[0122] After arranging the eigenvalues in descending order, calculate the first... Cumulative information retention rate in each independent fluctuation direction : ;
[0123] In the formula, the denominator is all The sum of the eigenvalues represents the total information content of the original data; the numerator is the sum of the first eigenvalues. The sum of the largest eigenvalues represents the sum of the first few eigenvalues selected. The amount of information retained in each independent fluctuation direction; For the front The cumulative information retention rate in each direction, with values between 0 and 1. A preset energy retention threshold. ,from Start by incrementing sequentially, and calculate the corresponding... until satisfaction is found The smallest Value, denoted as ,like There is already ,but Only retain the independent fluctuation direction with the largest proportion of information; if Increment to hour ,but All directions are preserved. In the scenario of multi-organ data in routine physical examinations, Typically significantly smaller than Before selection Each independent fluctuation direction is taken as the core feature direction set, and the remaining... The information carried by each direction is less than 15%, and the fluctuation information it carries has a small statistical variance. It may contain noise or minor physiological variation information, so it is stripped as a redundant direction and the projection data of these directions will not be used in subsequent steps.
[0124] Based on the core feature direction set and the baseline zero-return fluctuation matrix, the baseline zero-return fluctuation vector of each sampling time segment is projected onto each core feature direction one by one. The projection values of each time segment onto each core feature direction constitute the core health status feature set, which specifically includes:
[0125] Zeroing the baseline fluctuation matrix To the selected front Projection of a subspace composed of eigenvectors. Let the projection matrix be... From the beginning The feature vectors are arranged in columns and have a dimension of . , Each column represents a core feature direction vector, and the projected score matrix... , dimension Calculate using the following formula: ;
[0126] In the formula, The Line number The column element is the first The time section at the ... The projection score along the direction of the core feature, the physical meaning of which represents the projection score of the first core feature. After removing baseline offset, the multi-organ risk status at each time segment, on the [number]th [time segment], was [statistic]. The degree of deviation in the direction of the core feature is considered. The larger the absolute value of the score, the more significant the deviation of the cross-section from the mean of all time cross-sections in that core feature direction; a score of zero indicates that the cross-section is at the average level in that direction. Each column represents a core health status feature dimension. These dimensions are linearly uncorrelated due to their projection directions originating from pairwise orthogonal feature vectors, eliminating linear information redundancy. Each time segment consists of... The vector composed of the projection scores is the core health status feature vector of this section, and the score matrix is... of The core health status feature set is composed of column vectors, with a dimension number of... Less than the original number of dimensions The stripped dimensions correspond to independent fluctuation directions with smaller feature values, and the core health state feature vector for each time segment is composed of... It consists of linearly uncorrelated projection scores.
[0127] In this embodiment of the invention, by performing principal component extraction on the multidimensional risk accumulation matrix, the technical problems of feature coupling affecting classification accuracy due to redundancy of high-dimensional physiological data and lack of a unified quantitative integration framework for multi-source information at different time sections are overcome, achieving the technical effects of independent extraction of core risk features and effective information compression, automated and standardized generation of health rating labels.
[0128] In a preferred embodiment of the present invention, based on a core health status feature set, combined with continuously transmitted age and tumor markers, a tree-structured risk grading network is input to extract the risk level corresponding to the highest probability value, generating an initial health rating label, including:
[0129] Based on the core health status feature set, the projected scores of each independent core feature dimension are extracted and horizontally concatenated with the continuously transmitted user age value and tumor marker detection values to generate a multi-dimensional classification input vector, specifically including:
[0130] Based on the core health status feature set, in the form of a score matrix ,common OK The column extracts the user age values that have been acquired and stored in this step from the initial medical indicator sequence. Alpha-fetoprotein levels compared to tumor marker values With carcinoembryonic antigen (CEA) values For each sampling time segment, this segment is plotted in the scoring matrix. The corresponding The projected score is extracted and horizontally concatenated with the user's age value and the values of two tumor markers to generate a multidimensional classification input vector for that cross-section. In special cases where the information content of all principal components is below the preset threshold and the data variance is extremely small or close to a constant, the initial health rating label for that time segment is directly set to the preset minimum risk level, and the subsequent classification process is terminated. The core health status feature vector of each cross section is The concatenated multidimensional classification input vector for: ;
[0131] In the formula, The dimension is ; to For the first A time section in Projection scores along the directions of each core feature; The user's age value; This refers to the alpha-fetoprotein (AFP) test result; For carcinoembryonic antigen (CEA) detection values, in cases where tumor marker values are missing at a certain sampling time segment, the corresponding... or The field is filled with the median value of the marker across all sampling sections as the default fill value. This vector integrates three types of information: core physiological state characteristics decoupled by principal components, demographic parameters reflecting age-related baseline risk, and quantitative indicators of serum tumor markers reflecting lesion activity.
[0132] Based on the multidimensional classification input vector, it is input into a risk classification network composed of multiple sequentially cascaded decision trees to generate a comprehensive risk score, specifically including:
[0133] Will Input a pre-trained tree-structured risk classification network, which is a gradient boosting decision tree ensemble model, internally consisting of... The decision trees are constructed by sequentially cascading. The value is 200. The training dataset comes from physical examination data of a tertiary hospital's health management center from 2018 to 2023, containing a total of 15,623 subject samples. The feature vector dimension of each sample is [missing value]. The model includes core health status projection scores, user age, and tumor marker values. Risk level labels were independently assigned by two clinical physicians with associate chief physician or higher titles, based on the examinee's new disease diagnoses, hospitalization records, and all-cause mortality events within 3 years after the physical examination. The labeling consistency Kappa coefficient was 0.85. The model uses a log loss function (LogLoss) as the optimization objective and employs a learning rate of 0.05, a maximum tree depth of 6, a subsampling ratio of 0.8, and a minimum number of leaf node samples of 20 as the GBDT superposition parameters. The parameters are iteratively trained tree by tree, with each new decision tree fitting the prediction residuals of the preceding tree set. Training stops when the number of decision trees reaches 200 or the log loss on the validation set no longer decreases after 10 consecutive rounds. Five-fold cross-validation is used to evaluate the model performance. The validation results on the independent test set are: macro-average F1-score of 0.798, AUC of 0.834, and single-class recall rates for each risk level are 0.82 for level I, 0.79 for level II, 0.76 for level III, and 0.81 for level IV. The weight coefficients of the Softmax layer are... and bias terms These are parameters learned synchronously with the decision tree parameters during the model training phase. After training is complete, they are fixed into the model file along with the decision tree parameters, representing each risk level. satisfy The monotonically increasing relationship ensures the comprehensive risk score The higher the number, the greater the probability of a high-risk level.
[0134] After the model is trained, the split threshold of each node within the decision tree, the output value of the leaf nodes, and the softmax layer are all considered. All parameters are embedded in the model file and loaded directly during the inference phase. The network processing flow is as follows:
[0135] The first decision tree receives As input, starting from the root node, at each internal node, based on the comparison result of a certain feature dimension value with a preset splitting threshold (values greater than the threshold go to the right child node, values less than or equal to the threshold go to the left child node), the path is routed layer by layer downwards along the tree structure until a leaf node is reached. The initial risk score is output from the leaf node, and the second decision tree receives the same input. Using the residuals output by the preceding tree as the fitting target, residual correction scores are output according to the same decision routing method. Subsequent trees continue to fit based on the residuals of the preceding tree set. Tree output score increment After all After calculating the scores for each tree in sequence, the score increments output by each tree are summed to obtain the comprehensive risk score. : ;
[0136] In the formula, This represents the total number of decision trees in the ensemble model, with a value of 200. For the first The score increment of each decision tree output; For comprehensive risk scoring, A higher value indicates a higher overall risk quantification assessment result for the multidimensional classification input vector under the current model. The range of values for is determined by the label distribution of the training data, and after optimization with log loss, it usually falls within . Interval.
[0137] Based on the comprehensive risk score, and combined with the weighting coefficients and bias terms corresponding to each preset risk level, an exponential operation is performed and normalized to generate a probability distribution for each preset risk level. The risk level number with the highest probability value in the probability distribution is extracted to generate an initial health rating label, which specifically includes:
[0138] Comprehensive risk score The probability distribution of each preset risk level is mapped using the Softmax function. The preset risk levels are divided into four levels: Level I for low risk, Level II for medium risk, Level III for relatively high risk, and Level IV for high risk. The probability of each risk level Calculate using the following formula: ;
[0139] In the formula, To sum the index scores of all four risk levels and use them as a normalization factor, its function is to normalize the scores of all levels. The sum of the values equals 1, thus outputting a valid probability distribution. For the circular index in the summation formula, For the first Weighting coefficients for each risk level; For the first Bias items for each risk level; and All parameters are learned synchronously with the decision tree parameters during the model training phase, and are then stored in the model file after training is complete. They can be used directly during the inference phase, and there are four risk levels. satisfy The monotonically increasing relationship ensures the comprehensive risk score The higher the risk level, the greater the probability of a high-risk level and the smaller the probability of a low-risk level. The denominator is the sum of all four risk levels after exponential transformation, used to normalize the output into a probability distribution, so that... , The subjects at the current time segment are divided into the first... The probability of each risk level is calculated, with values ranging from 0 to 1. After obtaining the probability distribution of the four risk levels, the risk level number corresponding to the highest probability value is extracted. : ;
[0140] In the formula, The risk level number corresponding to the highest probability value. If two or more risk levels have the same probability value and are all at their maximum value, the level with the larger number shall be selected as the risk level. The initial health rating label is output as level I, II, III or IV for that time segment. The initial health rating label compresses the multi-source health information contained in the multi-dimensional classification input vector into a discrete risk level identifier.
[0141] like Figure 2 As shown, in a preferred embodiment of the present invention, based on the initial health rating label, the first and second underwriting tiers are divided by comparing the underwriting stratification thresholds, the corresponding set of coverage liability configuration parameters is associated, and a data labeling system is output, including:
[0142] Based on the initial health rating labels, the included disease risk level values and core physiological state dimension weights are analyzed to generate an underwriting assessment parameter set, specifically including:
[0143] Based on the initial health rating label, the label value is either Level I, Level II, Level III, or Level IV. This label is a discrete risk level identifier obtained by extracting the category corresponding to the maximum probability value after the multidimensional classification input vector is softmax mapped by the tree-structured risk grading network mentioned earlier. It maps the initial health rating label to a risk level value that can participate in numerical comparison and interval determination. The mapping rule is:
[0144] Level I Correspondence Level II corresponds to Level III corresponds to Level IV corresponds to The numerical risk levels increase from 1 to 4, with higher values indicating higher health risks. The natural order of the values corresponds to the progressive logic of clinical risk stratification. The risk level values are obtained by mapping the labels. Subsequently, a single discrete risk level is insufficient to fully characterize the risk distribution features of users across various core physiological dimensions. Different users at the same risk level may exhibit drastically different risk profiles across different dimensions such as the liver, kidneys, or cardiovascular system. Therefore, it is necessary to review all the data stored during the dimensional co-distribution matrix feature decomposition stage. eigenvalues Extract the top-ranked components that have been filtered by the energy retention threshold. Each principal component is calculated using its eigenvalues and the variance contribution rate is used as its dimensional weight. Variance contribution rate of each principal component Calculate using the following formula: ;
[0145] In the formula, For the first The eigenvalues corresponding to each principal component represent the magnitude of the projection variance of the original data along the direction of that principal component; the denominator For all The sum of the eigenvalues is the total variance of the original multidimensional risk cumulative matrix; For the first The proportion of information independently carried by each principal component to the total original information, with values ranging from 0 to 1. The larger the value, the greater the share of independent variation captured by the principal component in the overall fluctuation of multi-organ risk data, and the stronger the discriminative power of the corresponding physiological state dimension in distinguishing different health levels. After obtaining the variance contribution rate weights of each dimension, As the discriminative importance weights of each principal component dimension, and the core health status feature vector at the current time segment. Together they constitute the underwriting assessment parameter set. Indicates the first The proportion of information of each principal component in all fluctuation directions is used to characterize the discriminative importance of this dimension in distinguishing different health levels; For this time section at the 1st The deviation scores in each principal component direction are output as independent evaluation dimensions, without multiplication, in order to maintain the clarity of the physical meaning of each dimension.
[0146] Extracting user age from the initial medical indicator sequence, which has been acquired and continuously transmitted throughout the entire process. Alpha-fetoprotein levels and carcinoembryonic antigen (CEA) levels These three parameters were not involved in dimensionality compression in the preceding semantic parsing, fluctuation tracking, feature fusion, and principal component decoupling steps. They were retained as independent, uncompressed risk factors. These parameters were organized into an underwriting assessment parameter set, the internal structure of which is as follows:
[0147] Risk level value ; Group dimension parameters, each group contains dimension scores. Weighted by variance contribution rate User age Alpha-fetoprotein (AFP) levels Carcinoembryonic antigen (CEA) level The underwriting assessment parameter set takes the risk level of this time segment as the core anchor point, each weighted physiological dimension as the risk composition details, and independent risk factors as the supplementary verification basis, providing a complete input that can be verified item by item for the next step of progressive stratification.
[0148] Based on the underwriting assessment parameter set, and comparing it with the preset health condition admission threshold range, users whose disease risk level values fall into the low-risk range are classified into the first underwriting tier, and users whose disease risk level values fall into the medium-high-risk range are classified into the second underwriting tier, generating an underwriting tier result set, specifically including:
[0149] The underwriting assessment parameter set already carries complete assessment information for the current time segment, mapping the initial health rating label to a risk level value. Simultaneously, it binds the weighted risk contribution of each dimension with independent risk factors, extracting risk level values from the underwriting assessment parameter set. The preset health status threshold range is:
[0150] low-risk area Medium- and high-risk areas The basis for dividing the interval is:
[0151] The risk level probability distribution output by the tree-like risk grading network described above reflects the comprehensive risk quantification assessment of multidimensional physiological characteristics under the current model. Level I, corresponding to the low-risk range, indicates that the user's core physiological characteristics are most closely similar to the statistical pattern of low-risk individuals. Level II and above correspond to the medium-to-high-risk range, indicating that at least in some key dimensions, the user has deviated from the low-risk pattern. This indicates that the risk level has exceeded the low-risk range, and the user is directly directed to the second-level judgment process for review.
[0152] like The risk level falls into the low-risk range, but a single risk level label is a scalar compression result of multi-dimensional features by the network, and there is a possible scenario:
[0153] The user performed normally across most dimensions, resulting in an overall risk level of [missing information]. However, outliers exceeding the normal statistical variation range appeared in certain dimensions. To exclude such cases where the risk level is inconsistent with the dimensional state, [further details are needed]. Outlier detection is performed on each of the core physiological state dimensions. Each dimension, from the score matrix Extract that dimension from all Score sequence at each time section Calculate the standard deviation of the sequence. It describes the normal fluctuation range of the score of this dimension across all observation periods.
[0154] For the Each core physiological state dimension, from the score matrix Extract that dimension from all Score sequence at each time section The 95th and 5th percentiles are statistical values calculated based on the score distribution of all samples. These values are fixed in the model as warning limits. For outlier detection of new users, these fixed thresholds are used directly for judgment without recalculation. The 95th percentile of the sequence is then calculated as the upper limit of the warning. The 5th percentile as the lower limit of the warning : ; ;
[0155] The percentiles are calculated using a weighted average method, and the score sequence is arranged in ascending order. The percentile position is: ;
[0156] If the position is not an integer, then the weighted average of the scores of the two adjacent positions is taken. When the confidence interval of the percentile is estimated using the Bootstrap method (1000 resampling times), the median value estimated by Bootstrap is taken as the warning limit. When the sample size is too small, both percentile estimation and the Bootstrap method are unstable. In this case, outlier detection is not performed, and the dimensional warning status of this time segment is marked as all normal by default to correct the bias in percentile estimation under small sample conditions. For each dimension of the current time segment... Compare its current score With warning upper limit and the lower limit of the warning The size relationship, if If the dimension is within the normal range, it is marked as normal; if If the dimension exceeds the warning limit, it is marked as a positive anomaly; if If the value of this dimension is below the warning threshold, it is marked as a negative anomaly. If all... If all dimensions are within the normal range, the dimension warning status is marked as all normal; if any dimension is positively or negatively abnormal, that dimension is marked as abnormal, and the abnormal dimension number and deviation direction are recorded. The above-mentioned quantile-based warning limit setting method does not rely on the normal distribution assumption, is applicable to data with any distribution shape, and is based on the distribution of all samples rather than the individual's own historical data, which can effectively identify the deviation of an individual from the reference range of the healthy population. The result of the first-level judgment determines the path to the second-level judgment, and all three paths lead to the entry point of the second-level judgment:
[0157] Path 1, That is, the risk level itself is medium to high risk; Path two, However, the dimension warning status is marked as abnormal; Path 3, Furthermore, since the dimensional warning status is all normal, users in this path are close to meeting the conditions for the first underwriting tier, but still need to undergo a review of the high-risk supplementary factors in the second-level determination for final confirmation. The second-level determination involves extracting the user's age from the underwriting assessment parameter set. Alpha-fetoprotein levels and carcinoembryonic antigen (CEA) levels Each of the three factors is then compared to a preset supplementary risk threshold. Age supplementary threshold Values Age is a baseline risk factor independent of current health checkup indicators. Even if current indicators are normal, older individuals still have a significantly higher risk of declining organ reserve and degenerative diseases than younger individuals. Therefore, a separate age-related supplementary factor is necessary. High-risk age triggers supplemental factors, denoted as ;otherwise Upper reference limit for alpha-fetoprotein (AFP) Values Alpha-fetoprotein (AFP) at ng / mL is a serum marker for hepatocellular carcinoma and germ cell tumors. Elevated AFP levels are closely associated with the risk of malignant liver disease. Alpha-fetoprotein (AFP) high-risk factor supplementation trigger, denoted as ;otherwise Upper reference limit for carcinoembryonic antigen (CEA) Values Carcinoembryonic antigen (CEA) at ng / mL is a broad-spectrum tumor marker for colorectal cancer and other adenocarcinomas. Elevated levels suggest the need for further investigation to rule out the risk of epithelial-derived malignancies. High-risk supplementation factor for carcinoembryonic antigen (CEA) triggers, denoted as ;otherwise The comprehensive assessment results of the three high-risk supplementary factors Logical OR of the three:
[0158] if only , , Any one of them is , That is This indicates that at least one independent supplementary risk factor has been triggered; only when all three are... hour, for The results of the first-level and second-level assessments are combined to form the final classification rules. The admission criteria for the first underwriting tier are that all three conditions must be met simultaneously:
[0159] Risk level value The dimensional warning status is all normal; and the comprehensive supplementary factors are also in place. The triple conditions ensure that users entering the first underwriting tier have no abnormal signals in terms of comprehensive risk level, status of each core physiological dimension, and independent risk factors. Their health status is cross-validated from multiple angles, all supporting a low-risk conclusion. The second underwriting tier covers all other situations that do not meet the conditions of the first underwriting tier, including three typical situations: the risk level itself is medium to high risk ( Users who meet the following criteria will be categorized into underwriting tiers and risk levels: those with a low risk level but exhibiting outlier fluctuations in dimensionality; and those triggered by any high-risk supplementary factor. Detailed warning status for each dimension, and three independent supplementary factor marker values. and comprehensive supplementary factors It is integrated into a single underwriting tiered record.
[0160] Based on the underwriting stratification results set, the first coverage parameter group is configured for the first underwriting stratum, and the second coverage parameter group is configured for the second underwriting stratum, generating a differentiated coverage liability mapping table, specifically including:
[0161] Based on the underwriting stratification results set, each record clearly indicates the user's underwriting stratum at that time point. The first underwriting stratum corresponds to low-risk users who meet all three criteria, while the second underwriting stratum corresponds to users with any risk signal. The two strata have fundamentally different risk profiles, therefore, differentiated protection liability parameter groups need to be matched to ensure that the allocation of protection resources corresponds to the user's health risk status. Based on this logic, two sets of protection liability parameter groups with different coverage characteristics are predefined. The first protection parameter group corresponds to the first underwriting stratum and contains four protection liability parameters, each consisting of two fields: activation status and coverage type.
[0162] parameter This indicates that the enabled status is "on" and the coverage type is "inpatient care"; parameters This indicates that the enabled status is "on" and the coverage type is "outpatient medical care"; parameters This indicates that the activation status is enabled, and the coverage type is intensive care medical; parameters This indicates that the activation status is "On", the coverage type is "Annual Health Check-up", and the second protection parameter group corresponds to the second underwriting level. It contains four protection liability parameters, each of which is also composed of the activation status and coverage type:
[0163] parameter This indicates that the activation status is enabled, and the coverage type is malignant tumor-specific medical care; parameters This indicates that the enabled status is "on" and the coverage type is "inpatient care"; parameters This indicates that the activation status is enabled, and the coverage type is intensive care medical; parameters The activation status is "on" and the coverage type is "outpatient medical care". The difference between the two sets of parameters is reflected in the composition of the coverage type. The first set is characterized by comprehensive coverage of four dimensions: inpatient medical care, outpatient medical care, intensive care medical care, and annual health check-up. It has a wide coverage and is suitable for users who have been confirmed to be in a low-risk state through multi-dimensional cross-validation. The second set, in addition to the three basic coverages of inpatient medical care, intensive care medical care, and outpatient medical care, replaces the fourth item with malignant tumor-specific medical care. Users in the second underwriting tier have at least one risk signal among medium-to-high risk level, dimensional outlier fluctuation, or abnormal tumor markers. Their potential risk in the direction of malignant tumors is higher than that of users in the first tier. Therefore, a targeted tumor-specific protection is added to the protection configuration.
[0164] Match each stratification record in the underwriting stratification result set with its corresponding coverage parameter group. Iterate through the stratification records; if the underwriting stratum of a record is the first underwriting stratum, match it with the first coverage parameter group. The activation status value and coverage type value of each item are copied field by field to the new parameter field of this record; if the underwriting level is the second underwriting level, the second protection parameter group is copied. After copying field by field and matching, each record has eight protection responsibility parameter fields appended to the end of the original stratification determination field. Each of the four protection responsibilities corresponds to two sub-fields: activation status and coverage type. All records that have been matched and expanded are organized by user ID and timestamp to generate a differentiated protection responsibility mapping table. The mapping table uses user ID and timestamp as a joint index. The field structure of each record, from front to back, is the underwriting stratification result field and the protection responsibility parameter field. This structure creates a one-way definite mapping relationship between the underwriting stratification and the protection responsibility parameters.
[0165] Based on the differentiated coverage responsibility mapping table, the user's initial medical indicator sequence and underwriting stratification results are integrated to generate a data tagging system, specifically including:
[0166] Using the user identifier and sampling timestamp as the joint query key, the algorithm iterates backward from the first sampling node of the initial medical indicator sequence. For each sampling node, it retrieves all stratification records in the underwriting stratification result set that have the same user identifier as the current node and whose timestamp is no later than the current sampling timestamp. The record with the largest timestamp value is then selected from the search results set; this record represents the user's latest underwriting stratification determination as of the current sampling time. All fields in this stratification record, including underwriting stratification and risk level values, are then analyzed. Detailed warning status for each dimension and high-risk supplementary factor marker values Each field is attached to the current sampling node in turn. The attachment method is to append the above fields to the end of the existing original indicator fields of the sampling node in sequence, while keeping the original name and data type of each field unchanged.
[0167] For cases where the search results are empty, meaning the user has not generated any underwriting stratification records before the current sampling time, a null value is entered in the corresponding underwriting decision field of the sampling node, indicating that the underwriting assessment at that time has not yet been completed. This situation may occur when the user's first sampling timestamp is earlier than the start time of the first underwriting assessment cycle. After completion, each sampling node in the initial medical indicator sequence is expanded with a set of underwriting decision fields. The objective measurement values of the original indicators and the judgment conclusions of the underwriting decisions coexist in the same data row. The link between the two is the joint key of the user identifier and the sampling timestamp. Using the user identifier in the expanded sampling node as the query key, all records of the user are retrieved in the differentiated protection responsibility mapping table. The record with the latest timestamp is selected, and all contents of the protection responsibility parameter fields in that record are extracted, namely the activation status and coverage type of the four protection responsibilities in the first protection parameter group or the second protection parameter group, a total of eight sub-fields, which are appended to the end of the existing fields of the current sampling node. For the specific values of the protection responsibility parameter fields, their original values when matched in the previous text are retained:
[0168] If the user's latest record in the differentiated protection liability mapping table is associated with the first protection parameter group, then the sampling node obtains... The parameters; if associated with the second protection parameter group, then obtain the parameters. Each sampling node simultaneously carries the original medical indicators, underwriting decision results, and coverage configuration parameters. The original medical indicators are updated with each physical examination sampling, the underwriting stratification results are updated with each operation of the risk grading network, and the coverage configuration parameters are updated with changes in the underwriting stratification level. When attaching, the latest valid value up to the sampling timestamp is taken to ensure that the tag records are consistent in time sequence.
[0169] All mounted sampling nodes are grouped by user identifier, and within each group, they are sorted in ascending order by sampling timestamp. The data is output in the form of a structured data table. Each row in the data tagging system corresponds to a single sampling node of a user, and the fields are divided into three data layers according to their source:
[0170] The raw medical indicator data layer includes sampling timestamps, systolic blood pressure values, diastolic blood pressure values, and various physiological data.
[0171] The data includes chemical index values, lesion volume values for various organs, alpha-fetoprotein (AFP) values, and carcinoembryonic antigen (CEA) values. These are objective measurements directly extracted from the original text of the medical examination report or calculated through the accumulation of voxels in a three-dimensional geometric envelope space. The underwriting decision data layer includes underwriting level and risk grade values. , Projected scores for each core physiological state dimension Weighted by variance contribution rate Warning limits and judgment results for each dimension, and high-risk supplementary factor marker values. It fully records the reasoning chain and quantitative basis from the initial health rating label to the final underwriting level; the protection configuration data layer includes the activation status and coverage type of each of the four protection liabilities, and records the protection plan uniquely bound to the user's underwriting level.
[0172] The three layers of fields are separated by data source and processing stage. Field names are globally unique, and field meanings and value rules are consistent within each layer. Downstream business systems can query the evolution trajectory of a user's tags over time through user identifiers, query the distribution of the health status of all users at a certain moment through sampling timestamps, and filter a subset of users under a specific level through underwriting level fields to trace their original medical indicators and risk assessment basis, thus completing the data tagging system.
[0173] In this embodiment of the invention, by mapping the initial health rating label to a numerical risk level and combining it with the variance contribution rate weight of the core physiological dimension to construct an underwriting assessment parameter set, the technical problems of underwriting stratification relying on a single risk level that cannot reflect dimensional heterogeneity, the separation of health rating and coverage configuration, and the lack of time-series traceability of the label system are overcome. This achieves the technical effects of multi-dimensional cross-validation of underwriting stratification, accurate matching of coverage liability and risk status, and time-series traceability and stratified screening of label data.
[0174] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0175] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0176] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. An automated data tagging system construction system, characterized in that, include: The acquisition module is used to acquire physical examination reports and locate biochemical and imaging segments, extract time-series blood pressure, biochemical values, lesion volume, age and tumor markers, and generate an initial medical indicator sequence containing sampling nodes by sorting by timestamp; The parsing module is used to parse semantic associations and identify organ attributes and suspicious occupants based on the initial medical indicator sequence, and generate a structured entity relationship graph. The calculation module is used to extract time-series blood pressure based on the structured entity relationship graph, measure the rate of change of blood pressure between adjacent sampling nodes over time intervals, and generate a feature vector of index fluctuation trend. The extraction module is used to extract rate differences and direction identifiers based on the indicator fluctuation trend feature vector, filter sampling nodes with rising, falling and reversing characteristics and extract time sequence position identifiers to generate local fluctuation difference sequences. The comparison module is used to compare the frequency of biochemical abnormalities with the medical reference range based on the local fluctuation difference sequence and the structured entity relationship map, summarize and calculate the total abnormal load and lesion volume of each organ across the entire domain, perform feature splicing, and generate a multidimensional risk accumulation matrix. The transmission module is used to perform attribute decoupling operations to remove redundant dimensions based on the multidimensional risk accumulation matrix and generate a core health status feature set. Based on the core health status feature set, combined with the continuously transmitted age and tumor markers, the module is input into a tree-like risk grading network to extract the risk level corresponding to the maximum probability value and generate an initial health rating label. The segmentation module is used to divide the first and second underwriting tiers based on the initial health rating labels and the underwriting tier thresholds, associate the corresponding set of coverage liability configuration parameters, and output the data label system.
2. The automated data tagging system construction system according to claim 1, characterized in that, Obtain the physical examination report and locate the biochemical and imaging segments. Extract time-series blood pressure, biochemical values, lesion volume, age, and tumor markers. Sort by timestamp to generate an initial medical indicator sequence containing sampling nodes, including: Obtain the original text of the physical examination report and locate the biochemistry paragraph from the laboratory department and the description paragraph from the radiology department to generate a multi-source text paragraph set; Based on a collection of multi-source text paragraphs, extract blood pressure measurement values and their associated date and time strings, and arrange them in chronological order to generate a time-series blood pressure index group; Based on a multi-source text paragraph set, identify the biochemical item name and its corresponding detection value, and generate a biochemical indicator value group. Based on a multi-source text paragraph set, extract age and tumor marker values to generate user age and tumor marker value sets; Based on a multi-source text segment set, extract the diameter values and the names of the organs and sites to which the lesions belong from the lesion description segments, calculate the lesion volume and store them in pairs, and generate image lesion volume groups. Based on the time-series blood pressure index group, biochemical index value group, user age value, tumor marker value group, and imaging lesion volume group, the data are aggregated using date and time strings as the primary key to generate sampling nodes, and the initial medical index sequence is generated by arranging them in ascending order of time.
3. The automated data tagging system construction system according to claim 2, characterized in that, Based on a multi-source text segment set, the diameter values and associated organ names are extracted from the lesion description segments. The lesion volume is calculated and paired for storage, generating an image lesion volume group, including: Based on the image text subset in the multi-source text paragraph set, the lesion description fragments are scanned sentence by sentence and the number of diameter values contained therein is counted to generate diameter dimension determination results. Based on the results of the radial dimension determination, when there are at least two radial values, the largest radial value is extracted as the major axis parameter and the smallest radial value is extracted as the minor axis parameter. The arithmetic mean of the major axis parameter and the minor axis parameter is calculated as the vertical axis parameter. A basic three-dimensional geometric envelope is constructed based on the major axis parameter, the minor axis parameter and the vertical axis parameter. From the image text subset, extract the morphological modification phrases adjacent to the lesion description fragments, and map the morphological modification phrases to surface morphology compensation factors; Based on the surface morphology compensation factor, local spatial deformation compensation is performed on the surface of the basic three-dimensional geometric envelope to generate a compensated three-dimensional lesion envelope. Spatial voxel accumulation calculation is then performed on the compensated three-dimensional lesion envelope to generate the first lesion volume. Based on the results of the radial dimension determination, when the radial value is a single value, the single radial value is extracted as the diameter parameter. A basic spherical geometric envelope is constructed based on the diameter parameter, and spatial voxel accumulation calculation is performed to generate the second lesion volume. Based on the volume of the first or second lesion, and combined with the name of the attributing organ extracted backward from the sentence containing the lesion description fragment, a pairing and binding is performed to generate an image lesion volume group.
4. The automated data tagging system construction system according to claim 3, characterized in that, Based on the initial medical indicator sequence, semantic associations are analyzed and organ attributes and suspected lesions are identified, generating a structured entity relationship graph, including: Organ site names are extracted from the image lesion volume group of the initial medical indicator sequence to create entity nodes, and lesion modification phrases adjacent to each organ site are extracted from the image text subset of the multi-source text paragraph set and bound to the corresponding entity nodes to generate entity-modification attribute mapping. The values of biochemical indicators in the initial medical indicator sequence are compared with the preset medical reference interval to determine the abnormality. The associated organ parts are recorded according to the indicator-organ attribution relationship to generate an entity-abnormal association mapping. Entity nodes that have both modified records and abnormality pointing records are marked as carrying suspicious placeholder attributes to generate a basic entity node set. Based on each entity node in the basic entity node set, the corresponding biochemical index values, lesion volume, blood pressure index and tumor marker values in the initial medical index sequence are written to the timestamp, generating an entity node set with quantitative attributes. According to the organ hierarchical classification rules, entity nodes with quantitative attributes are assigned to the superior classification nodes and hierarchical classification relationship edges are established. For two entity nodes under the same organ location that carry suspicious occupancy attributes and have biochemical abnormality, cross-index association relationship edges are established. A structured entity relationship graph is constructed using all entity nodes as graph nodes and hierarchical affiliation edges and cross-index association edges as graph edges.
5. The automated data tagging system construction system according to claim 4, characterized in that, Based on the structured entity relationship graph, time-series blood pressure is extracted, and the rate of change of blood pressure between adjacent sampling nodes over time is calculated point by point to generate a feature vector of index fluctuation trends, including: Based on the structured entity relationship graph, extract the time-series blood pressure indicators bound to the timestamps and arrange them in ascending order of sampling time to generate a set of blood pressure time-series nodes; Based on the blood pressure time series node set, extract the blood pressure value difference and time interval length between adjacent sampling nodes to generate a node difference parameter group; Based on the node difference parameter group, the rate of change of blood pressure values between adjacent sampling nodes with the length of time interval is calculated point by point to generate an interval change rate sequence. Based on the interval change rate sequence, the direction of rate change and the rate amplitude at each sampling node are extracted to generate a feature vector of index fluctuation trend.
6. The automated data tagging system construction system according to claim 5, characterized in that, Based on the indicator fluctuation trend feature vector, the rate difference and direction identifier are extracted, sampling nodes with rising, falling, and reversing characteristics are selected and their time-series position identifiers are extracted to generate a local fluctuation difference sequence, including: Based on the indicator fluctuation trend feature vector, extract the interval change rate value and rate change direction identifier corresponding to each continuous sampling node to generate a node rate feature set. Based on the node rate feature set, the rate change between adjacent sampling nodes is calculated point by point and the direction of rise and fall is identified to generate a node rate change parameter set. Based on the reversal of the rising and falling directions in adjacent intervals of the node rate change parameter set, sampling nodes where the blood pressure fluctuation trend reverses are selected, and an initial state turning point node set is generated. Based on the initial state transition node set, extract the sampling timestamp and rate change of each transition node in the blood pressure time series node set to generate a local fluctuation difference sequence.
7. The automated data tagging system construction system according to claim 6, characterized in that, Based on the local fluctuation difference sequence and structured entity relationship map, the frequency of biochemical abnormalities is calculated by comparing with medical reference ranges. The total abnormal load and lesion volume of each organ are calculated across the entire domain. Feature splicing is then performed to generate a multidimensional risk accumulation matrix, including: Based on the structured entity relationship graph, the values of each biochemical indicator are extracted and compared with the preset medical reference interval to calculate the number of times the indicator exceeds the normal range, and the abnormal frequency of biochemical indicators is generated. Based on the frequency of abnormal biochemical indicators and the volume of lesions described by images, the total abnormal load of each organ is calculated and a global organ load feature vector is generated. Based on the local fluctuation difference sequence, the coefficient of variation of blood pressure values in each fluctuation segment is calculated and the cumulative mean is calculated to generate a cumulative blood pressure variation sequence. Based on the global organ load feature vector and the cumulative blood pressure variation sequence, a multidimensional attribute weighted splicing operation is performed to reorganize the physiological state dimension, resulting in a multidimensional risk accumulation matrix.
8. The automated data tagging system construction system according to claim 7, characterized in that, Based on the multidimensional risk accumulation matrix, attribute decoupling operations are performed to remove redundant dimensions, generating a core health status feature set, including: Based on the multidimensional risk accumulation matrix, the central reference value of each physiological risk dimension is calculated column by column across all sampling time sections to generate a baseline zero-return fluctuation matrix; based on the baseline zero-return fluctuation matrix, the degree of cooperative fluctuation between each pair of physiological risk dimensions is calculated to generate a dimension cooperative distribution matrix. Based on the dimensional cooperative distribution matrix, the independent fluctuation direction and the proportion of independent information it carries are calculated. The independent fluctuation direction sequence is generated by arranging them in descending order of the proportion of independent information. Based on the independent fluctuation direction sequence, starting from the first item, the proportion of independent information in each direction is accumulated until the accumulated value reaches the preset energy retention threshold. The independent fluctuation directions within the accumulation range are then extracted as the core feature direction set. Based on the core feature direction set and the baseline zero-return fluctuation matrix, the baseline zero-return fluctuation vector of each sampling time section is projected onto each core feature direction one by one, and the projection values of each time section on each core feature direction constitute the core health status feature set.
9. The automated data tagging system construction system according to claim 8, characterized in that, Based on the core health status feature set, combined with continuously transmitted age and tumor markers, the risk level corresponding to the highest probability value is extracted by the tree-structured risk grading network, generating an initial health rating label, including: Based on the core health status feature set, the projection scores of each independent core feature dimension are extracted and horizontally concatenated with the continuously transmitted user age value and tumor marker detection value to generate a multi-dimensional classification input vector. Based on the multidimensional classification input vector, it is input into a risk classification network composed of multiple decision trees cascaded in sequence to generate a comprehensive risk score. Based on the comprehensive risk score, the weight coefficients and bias terms corresponding to each preset risk level are combined to perform exponential calculation and normalization to generate the probability distribution of each preset risk level. The risk level number with the highest probability value in the probability distribution is extracted to generate the initial health rating label.
10. The automated data tagging system construction system according to claim 9, characterized in that, Based on the initial health rating labels, the first and second underwriting tiers are divided according to the underwriting stratification thresholds. The corresponding coverage configuration parameter sets are then associated, and a data labeling system is output, including: Based on the initial health rating labels, the included disease risk level values and core physiological state dimension weights are analyzed to generate an underwriting assessment parameter set; Based on the underwriting assessment parameter set, and compared with the preset health condition admission threshold range, users whose disease risk level values fall into the low-risk range are classified into the first underwriting tier, and users whose disease risk level values fall into the medium-high-risk range are classified into the second underwriting tier, thus generating an underwriting tier result set. Based on the underwriting stratification result set, configure the first protection parameter group for the first underwriting stratum and the second protection parameter group for the second underwriting stratum, and generate a differentiated protection liability mapping table; Based on the differentiated coverage responsibility mapping table, the user's initial medical indicator sequence and underwriting stratification result set are integrated to generate a data tagging system.