Pancreatic cancer prognosis analysis system based on big data mining

By constructing the rate of change difference and third-order response sequence, extracting abrupt trend feature points, and combining grade span and direction determination, stable samples are screened and the instability of the offset direction is quantified. This solves the problem of insufficient accuracy of traditional pancreatic cancer prognostic analysis systems in dynamic state identification, and realizes accurate classification and individualized intervention strategies for high-risk variant stages.

CN120895259AInactive Publication Date: 2025-11-04AFFILIATED HOSPITAL OF NANTONG UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510922384.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-11-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional pancreatic cancer prognostic analysis systems lack structured expression of the magnitude and direction of indicator transitions when dealing with dramatic temporal fluctuations or abrupt changes in the rate of indicator change. This results in insufficient accuracy in identifying dynamic states and susceptibility to interference from outliers or low-quality samples, affecting the clinical guidance value of the prediction system.

Method used

The pancreatic cancer prognostic analysis system based on big data mining constructs a rate of change difference and a third-order response sequence through a jump trend extraction module, an indicator level mapping module, a sample stability screening module, and a path node offset identification module. It extracts jump trend feature points, combines level span and direction determination, screens stable sample sets, quantifies the instability of offset direction, and achieves accurate classification and labeling.

Benefits of technology

It improves the accuracy of classifying high-risk variant stages in the prognostic pathway for pancreatic cancer patients, enhances the sensitivity of risk identification and the scientific basis of stratified intervention, provides a reference for individualized intervention strategies, and strengthens the clinical guidance value of the prediction system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120895259A_ABST
    Figure CN120895259A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of prognosis analysis, in particular to a pancreatic cancer prognosis analysis system based on big data mining, which comprises a jump trend extraction module, an index grade mapping module, a sample stability screening module, a path node offset identification module and a prognosis stage classification module. According to the method, a change rate difference value and a third-order response sequence are constructed for continuous three-stage pancreatic cancer biological index sequences, jump trend feature points are extracted, and a risk grade transition sequence is formed in combination with grade span and direction judgment; multi-dimensional normalized parameters of a variable coefficient, a survival score difference value and an organ function score standard deviation are introduced to perform stable sample screening, interference of data disturbance on a path judgment result is reduced, and path mutation node positions are screened in combination with a path stability threshold value; precise classification and marking of high-risk variation stages in prognosis paths of pancreatic cancer patients are realized, stage reference is provided for individualized intervention strategies, and risk identification sensitivity and layered intervention scientificity are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of prognosis analysis, in particular to a pancreatic cancer prognosis analysis system based on big data mining. BACKGROUND

[0002] The technical field of prognosis analysis involves modeling and predicting outcomes such as disease progression, treatment response, and survival time through multidimensional medical data, mainly applying statistical modeling, machine learning, and clinical data mining techniques to quantitatively evaluate the future health status of patients. The technical foundation of this field includes but is not limited to survival analysis models, risk score system construction, recursive prediction algorithms based on time series data, and multi-source medical data fusion analysis methods. Prognosis analysis systems are widely used in oncology, chronic disease management, postoperative rehabilitation evaluation, and other medical scenarios, aiming to optimize treatment options, improve disease management, and support clinical individualized decision-making.

[0003] Among them, the pancreatic cancer prognosis analysis system is an intelligent decision support system for quantitatively predicting the survival risk, recurrence probability, and treatment response trend of pancreatic cancer patients. The system collects multiple basic data of patients, including pathological stage, tumor marker expression, postoperative complication situation, treatment method, etc., and combines survival analysis to realize individualized prognosis evaluation of pancreatic cancer. Its use is to assist oncologists in developing differentiated follow-up cycles, determining whether adjuvant therapy intervention is appropriate, and providing scientific and quantitative basis for patient communication of expected disease course, thereby improving clinical management efficiency and patient quality of life.

[0004] Traditional analysis systems are based on single indicator trend modeling or static time point parameters as the basis for evaluation. In the case of dramatic fluctuations in time series or sudden changes in indicator rates, there is a lack of structured expression of indicator transition amplitude and direction, resulting in insufficient accuracy in dynamic state recognition. For example, when a patient's CA19-9 suddenly rises in the short term but there is no matching survival score update, the traditional system cannot recognize the impact of this mutation on the prognosis stage. In addition, the lack of stability evaluation mechanism based on sample volatility indicators makes it easy for data outliers or low-quality samples to interfere with the overall model training results, leading to fuzzy prognosis stage classification and inaccurate subsequent treatment strategy development, affecting the clinical guidance value of the prediction system. SUMMARY

[0005] The purpose of the present application is to solve the shortcomings in the prior art, and a pancreatic cancer prognosis analysis system based on big data mining is proposed.

[0006] To achieve the said purpose, the present application adopts the following technical solution: a pancreatic cancer prognosis analysis system based on big data mining, the system comprising:

[0007] The sudden trend extraction module obtains three continuous detection records of a pancreatic cancer patient, calculates the change rate of the two periods before and after respectively, constructs a change rate difference sequence, performs difference calculation between terms based on the sequence, extracts the point with the maximum amplitude and the corresponding detection cycle position from the third-order response sequence as the abnormal trend point, marks the sudden change characteristic type according to the change rate direction, and generates a sudden trend mark set;

[0008] The index grade mapping module classifies and matches each index value with the grade interval in the pancreatic cancer risk assessment reference table based on the sudden trend mark set, counts the number of grade span of each sudden point, forms a structured index grade transition item in combination with the sudden direction, and generates a risk grade transition sequence;

[0009] The sample stability screening module calculates the confidence weight grade according to the risk grade transition sequence, judges the score difference between adjacent two samples, identifies the sample paragraph with a continuous difference greater than the screening threshold, and obtains a stable sample set after screening;

[0010] The path node offset identification module calls the stable sample set after screening, calculates the sequence instability value according to the prognosis prediction time sequence, screens the offset direction mutation node position, and obtains a path offset node index table by summarizing to the path structure.

[0011] The sudden trend mark set specifically includes a sudden detection cycle position, a third-order mutation response value and a mutation direction type, the risk grade transition sequence includes a risk grade span value, an index corresponding grade number and a grade change direction, the stable sample set after screening specifically includes a reserved sample number, a sample confidence grade value and a removal paragraph start and end position, and the path offset node index table includes a node index number, a sliding offset direction value and a mutation trigger flag.

[0012] The sudden trend extraction module includes:

[0013] The change rate calculation submodule obtains three continuous detection records of a pancreatic cancer patient, extracts the original value sequence of three biological indexes CA19-9, CEA and CRP, calculates the change rate of each index between the first period and the second period and between the second period and the third period, obtains the change rate sequence, and then performs difference calculation between terms on adjacent two items in each group of sequences to obtain the change rate difference sequence of each index, and establishes change rate difference change trend information;

[0014] The acceleration sequence generation submodule calls the change rate difference sequence of each index based on the change rate difference change trend information, performs difference calculation between terms on the sequence, constructs a third-order response sequence reflecting the acceleration degree of the trend, extracts a group of third-order responses with the maximum value and the corresponding detection cycle position in each index, and generates a sudden acceleration change parameter group;

[0015] The abnormal trend identification submodule judges the positive and negative signs of the acceleration change directions of three indexes according to the group of sudden change acceleration change parameters, combines the position of the maximum value to mark the corresponding detection period, forms a sudden change feature combination of the acceleration direction and the position, is classified into a sudden change trend category, and establishes a sudden change trend marker set.

[0016] The index level mapping module comprises:

[0017] The index classification submodule calls the CA19-9, CEA and CRP values in the detection period corresponding to each sudden change point based on the sudden change trend marker set, matches the values of the three indexes with the level intervals in the pancreatic cancer risk assessment reference table respectively, labels the level numbers corresponding to the indexes according to the matching results, and generates level number attribution information;

[0018] The level span calculation submodule obtains the number difference between the current period level and the previous period level of each index according to the level number attribution information, counts the level change level numbers of each of the three indexes, marks the change direction as rising or falling, generates a multi-index change summary item in combination with the level number and the change direction, and obtains a group of level change level parameter.

[0019] The level transition generation submodule calls the group of level change level parameters, judges the consistency of the level change direction of each index and the sudden change direction, selects the index items with consistent directions, integrates the level span and the change direction to construct a level change structure combination in the sudden change period, and establishes a risk level transition sequence.

[0020] The sample stability screening module comprises:

[0021] The confidence parameter normalization submodule obtains the CA19-9, CEA and CRP detection values of each sample in the corresponding sample set according to the risk level transition sequence, calculates the coefficient of variation, the difference value of the survival function score, and the standard deviation of the organ function score, respectively groups the indexes for range standard normalization processing, obtains three normalized parameters, and establishes confidence factor information.

[0022] The weight score calculation submodule extracts the normalized value of the coefficient of variation, the normalized value of the difference value of the survival function score, and the normalized value of the standard deviation of the organ function score in each sample based on the confidence factor information, combines the total amount of the responses of the three indexes and the change rate of the organ function score of each sample in the detection period, calculates the sample confidence weight level, arranges the samples in descending order of the confidence weight level value, and generates a sample confidence level sequence.

[0023] The abnormal paragraph elimination submodule calls the sample confidence level sequence, calculates the absolute value of the score difference between each two adjacent samples after sorting, marks the paragraph with three consecutive score difference values greater than the screening threshold, eliminates the samples corresponding to the paragraph, and establishes a stable sample set after screening.

[0024] The path node offset identification module comprises:

[0025] The difference extraction submodule calls the stable sample set after screening, extracts the CA19-9, CEA and CRP combined values of three consecutive nodes according to the prognosis prediction time sequence of each patient, calculates the absolute difference sequence of the combined values of two adjacent nodes, generates a Manhattan distance difference sequence in time sequence, and establishes path difference sequence information;

[0026] The instability degree value calculation submodule constructs a sliding window structure according to each difference sequence based on the path difference sequence information, extracts the number of sign changes of the Manhattan distance difference in each window, calculates the offset direction instability degree value of the node, generates a node instability degree value sequence, and establishes a path offset node index table according to the node instability degree value sequence.

[0027] The offset node marking submodule compares the node score with the pre-set path stability threshold according to the node instability degree value sequence, marks the node as an offset trigger node when the score exceeds the threshold interval, and collects the marks according to the original sequence index position to establish a path offset node index table.

[0028] The system further comprises:

[0029] The prognosis stage classification module obtains each biological index combination value and organ function score under the corresponding node according to the path offset node index table, constructs a feature vector, calls a stage label vector set, performs Euclidean distance calculation between the feature vector and each class label center point, attributes to the class item with the smallest distance, and generates a prognosis stage classification list.

[0030] The prognosis stage classification list specifically refers to the classification stage number, stage label name and classified node position.

[0031] The prognosis stage classification module comprises:

[0032] The feature construction submodule obtains the CA19-9, CEA and CRP values and organ function score values corresponding to each node according to the path offset node index table, arranges the values in time sequence, constructs a four-dimensional index feature vector of each node, and establishes a node index feature set.

[0033] The distance calculation sub-module calls a standard stage label center point vector set based on the node index feature set, calculates the Euclidean distance between each node feature vector and each label vector, obtains the minimum weighted distance value of each node corresponding stage label, and obtains a label distance minimum value sequence;

[0034] The stage labeling sub-module selects the label category item with the minimum distance of each node as the belonging stage identifier according to the label distance minimum value sequence, maps and integrates each node and the corresponding category into structured labeling information, and establishes a prognosis stage classification list.

[0035] Compared with the prior art, the advantages and positive effects of the present application are that:

[0036] In the present application, by constructing the change rate difference and the third-order response sequence of the continuous three-stage pancreatic cancer biological index sequence, extracting the sudden trend feature point, combining the grade span and direction determination, forming the risk grade transition sequence, introducing the multi-dimensional normalization parameters of the coefficient of variation, the survival score difference and the organ function score standard deviation for stable sample screening, making the subsequent analysis based on the high stability sample set, reducing the interference of data disturbance on the path determination result, using the Manhattan distance difference and the sliding sign variation frequency to quantify the direction instability, combining the path stability threshold to screen the path mutation node position, determining the minimum distance stage belonging of the index combination by the Euclidean distance, the accurate classification and marking of the high-risk variation stage in the prognosis path of the pancreatic cancer patient are realized, the staging reference for individualized intervention strategy is provided, and the scientificity of risk identification sensitivity and stratified intervention is improved. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 The system flowchart of the present application;

[0038] Figure 2 The flowchart of the sudden trend extraction module of the present application;

[0039] Figure 3 The flowchart of the index grade mapping module of the present application;

[0040] Figure 4 The flowchart of the sample stability screening module of the present application;

[0041] Figure 5 The flowchart of the path node offset identification module of the present application;

[0042] Figure 6 The flowchart of the prognosis stage classification module of the present application. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and not to limit the present application.

[0044] In the description of the present application, it should be understood that the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, in the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise explicitly and specifically limited.

[0045] Please refer to Figure 1 The present application provides a technical scheme: a pancreatic cancer prognosis analysis system based on big data mining, the system comprising:

[0046] The sudden change trend extraction module obtains the continuous three-period detection records of the pancreatic cancer patient, extracts the original numerical value sequence of the three biological indicators CA19-9, CEA and CRP, calculates the change rate of the two periods before and after respectively, calculates the difference between the first-order items of the change rate sequence, constructs the change rate difference sequence, calculates the difference between the items based on the sequence, obtains the third-order response sequence reflecting the acceleration change, extracts the amplitude maximum point and the corresponding detection cycle position as the abnormal trend point according to the third-order response sequence, and generates a sudden change trend label set by marking the sudden change characteristics type in combination with the change rate direction with the abnormal trend point as a reference;

[0047] The index grade mapping module, based on the sudden change trend label set, calls the three original values of CA19-9, CEA and CRP under the corresponding detection cycle, classifies and matches each index value with the grade interval in the pancreatic cancer risk assessment reference table, calculates the grade span of the sudden change point and the previous cycle grade, counts the grade span layer of each sudden change point, forms a structured index grade transition item in combination with the sudden change direction, and generates a risk grade transition sequence;

[0048] The sample stability screening module, according to the risk grade transition sequence, obtains the tumor marker variation coefficient, the postoperative survival score difference, and the three-cycle organ function score standard deviation before and after treatment of each sample in the corresponding sample set, respectively normalizes the three parameters, calculates the confidence weight grade, sorts each sample according to the score, judges the score difference between the adjacent two samples, identifies the sample paragraph with a continuous difference greater than the screening threshold, eliminates the corresponding segment sample, and obtains a stable sample set after screening;

[0049] The path node offset identification module calls the stable sample set after screening, extracts the biological index combination difference value of the continuous three nodes in the patient path according to the prognosis prediction time sequence, calculates the sliding offset direction of the combination difference value sequence, obtains the Manhattan distance difference value sequence of each sequence, calculates the instability degree value, compares the instability degree value with the path stability threshold value extracted in the training stage of the pancreatic cancer prognosis path model, screens the mutation node position of the offset direction, and summarizes and marks to the path structure to obtain the path offset node index table;

[0050] The prognosis stage classification module obtains each biological index combination value and organ function score under the corresponding node according to the path offset node index table, constructs a feature vector, and calls a stage label vector set to perform Euclidean distance calculation between the feature vector and each class label center point, and belongs to the class item with the smallest distance, summarizes the belonging stage class of each node, and generates a prognosis stage classification list;

[0051] The sudden jump trend marker set specifically includes the sudden jump detection period position, the third-order mutation response value and the mutation direction type, the risk level transition sequence includes the risk level span value, the index corresponding level number and the level change direction, the stable sample set after screening specifically includes the reserved sample number, the sample confidence level value and the removed paragraph start and end position, the path offset node index table includes the node index number, the sliding offset direction value and the mutation trigger flag, and the prognosis stage classification list specifically refers to the classification stage number, the stage label name and the classification node position.

[0052] Please refer to Figure 2 The sudden jump trend extraction module includes:

[0053] The change rate calculation submodule obtains the continuous three period detection records of the pancreatic cancer patient, extracts the original value sequence of the three biological indexes CA19-9, CEA and CRP, calculates the change rate of each index between the first and second periods and between the second and third periods, obtains the change rate sequence, and then calculates the difference value between each adjacent item in each sequence to obtain the change rate difference value sequence of each index, and establishes the change trend information of the change rate difference value;

[0054] The continuous three period detection records of the pancreatic cancer patient are obtained, and the original detection values of the three biological indexes CA19-9, CEA and CRP in the three monitoring periods T1, T2 and T3 with a time interval of 10 days are extracted, the relative difference of the numerical difference between T1 and T2 and T2 and T3 is calculated to obtain the change rate sequence, for example, the detection values of CA19-9 in the three periods are 74 U / mL, 98 U / mL and 142 U / mL, respectively, and the change rate is The change rate sequence {0.324, 0.449} is constructed, and the difference between adjacent two items in the sequence is calculated to obtain the first level trend change in three periods, and the result is 0.449-0.324=0.125, which is the change rate difference, and the change rate difference sequence {0.125} is constructed. Similarly, the change rate sequence and the difference sequence of CEA and CRP are constructed by using the above method, and examples are shown as follows: the values of CRP in three periods are 3.2 mg / L, 3.9 mg / L and 5.4 mg / L, and the corresponding change rates are {0.219, 0.385}, and the difference is 0.166, thereby establishing the change trend information of the change rate difference, as shown in Table 1:

[0055] Table 1: Example table of change rate and change rate difference

[0056] Indicator T1 value T2 value T3 value R1 (T1-T2) R2 (T2-T3) Difference D CA19-9 74 98 142 0.324 0.449 0.125 CEA 5.2 6.1 6.9 0.173 0.131 -0.042 CRP 3.2 3.9 5.4 0.219 0.385 0.166

[0057] As shown in Table 1, the change rate difference of the three indexes is processed by the standard difference value calculation rule, which is used to reflect the similarities and differences of the change rates of different indexes in consecutive periods, thereby forming the basis trend track.

[0058] The acceleration sequence generation submodule generates the jump acceleration change parameter group based on the change trend information of the change rate difference, calculates the difference between items in the change rate difference sequence of each index, constructs the third-order response sequence reflecting the acceleration degree of the trend, and extracts the group of the maximum third-order response value and the corresponding detection period position of each index to generate the jump acceleration change parameter group;

[0059] Based on the change trend information of the change rate difference, the difference sequence corresponding to each index is selected, the difference between consecutive items in the sequence is calculated to form a second-order difference processing, for example, in another patient, the change rate difference sequence of CA19-9 is {0.08, 0.14, 0.29}, and the third-order response sequence is {0.06, 0.15}. The sequence is used to represent the acceleration information of the jump trend and reflects the trend change speed of the index in a unit period. The maximum item of the third-order response value and the corresponding period number are extracted as the basic quantity item of the jump parameter group. If the highest response value is 0.15 and the period is T3, the jump item is (0.15, T3). To determine the significance of the jump, a jump acceleration reference value is introduced and set to θ=0.12. The value is calculated from the median distribution value of the maximum third-order response of 500 diagnosed samples in the pancreatic cancer case data set. θ is defined as the 75th percentile value of the median of the maximum third-order response of each sample. For example, if the 375th item of the sorted maximum third-order response of 500 samples is 0.12, then θ=0.12 is set as the reference. If the response value of a certain index is greater than θ, it is marked as a jump point. The threshold value of the reference is divided by the 75th percentile line, which can improve the difference of the jump screening. Finally, the maximum value of the acceleration of each of the three indexes and the detection period are selected to form the jump acceleration change parameter group.

[0060] The abnormal trend identification submodule judges the positive and negative signs of the acceleration change direction of the three indexes according to the jump acceleration change parameter group, marks the corresponding detection period in combination with the position of the maximum value, forms a jump feature combination of the acceleration direction and the position, and is classified into a jump trend category to establish a jump trend marker set;

[0061] According to the jump acceleration change parameter group, the period in which the maximum response value of each index is located is obtained, and the positive and negative directions are extracted. By identifying the sign of the response value, it is judged whether the index in the period is a positive jump or a negative jump. If the response value is positive, the direction is marked as positive, and vice versa. Then, a trend identification marker is constructed according to the jump direction. For example, the CA19-9 response value is +0.15 in T3 period, the CEA response value is -0.04 in T2 period, and the CRP response value is +0.09 in T3 period. The combination feature {(CA19-9, T3, +), (CEA, T2, -), (CRP, T3, +)} is constructed. On this basis, priority weights are applied to the three indexes to determine the trend intensity. The weights are set as follows: ω1 = 0.5 (CA19-9), ω2 = 0.3 (CRP), and ω3 = 0.2 (CEA). The weights are given according to the sensitivity report of early diagnosis of pancreatic cancer. The weight value is equal to the normalized sensitivity of the index. If there is a positive jump in the main weight item (i.e. ω > 0.4) and the total positive jump index cumulative weight value is greater than 0.6, it is classified as "strong jump type". In the example, the positive jump items are CA19-9 and CRP, and the cumulative weight is 0.5 + 0.3 = 0.8 > 0.6, which constitutes a strong jump. If all are negative jump items, it is classified as "inhibition trend". The purpose of setting this classification threshold is to extract the dominant jump signal in the presence of multiple index interference. The final combination result constitutes a jump trend marker set, which is used for subsequent risk level mapping and disease course path updating.

[0062] Please refer to Figure 3 , the index level mapping module includes:

[0063] The index classification submodule calls the CA19-9, CEA, and CRP values in the detection period corresponding to each jump point based on the jump trend marker set, matches the three index values with the level intervals in the pancreatic cancer risk assessment reference table, and labels the corresponding level number of the index according to the matching result to generate level number attribution information.

[0064] Based on the sudden trend marker set, the specific values of CA19-9, CEA and CRP in the detection period corresponding to the sudden point are extracted, and the pancreatic cancer risk assessment reference table is called to match each index value to the risk level interval range to complete the interval attribution determination. In actual implementation, the pancreatic cancer risk assessment reference table is usually constructed based on existing epidemiological and clinical data, and the index level is divided into five interval levels, which are represented by level numbers 1 to 5. For example, the reference interval of CA19-9 can be set as: level 1 (0-37 U / mL), level 2 (38-74 U / mL), level 3 (75-150 U / mL), level 4 (151-300 U / mL), and level 5 (301+U / mL). If the CA19-9 in the sudden point is 98 U / mL, it is attributed to level 3. The intervals of CEA and CRP are set as follows: CEA level 1 (0-3 ng / mL), level 2 (4-5 ng / mL), level 3 (6-7 ng / mL), level 4 (8-10 ng / mL), and level 5 (11+ng / mL); and CRP level 1 (0-2 mg / L), level 2 (2.1-4 mg / L), level 3 (4.1-6 mg / L), level 4 (6.1-8 mg / L), and level 5 (8.1+mg / L). If the data of a sudden point is: CA19-9 is 98 U / mL, CEA is 6.4 ng / mL, and CRP is 5.3 mg / L, it is attributed to level numbers 3, 3 and 3 respectively. The attribution operation can generate three index level number record information of each sudden point to form the level number attribution information set. The interval division standard is based on the standard pancreatic cancer pathological reference interval to ensure consistency with the actual pathological distribution. The attribution operation is a direct comparison method, that is, if the index value falls between the boundaries of a certain interval, the corresponding interval number is assigned as the index level number to form a matching control table structure. This process establishes the index level number attribution information;

[0065] The level span calculation submodule obtains the number difference between the current period level and the previous period level of each index according to the level number attribution information, counts the number of level changes of each index in the three indexes, and marks the change direction as rising or falling. The level span parameter group is obtained by combining the number of levels and the change direction to generate a multi-index change summary item.

[0066] According to the level number attribution information, the level numbers of each index in the current detection period and the previous period of the sudden point are called to calculate the difference value of each index level number. The absolute value of the difference value is the level span, and the signed difference value is used for direction determination. If the difference value is positive, it means the level rises, and if it is negative, it means the level falls. For example, the CA19-9 level in the previous period is 2 and the current period is 3, which changes by 1 level. The CRP level decreases from 4 to 2, which changes by 2 levels. The values and directions are recorded respectively to generate the level difference setCA19-9 = +1, delta CEA = 0, delta CRP = -2}, in the process of constructing the difference, the single indicator is allowed to be unchanged (i.e. the difference is 0), in which case it is considered to be a neutral trend and does not participate in the directionality summary judgment, the difference set is combined to construct a level change structure, which is used for subsequent trend consistency judgment and transition identification, the number of levels is used as a quantitative standard of risk fluctuation amplitude, and the direction flag indicates the change main direction, if the number of indicators with the same direction change is greater than 2 and the total level change number is greater than or equal to 3, it is considered that there is a trend transition phenomenon, a multi-index change summary item is generated, and is further integrated into a level change level parameter group, and in the example, it is {(+1), (0), (-2)}.

[0067] The level transition generation submodule calls the level change level parameter group, judges the consistency of the level change direction and the transition direction of each indicator, and selects the indicators with the same direction, and integrates the level span and the change direction to construct a level change structure combination under the transition period, and establishes a risk level transition sequence.

[0068] The level change level parameter group is called, the consistency of the level change direction and the transition direction of each indicator is judged, i.e. if an indicator shows positive transition in the transition trend, and its level number increases from the last period, it is determined as a direction consistent item, otherwise it is determined as a direction conflict item, only the positively correlated indicators are retained in the consistency judgment, such as in the example, if the transition direction of CA19-9 is positive, it is consistent with the level change +1 direction, the level of CEA is neutral, and the transition direction of CRP is negative, which is consistent with the level decrease direction, therefore, the consistent items are marked as CA19-9 and CRP, and the level span and direction flag of the two are extracted, respectively, to construct a structure combination item, such as (CA19-9, +1, +) and (CRP, -2, -), which is used to establish a risk level transition sequence, and the final summary result is used as a risk evolution record of the transition period, the judgment basis only depends on the sign change of the level number and the consistency of the transition trend, without involving the specific values of the indicators, avoiding the influence of value domain heterogeneity on the stability of the judgment result, and the process constructs a complete transition structure. The result shows that in the process of establishing the risk level transition structure, the direction consistency of the indicators is used as the main line of judgment, the cross-period level difference and the trend direction are combined to form the transition expression, which has strong structure recognition ability.

[0069] Please refer to Figure 4 The sample stability screening module includes:

[0070] The confidence parameter normalization submodule obtains CA19-9, CEA and CRP detection values of each sample in the corresponding sample set according to the risk level transition sequence, calculates the coefficient of variation, the difference value of the survival function score and the standard deviation of the organ function score, respectively groups according to the indexes to perform range standard normalization processing, obtains three normalized parameters, and establishes confidence factor information;

[0071] According to the risk level transition sequence, the sample number of each jump point is extracted, the corresponding three biological index values, i.e. CA19-9, CEA and CRP three-period detection values, are called, and the survival function score and the organ function score record of the sample in the jump period and the adjacent period are combined to calculate three confidence factor basic indexes, wherein the coefficient of variation CV is defined as the ratio of the standard deviation to the mean value. For example, the CA19-9 three-period detection values of a sample are 112 U / mL, 137 U / mL and 161 U / mL, the mean value is calculated as The standard deviation is σ≈20.1, and The survival function score adopts the Karnofsky scoring system, the difference value of the three-period score value is -10 and -10, and the total difference value is ΔK=-20. The organ function score adopts the standard deviation of the difference of the three-period value of the SOFA score. For example, the SOFA scores are 6, 8 and 7, the standard deviation is about 0.816, and three original quantities are constructed. Then, the range normalization operation is performed on each index according to the range normalization of all samples, and the range normalization value is defined as Wherein x is the original value of a certain index of the current sample, x min and x max are the minimum value and the maximum value of the index in the whole sample range. The normalized values are defined as A1, A2 and A3, respectively, to construct the confidence factor information in a unified scale. The range normalization rule ensures that the normalized value is distributed in the range of 0 to 1, which is conducive to subsequent weighted integration processing and avoids error accumulation caused by inconsistent dimensions. The standard deviation, difference value calculation and mean value processing are standard statistical processing procedures, which can be directly completed through three-period data.

[0072] The weight score calculation submodule extracts the coefficient of variation normalized value, the survival function score difference normalized value and the organ function score standard deviation normalized value of each sample based on the confidence factor information, combines the three index response amplitude total amount and the organ function score change rate of each sample in the detection period, and adopts the formula:

[0073]

[0074] The operation obtains the sample confidence weight level, arranges the samples in descending order according to the confidence weight level value, and generates a sample confidence level sequence.

[0075] Wherein, A1 represents the normalized value of the coefficient of variation, A2 represents the normalized value of the difference in survival function score, A3 represents the normalized value of the standard deviation of organ function score, R represents the response amplitude of the three biological indicators, and AF represents the total amount of change in organ function score in the last three cycles, and S represents the confidence level value of the sample.

[0076] The confidence level value represents the reliability of each patient sample in the prognosis analysis of pancreatic cancer, reflecting the stability and consistency of the key indicators. The higher the score, the more the sample data conforms to the expected index change pattern of the system, and the data is more valuable.

[0077] Technical effects:

[0078] Excluding non-representative samples caused by accidental fluctuations, detection errors or extreme changes;

[0079] Improve the stability of the training set and analysis data, and reduce algorithm noise interference;

[0080] Provide a reliable data basis for subsequent path node offset identification and stage classification.

[0081] Calculation logic:

[0082] A1: Normalized value of the coefficient of variation (measures the intensity of fluctuation of the detection index);

[0083] A2: Normalized value of the difference in survival function score (measures the degree of function decline);

[0084] A3: Normalized value of the standard deviation of organ function score (measures the fluctuation of organ score);

[0085] R: Total amount of response amplitude of the three biological indicators (CA19-9, CEA, CRP) in three cycles;

[0086] AF: Change in organ function score in three cycles; S: Final confidence level.

[0087] The calculation principle embodies the following two points: comprehensive multi-dimensional index difference, construct a weighted combination reflecting the fluctuation of physiological state; Through variability, score difference and response amplitude, a mechanism for identifying abnormal samples is constructed.

[0088] Based on the confidence factor information, the normalized three index values in each sample are extracted in turn, and are denoted as A1, A2 and A3. The sum of the maximum and minimum differences of the CA19-9, CEA and CRP index values in each sample in the detection period is obtained, and is denoted as R. For example, the maximum and minimum differences of the three index values in a sample in three periods are CA19-9: 161-112=49, CEA: 6.9-5.1=1.8 and CRP: 5.2-2.6=2.6, and R=49+1.8+2.6=53.4. Meanwhile, the maximum difference sum ΔF of the organ function score is obtained. For example, the SOFA scores are 6, 8 and 7, and ΔF=|8-6|+|8-7|=2+1=3. The innovative score formula is:

[0089]

[0090] In the formula, |A1+A2-A3| represents the magnitude result of the cooperative variation direction in the confidence factor, which is used to evaluate the coordination of the index fluctuation. R represents the amplification strength of the biological response. The ΔF term represents the dynamics of the organ function variation. The denominator term adds |A1-A3| to prevent the weight polarization caused by the extreme value deviation. The square root and the fraction structure represent the overall harmony of the trend outbreak and convergence. The absolute value operation in the overall structure avoids direction interference. The square root narrows the difference, and improves the resolution of the score distribution density interval. Assuming that A1=0.42, A2=0.35, A3=0.27, R=53.4 and ΔF=3, the calculation is as follows:

[0091]

[0092] The results show that the score result S=7.78 is within the upper limit 10. According to the standardization interval statistics of the previous data, the sample rejection score threshold is set to 8.5. Therefore, the current sample can be retained for subsequent analysis. The beneficial aspects of the formula are that the non-linear harmony structure between the organ dynamic score term and the reduction factor is introduced, the response sensitivity to the uncoordinated variation of multiple indexes is improved, the individual fluctuation adjustment is placed in the risk amplification logic framework, and the confidence resolution of sample screening is improved.

[0093] The abnormal paragraph rejection submodule calls the sample confidence level sequence, calculates the absolute value of the score difference between each two adjacent samples after sorting, marks the paragraph with a score difference greater than the screening threshold for three consecutive groups, rejects the samples corresponding to the paragraph, and establishes a stable sample set after screening.

[0094] The sample confidence level sequence is called, and the absolute value operation is performed on the score difference between adjacent samples in the sorted sample sequence in sequence. For example, the sample S1 score is 8.2, S2 is 7.4, S3 is 6.9, and S4 is 5.2. The absolute value of the score difference is calculated as |S1-S2|=0.8, |S2-S3|=0.5, and |S3-S4|=1.7. If the screening threshold is set to 1.2, the threshold is set according to 1.5 times the standard deviation of the overall score value, that is, the score mean standard deviation is taken as a reference. For example, if the score sample standard deviation is 0.8, the threshold is 1.5x0.8≈1.2. If two of the three consecutive score differences are greater than 1.2, the point is marked as a screening breakpoint, and the sample sequence in which the direction of the score after the breakpoint is opposite to that before the breakpoint is removed. A stable sample set after screening is established, which is an important training input for subsequent prediction path reconstruction and classification learning, ensuring the continuity and score analyzability of the subsequent modeling sequence. The results show that the jump amplitude in the score difference sequence can be extracted and removed by absolute value comparison. The stable sample set after screening has usability in continuity and confidence. The numerical breakpoint as a key cutting point provides precision support for abnormal removal.

[0095] Please refer to Figure 5 The path node offset identification module includes:

[0096] The difference extraction submodule calls the stable sample set after screening, extracts the CA19-9, CEA, and CRP combined values of the three consecutive nodes according to the prognosis prediction time sequence of each patient, calculates the absolute difference sequence of the adjacent two node combinations, arranges the Manhattan distance difference sequence in time sequence, and establishes the path difference sequence information;

[0097] After calling the stable sample set for screening, all node information in the prognosis path of each pancreatic cancer patient is read in turn, and the combination of CA19-9, CEA and CRP index values at each of the three consecutive time nodes is extracted as a three-dimensional numerical combination, for example, the corresponding index values of a patient at the three consecutive nodes are node 1: (120, 5.8, 4.6), node 2: (134, 6.5, 6.1), and node 3: (118, 5.9, 5.3). The absolute difference value of each index between the adjacent two nodes is calculated according to the combination dimension, that is, the difference value of the combination of node 1 and node 2 is |134-120|+|6.5-5.8|+|6.1-4.6|=14+0.7+1.5=16.2, the difference value of the combination of node 2 and node 3 is |118-134|+|5.9-6.5|+|5.3-6.1|=16+0.6+0.8=17.4, and the two difference values are arranged in time sequence as a difference value sequence [16.2, 17.4], which is the Manhattan distance difference value sequence of the combined index between the path nodes. By sliding processing all the three nodes, the complete path difference sequence information is generated by traversing all the path segments.

[0098] The unstable degree value calculation submodule is based on the path difference sequence information, constructs a sliding window structure according to each difference value sequence, extracts the number of sign changes of the Manhattan distance difference value in each window, and uses the formula:

[0099]

[0100] The operation obtains the offset direction unstable degree value corresponding to the node, and generates a node unstable degree value sequence.

[0101] Wherein, F i represents the number of sign changes in the i-th sliding window, D ij represents the absolute value of the j-th Manhattan distance difference value in the i-th window, represents the cumulative result of the absolute values of the n difference values in the i-th window, λ i represents the attenuation exponential factor of the i-th window, is the attenuation weight coefficient corresponding thereto, M i represents the maximum difference value in the i-th window, U i is the offset direction unstable degree value calculated for the i-th window, and e is the base of natural logarithm.

[0102] The offset direction unstable degree value is used to measure the discontinuity of the index change direction of a specific detection node in the prognosis path, and reflects whether the pancreatic cancer development trend at the node appears a significant turning point or mutation risk.

[0103] Technical effects:

[0104] Label potential key course variation nodes;

[0105] Provide quantitative indicators for mutation detection in path analysis;

[0106] Provide scoring basis for path deviation node index.

[0107] Calculation logic:

[0108] F i : The number of times of sign change of the Manhattan distance difference value in the i-th sliding window (representing the frequency of fluctuation direction change);

[0109] D ij : The absolute value of the j-th difference value in the i-th window;

[0110] λ i : The attenuation factor of the current window (used to control the weight);

[0111] M i : The maximum value of the difference value in the window;

[0112] The attenuation weight of the difference value;

[0113] U i : The final instability value.

[0114] Principle: Use sliding window analysis, combine the direction change frequency F i of the distance sequence, introduce an exponential attenuation mechanism to weaken the influence of early large differences, highlight the nearby fluctuations; the denominator term reflects the concentration of changes, and the numerator term reflects the frequency of fluctuations, and finally output the comprehensive stability score.

[0115] Based on the constructed path difference sequence information, a sliding window structure with a length of 3 is constructed for each difference sequence, and the sign change count is performed for the three difference items in each window. If the positive and negative changes occur before and after, it is counted as a change, and is recorded as the i-th window change frequency F i , and the absolute value of each difference value in the window is taken to form an absolute value sequence D ij , such as the difference sequence [16.2, 17.4, 15.1], the absolute value sequence is [16.2, 17.4, 15.1], and the sum of the sequence is , n=3, set the exponential attenuation factor λ i of the window = 0.12×i, such as the 3rd window, λ3=0.36, the corresponding attenuation weight e -0.36 ≈0.697, and the maximum item M i of the difference value absolute value in the window is obtained = 17.4, and the improved formula is as follows:

[0116]

[0117] The offset direction instability value U under the window is generated i = 0.238, the score reflects the frequency of the fluctuation direction of the combined indicators in the path, and the larger the value, the stronger the instability. The fluctuation frequency and fluctuation amplitude drive the score weight structure, and the exponential decay factor dynamically scales the contribution of the late window to improve the score time sensitivity;

[0118] The formula structure is as follows: F i The frequency of symbol changes is evaluated, indicating the degree of direction repetition in the window, The sum of the absolute values of the differences in the window after attenuation, the maximum value M i The fluctuation peak value in the difference sequence, the difference structure is formed by subtracting the extreme value and taking the absolute value, and this value is under the square root to suppress the linear expansion of the score by the high fluctuation sequence. The denominator structure ensures that it cannot be 0 based on 1, and the square root processing suppresses the influence of extreme difference items. Finally, the score U i The direction and amplitude double factors are fused to express the instability level, and the exponential term increases the dynamic decreasing effect of the score, making the score mechanism more sensitive to recent windows.

[0119] The benefit of this score is that it forms a harmonic structure by superimposing the frequency of direction changes and the normalized fluctuation peak difference, and the exponential factor dynamically adjusts the window contribution, effectively improving the response ability to irregular fluctuations in node time series identification, and having better mutation identification ability in path anomaly judgment.

[0120] The offset node marking submodule compares the node score with the pre-set path stability threshold according to the node instability value sequence, and marks the node as an offset trigger node when the score exceeds the threshold interval. The marked nodes are summarized according to the original sequence index position to establish a path offset node index table.

[0121] The offset direction instability value U corresponding to each node is obtained i The threshold is set by referring to the average value of the score sequence of all nodes plus 1.2 times the standard deviation, such as the average value of all scores in a batch of samples is 0.195, and the standard deviation is 0.065, then the path stability threshold is set to 0.195 + 1.2 x 0.065 = 0.273, such as a node corresponds to U i = 0.295 > 0.273, then the node is marked as an offset trigger node, and its original time sequence index position is recorded. All threshold nodes are summarized to establish a path offset node index table for subsequent structure analysis or stage division. As shown in Table 2.

[0122] Table 2 Path offset node index table

[0123]

[0124] As shown in Table 2, the scores of the nodes determined as the offset trigger nodes in some samples and the node positions are listed, which can be used for subsequent breakpoint selection and key segment marking in path reconstruction. The structure supports batch extraction of key nodes, enables specific entry points and clear definition of broken segments in path reconstruction, and helps subsequent phased modeling and classification task execution.

[0125] Referring to Figure 6 , the prognosis stage classification module includes:

[0126] The feature construction submodule obtains the CA19-9, CEA, CRP values and organ function score values corresponding to each node according to the path offset node index table, arranges the values in time sequence order, constructs a four-dimensional index feature vector of each node, and establishes a node index feature set;

[0127] According to the path offset node index table, the CA19-9, CEA, CRP three biological index values and organ function score values of each node listed in the table at the detection period corresponding to the node are extracted, and the four index values are arranged in chronological order to form a four-dimensional index structure vector of the node. The CA19-9, CEA, and CRP values directly come from the detection report, such as CA19-9 = 128 U / mL, CEA = 6.1 ng / mL, and CRP = 5.3 mg / L. The organ function score adopts a scale scoring mechanism, and the standard interval is 0 to 10 points. For example, if the corresponding node score is 7.5 points, the feature vector of the node can be represented as [128, 6.1, 5.3, 7.5]. After all the offset nodes are traversed, a node index feature set with a structure of [X a1 ,X a2 ,X a3 ,X a4 ] is established, all nodes are numbered in time sequence, and the node index feature set construction is completed.

[0128] The distance calculation submodule calculates the Euclidean distance between each node feature vector and each label vector based on the node index feature set and the standard stage label center point vector set, using the formula:

[0129]

[0130] The minimum weighted distance value of each node corresponding to the stage label is obtained, and the label distance minimum value sequence is obtained.

[0131] wherein E a represents the minimum weighted distance value between the a-th node and all label center point vectors, and ω aGlobal normalization indicator weight of the a-th node, γ a Offset rate factor of the a-th node, φ a Confidence callback factor of the a-th node, X ab Feature value of the a-th node in the b-th dimension, C ab Reference value of the b-th label center point in the dimension;

[0132] The minimum weighted distance indicates that the system calculates the Euclidean distance between the node feature vector and each standard stage label center point when classifying the node into the prognosis stage, and the label corresponding to the minimum value is the prognosis stage to which the node most likely belongs.

[0133] Technical effect: realize quantitative classification of the clinical prognosis stage to which the offset node belongs; improve the objectivity and interpretability of the disease course segmentation; support the stage prediction model based on the node feature vector.

[0134] Calculation logic:

[0135] 1. Construct a four-dimensional feature vector (CA19-9, CEA, CRP, organ function score) for each offset node; 2. Perform Euclidean distance operation with all label center point vectors:

[0136] x i : i-th indicator value of the node;

[0137] c ki : i-th numerical value of the k-th label center point;

[0138] w i : i-th indicator weight (can be set to 1 or weighted according to indicator sensitivity);

[0139] 3. Select the label corresponding to the minimum d k as the stage attribution.

[0140] Calculation principle: nearest neighbor principle in multi-indicator vector space; the judgment ability for important indicators (such as CA19-9) can be enhanced by adjusting the weight; it has interpretability and stability, and is suitable for segmentation classification problems.

[0141] Based on the constructed node indicator feature set, for the four-dimensional vector of each node, call the standard stage label center point vector set C ab , where each label is constructed by clinical experts according to the diagnosis result, and the four-dimensional reference mean value corresponds to the same four indicators, respectively, perform Euclidean distance calculation between each node and all label vectors, and introduce offset rate factor γ a , node indicator weight ω a , and confidence callback factor φ aThe comprehensive structure evaluation is performed to improve the sensitivity of label fitting to the difference of nonlinear signal and the hierarchical expression of node importance. Assuming that a certain node is numbered as a=3, the eigenvalue vector of the node is X 31 =134,X 32 =6.3,X 33 =5.8,X 34 =7.2, the offset rate factor γ3=0.25, the weight ω3=0.7, the confidence callback factor φ3=0.02 are set, and the corresponding label center point vector is [C 31 =120,C 32 =5.9,C 33 =5.4,C 34 =7.5], the weighted distance between the node and the label is:

[0142]

[0143] The dimensional items are calculated respectively:

[0144] Dimension 1: (134-120) 2 =196, |134-120|=14, ln(1+0.02·134)=ln(3.68)≈1.304;

[0145] Dimension 2: (6.3-5.9) 2 =0.16, |6.3-5.9|=0.4, ln(1+0.02·6.3)=ln(1.126)≈0.119;

[0146] Dimension 3: (5.8-5.4) 2 =0.16, |5.8-5.4|=0.4, ln(1+0.02·5.8)=ln(1.116)≈0.110;

[0147] Dimension 4: (7.2-7.5) 2 =0.09, |7.2-7.5|=0.3, ln(1+0.02·7.2)=ln(1.144)≈0.134;

[0148] The formula is calculated as:

[0149]

[0150] The minimum value of the weighted Euclidean distance E3≈11.95 is obtained, and the best label distance value corresponding to the node attribution will be used as the final classification basis;

[0151] In the formula structure, the square term evaluates the overall deviation of the index, the absolute value term weightedly expresses the local mutation, and the confidence term adjusts the signal sensitivity of the abnormally high amplitude in a logarithmic form. The overall structure integrates the triple adjustment mechanism to improve the adaptability of the index combination structure to the abnormal response of complex signals and has comprehensive discrimination ability for different types of index fluctuations.

[0152] The innovation of the formula lies in that the logarithmic suppression is adopted for the abnormal amplitude term by introducing the confidence callback, and the deviation rate and the weight term are combined to adjust the node response value in the local and global ranges, so that the response characteristics of each dimension obtain the structured adjustment capability, thereby enhancing the overall discriminability and node recognition reliability of the classification process.

[0153] The results show that after adding the structure adjustment factor, the classification index is no longer based on the absolute difference of the Euclidean distance, but integrates the structure position, node signal response capability, and abnormal amplitude dynamic callback mechanism, thereby providing a more stable matching sequence for subsequent label attribution.

[0154] The stage labeling sub-module selects the label category item with the minimum distance of each node as the attribution stage identifier according to the minimum label distance sequence, and maps and integrates each node and the corresponding category into structured labeling information to establish a prognosis stage classification list.

[0155] After obtaining the weighted distance sequence of all nodes, for each node, the minimum distance item is selected from multiple label distance values, and the corresponding label category is the stage type to which the node belongs. For example, the distances of node 3 to each label are [12.4, 11.95, 14.8], and the attribution stage is label 2. After processing all nodes, the node number and the corresponding stage label are arranged into a structured mapping table, such as the attribution category of node 3 is “intermediate development period” and the attribution category of node 7 is “late stage”. The structured mapping table is arranged in the original index order to establish a prognosis stage classification list, which is used for subsequent structure stability tracking and disease course trend modeling analysis.

[0156] The above is only a preferred embodiment of the present application, and does not limit the present application in other forms. Any skilled person in the art can modify or change the disclosed technical content to equivalent embodiments applied to other fields, but any simple modification, equivalent change and modification made on the basis of the technical essence of the present application to the above embodiments still belongs to the protection scope of the present application.

Claims

1. A prognosis analysis system for pancreatic cancer based on big data mining, characterized by, The system comprises: The system comprises: The index grade mapping module classifies and matches each index value with the grade interval in the pancreatic cancer risk assessment reference table based on the sudden trend marker set, counts the grade span of each sudden point, forms a structured index grade transition item in combination with the sudden direction, and generates a risk grade transition sequence; The sample stability screening module calculates a confidence weight grade according to the risk grade transition sequence, judges the score difference between adjacent two samples, identifies a sample paragraph with a continuously greater screening threshold, and obtains a stable sample set after screening; The path node offset identification module calls the stable sample set after screening, calculates a sequence instability value according to a prognosis prediction time sequence, screens a direction mutation node position, and obtains a path offset node index table. 2.The big data mining based pancreatic cancer prognosis analysis system according to claim 1, wherein, The sudden trend marker set specifically refers to a sudden detection cycle position, a third-order mutation response value, and a mutation direction type, the risk grade transition sequence includes a risk grade span value, an index corresponding grade number, and a grade change direction, the stable sample set after screening specifically refers to a retained sample number, a sample confidence grade value, and a removal paragraph start and end position, and the path offset node index table includes a node index number, a sliding offset direction value, and a mutation trigger flag. 3.The big data mining based pancreatic cancer prognosis analysis system according to claim 2, wherein, The sudden trend extraction module comprises: The change rate calculation submodule obtains the continuous three-phase detection records of the pancreatic cancer patient, extracts the CA19-9, CEA, and CRP three-item biological index original value sequences, calculates the change rate of each index between the first phase and the second phase and between the second phase and the third phase, obtains the change rate sequence, performs inter-item difference calculation on adjacent two items in each group of sequences to obtain the change rate difference sequence of each index, and establishes change rate difference trend information; The acceleration sequence generation submodule calculates the inter-item difference of each index change rate difference sequence based on the change rate difference trend information, constructs a third-order response sequence reflecting the acceleration degree of the trend, extracts a group of three-order responses with the maximum value and the corresponding detection cycle position in each index, and generates a sudden acceleration change parameter group; The abnormal trend identification submodule judges the positive and negative signs of the acceleration change direction of the three indexes according to the sudden acceleration change parameter group, labels the corresponding detection cycle based on the position of the maximum value, forms a sudden feature combination of the acceleration direction and the position, and is classified into the sudden trend category to establish the sudden trend marker set. 4.The big data mining based pancreatic cancer prognosis analysis system according to claim 3, wherein, The index grade mapping module comprises: The index classification submodule calls the CA19-9, CEA and CRP values in the detection period corresponding to each inflection point based on the inflection trend marker set, matches the three index values with the interval in the pancreatic cancer risk assessment reference table respectively, labels the corresponding grade number of the index according to the matching result, and generates grade number attribution information; The grade span calculation submodule obtains the number difference between the current period grade and the previous period grade of each index according to the grade number attribution information, counts the number of grade changes of each of the three indexes, and marks the change direction as rising or falling, generates a multi-index change summary item combining the number of levels and the change direction, and obtains a grade change level parameter group; The grade transition generation submodule calls the grade change level parameter group, judges the consistency of the grade change direction of each index and the inflection direction, selects the index items with consistent directions, integrates the grade span and the change direction to construct a grade change structure combination in the inflection period, and establishes a risk grade transition sequence. 5.The big data mining based pancreatic cancer prognosis analysis system according to claim 4, wherein, The sample stability screening module comprises: The confidence parameter normalization submodule obtains the CA19-9, CEA and CRP detection values of each sample in the corresponding sample set according to the risk grade transition sequence, calculates the coefficient of variation, the difference in survival function score, and the standard deviation of organ function score, respectively groups them according to the index, performs range standard normalization processing, obtains three normalized parameters, and establishes confidence factor information; The weight score calculation submodule extracts the normalized value of the coefficient of variation, the normalized value of the difference in survival function score, and the normalized value of the standard deviation of organ function score in each sample based on the confidence factor information, combines the total amount of three index response amplitudes and the change rate of organ function score of each sample in the detection period, calculates the sample confidence weight grade, arranges the samples in descending order of confidence weight grade value, and generates a sample confidence grade sequence; The abnormal paragraph elimination submodule calls the sample confidence grade sequence, calculates the absolute value of the score difference between each two adjacent samples after sorting, marks the paragraph where the absolute value of the score difference of three consecutive groups is greater than the screening threshold, eliminates the corresponding sample, and establishes a stable sample set after screening. 6.The big data mining based pancreatic cancer prognosis analysis system according to claim 5, wherein, The path node offset identification module comprises: The difference extraction submodule calls the stable sample set after screening, extracts the combined values of CA19-9, CEA and CRP of three consecutive nodes according to the prognosis prediction time sequence of each patient, calculates the absolute difference sequence of the combination of two adjacent nodes, arranges the Manhattan distance difference sequence in time sequence to establish path difference sequence information; The unstable degree value calculation submodule constructs a sliding window structure according to each difference sequence based on the path difference sequence information, extracts the number of sign changes of the Manhattan distance difference in each window, calculates the offset direction unstable degree value of the node, and generates a node unstable degree value sequence; The offset node marking submodule compares the node score with the pre-set path stability threshold according to the node unstable degree value sequence, marks the node as an offset trigger node when the score exceeds the threshold interval, and arranges the marked nodes according to the original sequence index position to establish a path offset node index table. 7.The big data mining based pancreatic cancer prognosis analysis system according to claim 6, wherein, The system further comprises: The prognosis stage classification module obtains the biological index combination value and the organ function score of each node according to the path offset node index table, constructs a feature vector, calls a stage label vector set, performs Euclidean distance calculation between the feature vector and each type of label center point, is attributed to the category item with the smallest distance, and generates a prognosis stage classification list; The prognosis stage classification list specifically refers to a classification stage number, a stage label name, and a classification node position. 8.The big data mining based pancreatic cancer prognosis analysis system according to claim 7, wherein, The prognosis stage classification module comprises: A feature construction submodule obtains the CA19-9, CEA, CRP values and organ function score values corresponding to each node according to the path offset node index table, arranges the values in time sequence order, constructs a four-dimensional index feature vector of each node, and establishes a node index feature set; A distance calculation submodule calls a standard stage label center point vector set based on the node index feature set, calculates the Euclidean distance between each node feature vector and each type of label vector, obtains the minimum weighted distance value of each node corresponding to the stage label by calculation, and obtains a label distance minimum value sequence; A stage labeling submodule selects the label category item with the smallest distance of each node as an attribution stage identifier according to the label distance minimum value sequence, maps and integrates each node and the corresponding category into structured labeling information, and establishes a prognosis stage classification list.

Citation Information

Cited By

  • Breast dredging treatment course tracking and management method for postpartum care

    CN121393908A

  • A postpartum care breast unblocking treatment course tracking and management method

    CN121393908B

  • Ganoderma lucidum spore powder extracellular vesicle experiment result analysis method and system

    CN121963864A

  • Whole treatment course management system based on hepatitis B serology and virology index fusion analysis

    CN122158159A

  • A comprehensive treatment management system based on the fusion analysis of hepatitis B serological and virological indicators.

    CN122158159B