Clinical gene-disease risk dynamic prediction method and application thereof

By combining ROC curve analysis, principal component regression and time series modeling, user permissions and data labels are dynamically adjusted, which solves the problem of inaccurate and low flexibility in medical operation data query, and realizes the flexibility and timeliness of data query.

CN120376156APending Publication Date: 2025-07-25BEIJING KANGXU MEDICAL LAB CO LTD

Patent Information

Application Number
CN202510893187.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, after medical operation data is classified according to data sensitivity, it cannot be dynamically adjusted, resulting in untimely user information query and low flexibility in operation data query.

Method used

A multi-dimensional classification and dynamic prediction algorithm combining ROC curve analysis, principal component regression, logistic regression and time series modeling is used to analyze user rights and data importance, and dynamically adjust user rights and data labels by collecting medical operation information and user data.

Benefits of technology

It improves the flexibility and timeliness of medical operation data query, ensuring accurate matching and updates when data sensitivity changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120376156A_ABST
    Figure CN120376156A_ABST
Patent Text Reader

Abstract

The invention discloses a clinical-oriented gene-disease risk dynamic prediction method and application thereof, relates to the technical field of data prediction updating, is used for solving the problems that user information query is not timely and operation data query is low in flexibility, and analyzes user permission by collecting medical operation information data and user data and utilizing the user data to obtain a medical operation risk dynamic prediction result. And classifying the users according to the user permissions. The medical operation data importance is analyzed according to the medical operation information data, different labels are marked on different medical operation data according to the medical operation data importance, the label updating time is calculated, and the medical operation data marked by the different labels are matched with users with different permissions. And the matching feedback data and the information sensitivity change data are collected, and the label updating time and the user permission are adjusted by integrating the matching feedback data and the information sensitivity change data, so that the medical operation data query flexibility is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data prediction and update. More specifically, the present invention relates to a dynamic prediction method for gene-disease risk for clinical use and its application. Background Art

[0002] In clinical scientific research and refined medical operation management, the collection and analysis of medical operation data and user behavior data are of great significance for improving the efficiency of auxiliary decision-making. With the large accumulation of multi-source heterogeneous data such as genomic data, electronic medical record data, and operation log data, how to achieve dynamic access and prediction based on user permissions and data importance while ensuring information security and privacy has become a key issue in information management technology.

[0003] In the prior art, after classifying the previous medical operation data according to data sensitivity, different encryption protections are carried out on the operation data of different sensitivities. When the data sensitivity and user needs change, it causes untimely query of user information and low flexibility in querying operation data. Therefore, a dynamic prediction method for gene-disease risk for clinical use and its application are proposed.

[0004] The above information disclosed in the background art section is only used to strengthen the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a dynamic prediction method for gene-disease risk for clinical use and its application, and solves the problems proposed in the above background art by using a multi-dimensional classification and dynamic prediction algorithm mechanism combining ROC curve analysis, principal component regression, logistic regression, and time series modeling.

[0006] To achieve the above object, the present invention provides the following technical solutions, a dynamic prediction method for gene-disease risk for clinical use and its application, including the following methods: Step S1, collect medical operation information data and user data; Step S2, analyze the importance of medical operation data according to the medical operation information data by using the ROC curve analysis method, classify the medical operation data according to the importance of the medical operation data, analyze the user permissions through the user data, and classify the users according to the user permissions; Step S3, analyze whether the user classification is accurate according to the user data by using the principal component regression method, classify and match the users determined to be accurately classified with the medical operation data of different classifications, and collect the matching feedback data and information sensitivity change data; Step S4: Update the user classification using the logistic regression method based on the user data and the matching feedback data, update the importance of the medical operation data using time series according to the information sensitivity change data, and adjust the medical operation data update time.

[0007] In a preferred embodiment, in step S1, the medical operation information data includes the proportion of repeated entries and the number of clicks on the classification keywords. The proportion of repeated entries is the proportion of the entries appearing in the medical entry database. The user data includes the search frequencies of different drug entries, the opening frequencies of the operation system, the opening duration of the operation software, and the number of system logins. The search frequencies of different drug entries are the frequencies of users searching different drug entries in the medical operation software. The opening frequency of the operation system is the frequency of users clicking into the operation system in the medical operation software. The opening duration of the operation software is the time period from the opening to the closing of the medical operation software.

[0008] In a preferred embodiment, in step S2, collect the proportions of repeated entries of N types of entries and combine them into an entry data set. Collect the number of clicks on N types of classification keywords, standardize them, and combine them into a keyword data set. Use the entry data set and the keyword coefficient data set to set the optimal entry data threshold and the optimal keyword coefficient data threshold using the ROC curve analysis method. The specific steps are as follows: Step 1: Calculate the data average value in the entry data set or the keyword coefficient data set as the classification benchmark to classify the data into positive example data or negative example data. Step 2: Calculate the true positive rate and the false positive rate of the entry data set or the keyword coefficient data set through multiple preset entry data thresholds or keyword coefficient data thresholds, and then screen out the optimal entry data threshold and the optimal keyword coefficient data threshold.

[0009] In a preferred embodiment, in step S2, if the proportion of the entries appearing in the current medical operation data exceeds the preset percentage of the optimal entry data threshold or the number of clicks on the classification keywords in the classification column corresponding to the current medical operation data exceeds the preset percentage of the optimal keyword coefficient data threshold, then determine that the current medical operation data is regular data and mark it. Otherwise, mark the current medical operation data as important data. The search frequencies of different drug entries and the opening frequency of the operation system are represented by the number of times users search different drug entries and the number of times users open the operation system. Collect the number of times different drug entries are searched and the number of times the operation system is opened by multiple users, and combine them into an entry search data set and a system opening data set respectively; calculate the data average value and the data standard deviation in the entry search data set and mark them as and , set the entry search threshold by calculating the data average value and data standard deviation obtained , where is the entry search threshold, a is the standard deviation proportionality coefficient and a > 0; If the number of searches for different drug entries by the current user exceeds the entry search threshold and the number of times the operation system is opened exceeds the system opening threshold, it is determined that the current user has high permissions, and the current user is marked as a high-permission user; otherwise, the current user is marked as a low-permission user.

[0010] In a preferred embodiment, in step S3, the accuracy coefficient of user permissions is calculated by using the principal component regression method based on the opening duration of the user's operation software and the number of system logins to evaluate the classification accuracy of the user. The specific steps are as follows: Select a sufficient number of users within a period of time as sample users, collect the opening duration of the operation software and the number of system logins of the sample users, perform normalization processing, and then merge them into a duration coefficient dataset and a login coefficient dataset; Arrange the data in the duration coefficient dataset or the login coefficient dataset from smallest to largest, select the middle data to set the principal component threshold ratio to determine the principal component threshold interval, set the principal component threshold interval ratio to M%, expand the middle data by M% to both sides to obtain the principal component threshold interval, mark the data in the duration coefficient dataset within the principal component threshold interval as principal component data, and mark other data as secondary component data; By setting component weights for the principal component data and the secondary component data and performing weighted calculations, the comprehensive duration coefficient and the comprehensive login coefficient are obtained, which are respectively marked as zs and zd. Take the average value of zs and zd as the accuracy coefficient of user permissions and set it as the accuracy coefficient threshold of user permissions, marked as .

[0011] In a preferred embodiment, in step S3, collect the opening duration of the operation software and the number of system logins of the current user, and calculate the comprehensive duration coefficient or the comprehensive login coefficient by using the principal component regression method; take the average value of the comprehensive duration coefficient and the comprehensive login coefficient of the current user as the accuracy coefficient of the current user's permissions and mark it as ; If , it is determined that the classification accuracy of the current user is high, and preset important data is matched for the current user; if , it is determined that the classification accuracy of the current user is low, and preset conventional data is matched for the current user.

[0012] In a preferred embodiment, in step S3, the matching feedback data is the number of times the matching data page is opened after data matching for the user, and the information sensitivity change data includes the change amount of the proportion of repeated entries, the change amount of the number of clicks on classification keywords, and the entry search change frequency; The number of times the matching data page is opened refers to the number of times the user views the matching data page within a preset time after data matching for the user; The change amount of the proportion of repeated entries is the difference between the proportion of an entry appearing in the medical entry database at a later time point and the proportion of the entry appearing in the medical entry database at a previous time point within a preset time period; The change amount of the number of clicks on classification keywords is the difference between the change amount of the number of clicks on keywords at a later time point and the change amount of the number of clicks on keywords at a previous time point within a preset time period; The change frequency of entry searches is the change speed of entry searches in the medical entry search system.

[0013] In a preferred embodiment, a period of time is selected as the user sample time. Within the user sample time, the sufficient number of times the user's matching data page is opened is collected and normalized and merged into an opening coefficient data set; The average value of the opening coefficient data is taken as the average opening coefficient. Combining with the user permission accuracy coefficient threshold, logistic regression is constructed to obtain the user update coefficient, and the user update coefficient calculated from the average value of the opening coefficient data is set as the user update threshold; A time with the same time interval as the user sample time is selected to collect the number of times the current user's matching data page is opened. After the same normalization process, it is used as the average opening coefficient. Combining with the user permission accuracy coefficient threshold, logistic regression is constructed to obtain the current user update coefficient; The current user update coefficient is compared with the user update threshold. If the current user update coefficient exceeds the user update threshold, it is determined that the change degree of the current user is high, and the current user is marked and switched; if the current user update coefficient is lower than the user update threshold, it is determined that the change degree of the current user is low, and the current user mark is maintained.

[0014] In a preferred embodiment, a period of time is selected as the data sample time. Within the data sample time, multiple adjacent time points are randomly selected and combined in pairs. The two time points within each combination are respectively used as the historical time point and the change time point; The proportion of repeated entries and the number of clicks on classification keywords are collected at the historical time point. The number of clicks on classification keywords is divided by the number of clicks on the keyword with the most clicks among the total classification keywords in the system for normalization to obtain the classification keyword coefficient. The proportion of repeated entries and the number of clicks on classification keywords are collected at the change time point, and the classification keyword coefficient is calculated based on the number of clicks on classification keywords; The difference between the proportion of repeated entries collected at the change time point and the proportion of repeated entries collected at the historical time point is used to obtain the change amount of the proportion of repeated entries. The difference between the classification keyword coefficient calculated at the change time point and the classification keyword coefficient collected at the historical time point is used to obtain the change amount of the number of clicks on classification keywords; The time series model adopted is the ARIMAX model. The specific steps for data-sensitive prediction analysis through the time series model are as follows: Step A1: Obtain the data for prediction; Step A2: Establish the ARIMAX model; Step A3: Use the maximum likelihood estimation (MLE) method to estimate the parameters of the ARIMAX model; Step A4: Verify the effect of the fitted model and check the goodness of fit of the model through the residual analysis method; Step A5: Use the fitted model to predict the future allocation plan, and take the maximum value of the prediction result as the highest level of importance of medical operation data in the future time period.

[0015] In a preferred embodiment, the exogenous variables in the ARIMAX model include the change amount of the proportion of repeated entries and the change amount of the number of clicks on the classification keywords. The basic form of the ARIMAX model is: ; In the formula, is the data sensitivity at the current time point, is the constant term, is the i-th order autoregressive parameter, and p is the order of the autoregressive term, is the j-th order moving average parameter, and q is the order of the moving average term, is the white noise term, representing the random error, is the exogenous variable is the coefficient of, and m is the lag order of the exogenous variable, is the exogenous variable lagged by k periods, is the i-th order autoregressive term of the target variable at the current time point t, is the prediction error term at the time point t−j; In step A3, 、 、 、 、 are calculated and obtained through the maximum likelihood estimation method; Adjust the update time of the medical operation data through the calculation result of multiplying the updated initial setting time by the change ratio.

[0016] The technical effects and advantages of the present invention: 1. The present invention collects medical operation information data and user data, analyzes user permissions using user data, classifies users according to user permissions. Analyzes the importance of medical operation data based on medical operation information data, assigns different tags to different medical operation data according to the importance of medical operation data and calculates the tag update time, matches the medical operation data marked with different tags with users of different permissions, collects matching feedback data and information sensitivity change data, and adjusts the tag update time and user permissions by comprehensively considering the matching feedback data and information sensitivity change data, thereby improving the flexibility of querying medical operation data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flowchart of the implementation of a clinical-oriented gene-disease risk dynamic prediction method of the present invention.

[0018] Figure 2 It is a schematic diagram of the steps of a clinical-oriented gene-disease risk dynamic prediction method of the present invention.

[0019] Figure 3 It is a flowchart of the application of a clinical-oriented gene-disease risk dynamic prediction method and its application of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] Embodiment 1 Please refer to Figures 1 to 2 , a clinical-oriented gene-disease risk dynamic prediction method and its application, and the specific operation process is as follows: Step S1, collect medical operation information data and user data.

[0022] Step S2, analyze the importance of medical operation data using the ROC curve analysis method based on medical operation information data, classify medical operation data according to the importance of medical operation data, analyze user permissions through user data, and classify users according to user permissions.

[0023] Step S3, analyze whether the user classification is accurate using the principal component regression method based on user data, classify and match the users determined to be accurately classified with medical operation data of different classifications, and collect matching feedback data and information sensitivity change data.

[0024] Step S4: Update the user classification using logistic regression based on the user data and the matching feedback data, and update the importance of the medical operation data using time series based on the information sensitivity change data and adjust the update time of the medical operation data.

[0025] Among them, the medical operation data includes genomic data, electronic medical record data, and operation log data; The specific implementation is as follows: In step S1, the medical operation information data includes the proportion of repeated terms and the number of clicks on the classification keywords. The proportion of repeated terms can be calculated by collecting term data from the medical term database, and the number of clicks on the classification keywords can be obtained by accessing the system server logs.

[0026] The proportion of repeated terms, that is, the proportion of a term appearing in the medical term database. The higher the frequency of a term appearing in the medical database, the more popular the term is, and it can be retrieved and queried by users, and its importance is relatively low.

[0027] The number of clicks on the classification keywords, that is, on the classification column of the medical operation data in the medical operation software or website. When the classification column is clicked, it will jump to the corresponding medical operation data content. The more times a certain classification keyword is clicked, the more popular the classification keyword is, and it can be retrieved and queried by users, and its importance is relatively low.

[0028] The user data includes the search frequencies of different drug terms, the opening frequencies of the operation system, the opening duration of the operation software, and the number of system logins. The above data can be directly obtained or calculated by accessing the system server logs.

[0029] The search frequencies of different drug terms refer to the frequencies of users searching different drug terms in the medical operation software. The higher the search frequencies of different drug terms, the more familiar the user is with the medical operation transmission, and the greater the user's permissions.

[0030] The opening frequency of the operation system refers to the frequency of users clicking into the operation system in the medical operation software. The higher the opening frequency of the user's operation system, the more familiar the current user is with the medical operation transmission, and the greater the user's permissions.

[0031] The opening duration of the operation software refers to the time period from the opening to the closing of the medical operation software. The longer the opening duration of the user's medical operation software, the greater the possibility that the user is a medical operation management personnel. The greater the user's permissions.

[0032] The number of system logins refers to the number of times the operation software is logged in. Medical operation management personnel use the medical operation software more frequently than ordinary users. The greater the number of system logins, the greater the possibility that the user is a medical operation management personnel. The greater the user's permissions.

[0033] It should be noted that the medical term database is a database system specifically for the medical and health field, which includes the collection of various terms and the proportion of various terms in the medical term database. The system server log refers to the recorded information generated during the operation of the operating system, application program or other software systems, which is used to record various activities and events of the system. The operation system startup frequency and user data can be obtained by using the system server.

[0034] In step S2, the proportion of duplicate terms obtained by accessing the medical term database is data in percentage form, and its value ranges from 0 to 1. The proportions of duplicate terms of N types of terms are collected and combined into a term dataset. The click counts of N types of classification keywords obtained by accessing the system server log are combined into a keyword dataset. Since the data in the keyword dataset are click counts, normalization processing is required, and its normalization formula can be: , where is the result after normalization of the data in the keyword dataset, is the normalized data in the keyword dataset, is the minimum data in the keyword dataset, is the maximum data in the keyword dataset. The N calculated normalized results are sequentially replaced and updated with the corresponding data in the keyword dataset, and the keyword coefficient dataset is obtained through the replacement and update.

[0035] The optimal term data threshold and the optimal keyword coefficient data threshold can be set by using the term dataset and the keyword coefficient dataset. The ROC curve analysis method can be used for setting. Taking the term dataset as an example, the specific steps are as follows: First, set a term data classification benchmark. The term data classification benchmark is a binary classification value, which can be set as the average value of the data in the term dataset. The data in the term dataset that exceeds the term data classification benchmark is marked as positive example data, and the data in the term dataset that is lower than the data classification benchmark is marked as negative example data. After the data in the term dataset are marked respectively, the true and false nature of the positive example data and negative example data are determined respectively by presetting the term data threshold multiple times, and the true positive rate and false positive rate are calculated. The optimal term data threshold is calculated by using the obtained true positive rate and false positive rate.

[0036] The positive example data and negative example data that exceed the term data threshold are used as true positive data and false negative data respectively and marked as TP and TN. The positive example data and negative example data that are lower than the term data threshold are used as false positive data and false negative data respectively and marked as FP and FN. The formulas for calculating the true interest rate and false positive rate are respectively: 、 , where TPR is the true positive rate and FPR is the false positive rate. By setting different entry data thresholds, different TPRs and FPRs are calculated. The multiple obtained TPRs are subtracted from largest to smallest in turn to obtain multiple true positive rate differences, which are respectively marked as TPR1, TPR2, TPR3, etc. Using the same operation, multiple false positive rate differences are calculated using the obtained multiple FPRs and are respectively marked as FPR1, FPR2, FPR3, etc. It should be noted that the true positive rate difference is obtained by subtracting the minimum value of the true positive rate TPR arranged from largest to smallest from 0, and the same applies to the false positive rate FPR.

[0037] Calculate the optimal entry data threshold through the formula using the obtained true positive rate difference and false positive rate difference. The formula is: , where c is the calculation result, i is the subscript of the true positive rate difference and the false positive rate difference, and n is the number of entry data thresholds set. The calculation result c is used as the optimal entry data threshold.

[0038] For example, there are 100 data in the entry data set. Among them, 60 data exceed the average value of the entry data set, and 40 data do not exceed the average value of the entry data set. That is, there are 60 positive example data and 40 negative example data. Set the first entry data threshold to 0.1. Then there are 54 true positive example data, 36 true negative example data, 6 false positive example data, and 4 false negative example data. That is, TP = 4, TN = 36, FP = 6, FN = 4. Calculate through the formula , ; Set the second entry data threshold to 0.2. Then TP = 448, TN = 32, FP = 12, FN = 8. The calculated , . When n = 2, that is, only two entry data thresholds are set. Calculate the true positive rate difference and the false positive rate difference , , , . Calculate the optimal entry data threshold .

[0039] It should be noted that the values of the entry data thresholds set and the number n of settings are not unique and can be adjusted in detail according to the actual situation.

[0040] The keyword coefficient data set calculates the optimal keyword coefficient data threshold using the same method as above and marks it as b. The larger the optimal entry data threshold c or the optimal keyword coefficient data threshold b, that is, the higher the proportion of repeated entries or the higher the click-through rate of classification keywords, the lower the importance of the corresponding medical operation data.

[0041] Classify the medical operation data by combining the percentage settings of the optimal entry data threshold and the optimal keyword coefficient data threshold. If the proportion of entries in the current medical operation data exceeds the percentage set by the optimal entry data threshold or the number of clicks on the classification keywords in the classification column corresponding to the medical operation data exceeds the percentage set by the optimal keyword coefficient data threshold, then determine that the current medical operation data is regular data and mark it. Otherwise, mark the current medical operation data as important data.

[0042] For example, if c = 0.8 and b = 0.7, then the percentage settings are 80% and 70% respectively. If the proportion of entries in the current medical operation data exceeds 80% of the entries or the number of clicks on the classification keywords in the classification column corresponding to the current medical operation data exceeds 70% of the number of clicks on the classification keywords, then determine that the current medical operation data is regular data.

[0043] The search frequencies of different drug entries and the opening frequencies of the operation system can be represented by obtaining the number of searches for different drug entries by users and the number of times the user operation system is opened by accessing the system server logs for a selected period of time.

[0044] By collecting a sufficient amount of the number of searches for different drug entries by users and the number of times the user operation system is opened, they are respectively combined into an entry search data set and a system opening data set. Calculate the data average value and data standard deviation in the entry search data set and mark them respectively as and , and set the entry search threshold through the calculated data average value and data standard deviation, which can be set as , where a is the standard deviation proportionality coefficient and a > 0, and the specific value of a can be set according to the actual situation. Similarly, use the system opening data set to set the system opening threshold and mark it as t.

[0045] Comprehensively judge the user permissions and classify the users based on the entry search threshold and the system opening threshold. If the number of searches for different drug entries by the current user exceeds the entry search threshold and the number of times the operation system is opened exceeds the system opening threshold, then determine that the current user has high permissions and mark the current user as a high-permission user; otherwise, mark the current user as a low-permission user.

[0046] It should be noted that the ROC curve analysis method is a commonly used tool for evaluating the performance of binary classification models. It can analyze the optimal classification threshold, and the optimal entry data threshold and the optimal keyword coefficient data threshold can be calculated using the ROC curve. The normalization formula used in the above steps and the setting methods of each threshold are not unique.

[0047] In step S3, before matching the users classified by permission analysis with the medical operation data classified by importance, the classification accuracy of users can be evaluated to further improve the matching accuracy. The classification accuracy of users can be evaluated by calculating the user permission accuracy coefficient using the principal component regression method based on the opening duration of the user's operation software and the number of system logins. The specific steps are as follows: Select a sufficient number of users within a period of time as sample users, and collect the opening duration of the operation software and the number of system logins of the sample users. In addition, to improve the quality of the sample users, the sample users do not include users with 0 system logins. It should be noted that some users do not log in when using the software system and can only use some functions of the software system. These users are called tourists and are excluded from the selection range of sample users.

[0048] After normalizing the opening duration of the operation software and the number of system logins of the sample users, they are merged into a duration coefficient dataset and a login coefficient dataset. The normalization method can use the Min-Max normalization method, which is not limited here.

[0049] Taking the duration coefficient dataset as an example, set the principal component threshold interval. Arrange the data in the duration coefficient dataset from smallest to largest, select the middle data, set the principal component threshold ratio to determine the principal component threshold interval, set the principal component threshold interval ratio to M%, and expand the middle data by M% to both sides to obtain the principal component threshold interval. The value of M% can be set according to the actual situation. For example, if the principal component threshold interval ratio is set to 40%, then expand from the middle data by 40% of the principal component threshold ratio to both sides, and the obtained principal component threshold interval covers 80% of the middle data. Mark the data in the duration coefficient dataset within the principal component threshold interval as principal component data, and other data as secondary component data.

[0050] The comprehensive duration coefficient is obtained by weighted calculation by setting the component weights for the principal component data and the secondary component data. The formula can be , where is the comprehensive duration coefficient, and are the principal component weight and the secondary component weight respectively, is the average value of the principal component data, is the average value of the secondary component data. Further, the principal component weight and the secondary component weight can be set according to the principal component threshold interval ratio multiplied by the scaling ratio. The formulas are respectively , , where e is the scaling ratio, which can be adjusted according to the actual situation, and f is the principal component threshold interval ratio.

[0051] The same method can be used to calculate the comprehensive login coefficient using the login coefficient dataset and mark it as zd. The average value of the calculated comprehensive duration coefficient zs and the comprehensive login coefficient zd is used as the user permission accuracy coefficient and is set as the user permission accuracy coefficient threshold and marked as .

[0052] After collecting the operation software startup duration and system login times of the current user and performing the same normalization operation as obtaining the duration coefficient dataset and the login coefficient dataset, the normalized coefficients are regarded as the only principal component data , with the principal component weight being 1, and the secondary component data and the secondary component weights are both set to 0 to calculate the comprehensive duration coefficient or the comprehensive login coefficient. The average value of the current user's comprehensive duration coefficient and the comprehensive login coefficient is used as the current user permission accuracy coefficient and is marked as .

[0053] The larger the current user permission accuracy coefficient, the more accurate the classification of the current user. If the current user permission accuracy coefficient exceeds the user permission accuracy coefficient threshold, that is , it is determined that the user classification accuracy of the current user is high, and the current user is classified and matched; if the current user permission accuracy coefficient is lower than the user permission accuracy coefficient threshold, that is , it is determined that the user classification accuracy of the current user is low, and regular data is matched for the current user.

[0054] Classification matching means that if the current user is a high-privilege user, the current user is given the permission to open important data and is allowed to view the content of important data; if the current user is a low-privilege user, regular data is matched for the current user, and matching regular data means that only the user is allowed to view the content of regular data.

[0055] The matching feedback data is the number of times the matching data page is opened after data matching for the user, which can be obtained by accessing the system server log. The information sensitivity change data includes the change amount of the proportion of repeated entries, the change amount of the number of clicks on classification keywords, and the change frequency of entry searches. The information sensitivity change data can be calculated by accessing the system server log and the historical data of the system server log.

[0056] The number of times the matching data page is opened is the number of times the user views the matching data page within a certain period after data matching for the user. The more times the page is viewed, the more accurate the user classification.

[0057] The change amount of the proportion of repeated entries is the difference between the proportion of an entry appearing in the medical entry database at a later time point and the proportion of the entry appearing in the medical entry database at a previous time point within a certain time period. The importance change of the data content corresponding to the entry can be analyzed based on the change amount of the proportion of repeated entries.

[0058] The change in the number of clicks on classification keywords, that is, the difference between the change in the number of clicks on keywords at a later time point and the change in the number of clicks on keywords at a previous time point within a certain time period. The importance change of the data content corresponding to the keyword can be analyzed according to the change in the number of clicks on classification keywords.

[0059] The change frequency of term searches is the change speed of term searches in the medical term search system. The faster the change speed of term searches, the more necessary it is to reclassify medical operation data to ensure the classification accuracy of important data and regular data. The update time of medical operation data can be calculated according to the change frequency of term searches.

[0060] It should be noted that the method for analyzing the classification accuracy of users in the above steps is not unique, and the formulas involved in the analysis method can be adjusted according to the actual situation. The historical data of the system server log is the historical record of system acquisitions and events recorded on the server. The proportion of repeated terms and the number of clicks on classification keywords at a certain past time of the system can be obtained by accessing the historical data of the system server log.

[0061] In step S4, a period of time is selected as the user sample time. The number of times the user matching data page is opened is collected within the user sample time. After normalizing the number of times the user matching data page is opened, it is merged into an open coefficient data set. The normalization formula is not unique and can be selected by oneself. The logistic regression method is used to set the user update threshold. Take the average value of the open coefficient data as the average open coefficient and mark it as j, and comprehensively consider the user permission accuracy coefficient threshold Construct the following logistic regression formula: , where YH is the user update coefficient, e is the natural logarithm base, z is an intermediate variable, and its calculation formula is: , w is the coefficient weight and w is between 0 and 1. The specific value can be set according to the actual situation and will not be analyzed further here. Set the user update coefficient calculated using the average open coefficient data j as the user update threshold.

[0062] Select a time with the same time interval as the user sample time to collect the number of times the current user matching data page is opened. After performing the same normalization process on the number of times the current user matching data page is opened, it is used as the average open coefficient j, and comprehensively consider the current user permission accuracy coefficient Construct the following logistic regression formula: , where Yh is the current user update coefficient, u is an intermediate variable, and its calculation formula can be: , where w is the coefficient weight, and its value is the same as the coefficient weight used to calculate the user update threshold.

[0063] If the current user update coefficient exceeds the user update threshold, it is determined that the change degree of the current user is high. Perform a marking switch for the current user; if the update coefficient of the current user is lower than the user update threshold, it is determined that the change degree of the current user is low, and the current user marking is maintained. Marking switch means that if the current user is marked as a high-privilege user, the marking of the current user is switched to a low-privilege user; if the current user is marked as a low-privilege user, the marking of the current user is switched to a high-privilege user.

[0064] Select a period of time as the data sample time, randomly select multiple adjacent time points within the data sample time for pairwise combination, and use the two time points within each combination as the historical time point and the change time point respectively. Collect the proportion of repeated entries and the number of clicks on classification keywords at the historical time point. By normalizing the number of clicks on classification keywords collected at the historical time point, the classification keyword coefficient can be obtained by dividing the number of clicks on classification keywords by the number of clicks on the keyword with the most clicks in the total classification keywords of the system. Similarly, collect the proportion of repeated entries and the number of clicks on classification keywords at the change time point, and calculate the classification keyword coefficient according to the number of clicks on classification keywords.

[0065] Subtract the proportion of repeated entries collected at the change time point from the proportion of repeated entries collected at the historical time point to obtain the change amount of the proportion of repeated entries, and subtract the classification keyword coefficient calculated at the change time point from the classification keyword coefficient collected at the historical time point to obtain the change amount of the number of clicks on classification keywords.

[0066] Perform data-sensitive prediction analysis on the obtained change amount of the proportion of repeated entries and the change amount of the number of clicks on classification keywords through a time series model, and further refine and optimize the data-sensitive classification scheme; It should be noted that the time series model used in this embodiment is the ARIMAX model. Refining and optimizing the data-sensitive classification scheme means that on the basis of combining the change amount of the proportion of repeated entries and the change amount of the number of clicks on classification keywords, through further analysis and adjustment, the data-sensitive classification scheme is made more scientific and reasonable, so as to achieve the dynamic update of the importance level of medical operation data and the precise management of its update time scheduling; Furthermore, the specific steps of performing data-sensitive prediction analysis through a time series model are as follows: Step A1, obtain the data for prediction; Step A2, establish an ARIMAX model; Step A3, use the maximum likelihood estimation (MLE) method to estimate the ARIMAX model parameters; Step A4, verify the effect of the fitting model, and check the goodness of fit of the model through the residual analysis method; Step A5: Use the fitted model to predict the future allocation plan, and take the maximum value of the prediction result as the highest level of importance of medical operation data in the future time period.

[0067] Specifically, the data for prediction includes the change amount of the proportion of repeated entries corresponding to each time point, the change amount of the click-through times of classification keywords, and data sensitive data; among them, the change amount of the proportion of repeated entries corresponding to each time point and the change amount of the click-through times of classification keywords are not elaborated here as they have been exemplified above. Data sensitive data refers to historical data sensitive data, and its data serves as the main variable of the time series. Furthermore, the basic form of the ARIMAX model is: ; In the formula, is the data sensitivity at the current time point, is the constant term, is the autoregressive parameter of the i-th order, p is the order of the autoregressive term, is the moving average parameter of the j-th order, q is the order of the moving average term, is the white noise term, representing random error, is the exogenous variable 's coefficient, m is the lag order of the exogenous variable, is the exogenous variable lagged by k periods, is the i-th order autoregressive term of the target variable at the current time point t, is the prediction error term at time point t−j; It should be noted that the exogenous variable part can integrate the influence of other relevant variables (such as the change amount of the proportion of repeated entries and the change amount of the click-through times of classification keywords) on data sensitivity; In step A3, , , , , are calculated and obtained through the maximum likelihood estimation method; the specific steps are as follows: The error term follows a normal distribution, then the likelihood function is: ; Take the logarithm of the likelihood function to obtain the log-likelihood function: ; By maximizing the log-likelihood function, the parameter estimates , , , , ; Dynamic data-sensitive classification strategy by predicting the maximum, minimum, and average values of the prediction results The following is an example of this embodiment: The maximum value of the prediction result represents the highest level of importance of medical operation data in the future time period. An advanced classification strategy is formulated. At the peak of data sensitivity in the predicted data (i.e., the time point close to or equal to ), it is marked as important data; According to the lowest value of the prediction result or the low valley interval and the average value of the prediction result , a low-level allocation strategy is formulated. At the low valley period of data sensitivity and the average value of the prediction result (i.e., the time point close to or lower than ), it is marked as regular data; In addition to collecting the proportion of repeated entries and the number of clicks on classification keywords to update the classification of medical operation data at the historical time points and change time points selected within the data sample time, the number of types of entry searches is also collected to calculate the change frequency of entry searches to adjust the update time of medical operation data.

[0068] The update time of medical operation data can be adjusted by the calculation result of multiplying the initial set time by the change ratio. The adjustment formula can be: , where is the adjusted update time, g is the change ratio, is the initial set time of the update. The initial set time of the update can be obtained by accessing the system update log. The following is a method for setting a change ratio g in this example: Mark the number of types of entry searches collected at the historical time point as o, mark the data of the number of types of entry searches collected at the change time point as p, and count the number of repeated types of entry searches at the historical time point and the change time point as q. The formula for calculating the change ratio g can be: , when , that is, the new number of types of entry searches at the change time point is more than the number of types of entry searches at the historical time point, the change of the number of types of entry searches is fast, and the time for updating the classification of medical operation data needs to be shortened. At this time, g < 1, ; conversely, when , that is, the new number of types of entry searches at the change time point is less than the number of types of entry searches at the historical time point, the change of the number of types of entry searches is slow, and the time for updating the classification of medical operation data can be lengthened.

[0069] It should be noted that the selection methods of the user sample time, data sample time, historical time points, and change time points are set by professionals in the field according to the actual situation. The above-mentioned various coefficient calculations and threshold setting methods are only examples, and the analysis methods are not unique.

[0070] Example 2 Please refer to Figure 3 , a dynamic gene-disease risk prediction method for clinical use and its application In a real scenario deployed in a certain tertiary hospital, the gene-disease risk prediction system runs integrated with multiple terminals through the hospital data center

[0071] Gene data is collected by the Illumina NovaSeq sequencer in the hospital's molecular diagnostic center. After generating FastQ format files, they are uploaded to the bioinformatics analysis platform deployed on the GPU server (NVIDIA A100) in the data center

[0072] The data is transmitted to the central NAS storage array through a 10 Gigabit fiber optic network, aligned using BWA-MEM, variants are identified using GATK and a VCF file is output. Subsequently, it is parsed by a Python script to construct a feature vector containing variant sites and CADD scores

[0073] Electronic medical record data is transmitted from the HIS system to the structured processing module through the FHIR interface. Running on an Intel Xeon server, the Med-BERT model is used to perform entity extraction and ICD-10 coding conversion on the text chief complaint, and the results are stored in the MongoDB database in JSON format

[0074] Wearable devices, such as Huawei medical wristbands, send physiological data in real time to the bedside IoT relay device via Bluetooth 5.0. This device is modified based on the Raspberry Pi 4B and pushes the data to the edge computing node via the MQTT protocol. The latter runs the Node-RED platform for preprocessing and writes the data into the InfluxDB time series database

[0075] Laboratory test data is synchronized to the intermediate database daily by the LIS system. The system pulls from Oracle using an ETL job, unifies units and standardizes using a Pandas script and stores it in the PostgreSQL database. Above All the above data sources are aggregated nightly, and the Spark cluster performs multi-source data fusion processing to generate a unified feature vector Subsequently, the sklearn library in Python is used to calculate the ROC curve. By traversing the correspondence between different features as classification variables and the true disease labels, their TPR and FPR are calculated and the ROC curve is plotted. Features with an AUC value greater than 0.7 are retained, and its corresponding formula is , and this integral is numerically implemented by the Trapezoid method

[0076] Finally, the remaining important features will be used as input features for the next principal component regression and logistic regression models

[0077] After completing the feature importance screening, the system enters the modeling stage. Set the daily modeling time, and the hospital's private cloud platform schedules Airflow tasks to call the modeling script in the Spark cluster to start the principal component regression analysis on the retained features.

[0078] This process is implemented by the Spark MLlib module. First, the feature matrix X (with dimensions n×m) is Z-score standardized, then the covariance matrix Σ is calculated, and the principal component direction vectors are obtained using the eigenvalue decomposition method. The number of required principal components is determined by the cumulative contribution rate, and the first k principal components Z that achieve a contribution rate of more than 95% are retained. Spark then performs logistic regression modeling on Z, with the target variable being the structured disease label Y, and uses the sigmoid function to establish a binary classification model; Furthermore, the model training uses the L-BFGS optimizer to iteratively update the weight parameter β in parallel in a distributed environment. After each round of iteration, the cross-entropy loss function is calculated through the validation set, and if the convergence condition is met, the training ends.

[0079] The training results (model weights, principal component load matrix, feature mapping table) are stored in the Ceph distributed object storage in Pickle format and registered to the hospital's AI model management platform, providing RESTful interfaces through Flask services for the front-end system to call.

[0080] After the model training is completed and saved, the system enters the user permission feature fusion and risk inference stage. The hospital's clinical information system (HIS) pushes the daily user behavior data at 8:00 every day, including account ID, accessed modules, operation frequencies, etc., which are collected by the edge computing gateway (deployed at the outpatient floor switching node) and transmitted in real-time to the permission portrait construction module in the hospital's private cloud.

[0081] This module runs on the NVIDIA Jetson AGX Xavier hardware platform and uses a PyTorch lightweight model to extract features from the user behavior data, outputting a permission vector U (with dimensions 1×l), where each dimension represents the access stability and sensitive data access frequency of the user under different system function modules.

[0082] The permission vector U is then transmitted to the inference engine through the gRPC interface. This engine runs on the TensorRT deployment node in the hospital's core business cloud cluster, loads the previously trained principal component logistic regression model, and performs risk level inference.

[0083] During the inference process, first, the daily medical data features X' corresponding to the user are transformed into principal component vectors Z through principal component transformation, and then merged with the permission vector U to construct a fusion vector F.

[0084] The fusion vector F is fed into a logistic regression model with a weight adjustment layer: where R represents whether there is unauthorized or potential risk in the user's access to this type of data. The inferred probability value P will be compared with a set risk threshold T (this threshold is dynamically adjusted based on the historical behavior average); If P > T, it is considered that there is an abnormal access.

[0085] The inference result is pushed to the hospital data security monitoring platform through the MQTT protocol and stored in the MongoDB distributed document database in JSON format, supporting asynchronous auditing and triggering of risk warnings.

[0086] When the monitoring platform receives the risk inference result, the system enters the feedback adjustment and dynamic permission policy update stage.

[0087] The abnormal access records in the MongoDB database are written into the policy control module in real time. This module is deployed in the OpenShift container cloud of the hospital data center and runs the microservice logic built based on SpringBoot.

[0088] In this module, first, the historical access database (stored in TimescaleDB) is called, and the average access frequency, abnormal trigger times, and the usage popularity of the corresponding module of this user in the past 30 days are queried according to the user account ID to form the historical behavior vector H; Next, the policy adjustment function performs a difference analysis on the current behavior and the historical behavior, and calculates the adjustment factor through the following deviation formula: ; In the formula, is the i-th dimension of the current behavior vector U, is the historical average, is a small constant to avoid division by zero. If the adjustment factor exceeds the preset threshold, it is judged that there is a significant deviation in the user's permission usage pattern. At this time, the system will automatically trigger the permission policy update process.

[0089] The policy update process calls the policy generation function node scheduled by Kubernetes. Inside the node, based on the permission level mapping table (cached in Redis) and the user's current risk level P, the access weight of its module is adjusted and a new policy JSON structure is generated.

[0090] Subsequently, the hospital IAM (Identity and Access Management) system is called through the RESTful API to dynamically adjust the permission scope of the role to which this user account belongs, and the operation log is written into Elasticsearch and visually displayed in Kibana.

[0091] After the permission policy update is completed, the system enters the iterative update stage of the time series model to enable the permission determination model to adapt to the evolution of user behavior.

[0092] In this stage, the NVIDIA Jetson Xavier NX computing unit deployed on the edge node performs preliminary calculations, and the results are transmitted back to the AI cluster of the central inference platform through the hospital's internal 5G network. The cluster is based on the TensorFlow Serving architecture and runs on a GPU-accelerated VMware vSphere private cloud.

[0093] The Jetson side first normalizes the access behavior vector of the user in the latest hour to construct a normalized sequence.

[0094] The normalized sequence is encapsulated in Tensor format and transmitted to the central AI platform through the gRPC protocol, and is input into the behavior prediction model constructed based on the time series.

[0095] The model predicts the risk score at the next moment according to the historical time window, and calculates the difference with the determined result of the real feedback to form an error term: If the error continuously exceeds the preset threshold for more than three times, the system automatically triggers the model fine-tuning process.

[0096] This process is scheduled by the MLflow platform, performs incremental training on the LSTM model using the latest sampled data, and records the parameter version control information.

[0097] The weights of the fine-tuned model are automatically stored in the MinIO distributed object storage and pushed back to the Jetson edge node through the GitOps process for the next round of local inference.

[0098] Finally, this update process maintains the timeliness prediction ability of the model for user behavior, and completes a complete adaptive closed-loop of behavior modeling under the constraint that the overall response time of the system is less than 3 seconds.

[0099] The above formulas are all dimensionless and take their numerical calculations. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters in the formulas are set by technicians in this field according to the actual situation.

[0100] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more sets of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0101] It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0102] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0103] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0104] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0105] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0106] In addition, in each embodiment of the present application, the functional units can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0107] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the essence of the technical solution of the present application, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0108] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A dynamic prediction method for gene-disease risk oriented to clinical practice, characterized in that: including the following methods: Step S1, collect medical operation information data and user data; Step S2, analyze the importance of medical operation data using the ROC curve analysis method based on the medical operation information data, classify the medical operation data according to the importance of the medical operation data, analyze the user permissions through the user data, and classify the users according to the user permissions; Step S3, analyze whether the user classification is accurate using the principal component regression method based on the user data, classify and match the users judged to be accurately classified with the medical operation data of different classifications, and collect the matching feedback data and information sensitivity change data; Step S4, update the user classification using the logistic regression method based on the user data and the matching feedback data, update the importance of the medical operation data using the time series based on the information sensitivity change data, and adjust the medical operation data update time.

2. A gene-disease risk dynamic prediction method for clinical use according to claim 1, characterized in that: In step S1, the medical operation information data includes the proportion of repeated entries and the number of clicks on the classification keywords. The proportion of repeated entries is the proportion of the entries appearing in the medical entry database; The user data includes the search frequencies of different drug entries, the opening frequencies of the operation system, the opening duration of the operation software, and the number of system logins; The search frequencies of different drug entries are the frequencies of users searching for different drug entries in the medical operation software. The opening frequency of the operation system is the frequency of users clicking into the operation system in the medical operation software. The opening duration of the operation software is the time period from the opening to the closing of the medical operation software.

3. A gene-disease risk dynamic prediction method for clinical use according to claim 2, characterized in that: In step S2, collect the proportions of repeated entries of N types of entries and merge them into an entry data set. Collect the number of clicks on N types of classification keywords, standardize them, and merge them into a keyword data set; Use the entry data set and the keyword coefficient data set to set the optimal entry data threshold and the optimal keyword coefficient data threshold using the ROC curve analysis method. The specific steps are as follows: Step 1: Calculate the average value of the data in the entry data set or the keyword coefficient data set as the classification benchmark to classify the data into positive example data or negative example data; Step 2: Calculate the true positive rate and the false positive rate of the entry data set or the keyword coefficient data set through multiple preset entry data thresholds or keyword coefficient data thresholds, and then screen out the optimal entry data threshold and the optimal keyword coefficient data threshold.

4. A gene-disease risk dynamic prediction method for clinical use according to claim 3, characterized in that: In step S2, if the proportion of the entries appearing in the current medical operation data exceeds the preset percentage of the optimal entry data threshold or the number of clicks on the classification keywords in the classification column corresponding to the current medical operation data exceeds the preset percentage of the optimal keyword coefficient data threshold, then judge that the current medical operation data is regular data and mark it, otherwise judge and mark the current medical operation data as important data; The search frequencies of different drug entries and the opening frequency of the operation system are represented by the number of times users search for different drug entries and the number of times users open the operation system. Collect the search times of different drug entries and the opening times of the operation system of multiple users, and merge them into an entry search data set and a system opening data set respectively; The data mean and data standard deviation in the computational entry search dataset are respectively labeled as and , and the data mean and data standard deviation obtained through calculation are used to set the entry search threshold , where is the entry search threshold, a is the standard deviation proportionality coefficient and a > 0; If the search times of different drug entries of the current user exceed the entry search threshold and the opening times of the operation system exceed the system opening threshold, it is determined that the current user has high permissions, and the current user is marked as a high-permission user; otherwise, the current user is marked as a low-permission user.

5. The method for dynamically predicting gene-disease risk for clinical use according to claim 2, wherein: In step S3, the accuracy coefficient of user permissions is calculated by using the main component regression method based on the opening duration of the user's operation software and the number of system logins to evaluate the classification accuracy of users. The specific steps are as follows: Select a sufficient number of users within a period of time as sample users, collect the opening duration of the operation software and the number of system logins of the sample users, perform normalization processing, and merge them into a duration coefficient data set and a login coefficient data set; Arrange the data in the duration coefficient data set or the login coefficient data set from small to large, select the middle data, set the main component threshold ratio to determine the main component threshold interval, set the main component threshold interval ratio to M%, expand the middle data by M% to both sides to obtain the main component threshold interval, mark the data in the duration coefficient data set within the main component threshold interval as main component data, and mark other data as sub-component data; By setting component weights for the principal component data and the secondary component data and performing weighted calculations, the comprehensive duration coefficient and the comprehensive login coefficient are obtained, which are respectively marked as zs and zd. The average value of zs and zd is taken as the user permission accuracy coefficient and is set as the user permission accuracy coefficient threshold, marked as .

6. The method for dynamically predicting gene-disease risk for clinical use according to claim 5, wherein: In step S3, collect the operation software startup duration and system login times of the current user, and use the principal component regression method to calculate the comprehensive duration coefficient or the comprehensive login coefficient; take the average of the current user's comprehensive duration coefficient and comprehensive login coefficient as the current user's permission accuracy coefficient and label it as ; If , it is determined that the user classification accuracy of the current user is high, and preset important data is matched for the current user; if , it is determined that the user classification accuracy of the current user is low, and preset regular data is matched for the current user.

7. The method for dynamically predicting gene-disease risk for clinical use according to claim 1, wherein: In step S3, the matching feedback data is the number of times the matching data page is opened after data matching for the user. The information sensitivity change data includes the change amount of the proportion of repeated entries, the change amount of the number of clicks on classification keywords, and the entry search change frequency; The number of times the matching data page is opened is the number of times the user browses the matching data page within a preset time after data matching for the user; The change amount of the proportion of repeated entries is the difference between the proportion of an entry appearing in the medical entry database at the latter time point and the proportion of the entry appearing in the medical entry database at the former time point within a preset time period; The change amount of the number of clicks on classification keywords is the difference between the change amount of the number of clicks on keywords at the latter time point and the change amount of the number of clicks on keywords at the former time point within a preset time period; The entry search change frequency is the entry search change speed in the medical entry search system.

8. The method for dynamically predicting gene-disease risk for clinical use according to claim 7, wherein: Select a period of time as the user sample time, collect the number of times the matching data page is opened for a sufficient number of users within the user sample time, and perform normalization processing and merge them into an opening coefficient data set; Take the average value of the opening coefficient data as the average opening coefficient, construct a logistic regression by combining the user permission accuracy coefficient threshold to obtain the user update coefficient, and set the user update coefficient calculated by using the average value of the opening coefficient data as the user update threshold; Collect the number of times the current user's matching data page is opened at time intervals identical to those of the user sample time, and after performing the same normalization process, use it as the average opening coefficient. Combine it with the user permission accuracy coefficient threshold to construct a logistic regression to obtain the current user update coefficient; Compare the current user update coefficient with the user update threshold. If the current user update coefficient exceeds the user update threshold, it is determined that the degree of change of the current user is high, and the current user is marked for switching; If the current user update coefficient is lower than the user update threshold, it is determined that the degree of change of the current user is low, and the current user's mark is maintained.

9. A gene-disease risk dynamic prediction method for clinical use according to claim 8, characterized in that: Select a period of time as the data sample time. Randomly select multiple adjacent time points within the data sample time for pairwise combination, and use the two time points within each combination as the historical time point and the changing time point respectively; Collect the proportion of repeated entries and the number of clicks on classification keywords at the historical time point. Divide the number of clicks on classification keywords by the number of clicks on the keyword with the most clicks among the total system classification keywords for normalization to obtain the classification keyword coefficient. Collect the proportion of repeated entries and the number of clicks on classification keywords at the changing time point, and calculate the classification keyword coefficient based on the number of clicks on classification keywords; Subtract the proportion of repeated entries collected at the historical time point from the proportion of repeated entries collected at the changing time point to obtain the change amount of the proportion of repeated entries. Subtract the classification keyword coefficient collected at the historical time point from the classification keyword coefficient calculated at the changing time point to obtain the change amount of the number of clicks on classification keywords; The time series model used is the ARIMAX model. The specific steps for performing data sensitivity prediction analysis through the time series model are as follows: Step A1, obtain the data for prediction; Step A2, establish an ARIMAX model; Step A3, use the maximum likelihood estimation (MLE) method to estimate the parameters of the ARIMAX model; Step A4, verify the effect of the fitted model, and check the goodness of fit of the model through the residual analysis method; Step A5, use the fitted model to predict the future allocation plan, and use the maximum value of the prediction result as the highest level of importance of medical operation data within the future time period.

10. A gene-disease risk dynamic prediction method for clinical use according to claim 9, characterized in that: The exogenous variables in the ARIMAX model include the change amount of the proportion of repeated entries and the change amount of the number of clicks on classification keywords. The basic form of the ARIMAX model is: ; Wherein, is the data sensitivity at the current time point, is the constant term, is the i-th order autoregressive parameter, p is the order of the autoregressive term, is the j-th order moving average parameter, q is the order of the moving average term, is the white noise term, representing the random error, is the exogenous variable is the coefficient of, m is the lag order of the exogenous variable, is the exogenous variable lagged by k periods, is the i-th order autoregressive term of the target variable at the current time point t, is the prediction error term at time point t−j; In step A3, , , , , are obtained by calculation through the maximum likelihood estimation method. Adjust the update time of the medical operation data by the calculation result of multiplying the initially set time by the change ratio.

Citation Information

Patent Citations

  • Construction method and application of thoracic disease detection model

    CN108898595A

  • Medical information data query method and system

    CN118035319A

  • Method and system for managing medical health care records of user

    CN120126782A

  • Detecting Early Symptoms And Providing Preventative Healthcare Using Minimally Required But Sufficient Data

    US20210407686A1

Cited By

  • Data security risk identification method and system based on big data

    CN121365410A

  • A Data Security Risk Identification Method and System Based on Big Data

    CN121365410B