A method for tracing the sample information of a kidney disease biobank
By building an FP tree and obtaining the optimal support threshold based on the reasonableness of the support threshold, the problem of inaccurate process prediction of abnormal links caused by inaccurate support threshold in the prior art is solved, and accurate prediction and risk avoidance of abnormal links of kidney disease biological samples are achieved.
Patent Information
- Application Number
- CN202510097228.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The artificially set support threshold in the prior art is inaccurate, resulting in inaccurate prediction of abnormal link processes and ineffective in reducing abnormal phenomena in kidney disease biological samples.
By obtaining the flow of abnormal kidney disease biological samples and building an FP tree, the FP tree is processed for any support threshold in the preset support threshold set, the maximum frequent item set is obtained, and the reasonableness of the support threshold is obtained based on the reference weight of the maximum frequent item set, and finally obtaining the optimal support threshold based on the reasonableness, accurately predicting the abnormal process link of the kidney disease biological samples.
By automatically obtaining the optimal support threshold and accurately predicting abnormal processes, it effectively reduces abnormal phenomena in kidney disease biological samples and improves the ability to deal with and avoid abnormal risks.
Smart Images

Figure CN119517442B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of predicting abnormal process links, and particularly relates to a method for tracing the sample information of a kidney disease biobank. Background Art
[0002] Kidney disease is a disease that affects renal function and structure. In order to develop drugs for the effective treatment of kidney disease, it is necessary to study kidney disease biological samples. During the process of studying kidney disease biological samples, it is necessary to ensure the effectiveness and normality of kidney disease biological samples. In actual situations, abnormal phenomena such as loss and contamination are likely to occur during the process of kidney disease biological samples from collection to research. In order to avoid abnormalities in kidney disease biological samples, it is necessary to trace the whole process of kidney disease biological samples from collection, distribution, use, and destruction, analyze the abnormal conditions of specific link processes, timely increase risk control, and effectively avoid the reasons for causing abnormalities in kidney disease biological samples.
[0003] In the existing method, the process information of abnormal kidney disease biological samples is traced through the FP tree (Frequent Pattern Tree), and the abnormal link process is predicted according to the artificially set support threshold, and timely processing and risk avoidance are made for the relevant abnormal transfer processes of kidney disease biological samples, reducing the abnormal phenomena of kidney disease biological samples. However, in actual situations, the artificially set support threshold is inaccurate. During the process of tracing the process information of abnormal kidney disease biological samples through the FP tree (Frequent Pattern Tree), too large or too small support thresholds will lead to inaccurate prediction of the abnormal link process, resulting in the inability to accurately predict the transfer process that is likely to cause abnormalities in kidney disease biological samples, and thus unable to accurately and timely handle and avoid abnormal risks, and unable to effectively reduce the abnormal phenomena of kidney disease biological samples. Summary of the Invention
[0004] In order to solve the technical problem that the inaccuracy of the artificially set support threshold leads to inaccurate prediction of the abnormal link process, and thus unable to effectively reduce the abnormalities in kidney disease biological samples, the purpose of the present invention is to provide a method for tracing the sample information of a kidney disease biobank, and the specific technical solution adopted is as follows:
[0005] An embodiment of the present invention provides a method for tracing the sample information of a kidney disease biobank, and the method includes the following steps:
[0006] Obtain the process of each abnormal kidney disease biological sample transfer and construct an FP tree; for any support threshold in the preset support threshold set, process the FP tree through the support threshold to obtain the maximum frequent item set;
[0007] Obtain the reference weight of each maximum frequent item set based on the values and depths of each node on the path corresponding to each maximum frequent item set in the FP tree;
[0008] According to the reference weight, obtain the rationality degree of the support threshold;
[0009] Based on the rationality degree, obtain the optimal support threshold, and predict the abnormal process links of kidney disease biological samples according to the optimal support threshold.
[0010] Furthermore, the method for obtaining the reference weight is as follows:
[0011] For any maximum frequent item set, take the path corresponding to this maximum frequent item set as the target path, divide the nodes in the target path according to the value size, and obtain the first category and the second category; among them, the values of the nodes in the first category are larger, and the values of the nodes in the second category are smaller;
[0012] According to the numerical difference between the nodes in the first category and the second category, obtain the similarity degree of the nodes on the target path;
[0013] Take the node with the largest value in the second category as the splitting node of the target path, and according to the numerical difference between the splitting node and other nodes at the same depth in the FP tree, as well as the number of all nodes at the depth where the splitting node is located in the FP tree, obtain the first abnormal possibility degree of the first category;
[0014] Take the node with the smallest value in the second category as the end point of the target path, and according to the fluctuation of the values of the nodes in the second category, as well as the similarity between the values of the nodes at the next depth where the end point is located in the FP tree and the values of the nodes in the second category, obtain the second abnormal possibility degree of the second category;
[0015] According to the similarity degree, the first abnormal possibility degree and the second abnormal possibility degree, obtain the overall abnormal degree of this maximum frequent item set;
[0016] According to the depth of the end point in the FP tree and the values of each node in the target path, obtain the importance degree of this maximum frequent item set;
[0017] Take the product of the overall abnormal degree and the importance degree as the reference weight of this maximum frequent item set.
[0018] Furthermore, the method for obtaining the similarity degree is as follows:
[0019] Obtain the mean value of the values of all nodes in the first category as the first reference value;
[0020] Obtain the mean value of the values of all nodes in the second category as the second reference value;
[0021] The result of negatively correlating and normalizing the difference between the first reference value and the second reference value is used as the similarity degree of the nodes on the target path.
[0022] Furthermore, the method for obtaining the first degree of abnormal possibility is as follows:
[0023] Taking the depth of the splitting node in the FP tree as the target depth, and obtaining the mean value of the numerical values of all other nodes except the splitting node at the target depth as the first eigenvalue;
[0024] Taking the difference between the numerical value of the splitting node and the first eigenvalue as the reference degree of abnormal for the first category;
[0025] Taking the result of normalizing the product of the reciprocal of the number of all nodes at the target depth and the reference degree of abnormal as the first degree of abnormal possibility for the first category.
[0026] Furthermore, the method for obtaining the second degree of abnormal possibility is as follows:
[0027] Taking the result of negatively correlating and normalizing the standard deviation of the numerical values of the nodes in the second category as the initial degree of abnormal possibility for the second category;
[0028] Obtaining the variance of the numerical values of the nodes in the second category as the first variance;
[0029] Obtaining the variance of the numerical values of all nodes at the next depth of the depth where the end point is located in the FP tree and the numerical values of the nodes in the second category as the second variance;
[0030] Taking the result of normalizing the absolute value of the difference between the first variance and the second variance as the correction weight for the second category;
[0031] Taking the product of the correction weight for the second category and the initial degree of abnormal possibility as the second degree of abnormal possibility for the second category.
[0032] Furthermore, the method for obtaining the overall degree of abnormal is as follows:
[0033] Obtaining the mean value of the first degree of abnormal possibility and the second degree of abnormal possibility as the overall degree of abnormal possibility of this maximum frequent item set;
[0034] Taking the product of the similarity degree and the overall degree of abnormal possibility as the overall degree of abnormal of this maximum frequent item set.
[0035] Furthermore, the method for obtaining the importance degree is as follows:
[0036] The ratio of the mean value of the numerical values of all nodes in the target path to the depth of the end point in the FP tree is used as the importance degree of the maximum frequent item set.
[0037] Further, the method for obtaining the reasonable degree is as follows:
[0038] The mean value of the reference weights of all maximum frequent item sets is used as the reasonable degree of the support threshold.
[0039] Further, the method for obtaining the optimal support threshold is as follows:
[0040] The support threshold corresponding to the maximum reasonable degree is used as the optimal support threshold.
[0041] Further, the method for obtaining the abnormal process link is as follows:
[0042] The FP tree is processed by the optimal support threshold, and each maximum frequent item set obtained at this time is used as the target item set;
[0043] The process of each normal kidney disease biological sample transfer is obtained and an FP tree is constructed as the reference FP tree. The reference FP tree is processed by the optimal support threshold, and each maximum frequent item set obtained at this time is used as the reference item set;
[0044] Each target item set is compared with each reference item set, and the process link corresponding to the target item set that is not the same as the reference item set is used as the abnormal process link of the kidney disease biological sample.
[0045] The present invention has the following beneficial effects:
[0046] The present invention constructs an FP tree according to the process of abnormal kidney disease biological sample transfer, avoiding the interference of normal kidney disease biological samples, which is conducive to efficiently analyzing the process causing the abnormality of kidney disease biological samples. In order to obtain a reasonable support threshold to accurately predict the process links causing the abnormality of kidney disease biological samples, for any support threshold in the preset support threshold set, the FP tree is processed by this support threshold to obtain the maximum frequent item set, avoiding the analysis of repeated process links. In order to accurately analyze whether this support threshold is reasonable, based on the values and depths of each node of the path corresponding to each maximum frequent item set in the FP tree, the reference weight of each maximum frequent item set is obtained, accurately reflecting the possibility that the process link corresponding to each maximum frequent item set causes the abnormality of kidney disease biological samples, which is conducive to accurately analyzing whether this support threshold is reasonable subsequently. Therefore, the reasonable degree of this support threshold is accurately obtained according to the reference weight, accurately reflecting the rationality of this support threshold. Furthermore, the optimal support threshold is obtained based on the reasonable degree, making the finally analyzed process links causing the abnormality of kidney disease biological samples more accurate. Then, the abnormal process links of kidney disease biological samples are accurately predicted according to the optimal support threshold, which is conducive to timely handling and avoiding abnormal risks, and effectively reducing the abnormal phenomenon of kidney disease biological samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0048] Figure 1 Schematic flowchart of a method for tracing sample information of a kidney disease biological sample library provided by an embodiment of the present invention;
[0049] Figure 2 Flowchart of a method for obtaining a reference weight provided by an embodiment of the present invention;
[0050] Figure 3 Structural diagram of a system for tracing sample information of a kidney disease biological sample library provided by an embodiment of the present invention;
[0051] Figure 4 Schematic diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following specifically describes, in conjunction with the accompanying drawings and preferred embodiments, a method for tracing the sample information of a kidney disease biobank according to the present invention, including its specific implementation manner, structure, features, and effects, as follows. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs.
[0054] The following specifically describes, in conjunction with the accompanying drawings, the specific solution of a method for tracing the sample information of a kidney disease biobank provided by the present invention.
[0055] Embodiment 1:
[0056] The present invention proposes a method for tracing the sample information of a kidney disease biobank. Please refer to Figure 1 , which shows a schematic flowchart of a method for tracing the sample information of a kidney disease biobank provided by an embodiment of the present invention. The method includes the following steps:
[0057] Step S1: Obtain the flow of each abnormal kidney disease biological sample and construct an FP tree; for any support threshold in the preset support threshold set, process the FP tree through this support threshold to obtain the maximum frequent item set.
[0058] Specifically, kidney disease biological samples are first collected in a hospital, then transported to a dedicated storage center, and finally uniformly distributed to various research institutes or academies. The entire transportation process includes multiple process links. The storage center needs to perform abnormal detection on the kidney disease biological samples first to ensure that the kidney disease biological samples distributed to the research institutes or academies are normal, thus guaranteeing the significance of the research.
[0059] Since kidney disease biological samples may be damaged and become abnormal during transportation due to environmental factors such as temperature and humidity, as well as unreasonable transportation methods, in this embodiment, a batch of kidney disease biological samples is selected, and the transfer process of each kidney disease biological sample in this batch of kidney disease biological samples is traced. After arriving at the storage center, each kidney disease biological sample in this batch of kidney disease biological samples is subjected to abnormal detection. Kidney disease biological samples with obvious DNA or protein abnormalities are all marked as abnormal kidney disease biological samples. The number of kidney disease biological samples in this batch of kidney disease biological samples selected in this embodiment is set to 1000. The implementer can set the number of this batch of kidney disease biological samples according to the actual situation, and no limitation is made here. Among them, a unique identifier, that is, a unique ID, is assigned to each kidney disease biological sample.
[0060] In order to analyze the process links that are likely to cause abnormalities in kidney disease biological samples, and considering the association between different processes, in this embodiment, abnormal kidney disease biological samples are first obtained, and then according to each process through which each abnormal kidney disease biological sample flows, an FP tree is constructed by the FP algorithm (Frequent Pattern Mining). After that, by analyzing the FP tree, the process links that are likely to cause abnormalities in kidney disease biological samples are predicted. It should be noted that the process links that are likely to cause abnormalities in kidney disease biological samples are composed of multiple related processes. Among them, the FP algorithm (Frequent Pattern Mining) is a well-known technology and will not be traced back.
[0061] Obtaining frequent item sets by setting a support threshold for the FP tree is actually pruning the FP tree. Different support thresholds represent pruning different positions of the FP tree. Therefore, different support thresholds correspond to different frequent item sets, and judging the association between different processes depends on the frequent item sets, thereby predicting the process links that are likely to cause abnormalities in kidney disease biological samples. It is known that too large or too small support thresholds will result in inaccurate prediction of the process links that are likely to cause abnormalities in kidney disease biological samples. Therefore, in this embodiment, a set of preset support thresholds is first set, and then the most reasonable support threshold in the set of preset support thresholds is analyzed, and the FP tree is processed to more accurately predict the process links that are likely to cause abnormalities in kidney disease biological samples. It is known that the value range of the support threshold is from 0 to 1. In order to improve the efficiency of obtaining the most reasonable support threshold, the set of preset support thresholds in this embodiment is set to , and the implementer can set the set of preset support thresholds according to the actual situation, and no limitation is made here.
[0062] For clear analysis, in this embodiment, each support threshold in the preset support threshold set is analyzed separately. That is, for any support threshold in the preset support threshold set, the FP tree is processed by this support threshold to obtain the maximum frequent item set corresponding to this support threshold. The purpose of obtaining the maximum frequent item set in this embodiment is to avoid repeated analysis of the processes associated with abnormal kidney disease biological samples. Therefore, each maximum frequent item set is analyzed in this embodiment. Among them, the method for obtaining the maximum frequent item set is a well-known technology and will not be elaborated here. It should be noted that the maximum frequent item sets analyzed subsequently are all those corresponding to this support threshold.
[0063] Step S2: Based on the values and depths of each node in the path corresponding to each maximum frequent item set in the FP tree, obtain the reference weight of each maximum frequent item set.
[0064] It is known that in the FP tree, the leaf nodes indicate that a specific item appears only together with other items in the path of the corresponding frequent item set in the frequent item set. Therefore, the item corresponding to the leaf node, that is, the process appears less frequently in the processes through which the abnormal kidney disease biological sample flows. When the FP tree is pruned by the set support threshold, the leaf nodes will be pruned first, indirectly reflecting that the process corresponding to the node with a greater depth, that is, the node farther away from the root node of the FP tree, is less likely to be associated with other processes, and the number of times the corresponding process appears in the processes passed by the abnormal kidney disease biological sample is less, indirectly reflecting that the possibility of the corresponding process causing the abnormality of the kidney disease biological sample is smaller. Therefore, when the depth of the path corresponding to a certain maximum frequent item set in the FP tree is greater, that is, the farther away from the root node of the FP tree, the fewer times the process link of the path corresponding to this maximum frequent item set appears in the process links passed by the abnormal kidney disease biological sample, and the smaller the possibility that the process link of the path corresponding to this maximum frequent item set causes the abnormality of the kidney disease biological sample.
[0065] Considering that the larger the value of the node in the FP tree, the more times the process corresponding to this node appears in the processes passed by the abnormal kidney disease biological sample, indirectly reflecting that the possibility of the process corresponding to this node causing the abnormality of the kidney disease biological sample is greater. Therefore, when the values of the nodes in the path corresponding to a certain maximum frequent item set in the FP tree are all larger, the more times the process link of the path corresponding to this maximum frequent item set appears in the process links passed by the abnormal kidney disease biological sample, and the greater the possibility that the process link of the path corresponding to this maximum frequent item set causes the abnormality of the kidney disease biological sample.
[0066] Therefore, in this embodiment, based on the values and depths of each node in the path corresponding to each maximum frequent item set in the FP tree, the reference weight of each maximum frequent item set is obtained. Among them, the greater the reference weight, the more likely the process link corresponding to the path of the corresponding maximum frequent item set is the process link that causes abnormalities in kidney disease biological samples.
[0067] Preferably, in an implementable manner of this embodiment, for the method of obtaining the reference weight, please refer to Figure 2 , which shows a flowchart of a method for obtaining a reference weight provided by this embodiment. The method includes the following steps:
[0068] Step S201: For any maximum frequent item set, take the path corresponding to the maximum frequent item set as the target path, and divide the nodes in the target path according to the value size to obtain a first category and a second category; among them, the values of the nodes in the first category are larger, and the values of the nodes in the second category are smaller.
[0069] In this embodiment, the value of K in the K-means clustering algorithm is set to 2, and the nodes in the target path are divided according to the value size of the nodes through the K-means clustering algorithm to obtain a first category and a second category. Among them, the K-means clustering algorithm is a well-known technology and will not be elaborated. Among them, the larger values of the nodes in the first category indicate that the nodes in the first category are frequently shared in the path, that is, the processes corresponding to the nodes in the first category appear frequently in the processes passed by abnormal kidney disease biological samples; the smaller values of the nodes in the second category indicate that the processes corresponding to the nodes in the second category appear less frequently in the processes passed by abnormal kidney disease biological samples.
[0070] It should be noted that if the values of all nodes in the target path are the same, then according to the depths of the nodes in the target path, the nodes in the target path are classified into categories through the K-means clustering algorithm to obtain a first category and a second category. Among them, the depths of the nodes in the first category are smaller, and the depths of the nodes in the second category are larger.
[0071] Step S202: Obtain the similarity degree of the nodes on the target path according to the numerical difference between the nodes in the first category and the second category.
[0072] When the values of the nodes on the target path are more equal, it indicates that the correlation between the processes in the maximum frequent item set corresponding to the target path is stronger, indirectly indicating that the process links corresponding to the target path are the process links frequently passed by abnormal kidney disease biological samples; when the difference in the values of the nodes on the target path is greater, it indirectly indicates that the process links corresponding to the target path include some process links that abnormal kidney disease biological samples do not pass through, that is, they are not the process links frequently passed by abnormal kidney disease biological samples. In predicting the process links that cause abnormalities in kidney disease biological samples, the possibility that the process links corresponding to the target path cause abnormalities in kidney disease biological samples is smaller. Therefore, in this embodiment, according to the numerical difference between the nodes in the first category and the second category, the similarity degree of the nodes on the target path is obtained. The greater the similarity degree, the greater the possibility that the process link corresponding to the target path causes abnormalities in kidney disease biological samples.
[0073] Preferably, in an implementable manner of this embodiment, the method for obtaining the similarity degree is as follows: obtain the mean value of the values of all nodes in the first category as the first reference value; obtain the mean value of the values of all nodes in the second category as the second reference value; when the first reference value and the second reference value are more equal, it indicates that the values of the nodes on the target path are more equal, and the process corresponding to the nodes on the target path is more likely to be the process jointly passed by abnormal kidney disease biological samples, indirectly indicating that the possibility that the process link corresponding to the target path causes abnormalities in kidney disease biological samples is greater. Then, obtain the result of negative correlation and normalization of the difference between the first reference value and the second reference value as the similarity degree of the nodes on the target path. Among them, the difference between the first reference value and the second reference value must be greater than 0.
[0074] Among them, the calculation formula for the similarity degree is: ; in the formula, S is the similarity degree of the nodes on the target path; is the first reference value; is the second reference value; exp is the exponential function with the natural constant as the base.
[0075] Step S203: Use the node with the largest value in the second category as the splitting node of the target path, and obtain the first abnormal possibility degree of the first category according to the numerical difference between the splitting node and other nodes at the same depth in the FP tree, and the number of all nodes at the depth where the splitting node is located in the FP tree.
[0076] In actual situations, due to issues such as temperature, humidity, and sealing, abnormal kidney disease biological samples may occur in different process links. Therefore, when the similarity degree of nodes on the target path is small, some process links corresponding to the target path may also cause abnormalities in kidney disease biological samples, indirectly indicating that the process links corresponding to the target path may cause abnormalities in kidney disease biological samples. The values of the known nodes change in ascending order of path depth. Therefore, the first category and the second category in step S201 actually segment the target path. In order to more accurately analyze whether the process links corresponding to the target path have an impact on the abnormalities of kidney disease biological samples, in this embodiment, the first category and the second category are analyzed respectively to determine the possibility that the partial process links corresponding to each stage of the target path cause abnormalities in kidney disease biological samples.
[0077] In this embodiment, first, the node with the largest value in the second category is used as the segmentation node of the target path, that is, the target path is divided into two segments. When the value of the segmentation node is larger, it indicates that the values of the nodes in the first category are larger, indirectly reflecting that the process links corresponding to the first category of paths are more likely to be the process links frequently passed by abnormal kidney disease biological samples. Considering that if faced with abnormal kidney disease biological samples of different scales, it will lead to large differences in the values of the nodes, so there will be errors in only analyzing the values of the nodes. However, when the value of the segmentation node is greater than the values of other nodes at the same depth as the segmentation node in the FP tree, and the number of nodes at the depth where the segmentation node is located in the FP tree is smaller, it can accurately indicate that the path corresponding to the first category has a higher node sharing situation compared to other paths, indirectly reflecting that the process links corresponding to the first category of paths are more likely to be the process links frequently passed by abnormal kidney disease biological samples. Therefore, in this embodiment, according to the numerical difference between the segmentation node and other nodes at the same depth as it in the FP tree, and the number of all nodes at the depth where the segmentation node is located in the FP tree, the first abnormal possibility degree of the first category is obtained. The greater the first abnormal possibility degree, the greater the possibility that the process links corresponding to the first category of paths are the process links frequently passed by abnormal kidney disease biological samples.
[0078] Preferably, in an implementable manner of this embodiment, the method for obtaining the first anomaly possibility degree is as follows: taking the depth of the splitting node in the FP tree as the target depth, obtaining the average value of the numerical values of all other nodes except the splitting node at the target depth as the first eigenvalue; taking the difference between the numerical value of the splitting node and the first eigenvalue as the reference anomaly degree of the first category; the greater the reference anomaly degree, the higher the node sharing situation of the path corresponding to the first category compared to other paths. At the same time, when the number of all nodes at the target depth is smaller, it indirectly indicates that the path corresponding to the first category has a higher node sharing situation. Furthermore, taking the product of the reciprocal of the number of all nodes at the target depth and the reference anomaly degree and normalizing the result as the first anomaly possibility degree of the first category.
[0079] Among them, the calculation formula for the first anomaly possibility degree is: ; in the formula, q is the first anomaly possibility degree of the first category; n is the number of all nodes at the target depth; E is the numerical value of the splitting node; is the first eigenvalue; is the reference anomaly degree; norm is the normalization function.
[0080] Step S204: Taking the node with the smallest numerical value in the second category as the end point of the target path, and obtaining the second anomaly possibility degree of the second category according to the fluctuation of the numerical values of the nodes in the second category and the similarity between the numerical values of the nodes at the next depth of the depth where the end point is located in the FP tree and the numerical values of the nodes in the second category.
[0081] The process corresponding to the nodes in the second category can be defaulted to the process passed by some abnormal kidney disease biological samples. If the kidney disease biological samples are abnormal due to the process corresponding to the nodes in the second category, these processes should also be obtained through the FP tree; at this time, although the process corresponding to the nodes in the second category is not the process passed by all abnormal kidney disease biological samples, it also reflects the obvious abnormal characteristics of some abnormal kidney disease biological samples.
[0082] In order to analyze the possibility that the process links corresponding to some paths of the second category cause the abnormality of kidney disease biological samples and avoid missing the process links that cause the abnormality of kidney disease biological samples, first take the node with the smallest numerical value in the second category as the end point of the target path. When the numerical value of the end point has a high similarity with the numerical values of other nodes in the second category, it indicates that the process corresponding to the end point and the processes corresponding to other nodes in the second category always appear together. The stronger the correlation between the processes corresponding to some paths of the second category, the greater the possibility that the process links corresponding to some paths of the second category jointly cause the abnormality of kidney disease biological samples.
[0083] On the other hand, if the value of the next node of the end point has a high similarity with the values of the nodes in the second category, it indicates that there is also a high correlation between the process corresponding to the next node not included in the target path and the process of the target path. Indirectly, it shows that the determination of the associated processes obtained from the maximum frequent item set corresponding to the target path is incomplete, and further reflects that the process links corresponding to the path of the second category are less likely to be the process links causing abnormalities in kidney disease biological samples.
[0084] Therefore, in this embodiment, according to the fluctuation of the values of the nodes in the second category and the similarity between the values of the nodes at the next depth of the depth where the end point is located in the FP tree and the values of the nodes in the second category, the second abnormal possibility degree of the second category is obtained. The greater the second abnormal possibility degree, the greater the possibility that the process links corresponding to the path of the second category are the process links causing abnormalities in kidney disease biological samples.
[0085] Preferably, in an implementable manner of this embodiment, the method for obtaining the second abnormal possibility degree is as follows: The result of performing negative correlation and normalization on the standard deviation of the values of the nodes in the second category is used as the initial abnormal possibility degree of the second category; the greater the initial abnormal possibility degree, the more equal the values of the nodes in the second category, indirectly indicating that the processes corresponding to the path of the second category are more likely to jointly cause abnormalities in kidney disease biological samples. In order to analyze the integrity of the maximum frequent item set corresponding to the target path, the variance of the values of the nodes in the second category is obtained as the first variance; the variance of the values of all the nodes at the next depth of the depth where the end point is located in the FP tree and the values of the nodes in the second category is obtained as the second variance; when the first variance and the second variance are more equal, it indicates that the maximum frequent item set corresponding to the target path is less complete, indirectly indicating that the path corresponding to the second category is less complete, and the possibility that the process links corresponding to the path of the second category jointly cause abnormalities in kidney disease biological samples is smaller. Furthermore, in this embodiment, the result of normalizing the absolute value of the difference between the first variance and the second variance is used as the correction weight of the second category. The greater the correction weight, the more complete the path corresponding to the second category, and the greater the possibility that the process links corresponding to the path of the second category jointly cause abnormalities in kidney disease biological samples; furthermore, the product of the correction weight of the second category and the initial abnormal possibility degree is used as the second abnormal possibility degree of the second category.
[0086] Among them, the calculation formula for the second abnormal possibility degree is: ; in the formula, p is the second abnormal possibility degree of the second category; is the first variance; is the second variance; is the absolute value function; is the correction weight of the second category; is the standard deviation of the values of the nodes in the second category; The initial abnormal possibility degree for the second category; norm is the normalization function; exp is the exponential function with the natural constant as the base.
[0087] Step S205: Obtain the overall abnormal degree of this maximum frequent item set according to the similarity degree, the first abnormal possibility degree, and the second abnormal possibility degree.
[0088] It is known that the greater the similarity degree, the greater the possibility that the process link corresponding to the target path causes abnormalities in kidney disease biological samples; the greater the first abnormal possibility degree, the greater the possibility that the process link corresponding to the first category path is a frequently passed process link for abnormal kidney disease biological samples, indirectly indicating that the process link corresponding to the target path is more likely to cause abnormalities in kidney disease biological samples; the greater the second abnormal possibility degree, the greater the possibility that the process links corresponding to the second category path jointly cause abnormalities in kidney disease biological samples, also indirectly indicating that the process link corresponding to the target path is more likely to cause abnormalities in kidney disease biological samples; therefore, in this embodiment, the overall abnormal degree of this maximum frequent item set is obtained according to the similarity degree, the first abnormal possibility degree, and the second abnormal possibility degree. The greater the overall abnormal degree, the more likely the process in this maximum frequent item set is the process that causes abnormalities in kidney disease biological samples.
[0089] Preferably, in a feasible implementation manner of this embodiment, the method for obtaining the overall abnormal degree is: obtain the average value of the first abnormal possibility degree and the second abnormal possibility degree as the overall abnormal possibility degree of this maximum frequent item set; multiply the similarity degree by the overall abnormal possibility degree as the overall abnormal degree of this maximum frequent item set.
[0090] Among them, the calculation formula for the overall abnormal degree is: ; in the formula, r is the overall abnormal degree of this maximum frequent item set; S is the similarity degree of the nodes on the target path; q is the first abnormal possibility degree of the first category; p is the second abnormal possibility degree of the second category; is the overall abnormal possibility degree of this maximum frequent item set.
[0091] Step S206: Obtain the importance degree of this maximum frequent item set according to the depth of the end point in the FP tree and the values of each node in the target path.
[0092] When the depth of the target path is greater, it indicates that there are some processes with relatively low relevance in the maximum frequent item set. The processes in this maximum frequent item set cannot represent the processes commonly passed by abnormal kidney disease biological samples, and the importance of this maximum frequent item set is smaller. When the node value on the target path of this maximum frequent item set is larger, it indicates that the processes in this maximum frequent item set are more likely to be the processes commonly passed by abnormal kidney disease biological samples. The greater the possibility that the processes in this maximum frequent item set cause abnormalities in kidney disease biological samples, the greater the importance of this maximum frequent item set. Therefore, in this embodiment, according to the depth of the end point in the FP tree and the value of each node in the target path, the importance degree of this maximum frequent item set is obtained.
[0093] Preferably, in an implementable manner of this embodiment, the method for obtaining the importance degree is: taking the ratio of the mean value of the values of all nodes in the target path to the depth of the end point in the FP tree as the importance degree of this maximum frequent item set.
[0094] Step S207: Taking the product of the overall abnormality degree and the importance degree as the reference weight of this maximum frequent item set.
[0095] If there are fewer processes in this maximum frequent item set, that is, the target path corresponding to this maximum frequent item set is shorter. After classifying the target path, the number of nodes in each category is relatively small, so that the calculated overall abnormality degree of this maximum frequent item set cannot truly reflect the abnormal processes corresponding to abnormal kidney disease biological samples. Therefore, in this embodiment, in combination with the importance degree of this maximum frequent item set, that is, taking the product of the overall abnormality degree and the importance degree of this maximum frequent item set as the reference weight of this maximum frequent item set.
[0096] So far, the reference weight of each maximum frequent item set is obtained.
[0097] Step S3: According to the reference weight, obtain the rationality degree of this support threshold.
[0098] In order to comprehensively analyze the possibility that the process links corresponding to the maximum frequent item set obtained through this support threshold cause abnormalities in kidney disease biological samples, furthermore, taking the mean value of the reference weights of all maximum frequent item sets as the rationality degree of this support threshold. The greater the rationality degree, the more reasonable this support threshold is, and the more likely the process links corresponding to this support threshold and the maximum frequent item set are to cause abnormalities in kidney disease biological samples.
[0099] So far, the rationality degree of each support threshold is obtained.
[0100] Step S4: Based on the rationality degree, obtain the optimal support threshold, and predict the abnormal process links of kidney disease biological samples according to the optimal support threshold.
[0101] Specifically, in order to avoid missing the process links that may cause abnormalities in kidney disease biological samples, in this embodiment, the support threshold corresponding to the maximum reasonable degree is used as the optimal support threshold. The FP tree constructed for the process of abnormal transfer of kidney disease biological samples is processed by the optimal support threshold, and each maximum frequent item set obtained at this time is used as a target item set; among them, the process links corresponding to the paths of each target item set are the process links that may cause abnormalities in kidney disease biological samples.
[0102] In order to accurately determine the abnormal process links that cause abnormalities in kidney disease biological samples, in this embodiment, the process of transfer of each normal kidney disease biological sample in this batch of kidney disease biological samples is further obtained and an FP tree is constructed as a reference FP tree. The reference FP tree is processed by the optimal support threshold, and each maximum frequent item set obtained at this time is used as a reference item set; then each target item set is compared with each reference item set, and the process links corresponding to the target item sets that are not the same as the reference item sets are used as the abnormal process links of the kidney disease biological samples.
[0103] So far, the abnormal process links that are likely to cause abnormalities in kidney disease biological samples are accurately predicted. Before the kidney disease biological samples pass through the predicted abnormal process links, the environment such as the temperature and humidity of the kidney disease biological samples can be strictly controlled, and the behaviors that cause abnormalities in the kidney disease biological samples can be avoided. At the same time, the abnormal process links can be processed and risk-averted in advance, effectively reducing the occurrence of abnormalities in the kidney disease biological samples during transportation.
[0104] In summary, this embodiment obtains the process of abnormal transfer of kidney disease biological samples and constructs an FP tree; for any support threshold in the preset support threshold set, the FP tree is processed by this support threshold to obtain the maximum frequent item set; based on the values and depths of each node of the path corresponding to the maximum frequent item set in the FP tree, the reference weight of the maximum frequent item set is obtained, and then the reasonable degree of this support threshold is obtained; based on the reasonable degree, the optimal support threshold is obtained, and the abnormal process links of the kidney disease biological samples are predicted according to the optimal support threshold. The present invention accurately predicts the transfer process that is likely to cause abnormalities in kidney disease biological samples by obtaining the optimal support threshold, timely processes and avoids abnormal risks, and effectively reduces the occurrence of abnormal phenomena in kidney disease biological samples.
[0105] Embodiment 2:
[0106] The present invention also proposes a sample information traceability system for a kidney disease biological sample bank. Please refer to Figure 3, which shows the structural diagram of a sample information traceability system for a kidney disease biobank provided by an embodiment of the present invention. The system includes: a maximum frequent item set acquisition module 10, a reference weight acquisition module 20, a rationality degree acquisition module 30, and an abnormal process link acquisition module 40.
[0107] The maximum frequent item set acquisition module 10 is used to obtain the process of each abnormal kidney disease biological sample transfer and construct an FP tree; for any support threshold in the preset support threshold set, the FP tree is processed through the support threshold to obtain the maximum frequent item set.
[0108] The reference weight acquisition module 20 is used to obtain the reference weight of each maximum frequent item set based on the values and depths of each node on the path corresponding to each maximum frequent item set in the FP tree.
[0109] The rationality degree acquisition module 30 is used to obtain the rationality degree of the support threshold according to the reference weight.
[0110] The abnormal process link acquisition module 40 is used to obtain the optimal support threshold based on the rationality degree, and predict the abnormal process link of the kidney disease biological sample according to the optimal support threshold.
[0111] It should be noted that: for the system provided in the above embodiment, only the above division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, a sample information traceability system for a kidney disease biobank provided in the above embodiment and an embodiment of a sample information traceability method for a kidney disease biobank belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0112] Embodiment 3:
[0113] The present invention also proposes a sample information traceability device for a kidney disease biobank. The device includes a memory and a processor. Among them, an executable program code is stored in the memory, and the processor is used to call and execute the executable program code to execute a sample information traceability method for a kidney disease biobank provided in an embodiment of the present application. The device may specifically be a chip, a component or a module. The chip may include a connected processor and a memory; among them, the memory is used to store instructions. When the processor calls and executes the instructions, the chip can execute a sample information traceability method for a kidney disease biobank provided in the above embodiment.
[0114] In addition, an embodiment of the present application also protects a computer device. Please refer to Figure 4, the computer device includes a memory 401, a processor 402, and a computer program 403 stored in the memory 401 and running on the processor 402. When the processor 402 executes the computer program 403, the computer device can execute any one of the sample information traceability methods of the kidney disease biobank introduced above.
[0115] Example 4:
[0116] This embodiment also provides a computer-readable storage medium. Computer program code is stored in the computer-readable storage medium. When the computer program code runs on a computer, the computer is enabled to execute the above-related method steps to implement a sample information traceability method of a kidney disease biobank provided in the above embodiment.
[0117] Example 5:
[0118] This embodiment also provides a computer program product. When the computer program product runs on a computer, the computer is enabled to execute the above-related steps to implement a sample information traceability method of a kidney disease biobank provided in the above embodiment.
[0119] Among them, the device, computer-readable storage medium, computer program product, or chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.
[0120] It should be noted that: the above sequence of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0121] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments.
Claims
1. A method for tracing sample information of a kidney disease biological sample library, characterized in that: The method comprises the following steps: Obtain the flow process of each abnormal kidney disease biological sample and construct an FP tree; for any support threshold in the preset support threshold set, process the FP tree through the support threshold to obtain the maximum frequent item set; Based on the value and depth of each node in the FP tree corresponding to each maximum frequent item set, obtain the reference weight of each maximum frequent item set; According to the reference weight, obtaining the rationality of the support threshold; Obtaining the best support threshold based on the reasonableness, and predicting the abnormal process link of the kidney disease biological sample according to the best support threshold; The method for obtaining the reasonable degree is: The mean of the reference weights of all the maximum frequent itemsets is taken as the reasonableness of the support threshold; The method for obtaining the optimal support threshold is: The support threshold corresponding to the maximum reasonable degree is taken as the optimal support threshold; The method for obtaining the abnormal process link is: The FP tree is processed by the optimal support threshold, and each maximum frequent item set obtained at this time is used as the target item set; Obtain the flow process of each normal kidney disease biological sample and construct an FP tree as a reference FP tree, process the reference FP tree through the optimal support threshold, and use each maximum frequent item set obtained at this time as a reference item set; Each target item set is compared with each reference item set, and the process links of the paths corresponding to the target item sets that are different from the reference item sets are regarded as abnormal process links of the kidney disease biological samples.
2. The method for tracing the sample information of a kidney disease biological sample library according to claim 1, characterized in that: The method for obtaining the reference weight is: For any maximum frequent item set, the path corresponding to the maximum frequent item set is used as the target path, and the nodes in the target path are divided according to the value size to obtain the first category and the second category; wherein the value of the node in the first category is larger, and the value of the node in the second category is smaller; According to the numerical difference between the nodes in the first category and the nodes in the second category, the similarity of the nodes on the target path is obtained; The node with the largest value in the second category is used as the split node of the target path, and the first abnormality possibility degree of the first category is obtained according to the numerical difference between the split node and other nodes at the same depth in the FP tree, and the number of all nodes at the depth of the split node in the FP tree; The node with the smallest value in the second category is taken as the end point of the target path, and the second abnormal possibility degree of the second category is obtained according to the fluctuation of the value of the node in the second category and the similarity between the value of the node at the next depth of the end point in the FP tree and the value of the node in the second category; According to the similarity degree, the first abnormal possibility degree and the second abnormal possibility degree, obtaining the overall abnormality degree of the maximum frequent item set; Obtaining the importance of the maximum frequent itemset according to the depth of the end point in the FP tree and the value of each node in the target path; The product of the overall abnormality degree and the importance degree is used as the reference weight of the maximum frequent item set.
3. The method for tracing the sample information of a kidney disease biological sample library according to claim 2, characterized in that: The method for obtaining the similarity degree is: Obtaining the average value of the values of all nodes in the first category as a first reference value; Obtaining the average value of the values of all nodes in the second category as a second reference value; The result of negatively correlating and normalizing the difference between the first reference value and the second reference value is used as the similarity of the nodes on the target path.
4. The method for tracing the sample information of a kidney disease biological sample library according to claim 2, characterized in that: The method for obtaining the first abnormality possibility degree is: The depth of the split node in the FP tree is taken as the target depth, and the average value of the values of all nodes except the split node at the target depth is obtained as the first eigenvalue; The difference between the value of the segmentation node and the first characteristic value is used as a reference abnormality degree of the first category; The result of normalizing the product of the inverse of the number of all nodes at the target depth and the reference abnormality degree is taken as the first abnormality possibility degree of the first category.
5. The method for tracing the sample information of a kidney disease biological sample library according to claim 2, characterized in that: The method for obtaining the second abnormal possibility degree is: The result of negatively correlating and normalizing the standard deviation of the values of the nodes in the second category is used as the initial abnormal possibility degree of the second category; Obtain the variance of the values of the nodes in the second category as the first variance; Obtaining the variance of the values of all nodes at the next depth of the end point in the FP tree and the values of the nodes in the second category as the second variance; The result of normalizing the absolute value of the difference between the first variance and the second variance is used as the correction weight of the second category; The product of the revised weight of the second category and the initial abnormality possibility degree is taken as the second abnormality possibility degree of the second category.
6. The method for tracing the sample information of a kidney disease biological sample library according to claim 2, characterized in that: The method for obtaining the overall abnormality degree is: Obtaining the average of the first abnormality possibility degree and the second abnormality possibility degree as the overall abnormality possibility degree of the maximum frequent item set; The product of the similarity degree and the overall abnormality possibility degree is taken as the overall abnormality degree of the maximum frequent item set.
7. The method for tracing the sample information of a kidney disease biological sample library according to claim 2, characterized in that: The method for obtaining the importance is: The ratio of the average value of all nodes in the target path to the depth of the end point in the FP tree is used as the importance of the maximum frequent itemset.
Citation Information
Patent Citations
Time sequence anomaly detection method based on FP-Growth algorithm
CN117591973A
Analysis method and apparatus for operating data of network function virtualization device
WO2022083576A1