Laboratory quality control data management method
By conducting multi-dimensional credibility assessments and weighted hierarchical management of laboratory quality control data, the problem of insufficient intrinsic quality and credibility of the data has been solved, and the intelligentization and reliability improvement of data management have been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 黑龙江省粮食质量安全监测和技术中心
- Filing Date
- 2026-01-15
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies lack in-depth assessment of the intrinsic quality and reliability of data in laboratory quality control data management, resulting in a large volume of data with inconsistent quality, which affects the accuracy of data analysis and the reliability of decision-making.
By performing trend consistency analysis, category completeness analysis, and horizontal trend consistency analysis on quality control data, category credibility, completeness credibility, and horizontal credibility are obtained, and a weighted comprehensive credibility is calculated for hierarchical management.
It significantly improved the intrinsic quality and reliability of laboratory quality control data, enhanced the intelligence and trustworthiness of data management, and optimized the allocation of data storage and application resources.
Smart Images

Figure CN121998488A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data management technology, specifically to a method for managing laboratory quality control data. Background Technology
[0002] Laboratories generate massive amounts of heterogeneous, multi-source quality control data during routine grain quality monitoring. This data covers various grain types, storage locations, testing times, and multiple testing categories such as moisture, impurities, fatty acid values, and mycotoxins, forming a crucial basis for continuously tracking the quality and safety status of stored grains. However, faced with such a complex data stream, existing technologies primarily focus on data collection, input, structured storage, and basic query and statistical functions, lacking in-depth evaluation mechanisms for the intrinsic quality and reliability of the data itself. This results in a large volume of laboratory quality control data with inconsistent quality and unclear reliability, severely restricting the accuracy of subsequent data analysis, the timeliness of risk warnings, and the reliability of data-driven decision-making. Summary of the Invention
[0003] This application provides a laboratory quality control data management method to address the technical problem of ineffective quality control data management in the prior art.
[0004] In view of the above problems, this application provides a laboratory quality control data management method, the method comprising:
[0005] Obtain laboratory quality control data and classify it according to grain category to obtain multiple quality control datasets, wherein each quality control dataset includes multiple quality control data sequences;
[0006] Perform trend consistency analysis on multiple detection category data in the quality control data sequence to obtain category credibility;
[0007] Perform category integrity analysis on the quality control data sequence to obtain integrity reliability;
[0008] The quality control data sequences from multiple storage locations are selected and obtained. A horizontal trend consistency analysis is performed on the same category data of the multiple quality control data sequences to obtain the horizontal reliability.
[0009] The category credibility, the integrity credibility, and the horizontal credibility are weighted and calculated to obtain the comprehensive credibility. The credibility of the quality control data sequence is classified, and the data is stored hierarchically based on the credibility classification for laboratory quality control data management.
[0010] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0011] This application proposes a laboratory quality control data management method. By performing internal category trend consistency analysis, category integrity analysis, and horizontal trend consistency analysis on quality control data sequences, and then weighting and fusing these analyses into a comprehensive reliability score for hierarchical management, the method significantly improves the ability to quantitatively assess the intrinsic quality and reliability of laboratory quality control data from multiple dimensions. Compared with traditional methods, the technical solution provided in this application significantly enhances the intelligence level of the data management process and the reliability of the data management results, achieving the technical effect of improving data management effectiveness from the source. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart illustrating a laboratory quality control data management method provided in an embodiment of this application.
[0014] Figure 2 This is a schematic diagram illustrating the process of obtaining high-frequency category sets in a laboratory quality control data management method provided in this application embodiment. Detailed Implementation
[0015] This application provides a laboratory quality control data management method to address the technical problem of ineffective quality control data management in the prior art.
[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0017] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to these processes, methods, products, or devices.
[0018] Examples, such as Figure 1 As shown, this application provides a laboratory quality control data management method, wherein the method includes:
[0019] S10: Obtain laboratory quality control data and classify it according to grain category to obtain multiple quality control datasets, wherein each quality control dataset includes multiple quality control data sequences.
[0020] During the storage, transportation and processing of grains, routine monitoring generates massive amounts of data from a wide range of sources and in various formats. This data usually includes different grain categories such as wheat, rice and corn, and each category contains multiple testing time points, different storage locations, and numerous testing categories such as moisture, impurities, fatty acid value, and mycotoxins.
[0021] Step S10 in the method provided in this application embodiment includes:
[0022] Obtain laboratory quality control data, which includes grain type, grain storage location, testing time, and data on multiple testing categories;
[0023] Based on the grain category, the laboratory quality control data is classified to obtain multiple quality control datasets. Each quality control dataset includes multiple quality control data sequences, and each quality control data sequence includes a set of samples of a grain category in a single monitoring session, along with the grain category, grain storage location, and multiple testing categories.
[0024] In this embodiment of the application, laboratory quality control data is obtained and classified according to grain category to obtain multiple quality control datasets, wherein each quality control dataset includes multiple quality control data sequences.
[0025] Specifically, the first step is to acquire laboratory quality control data, which includes grain type, grain storage location, testing time, and data on multiple testing categories. For example, grain type may include rice, wheat, etc., grain storage location may include multiple grain warehouse numbers, testing time can be stored in year / month / day format, and testing category data may include wheat hardness index, protein content, etc.
[0026] Furthermore, based on grain categories, the laboratory quality control data is classified to obtain multiple quality control datasets. Each quality control dataset includes multiple quality control data sequences. Each quality control data sequence includes a set of samples of a grain category in a single monitoring session, specifying the grain category, storage location, and multiple testing categories. For example, a quality control data sequence might be: {Wheat, A01 grain storage, 2025 / 12 / 06, Protein content: 13.5%, Wheat hardness index: 65, Imperfect grains: 3.1%...}.
[0027] By automatically classifying the original laboratory quality control data according to grain categories, the data is initially organized to facilitate subsequent analysis.
[0028] S20: Perform trend consistency analysis on multiple detection category data in the quality control data sequence to obtain category credibility.
[0029] A single laboratory test report or a batch of sample quality control data contains multiple test categories, such as moisture content, fatty acid value, and total mold count. These categories do not exist in isolation; they usually imply inherent connections or consistent trends based on the patterns of biochemical changes in grains.
[0030] Step S20 in the method provided in this application embodiment includes:
[0031] Obtain the category trend analysis model;
[0032] The category trend analysis model includes:
[0033] Obtain sample quality control data sequences and obtain category trend conflict rules, wherein the category trend conflict rules include trend conflict category pairs and trend conflict importance;
[0034] Based on the category trend conflict rules, the sample quality control data sequence is labeled to obtain the sample conflict coefficient;
[0035] A category trend analysis model is constructed, with the sample quality control data sequence as input and the sample conflict coefficient as supervision, and the category trend analysis model is trained until convergence.
[0036] The quality control data sequence is input into the category trend analysis model to obtain the conflict coefficient. The category credibility is obtained by subtracting the conflict coefficient from 1.
[0037] In this embodiment of the application, trend consistency analysis is performed on multiple detection category data in the quality control data sequence to obtain category credibility.
[0038] Specifically, obtain the category trend analysis model.
[0039] First, obtain the sample quality control data sequence and category trend conflict rules, which include trend conflict category pairs and the importance of the trend conflict. Preferably, obtain the laboratory quality control data sequence from the historical database as the sample quality control data sequence. Further, domain experts pre-define category trend conflict rules based on knowledge of grain storage science. For example, a rule might state that for wheat varieties, mycotoxin content and moisture content constitute a trend conflict category pair. The logic is that under normal storage conditions, a significant increase in mycotoxins is usually accompanied by an abnormal increase in moisture content or its maintenance at a high level; if a data set shows a high mycotoxin content but a low moisture content, these two trends logically conflict, potentially indicating an error in the detection. Furthermore, each category trend conflict rule is assigned a trend conflict importance weight; for example, a weight of 0.8 for the mycotoxin-moisture conflict indicates it is relatively important, and a weight of 0.6 for the fatty acid value-broken grain ratio conflict indicates it is relatively important.
[0040] Furthermore, based on category trend conflict rules, the sample quality control data sequence is labeled to obtain the sample conflict coefficient. The sample quality control data sequence is then automatically checked based on these category trend conflict rules. For each sample quality control data sequence, all conflict rules are traversed. If the actual values of a conflicting category pair in the data show a contradictory trend as defined by the rule, a conflict score is accumulated according to the importance weight of the corresponding rule. Finally, through a preset normalization function, the accumulated conflict score is converted into a sample conflict coefficient between 0 and 1. A value of 1 indicates a serious trend conflict according to the rule, while 0 indicates no obvious conflict is found.
[0041] Furthermore, a category trend analysis model is constructed, using sample quality control data sequences as input and sample conflict coefficients as supervision, and trained until convergence. For example, a category trend analysis model is constructed based on a fully connected neural network. The number of neurons in the input layer equals the number of possible detectable categories in a quality control data sequence, such as 20. Each input neuron corresponds to one detectable category. The hidden layer includes 32 neurons activated using the ReLU function, and the output layer includes one neuron activated using the Sigmoid activation function, whose output value is a predicted conflict probability between 0 and 1. The gradient descent method in supervised learning is used to train the model. The sample quality control data sequence is used as input, and the corresponding sample conflict coefficients are used as supervision for training. The model calculates the predicted output under the current parameters, and then calculates the difference between the predicted value and the sample conflict coefficient using the mean squared error loss. Next, the gradient of the loss relative to the model parameters is calculated using the backpropagation algorithm, and the model parameters are updated along the gradient direction using the Adam optimizer to reduce the loss. This process is iterated repeatedly on all training data until the model's predicted loss no longer decreases significantly, i.e., the model converges.
[0042] Furthermore, the quality control data sequence is input into the category trend analysis model to obtain the conflict coefficient. The category credibility is then calculated by subtracting the conflict coefficient from 1. For example, if a conflict coefficient of 0.2 is obtained, then the category credibility of this quality control data sequence is 1 - 0.2 = 0.8.
[0043] The constructed category trend analysis model can automatically assess whether the trends reflected by the data of each category in the sequence conform to the general and known change patterns of that category under corresponding conditions. When the data performance of some categories deviates significantly from the trend predictions of other categories, the model can identify this inconsistency and output a lower category credibility quantification index, which can elevate the judgment of data credibility from a single threshold check to a level of dynamic correlation and logical verification.
[0044] S30: Perform category integrity analysis on the quality control data sequence to obtain integrity reliability.
[0045] In actual laboratory testing, due to reasons such as updates to testing standards, additions or subtractions of testing items, differences in specifications between different laboratories, or omissions in records, the resulting quality control data sequences often exhibit inconsistencies in the included testing categories.
[0046] Step S30 in the method provided in this application embodiment includes:
[0047] Based on multiple quality control data sequences, high-frequency category analysis is performed to obtain a high-frequency category set;
[0048] Among these methods, high-frequency category analysis is performed based on multiple quality control data sequences to obtain a high-frequency category set, such as... Figure 2 As shown, it includes:
[0049] Obtain a quality control data sequence as an initial control data sequence, and extract multiple detection categories from the initial control data sequence;
[0050] Randomly select one detection category from the initial control data sequence as the first detection category, traverse the remaining quality control data sequence, calculate the deviation of the first detection category, count the number of sequences whose deviation of the first detection category is less than the preset first deviation threshold, and calculate the first high frequency coefficient of the first detection category. The deviation of the first detection category includes the standard deviation of data acquisition and the data accuracy deviation.
[0051] When the first high-frequency coefficient is greater than the preset high-frequency coefficient threshold, a detection category is randomly selected from the initial control data sequence as the second detection category, the deviation of the second detection category is calculated and the second high-frequency coefficient is obtained, until the high-frequency coefficient calculation of all detection categories is completed.
[0052] When the first high-frequency coefficient is less than or equal to the preset high-frequency coefficient threshold, one is randomly selected from the remaining quality control data sequences as the initial control data sequence, and the first detection category is reacquired.
[0053] When all high-frequency coefficients in the initial control data sequence are greater than the preset high-frequency coefficient threshold, the detection category in the initial control data sequence is added to the high-frequency category set as a high-frequency category.
[0054] Based on the high-frequency category set, combined with the essential category set, a complete category set is obtained;
[0055] Calculate the percentage of completeness of the categories in the quality control data sequence relative to the complete category set, and obtain the completeness reliability.
[0056] In this embodiment of the application, category integrity analysis is performed on the quality control data sequence to obtain integrity reliability.
[0057] Specifically, firstly, based on multiple quality control data sequences, high-frequency category analysis is performed to obtain a high-frequency category set.
[0058] Among these, high-frequency category analysis is performed based on multiple quality control data sequences to obtain a high-frequency category set, including:
[0059] First, a quality control data sequence is obtained as the initial control data sequence, and multiple detection categories are extracted from the initial control data sequence. For example, the multiple detection categories of the initial control data sequence may include moisture content, protein content, wheat hardness index, and imperfect grains.
[0060] Further, a detection category is randomly selected from the initial control data sequence as the first detection category. The remaining quality control data sequences are traversed, the deviation of the first detection category is calculated, and the number of sequences with a first detection category deviation less than a preset first deviation threshold is counted. The first high-frequency coefficient of the first detection category is calculated. The first detection category deviation includes the data acquisition standard deviation and the data accuracy deviation. For example, moisture is randomly selected as the first detection category. All other sequences in the dataset, except the initial control data sequence, are traversed. For each traversed sequence, the first detection category deviation between that sequence and the initial control data sequence in the moisture category is calculated. The first detection category deviation consists of two parts: first, the data acquisition standard deviation, which determines whether the detection standards used by the two records are consistent, such as whether they are both based on the GB5009.3-2016 standard; second, the data format deviation, which determines whether the specific detection methods and reporting formats used by the two records for this category are consistent, such as whether they are both oven drying methods rather than vacuum drying methods, and whether the numerical reporting units are both "%" rather than other units. A first deviation threshold is preset. This threshold is used to evaluate whether the data collection standards and data format meet the consistency requirements. For example, the first deviation threshold can be set to 0.67, meaning that if the data collection standards and detection methods are consistent, the consistency requirements are considered met. The number of sequences with a deviation less than the first deviation threshold is counted across all traversed sequences, and this number is used as the first high-frequency coefficient for water classification. For example, if 1000 other sequences are traversed and 850 meet the deviation threshold condition, then the high-frequency coefficient for water is 0.85.
[0061] Furthermore, when the first high-frequency coefficient is greater than the preset high-frequency coefficient threshold, a detection category is randomly selected from the initial control data sequence as the second detection category. The deviation of the second detection category is calculated, and the second high-frequency coefficient is obtained, until the high-frequency coefficients of all detection categories are calculated. For example, if the high-frequency coefficient (85) of water is greater than the preset high-frequency coefficient threshold (e.g., 80), it indicates that the water category exhibits high frequency occurrence and uniform detection form in the dataset. Further, another category that has not yet been calculated is randomly selected from the initial control data sequence as the second detection category, and the traversal and calculation process is repeated to obtain the second high-frequency coefficient. This process is repeated until the high-frequency coefficients of all categories in the initial sequence are calculated.
[0062] Furthermore, when the first high-frequency coefficient is less than or equal to a preset high-frequency coefficient threshold, a random sequence is selected from the remaining quality control data sequences as the initial control data sequence, and the first detection category is re-acquired. If, when calculating a certain category (such as aflatoxin), its high-frequency coefficient is less than the high-frequency coefficient threshold, it indicates that the currently selected initial control data sequence may not be representative. In this case, the current initial control data sequence is discarded, and a new initial control data sequence is randomly selected from the remaining data sequences, and the calculation process restarts.
[0063] Furthermore, when all high-frequency coefficients in the initial control data sequence are greater than a preset high-frequency coefficient threshold, the detection categories in the initial control data sequence are added to the high-frequency category set as high-frequency categories. This set of detection categories in the initial control data sequence represents a stable and frequently co-occurring combination of categories. By repeatedly performing the above calculations and analyses, several potentially slightly different high-frequency category combinations can be collected; the union of these high-frequency category combinations constitutes the high-frequency category set.
[0064] Furthermore, based on the high-frequency category set and combined with the essential category set, a complete category set is obtained. The essential category set is a list of categories that must be tested for a specific grain category, pre-defined according to grain storage standards. For example, for rice, the essential category set might be {moisture, impurities, aflatoxin B1}. The complete category set is obtained by taking the union of the high-frequency category set and the essential category set. This set represents all the categories that a quality control data sequence needs to include to better characterize the grain quality under the current data environment and standard requirements.
[0065] Furthermore, the completeness ratio of categories in the quality control data sequence to the complete category set is calculated to obtain the completeness reliability. Completeness reliability = Number of intersection categories between the quality control data sequence and the complete category set / Total number of categories in the complete category set. A higher completeness reliability indicates a larger number of categories in the quality control data sequence, better completeness, and potentially higher reliability.
[0066] By conducting category completeness analysis, the completeness and applicability of acquired data in terms of information dimensions can be analyzed. Based on high-frequency category analysis and domain knowledge such as essential category sets, a reference standard for a complete category set for a specific grain category can be constructed. By calculating the matching ratio between the existing categories of the data sequence to be evaluated and this complete category set, a completeness credibility can be quantitatively output. This can effectively identify and mark data sequences lacking key information, preventing them from being misused in serious decision-making scenarios that require complete information.
[0067] S40: Filter and obtain multiple quality control data sequences from the same storage location, perform horizontal trend consistency analysis on the same category data of the multiple quality control data sequences, and obtain horizontal reliability.
[0068] Traditional laboratory quality control data management typically treats each test sequence as an independent, isolated record, focusing only on its own historical changes or internal logic. However, in the actual grain storage ecosystem, different batches of grain stored in the same storage location over similar time periods experience highly similar macro-environmental conditions, such as storage temperature and humidity. This commonality in environment may lead to comparability or group-wide commonalities in the trends of certain quality indicators.
[0069] Step S40 in the method provided in this application embodiment includes:
[0070] The quality control data sequences are filtered based on the storage location to obtain quality control data sequences from the same location.
[0071] Based on the detection time, the same-location quality control data sequences are filtered to obtain similar quality control data sequences;
[0072] Horizontal trend consistency analysis is performed on multiple similar quality control data sequences to obtain horizontal reliability.
[0073] Among these steps, a horizontal trend consistency analysis is performed on multiple similar quality control data sequences to obtain horizontal reliability, including:
[0074] Based on multiple similar quality control data sequences, common environmental characteristics are obtained;
[0075] Based on the common environmental characteristics, a common feature deviation analysis is performed on multiple similar quality control data sequences to obtain horizontal reliability. The horizontal reliability is obtained based on the category with the largest deviation of common features.
[0076] In this embodiment of the application, multiple quality control data sequences from the same storage location are selected and obtained. Then, a horizontal trend consistency analysis is performed on the same category data of the multiple quality control data sequences to obtain the horizontal reliability.
[0077] Specifically, firstly, the quality control data sequences are filtered based on the storage location to obtain quality control data sequences from the same location. Specifically, based on the storage location of each data record, such as grain warehouse A01, all quality control data sequences belonging to that location are selected to obtain quality control data sequences from the same location.
[0078] Furthermore, based on the detection time, the quality control data sequences from the same location are filtered to obtain similar quality control data sequences. For example, a time window, such as 30 days, is set, and based on the detection time of the currently evaluated control data sequence, all other data sequences from the same location whose detection time falls within that time window are selected to obtain similar quality control data sequences.
[0079] Furthermore, a horizontal trend consistency analysis was performed on multiple similar quality control data sequences to obtain horizontal reliability.
[0080] Specifically, firstly, common environmental characteristics are obtained based on multiple similar quality control data sequences. For example, the mean of multiple similar quality control data sequences is calculated as a common environmental characteristic. The calculation results of all common categories are summarized to constitute the common environmental characteristics, which describe the general level of various quality indicators of grain under that specific environment and time period.
[0081] Furthermore, based on common environmental characteristics, a common feature deviation analysis was performed on multiple similar quality control data sequences to obtain lateral confidence. The lateral confidence was obtained based on the category with the largest common feature deviation. Common feature deviation = |Target sequence value for that category - Average value of that category in common features| / Average value of that category in common features. Lateral confidence = 1 - Maximum deviation. For example, if the calculated common feature deviations of the sequence across multiple categories are 0.05, 0.12, and 0.03, then the maximum deviation is 0.12. Therefore, the lateral confidence of the target sequence = 1 - 0.12 = 0.88. This result indicates that, in terms of lateral consistency with other data from the same location and period, the sequence has high confidence. If a certain category deviates significantly, the lateral confidence will be low, suggesting that the data may be an outlier in this environment.
[0082] By screening data sequences from the same storage locations and conducting cross-sectional comparisons, it is possible to effectively capture and utilize the collective characteristics of data shaped by shared environmental factors. When a data sequence deviates significantly from the trend baseline or common range formed by other data sequences from the same location and period in key categories, it is given lower cross-sectional reliability due to its environmental inconsistency. This effect greatly enhances the ability to identify anomalous data that is highly concealed and difficult to detect from an individual perspective alone.
[0083] S50: Calculate the category credibility, the integrity credibility, and the horizontal credibility using weighted averages to obtain the comprehensive credibility. Then, classify the credibility of the quality control data sequence and store it hierarchically based on the credibility classification to manage laboratory quality control data.
[0084] Through the aforementioned steps, quantitative credibility indicators for data credibility can be obtained from three different but important dimensions: internal logic, structural integrity, and horizontal correlation. Current technology lacks an effective mechanism to scientifically integrate these multi-dimensional evaluation results into a comprehensive, actionable judgment.
[0085] Step S50 in the method provided in this application embodiment includes:
[0086] Based on expert evaluation, credibility weights are restructured.
[0087] The credibility weight reorganization is used to perform a weighted calculation on the category credibility, the integrity credibility, and the horizontal credibility to obtain a comprehensive credibility.
[0088] Based on the grain category, obtain the credibility grading threshold;
[0089] Based on the aforementioned credibility grading threshold, the credibility of the quality control data sequence is graded to obtain a credibility level.
[0090] Based on the trust level, the quality control data sequence is stored in a hierarchical manner and labeled with application scenarios.
[0091] In this embodiment of the application, the category credibility, integrity credibility, and horizontal credibility are weighted and calculated to obtain the comprehensive credibility. The credibility of the quality control data sequence is classified, and the data is stored hierarchically based on the credibility classification for laboratory quality control data management.
[0092] Specifically, firstly, based on expert evaluation, a credibility weight reassessment is obtained. For example, the Delphi method is used for expert evaluation, where category credibility reflects the internal logical consistency of the data, completeness credibility reflects the completeness of the data information, and horizontal credibility reflects the consistency of the data with data in the same environment. The mean value of each weight in the credibility weight reassessment given by the experts is obtained to obtain the credibility weight reassessment. For example, a credibility weight reassessment is 0.52, 0.30, and 0.18.
[0093] Furthermore, a credibility weighting reassessment is employed to weight the category credibility, completeness credibility, and horizontal credibility to obtain the overall credibility. Overall Credibility = (Category Credibility × Weight of Category Credibility in Weighting Reassessment) + (Completeness Credibility × Weight of Completeness Credibility in Weighting Reassessment) + (Horizontal Credibility × Weight of Horizontal Credibility in Weighting Reassessment). For example, the three credibility values of a quality control data sequence are: Category Credibility 0.7, Completeness Credibility 0.6, and Horizontal Credibility 0.9. Using the credibility weighting reassessment (0.52, 0.30, 0.18) from the example above, the overall credibility is calculated as follows: Overall Credibility = (0.7 × 0.52) + (0.6 × 0.30) + (0.9 × 0.18) = 0.364 + 0.18 + 0.162 = 0.706. A higher overall credibility indicates higher overall credibility.
[0094] Furthermore, based on the grain category, credibility grading thresholds are obtained. A credibility mapping table is then created, which is a pre-defined mapping table for grading credibility based on historical data analysis and domain experience. For example, for wheat, the grading thresholds are a high credibility lower limit of 0.8 and a medium credibility lower limit of 0.6. Different categories may have different thresholds; for example, for rice, the high credibility lower limit might be 0.75 and the medium credibility lower limit might be 0.55.
[0095] Furthermore, based on the confidence level threshold, the quality control data sequence is classified into confidence levels to obtain a confidence grade. For example, if the confidence level of a quality control data sequence for a wheat variety is 0.706, which is greater than the lower limit of medium confidence, then the confidence grade of the quality control data sequence is medium confidence.
[0096] Furthermore, based on the trust level, the quality control data sequences are stored in a tiered manner and labeled with application scenarios. For example, high-trust-level data sequences are stored in a high-performance, high-access-priority main database and labeled with the application scenario tag "can be used for automatic early warning and precise analysis." Medium-trust-level data sequences are stored in a standard quality control historical database and labeled with the application scenario tag "requires manual review or combination with other data before it can be used for statistical analysis." Low-trust-level data sequences are transferred to a dedicated, lower-performance archiving and review database for storage and labeled with the application scenario tag "only for data traceability or problem investigation reference, not involved in routine analysis and decision-making."
[0097] This step, through weighted fusion and hierarchical management, ultimately achieves a comprehensive quantification and value realization of data credibility. Multiple credibility indicators characterizing different aspects of data quality are weighted according to the importance of each dimension, resulting in a unified comprehensive credibility score. This score provides a concise summary of the overall reliability of the data, solving the problem of the difficulty in comprehensively applying multi-dimensional evaluation results. Furthermore, based on this comprehensive credibility score, data sequences are classified, and hierarchical storage strategies and differentiated application strategies are implemented according to these levels. This elevates data management from indiscriminate storage to value-based refined management, enabling data management methods to automatically identify and differentiate data quality, optimizing storage resource allocation and subsequent applications.
[0098] In summary, the embodiments of this application have at least the following technical effects:
[0099] This application proposes a laboratory quality control data management method. By performing internal category trend consistency analysis, category completeness analysis, and horizontal trend consistency analysis on quality control data sequences, and weighting and integrating these into a comprehensive reliability score for hierarchical management, it significantly improves the ability to quantitatively assess the intrinsic quality and reliability of laboratory quality control data from multiple dimensions. Specifically, firstly, category trend consistency analysis can automatically identify potential logical conflicts or abnormal trends between different categories of data in the same test sequence, thereby effectively filtering unreliable data caused by operational errors, instrument malfunctions, or sample contamination, and improving the internal logical consistency of single test results. Secondly, through completeness assessment based on high-frequency and essential category analysis, it can adaptively determine the completeness of data sequences in terms of information dimensions, ensuring that data used for key decisions has the necessary category support, and avoiding analytical bias and decision risks caused by missing information. Furthermore, by introducing horizontal trend consistency analysis of data from similar time periods at the same storage location, isolated abnormal data deviating from the commonalities of the group can be discovered, enhancing the objectivity of data quality judgment and the environmental supporting evidence. Ultimately, by weightedly integrating the aforementioned multi-dimensional credibility indicators to form a comprehensive credibility score and implementing tiered storage, this application achieves refined management of data value. This allows high-credibility data to be prioritized for key scenarios such as precise analysis and risk warning, while low-credibility data is restricted in use or marked for review, optimizing the allocation efficiency of data storage and computing resources. Compared to traditional methods, the technical solution provided in this application significantly enhances the intelligence level of the data management process and the reliability of data management results, achieving a technical effect of improving data management effectiveness from the source.
[0100] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0101] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0102] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.
Claims
1. A method for managing laboratory quality control data, characterized in that, include: Obtain laboratory quality control data and classify it according to grain category to obtain multiple quality control datasets, wherein each quality control dataset includes multiple quality control data sequences; Perform trend consistency analysis on multiple detection category data in the quality control data sequence to obtain category credibility; Perform category integrity analysis on the quality control data sequence to obtain integrity reliability; The quality control data sequences from multiple storage locations are selected and obtained. A horizontal trend consistency analysis is performed on the same category data of the multiple quality control data sequences to obtain the horizontal reliability. The category credibility, the integrity credibility, and the horizontal credibility are weighted and calculated to obtain the comprehensive credibility. The credibility of the quality control data sequence is classified, and the data is stored hierarchically based on the credibility classification for laboratory quality control data management.
2. The laboratory quality control data management method according to claim 1, characterized in that, Obtain laboratory quality control data and classify it according to grain category to obtain multiple quality control datasets. Each quality control dataset includes multiple quality control data sequences, including: Obtain laboratory quality control data, which includes grain type, grain storage location, testing time, and data on multiple testing categories; Based on the grain category, the laboratory quality control data is classified to obtain multiple quality control datasets. Each quality control dataset includes multiple quality control data sequences, and each quality control data sequence includes a set of samples of a grain category in a single monitoring session, along with the grain category, grain storage location, and multiple testing categories.
3. The laboratory quality control data management method according to claim 1, characterized in that, Perform trend consistency analysis on multiple detection category data in the quality control data sequence to obtain category credibility, including: Obtain the category trend analysis model; The quality control data sequence is input into the category trend analysis model to obtain the conflict coefficient. The category credibility is obtained by subtracting the conflict coefficient from 1.
4. The laboratory quality control data management method according to claim 3, characterized in that, Obtain the category trend analysis model, including: Obtain sample quality control data sequences and obtain category trend conflict rules, wherein the category trend conflict rules include trend conflict category pairs and trend conflict importance; Based on the category trend conflict rules, the sample quality control data sequence is labeled to obtain the sample conflict coefficient; A category trend analysis model is constructed, using sample quality control data sequences as input and sample conflict coefficients as supervision, and the model is trained until convergence.
5. The laboratory quality control data management method according to claim 1, characterized in that, Perform category integrity analysis on the quality control data sequence to obtain integrity reliability, including: Based on multiple quality control data sequences, high-frequency category analysis is performed to obtain a high-frequency category set; Based on the high-frequency category set, combined with the essential category set, a complete category set is obtained; Calculate the percentage of completeness of the categories in the quality control data sequence relative to the complete category set, and obtain the completeness reliability.
6. The laboratory quality control data management method according to claim 5, characterized in that, Based on multiple quality control data sequences, high-frequency category analysis is performed to obtain a high-frequency category set, including: Obtain a quality control data sequence as an initial control data sequence, and extract multiple detection categories from the initial control data sequence; Randomly select one detection category from the initial control data sequence as the first detection category, traverse the remaining quality control data sequence, calculate the deviation of the first detection category, count the number of sequences whose deviation of the first detection category is less than the preset first deviation threshold, and calculate the first high frequency coefficient of the first detection category. The deviation of the first detection category includes the standard deviation of data acquisition and the data accuracy deviation. When the first high-frequency coefficient is greater than the preset high-frequency coefficient threshold, a detection category is randomly selected from the initial control data sequence as the second detection category, the deviation of the second detection category is calculated and the second high-frequency coefficient is obtained, until the high-frequency coefficient calculation of all detection categories is completed. When the first high-frequency coefficient is less than or equal to the preset high-frequency coefficient threshold, one is randomly selected from the remaining quality control data sequences as the initial control data sequence, and the first detection category is reacquired. When all high-frequency coefficients in the initial control data sequence are greater than the preset high-frequency coefficient threshold, the detection category in the initial control data sequence is added to the high-frequency category set as a high-frequency category.
7. The laboratory quality control data management method according to claim 1, characterized in that, The quality control data sequences from multiple storage locations are selected and analyzed for horizontal trend consistency to obtain horizontal reliability, including: The quality control data sequences are filtered based on the storage location to obtain quality control data sequences from the same location. Based on the detection time, the same-location quality control data sequences are filtered to obtain similar quality control data sequences; A horizontal trend consistency analysis was performed on multiple similar quality control data sequences to obtain horizontal reliability.
8. The laboratory quality control data management method according to claim 7, characterized in that, A cross-sectional trend consistency analysis was performed on multiple similar quality control data sequences to obtain cross-sectional reliability, including: Based on multiple similar quality control data sequences, common environmental characteristics are obtained; Based on the common environmental characteristics, a common feature deviation analysis is performed on multiple similar quality control data sequences to obtain horizontal reliability. The horizontal reliability is obtained based on the category with the largest deviation of common features.
9. A laboratory quality control data management method according to claim 1, characterized in that, The overall credibility is obtained by weighted calculation of the category credibility, the completeness credibility, and the horizontal credibility, including: Based on expert evaluation, credibility weights are restructured. The credibility weight reorganization is used to perform a weighted calculation on the category credibility, the integrity credibility, and the horizontal credibility to obtain a comprehensive credibility.
10. A laboratory quality control data management method according to claim 1, characterized in that, Based on credibility levels, tiered storage is implemented for laboratory quality control data management, including: Based on the grain category, obtain the credibility grading threshold; Based on the aforementioned credibility grading threshold, the credibility of the quality control data sequence is graded to obtain a credibility level. Based on the trust level, the quality control data sequence is stored in a hierarchical manner and labeled with application scenarios.