An antibacterial peptide experimental data management method and system

CN121705658BActive Publication Date: 2026-09-18XIAN MEDICAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610101109.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-26
Publication Date
2026-09-18
Estimated Expiration
2046-01-26

AI Technical Summary

Technical Problem

[0004]本发明针对现有技术中抗菌肽实验数据因环境条件变化导致数据质量不稳定、难以准确管理实验数据的技术问题,提供一种抗菌肽实验数据管理方法及系统来解决

Benefits of technology

获取多个实验室内的抗菌肽实验数据集和实验条件数据集,根据多个实验条件数据集进行数据置信度分析,获得多个数据置信度集,从而通过云平台分布式架构,建立实验条件与数据可信程度的对应关系,为云平台多节点数据筛选提供量化依据。根据多个数据置信度集,分别在多个抗菌肽实验数据集内选取训练数据集和测试数据集,采用多个训练数据集集成训练多个基础分析器,对多个测试数据集进行测试,获得多个测试精度集,从而利用云平台强大的计算资源验证不同置信度数据的实际分析性能,为环境影响评估提供客观的准确性指标。对多个测试精度集和多个测试数据集的数据置信度进行验证,获得多个环境影响可信度,并对多个实验条件数据集进行波动异常分析并校正多个环境影响可信度,获得多个校正环境可信度,从而在云平台统一处理框架下消除偶然性环境变化和传感器误差对环境影响评估的干扰,提高可信度评估的准确性。根据多个校正环境可信度,选取获得最优分析器,并在抗菌肽实验数据集内选取获得可信实验数据集,进行实验数据管理,确保使用最可靠的分析模型和最高质量的实验数据,从而提升整体数据管理的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121705658B_ABST
    Figure CN121705658B_ABST
Patent Text Reader

Abstract

The application provides an antibacterial peptide experimental data management method and system, and belongs to the field of cloud-edge collaborative data management. The method comprises the following steps: obtaining antibacterial peptide experimental data sets and experimental condition data sets in multiple laboratories; performing data confidence analysis according to the multiple experimental condition data sets to obtain multiple data confidence sets; training multiple basic analyzers, testing the multiple test data sets, and obtaining multiple test accuracy sets; verifying the data confidence of the multiple test accuracy sets and the multiple test data sets to obtain multiple environmental influence credibility, and obtaining multiple corrected environmental credibility; selecting the optimal analyzer, and selecting the credible experimental data set in the antibacterial peptide experimental data set to perform experimental data management. Through data confidence analysis and environmental influence correction, the credible experimental data is accurately screened, and the accuracy of antibacterial peptide experimental data management is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cloud-edge collaborative data management, and in particular to a method and system for managing experimental data of antimicrobial peptides. Background Technology

[0002] Antimicrobial peptides, as a class of natural biomolecules with broad-spectrum antimicrobial activity, have significant application value in fields such as antibiotic alternatives, food preservation, and medical devices. With the in-depth development of antimicrobial peptide research, laboratories in different geographical locations have generated a large amount of experimental data, including antimicrobial activity test data, structural analysis data, and stability test data. The effective management and analysis of this experimental data is of great significance for the research and application of antimicrobial peptides.

[0003] However, antimicrobial peptide experiments involve various environmental factors, such as temperature, humidity, pH, and light conditions. Changes in these environmental conditions directly affect the activity of antimicrobial peptides and the accuracy of experimental results. Current experimental data management methods typically employ simple data storage and classification approaches, lacking effective monitoring of changes in experimental environmental conditions and data quality assessment mechanisms. Consequently, in practical applications, due to the complexity and uncertainty of the experimental environment, antimicrobial peptide experimental data often suffers from unstable quality, making it difficult to accurately identify which data have been affected by changes in environmental conditions. This leads to insufficient accuracy in experimental data management, impacting subsequent data analysis and application effectiveness. Summary of the Invention

[0004] This invention addresses the technical problem in the prior art where the quality of antimicrobial peptide experimental data is unstable due to changes in environmental conditions, making it difficult to accurately manage experimental data. It provides a method and system for managing antimicrobial peptide experimental data to solve this problem.

[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:

[0006] In a first aspect, the present invention provides a method for managing experimental data of antimicrobial peptides, comprising: acquiring multiple experimental datasets and experimental condition datasets of antimicrobial peptides in multiple laboratories; performing data confidence analysis on the multiple experimental condition datasets to obtain multiple data confidence sets; selecting training datasets and test datasets from the multiple antimicrobial peptide experimental datasets according to the multiple data confidence sets; integrating and training multiple basic analyzers using the multiple training datasets; testing the multiple test datasets to obtain multiple test accuracy sets; verifying the data confidence of the multiple test accuracy sets and multiple test datasets to obtain multiple environmental influence confidence levels; performing fluctuation anomaly analysis on the multiple experimental condition datasets and correcting the multiple environmental influence confidence levels to obtain multiple corrected environmental confidence levels; selecting the optimal analyzer based on the multiple corrected environmental confidence levels; and selecting a reliable experimental dataset from the antimicrobial peptide experimental datasets for experimental data management.

[0007] Secondly, the present invention provides an antimicrobial peptide experimental data management system, comprising: a data confidence analysis module, used to acquire multiple antimicrobial peptide experimental datasets and experimental condition datasets from multiple laboratories, perform data confidence analysis based on the multiple experimental condition datasets, and obtain multiple data confidence sets; a test accuracy evaluation module, used to select training datasets and test datasets from the multiple antimicrobial peptide experimental datasets based on the multiple data confidence sets, integrate and train multiple basic analyzers using the multiple training datasets, and test the multiple test datasets to obtain multiple test accuracy sets; a credibility verification and correction module, used to verify the data confidence of the multiple test accuracy sets and multiple test datasets to obtain multiple environmental influence credibility, and perform fluctuation anomaly analysis on the multiple experimental condition datasets and correct the multiple environmental influence credibility to obtain multiple corrected environmental credibility; and an experimental data management module, used to select the optimal analyzer based on the multiple corrected environmental credibility, and select a reliable experimental dataset from the antimicrobial peptide experimental datasets for experimental data management.

[0008] The beneficial effects of this invention are: This study acquires multiple experimental datasets of antimicrobial peptides and experimental conditions from various laboratories. Based on these datasets, it performs data confidence analysis to obtain multiple confidence sets. Through a distributed cloud platform architecture, it establishes a correlation between experimental conditions and data confidence levels, providing a quantitative basis for multi-node data screening on the cloud platform. Based on these confidence sets, training and test datasets are selected from the multiple antimicrobial peptide experimental datasets. Multiple basic analyzers are trained using the integrated training datasets, and tested on multiple test datasets to obtain multiple test accuracy sets. This leverages the powerful computing resources of the cloud platform to verify the actual analytical performance of data at different confidence levels, providing objective accuracy indicators for environmental impact assessment. The confidence levels of the multiple test accuracy sets and test datasets are validated to obtain multiple environmental impact confidence levels. Furthermore, fluctuation anomaly analysis is performed on the multiple experimental condition datasets to correct multiple environmental impact confidence levels, resulting in multiple corrected environmental confidence levels. This eliminates the interference of accidental environmental changes and sensor errors on environmental impact assessment within the unified processing framework of the cloud platform, improving the accuracy of confidence assessment. Based on the credibility of multiple calibration environments, the optimal analyzer is selected, and a reliable experimental dataset is selected within the antimicrobial peptide experimental dataset for experimental data management. This ensures the use of the most reliable analysis model and the highest quality experimental data, thereby improving the accuracy of overall data management.

[0009] By using the above technical solution, multiple laboratories are used as edge data acquisition nodes to achieve distributed data aggregation and collaborative processing. The aggregated data is processed uniformly through a cloud platform. Based on the distributed data processing capabilities of the cloud platform, reliable antimicrobial peptide experimental data can be accurately identified and screened, effectively eliminating the adverse effects of changes in environmental conditions on data quality, and achieving the technical effect of improving the accuracy of antimicrobial peptide experimental data management. Attached Figure Description

[0010] Figure 1 A flowchart illustrating an antimicrobial peptide experimental data management method provided by the present invention; Figure 2 This is a schematic diagram of the structure of an antimicrobial peptide experimental data management system provided by the present invention.

[0011] In the attached diagram, the components represented by each number are as follows: Data confidence analysis module 11, test accuracy evaluation module 12, confidence verification and correction module 13, experimental data management module 14. Detailed Implementation

[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0013] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0014] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.

[0015] Example 1, as Figure 1 As shown, this embodiment of the invention provides a method for managing experimental data of antimicrobial peptides, including: S1. Obtain multiple experimental datasets of antimicrobial peptides and experimental conditions from multiple laboratories, and perform data confidence analysis based on multiple experimental condition datasets to obtain multiple data confidence sets.

[0016] Specifically, based on a distributed cloud platform architecture, an antimicrobial peptide data management system centered on the cloud platform is constructed. Each laboratory acts as an edge data acquisition node, transmitting local experimental data to the cloud platform for centralized storage in real time via standardized interface protocols. The cloud platform provides unified data format specifications and quality control mechanisms to ensure the consistency and comparability of data from laboratory nodes in different geographical locations. First, antimicrobial peptide-related data is collected from multiple laboratories. The acquired data consists of two components: an antimicrobial peptide experimental dataset and an experimental condition dataset. The antimicrobial peptide experimental dataset contains molecular characteristic information of the antimicrobial peptides and the corresponding minimum inhibitory concentration (MIC) determination results. This data reflects the antimicrobial effects of different antimicrobial peptide molecules under specific experimental conditions and forms the core data foundation for subsequent data analysis. The experimental condition dataset records the environmental parameters used to obtain the above experimental data, including key environmental factors such as pH, ion concentration, and temperature during the experiment. These experimental condition parameters have a significant impact on the activity of antimicrobial peptides; different experimental conditions may lead to different antimicrobial effects for the same antimicrobial peptide.

[0017] For the collected datasets of experimental conditions, a data confidence analysis was performed to evaluate the reliability and consistency of the experimental data. Specifically, the confidence level of the corresponding experimental data was quantified by calculating the similarity between the experimental condition data of each experiment in the multiple experimental condition datasets and the preset standard experimental condition data. The closer the experimental condition data is to the standard experimental condition data, the higher the confidence level of the corresponding experimental data; conversely, the experimental data corresponding to experiments whose experimental condition data deviates significantly from the standard experimental condition data are assigned a lower confidence level.

[0018] Through the above analysis and processing, data confidence sets corresponding to each antimicrobial peptide experimental dataset are obtained, resulting in multiple data confidence sets, which provide a quantitative basis for confidence assessment for subsequent data screening, analyzer training and quality control.

[0019] S2. Based on the multiple data confidence sets, select training datasets and test datasets from multiple antimicrobial peptide experimental datasets respectively, integrate and train multiple basic analyzers using multiple training datasets, and test multiple test datasets to obtain multiple test accuracy sets.

[0020] Specifically, firstly, a differentiated data selection strategy is implemented across multiple antimicrobial peptide experimental datasets based on multiple data confidence sets. Specifically, antimicrobial peptide experimental data with confidence levels above a pre-set confidence threshold are categorized as high-quality data and prioritized for use as the training dataset; correspondingly, experimental data with confidence levels below this threshold are used as the test dataset. This tiered selection ensures that the training phase uses samples with relatively standard experimental conditions and relatively reliable data quality, while data of relatively lower quality or with experimental conditions deviating from the standard are used for generalization ability verification tests.

[0021] Then, using an ensemble learning approach, multiple fundamental analyzers with identical structures but slightly different training data were constructed. Each fundamental analyzer was built based on a machine learning algorithm, using antimicrobial peptide features from antimicrobial peptide experimental data as input variables and minimum inhibitory concentration (MIC) as the prediction target output. These fundamental analyzers were trained under supervised supervision using multiple training datasets, and each fundamental analyzer learned the intrinsic correlation between antimicrobial peptide molecular features and their antimicrobial activity during the training process. The training process continued until the loss functions of each analyzer converged, ensuring stable prediction performance and resulting in multiple fundamental analyzers.

[0022] After training, the performance of each base analyzer was evaluated using multiple test datasets. Each analyzer predicted the minimum inhibitory concentration (MIC) for the test samples. By calculating the similarity or error index between the predicted results and the true values, the prediction accuracy of each analyzer on different test datasets was quantified, thus obtaining multiple test accuracy sets. These multiple test accuracy sets provide crucial performance evaluation criteria for subsequent environmental impact credibility analysis and optimal analyzer selection.

[0023] S3. Verify the confidence levels of multiple test precision sets and multiple test datasets to obtain multiple environmental influence confidence levels. Perform fluctuation anomaly analysis on multiple experimental condition datasets and correct the multiple environmental influence confidence levels to obtain multiple corrected environmental confidence levels.

[0024] Specifically, firstly, the data confidence scores corresponding to multiple test datasets are extracted and validated against multiple test precision sets. Specifically, the similarity between the data confidence scores and corresponding test precisions of each test dataset is calculated, and the overall environmental impact confidence score is obtained by averaging these scores, thus yielding multiple environmental impact confidence scores. The environmental impact confidence score reflects the actual impact of experimental condition deviations on the prediction accuracy of the basic analyzer; for experimental data with lower confidence scores, i.e., data whose experimental conditions deviate significantly from standard experimental conditions, a lower test precision should be correspondingly applied, and there should be a positive correlation between the two.

[0025] Considering that variations in experimental conditions may stem from various factors, including equipment sensor errors, accidental environmental fluctuations, or differences in human operation, fluctuation anomaly analysis is performed on multiple experimental condition datasets. Specifically, firstly, the magnitude of differences between adjacent experimental condition data within each dataset is statistically analyzed. When the magnitude of the difference reaches a preset condition difference threshold, it is identified as an abnormal fluctuation event. By counting the number of abnormal fluctuation events in each experimental condition dataset, the number of fluctuations for each is obtained, resulting in multiple fluctuation quantities. Subsequently, the ratio of each fluctuation quantity to the mean of the multiple fluctuation quantities is calculated as the fluctuation frequency, which is defined as the experimental condition fluctuation coefficient. This coefficient reflects the stability of each experimental environment and the reliability of data acquisition: a higher experimental condition fluctuation coefficient indicates less stable condition control in the corresponding experimental environment; a lower coefficient indicates a relatively stable experimental environment and more reliable data quality.

[0026] Subsequently, based on the analysis results of the experimental condition fluctuation coefficient, the initially obtained environmental impact credibility was corrected. The correction process directly multiplies the environmental impact credibility by the corresponding experimental condition fluctuation coefficient. Through this correction calculation, environmental impact assessments caused by equipment errors or accidental environmental fluctuations are effectively filtered out, resulting in more accurate and reliable corrected environmental credibility values. These corrected environmental credibility values ​​can more realistically reflect the essential impact of changes in experimental conditions on the antimicrobial peptide activity assay results, providing a more accurate evaluation basis for subsequent optimal analyzer selection.

[0027] S4. Based on the credibility of the multiple calibration environments, select the optimal analyzer and select a reliable experimental dataset within the antimicrobial peptide experimental dataset for experimental data management.

[0028] Specifically, based on the aforementioned assessment results of the credibility of the calibration environment, the optimal analyzer is selected and a credible experimental dataset is constructed to achieve high-quality management of antimicrobial peptide experimental data.

[0029] First, the validation accuracy of the constructed basic analyzers during training is obtained as the analytical accuracy of each analyzer. Then, the analytical accuracy of each analyzer is weighted and calculated with the corresponding calibration environment reliability to obtain the comprehensive analytical precision. This analytical precision comprehensively considers both the predictive ability of the basic analyzer itself and the impact of environmental conditions on data quality. By comparing the analytical precision of each basic analyzer, the basic analyzer with the highest analytical precision is selected as the optimal analyzer. This optimal analyzer not only possesses excellent antimicrobial peptide activity prediction capabilities but also maintains high predictive reliability even under changing environmental conditions.

[0030] The selected optimal analyzer was used to iterate through multiple antimicrobial peptide experimental datasets. For each dataset, the optimal analyzer predicted the corresponding minimum inhibitory concentration (MIC) based on its antimicrobial peptide characteristics and compared the predicted result with the experimentally measured MIC to calculate the matching accuracy. When the iterative testing accuracy of a dataset exceeded a preset accuracy threshold, the dataset was considered high-quality and reliable. Through this screening process, all antimicrobial peptide experimental datasets with acceptable testing accuracy were compiled to form a reliable experimental dataset. This reliable dataset, validated by the optimal analyzer, ensured data reliability and predictive consistency, and can serve as a high-quality data resource for subsequent antimicrobial peptide activity analysis, data mining, and predictive modeling research, achieving effective management and quality control of antimicrobial peptide experimental data.

[0031] The above technical solutions effectively identify and screen high-quality antimicrobial peptide experimental data, solving the problem of inconsistent data quality caused by differences in experimental conditions in traditional data management. Through data confidence analysis and environmental impact correction, interference from equipment errors and accidental environmental factors is effectively eliminated, improving the accuracy of data quality assessment.

[0032] Furthermore, multiple experimental datasets of antimicrobial peptides and experimental conditions from various laboratories were obtained. Data confidence analysis was performed on these datasets to obtain multiple data confidence sets, including: S11. Obtain multiple antimicrobial peptide experimental datasets and multiple experimental condition datasets from multiple laboratory experimental records. Each experimental condition dataset includes pH, ion concentration, and temperature, and each antimicrobial peptide experimental dataset includes antimicrobial peptide characteristics and minimum inhibitory concentration. S12. Perform data confidence analysis based on multiple experimental condition datasets to obtain multiple data confidence sets.

[0033] In a preferred embodiment, experimental data is first obtained from multiple laboratories, with each laboratory providing a complete antimicrobial peptide experimental dataset and its corresponding experimental condition dataset. Specifically, each antimicrobial peptide experimental dataset contains multiple antimicrobial peptide experimental data points tested by that laboratory. Each antimicrobial peptide experimental data point includes two parts: first, the antimicrobial peptide characteristics, covering structural features such as molecular weight, amino acid sequence, charge distribution, and hydrophobicity index; second, the minimum inhibitory concentration (MIC), which is the lowest concentration required for the antimicrobial peptide to inhibit the target pathogen. The corresponding experimental condition dataset for each antimicrobial peptide experimental dataset details the experimental environmental parameters used to obtain each antimicrobial peptide experimental data point. Each experimental condition dataset includes pH, ion concentration, and temperature. The pH value reflects the acidity or alkalinity of the experimental system, the ion concentration represents the ionic strength in the experimental buffer solution, and the temperature reflects the ambient temperature during the experiment.

[0034] Then, data confidence assessment was performed on multiple experimental condition datasets collected from various laboratories. The analysis process first established standard experimental condition data as a reference benchmark, and then calculated the similarity between each experimental condition dataset and the standard experimental condition data. The closer the experimental condition data of a certain experiment is to the standard experimental condition data, the higher the confidence level of the antimicrobial peptide experimental data obtained in that experiment is considered; conversely, the further the experimental condition data deviates from the standard experimental condition data, the lower the confidence level of the corresponding data.

[0035] Through the above confidence analysis, a corresponding confidence score sequence was generated for the antimicrobial peptide experimental dataset of each laboratory, resulting in multiple data confidence sets. This provides a quantitative reliability assessment basis for subsequent data stratification processing, analyzer training, and quality control.

[0036] Furthermore, data confidence analysis was performed on multiple experimental condition datasets to obtain multiple data confidence sets, including: S121. Based on multiple experimental condition datasets, statistical processing is performed to obtain standard experimental condition data; S122. Calculate the similarity between each experimental condition data and the standard experimental condition data, and use it as the data confidence score to obtain multiple data confidence score sets.

[0037] In one feasible implementation, firstly, statistical analysis is performed on all experimental condition datasets collected from multiple laboratories to establish standard experimental condition data as an evaluation benchmark. Specifically, three parameters—pH, ion concentration, and temperature—are extracted from each experimental condition dataset, and statistical analysis is performed on each parameter to obtain standard experimental condition data. For example, by calculating the median value of each parameter across all experimental condition data, standard values ​​for pH, ion concentration, and temperature are determined, and these three standard values ​​are combined to form standard experimental condition data. This standard experimental condition data represents typical condition parameters commonly used by multiple laboratories in antimicrobial peptide experiments, serving as a unified reference benchmark for subsequent data confidence assessment. The establishment of standard conditions eliminates differences in operating habits between different laboratories, providing a consistent evaluation standard for data quality assessment.

[0038] Then, based on the established standard experimental conditions data, the similarity between each experimental condition data and the standard experimental conditions data is calculated one by one. Specifically, firstly, the three parameters of pH value, ion concentration, and temperature in each experimental condition data and the standard experimental conditions data are normalized, mapping each parameter value to the [0,1] interval to eliminate the influence of dimensional differences between different parameters, resulting in normalized experimental condition data and standard experimental condition data. Subsequently, the Euclidean distance between the normalized experimental condition data and the standard experimental conditions data is calculated, obtaining the distance value between each experimental condition data and the standard experimental conditions data. Next, using a pre-set distance-similarity mapping table, the calculated distance value is used as input to directly obtain the corresponding similarity. This distance-similarity mapping table was pre-established by experts in the field of antimicrobial peptide research based on their experience regarding the impact of experimental condition deviations on the results of antimicrobial peptide activity determination. The similarity is directly used as the data confidence level of the corresponding antimicrobial peptide experimental data: the higher the similarity, the closer the experimental conditions are to the standard conditions, and the higher the reliability of the corresponding experimental data; the lower the similarity, the more the experimental conditions deviate from the standard, and the lower the data confidence level. By calculating the experimental conditions data of each laboratory one by one, a corresponding data confidence set is generated for each antimicrobial peptide experimental dataset, and finally multiple data confidence sets are obtained.

[0039] Through the above steps, a standard experimental condition benchmark is objectively established, effectively quantifying the reliability level of data under each experimental condition. At the same time, through normalization processing and the application of mapping tables, the consistency and accuracy of data quality assessment among different laboratories are ensured, providing a reliable quality control basis for subsequent data screening and analyzer training.

[0040] Furthermore, based on the multiple data confidence sets, training and test datasets are selected from multiple antimicrobial peptide experimental datasets, and multiple basic analyzers are trained using the multiple training datasets, including: S21. Based on the multiple data confidence sets, filter the antimicrobial peptide experimental data that are greater than or equal to and less than the data confidence threshold respectively to obtain multiple first antimicrobial peptide experimental datasets and multiple second antimicrobial peptide experimental datasets. S22. Randomly select antimicrobial peptide experimental data from the multiple first antimicrobial peptide experimental datasets and multiple second antimicrobial peptide experimental datasets to obtain multiple training datasets and multiple test datasets. S23. Based on machine learning, multiple basic analyzers with identical structures are constructed using antimicrobial peptide experimental data, including antimicrobial peptide characteristics and minimum inhibitory concentration, as inputs and outputs. S24. Supervised training and validation of the multiple basic analyzers are performed using the multiple training datasets respectively, and the training is completed after convergence.

[0041] In a preferred embodiment, the antimicrobial peptide experimental datasets are stratified and filtered based on multiple obtained data confidence sets. Specifically, a predetermined data confidence threshold is set as a data quality classification standard. This data confidence threshold can be set by experts according to the actual data distribution and application requirements, typically using the median or a specific percentile of the data confidence score as the threshold standard. Specifically, each antimicrobial peptide experimental dataset is traversed, and the corresponding data confidence set is obtained from multiple data confidence sets. For the antimicrobial peptide experimental data within each dataset, the data is classified by comparing its corresponding data confidence set with the data confidence threshold. For each antimicrobial peptide experimental dataset, the antimicrobial peptide experimental data with a data confidence score greater than or equal to the data confidence threshold are selected to form a corresponding first antimicrobial peptide experimental dataset; the antimicrobial peptide experimental data with a data confidence score less than the data confidence threshold are selected to form a corresponding second antimicrobial peptide experimental dataset. By processing each antimicrobial peptide experimental dataset individually, multiple first antimicrobial peptide experimental datasets and multiple second antimicrobial peptide experimental datasets were obtained, realizing hierarchical data organization based on data quality, laying the foundation for subsequent differentiated applications.

[0042] After data stratification, random sampling was performed from the obtained first and second antimicrobial peptide experimental datasets to obtain multiple training and test datasets. Specifically, each first antimicrobial peptide experimental dataset was traversed, and a predetermined number or proportion of antimicrobial peptide experimental data was randomly selected from each dataset to form a training dataset. Since the first antimicrobial peptide experimental datasets contain high-confidence data, the resulting training datasets have high data quality and can provide reliable learning samples for analyzer training. Simultaneously, each second antimicrobial peptide experimental dataset was traversed, and a predetermined number or proportion of antimicrobial peptide experimental data was randomly selected from each dataset to form a test dataset. The data in the second antimicrobial peptide experimental datasets have relatively low confidence; using them as test data can verify the model's predictive robustness under experimental condition deviations. Through this random sampling process, multiple training and test datasets were obtained, providing a data foundation for the subsequent training and performance evaluation of the basic analyzer.

[0043] Then, based on machine learning techniques, multiple fundamental analyzers with identical structures were constructed. Each fundamental analyzer employs the same neural network architecture, including an input layer, hidden layers, and an output layer. The input layer receives antimicrobial peptide features, such as molecular weight, amino acid sequence features, charge distribution, and hydrophobicity index, while the output layer predicts the corresponding minimum inhibitory concentration (MIC). Although the fundamental analyzers have the same structure, they are trained using different training datasets, allowing them to learn different feature patterns in the data. Then, supervised training of the corresponding fundamental analyzers is performed using multiple training datasets. For each fundamental analyzer, parameters are optimized using its corresponding training dataset, and the network weights are continuously adjusted through backpropagation to gradually reduce the error between the predicted output and the actual MIC. During training, model validation is performed simultaneously, monitoring the trend of the loss function and the improvement in prediction accuracy. When the loss function value stabilizes and the prediction accuracy no longer significantly improves in multiple training rounds, the fundamental analyzer is considered to have reached convergence, completing the training process.

[0044] Through the above training methods, each basic analyzer can fully learn the intrinsic correlation between antimicrobial peptide characteristics and antimicrobial activity in high-quality training data, ensuring that the model has accurate antimicrobial peptide activity prediction capabilities, and providing a reliable model foundation for subsequent test accuracy evaluation and optimal analyzer selection.

[0045] Furthermore, tests were conducted on multiple test datasets to obtain multiple test accuracy sets, including: S25. Using the multiple test datasets, test multiple basic analyzers respectively to obtain multiple sets of output; S26. Calculate the similarity between multiple test minimum inhibitory concentration sets and the true minimum inhibitory concentrations in multiple test datasets to obtain multiple test accuracy sets.

[0046] In a preferred embodiment, after each base analyzer has completed training, its performance is tested using multiple test datasets. Specifically, each test dataset is input into its corresponding base analyzer, and each base analyzer calculates the minimum inhibitory concentration (MIC) prediction based on the antimicrobial peptide characteristics of the test samples in the test dataset. Through this testing process, each base analyzer outputs a set of predicted MIC values ​​for its corresponding test dataset, resulting in multiple test MIC sets. These prediction results reflect the prediction performance of each base analyzer on relatively low-confidence data, providing foundational data for subsequent accuracy evaluation.

[0047] Then, based on the obtained multiple sets of minimum inhibitory concentrations (MICs), a quantitative evaluation of prediction accuracy is performed. Specifically, multiple basic analyzers are traversed, and one basic analyzer is selected as the first basic analyzer at a time. Then, the corresponding set of minimum inhibitory concentrations (MICs) is extracted from the multiple sets of MICs, denoted as the first set of MICs, and the corresponding test dataset is extracted from the test dataset, denoted as the first test dataset. Subsequently, the MICs in the first set of MICs are compared one by one with the actual measured MICs in the first test dataset. The relative error of each pair of MICs is calculated, where the relative error = |MIC - MIC| / MIC. The test accuracy of the corresponding sample is obtained by subtracting the relative error from 1, thus obtaining the test accuracy set corresponding to the first basic analyzer. By calculating the accuracy of each basic analyzer, multiple test accuracy sets are finally obtained, providing a quantitative performance evaluation basis for subsequent environmental impact credibility analysis and optimal analyzer selection.

[0048] The above testing process allows for an objective evaluation of the predictive stability of each basic analyzer under experimental condition deviations, providing support for the reliability assessment of the basic analyzers.

[0049] Furthermore, the confidence levels of multiple test precision sets and multiple test datasets are verified to obtain multiple environmental impact confidence sets. Fluctuation anomaly analysis is then performed on multiple experimental condition datasets, and the multiple environmental impact confidence sets are corrected to obtain multiple corrected environmental confidence sets, including: S31. Obtain multiple data confidence sets corresponding to multiple test datasets, calculate the similarity with the multiple test precision sets, and calculate the mean to obtain multiple environmental influence confidence levels. S32. Perform fluctuation anomaly analysis on the multiple experimental condition datasets to obtain multiple experimental condition fluctuation coefficients, correct and calculate the credibility of the multiple environmental influences, and obtain multiple corrected environmental credibility.

[0050] In a preferred embodiment, firstly, the corresponding data confidence sets for each test dataset are obtained. Specifically, for each test dataset, the corresponding data confidence set is found within the multiple data confidence sets obtained. Then, the similarity between the data confidence set and the corresponding test precision set for each test dataset is calculated. The similarity calculation uses a relative difference method. Specifically, for each test dataset, its data confidence set and test precision set are paired one by one, and the absolute difference between each pair of data confidence and test precision is calculated, i.e., |data confidence - test precision|. Subsequently, the average of all absolute differences corresponding to the test dataset is calculated to obtain the mean absolute difference. 1 - mean absolute difference is taken as the environmental impact confidence level corresponding to that test dataset. By calculating for each of all test datasets, multiple environmental impact confidence levels are obtained, with each environmental impact confidence level corresponding to one test dataset.

[0051] Then, fluctuation anomaly analysis was performed on multiple experimental condition datasets to identify non-substantial condition changes. Specifically, each experimental condition dataset was traversed, and the difference magnitude between adjacent experimental condition data within each dataset was calculated. When the difference magnitude exceeded a preset condition difference threshold, the change was identified as an abnormal fluctuation event. The frequency of abnormal fluctuation events in each experimental condition dataset was counted to calculate the fluctuation frequency, resulting in multiple experimental condition fluctuation coefficients, each corresponding to one experimental condition dataset. Subsequently, the obtained environmental impact confidence level was multiplied by the corresponding experimental condition fluctuation coefficient, and a correction calculation was performed to obtain multiple corrected environmental confidence levels. This correction process effectively eliminates the interference of equipment errors and accidental environmental fluctuations on the environmental impact assessment.

[0052] Through the above verification and correction process, the actual environmental impact effects can be accurately identified, the interference of accidental factors can be eliminated, and a reliable evaluation basis can be provided for the subsequent selection of the optimal analyzer.

[0053] Furthermore, fluctuation anomaly analysis is performed on the multiple experimental condition datasets to obtain multiple experimental condition fluctuation coefficients. The confidence levels of the multiple environmental influences are then corrected to obtain multiple corrected environmental confidence levels, including: S321. Count the number of adjacent experimental condition data within the multiple experimental condition datasets that reach the condition difference threshold, and obtain multiple fluctuation quantities. S322. Calculate multiple fluctuation frequencies based on multiple fluctuation quantities, and use them as fluctuation coefficients for multiple experimental conditions; S323. Based on the fluctuation coefficients of the multiple experimental conditions, the confidence levels of multiple environmental influences are corrected and calculated to obtain multiple corrected environmental confidence levels.

[0054] In a preferred embodiment, firstly, multiple experimental condition datasets are traversed, and fluctuation anomaly statistics are performed on each dataset. Specifically, for each experimental condition dataset, the difference magnitude between adjacent experimental condition data is calculated one by one. For each pair of adjacent experimental condition data, the difference values ​​of three parameters—pH value, ion concentration, and temperature—are calculated, and the difference magnitude is obtained by calculating the Euclidean distance for that pair of data. When the difference magnitude reaches or exceeds a preset condition difference threshold, the change is recorded as an abnormal fluctuation event. The condition difference threshold is a critical standard for determining whether abnormal fluctuations have occurred in experimental conditions; this threshold is preset based on the condition stability requirements of the antimicrobial peptide experiment. The number of abnormal fluctuation events in each experimental condition dataset is counted, obtaining multiple fluctuation counts, with each fluctuation count corresponding to one experimental condition dataset.

[0055] Subsequently, based on the obtained multiple fluctuation counts, the fluctuation frequency of each experimental condition dataset is calculated. Specifically, the mean of all fluctuation counts is calculated as a baseline value, and then the ratio of each fluctuation count to this mean is calculated; this ratio represents the fluctuation frequency of the corresponding experimental condition dataset. Each fluctuation frequency is used as the corresponding experimental condition fluctuation coefficient, resulting in multiple experimental condition fluctuation coefficients.

[0056] Subsequently, based on the obtained experimental condition fluctuation coefficients, the obtained environmental impact confidence levels were corrected. Specifically, each environmental impact confidence level was multiplied by its corresponding experimental condition fluctuation coefficient, and the correction calculation was performed through multiplication to obtain multiple corrected environmental confidence levels.

[0057] Through the above fluctuation analysis and correction, we can effectively identify and compensate for the deviations in data quality assessment caused by equipment errors and accidental environmental fluctuations, thereby improving the accuracy and reliability of environmental impact assessment.

[0058] Furthermore, based on the credibility of the multiple calibration environments, the optimal analyzer is selected, and a reliable experimental dataset is selected within the antimicrobial peptide experimental dataset for experimental data management, including: S41. Obtain the analysis accuracy of multiple basic analyzers, and calculate the analysis precision of multiple basic analyzers by combining the confidence levels of multiple calibration environments. S42. Select the basic analyzer with the highest analysis accuracy as the optimal analyzer; S43. Using the aforementioned optimal analyzer, perform traversal testing on multiple antimicrobial peptide experimental datasets, and use the antimicrobial peptide experimental data with traversal test accuracy greater than the test accuracy threshold as the reliable experimental dataset for experimental data management.

[0059] In a preferred embodiment, firstly, multiple basic analyzers are traversed, and the validation accuracy of each basic analyzer during the training process is obtained as the corresponding basic analyzer analysis accuracy. This analysis accuracy is obtained through prediction performance on the validation set and reflects the intrinsic prediction performance of each basic analyzer under training conditions. Subsequently, the analysis accuracy of each basic analyzer is weighted and calculated with the obtained corresponding calibration environment confidence. Specifically, for each basic analyzer, its analysis accuracy and corresponding calibration environment confidence are weighted and combined to obtain the analysis precision of that basic analyzer. For example, when more emphasis is placed on the performance of the model itself, the weight of the analysis accuracy can be set to 0.6, and the weight of the calibration environment confidence can be set to 0.4; if the analysis accuracy of a certain basic analyzer is 0.85 and the corresponding calibration environment confidence is 0.92, then the analysis precision of the basic analyzer is 0.85×0.6+0.92×0.4=0.878. This analysis precision considers both the prediction accuracy of the basic analyzer itself and the degree of influence of changes in environmental conditions on prediction reliability. The analytical accuracy of multiple fundamental analyzers is obtained by calculating each of the fundamental analyzers one by one.

[0060] Then, the analysis precision of multiple basic analyzers is iterated, and the maximum analysis precision is found through numerical comparison. The corresponding basic analyzer is selected as the optimal analyzer. This optimal analyzer performs best in the comprehensive evaluation, possessing both excellent antimicrobial peptide activity prediction capability and maintaining high prediction stability and reliability even when experimental conditions are biased. Subsequently, the selected optimal analyzer is used to iterate through all original antimicrobial peptide experimental datasets. Specifically, for each antimicrobial peptide experimental dataset, the antimicrobial peptide features of each antimicrobial peptide experimental data are input into the optimal analyzer to obtain the predicted minimum inhibitory concentration (MIC). Then, the predicted MIC is compared with the actual MIC of the antimicrobial peptide experimental data, and the iterative test precision of the antimicrobial peptide experimental data is calculated using the relative error method, where the iterative test precision = 1 - |predicted MIC - actual MIC| / actual MIC. When the iterative test precision of an antimicrobial peptide experimental data is greater than the preset test precision threshold, the antimicrobial peptide experimental data is marked as reliable experimental data. All experimental data of antimicrobial peptides that passed the accuracy test were compiled to form a reliable experimental dataset.

[0061] The acquired reliable experimental datasets have undergone rigorous quality verification by the optimal analyzer, ensuring high data reliability and predictive consistency. This provides a reliable data foundation for subsequent applications such as antimicrobial peptide research, drug development, and database construction, enabling high-quality management of antimicrobial peptide experimental data.

[0062] Furthermore, the analytical accuracy of multiple fundamental analyzers is obtained, and combined with the confidence levels of multiple calibration environments, the analytical precision of multiple fundamental analyzers is calculated, including: S411. The accuracy verified during the training of multiple basic analyzers is used as the multiple analysis accuracy. S412. The analytical accuracy of multiple basic analyzers is obtained by weighted calculation based on multiple analytical accuracy rates and multiple calibration environment confidence levels.

[0063] In a preferred embodiment, firstly, during the supervised training of the basic analyzers, each basic analyzer obtains a validation accuracy by calculating the degree of matching between the predicted results and the true values ​​during the training validation process. The validation accuracy of each basic analyzer at training convergence is then used as the corresponding analytical accuracy, resulting in multiple analytical accuracies. These analytical accuracies reflect the inherent predictive performance level of each basic analyzer under standard training conditions.

[0064] Then, a weighted comprehensive calculation is performed based on the obtained multiple analytical accuracies and multiple calibration environment confidence scores. Specifically, for each basic analyzer, its analytical accuracy and corresponding calibration environment confidence score are linearly combined according to preset weights. The weighting coefficients are determined according to actual application requirements. For example, when more emphasis is placed on the predictive performance of the model itself, the analytical accuracy weight can be set to 0.6 and the calibration environment confidence score weight to 0.4; when more emphasis is placed on environmental adaptability, the analytical accuracy weight can be set to 0.4 and the calibration environment confidence score weight to 0.6. By performing weighted calculations on all basic analyzers one by one, the analytical accuracy of multiple basic analyzers is obtained.

[0065] The above calculation method enables a comprehensive evaluation of the model's predictive ability and environmental adaptability, providing a quantitative basis for selecting the optimal analyzer.

[0066] Example 2, as Figure 2 As shown, based on the same inventive concept as the antimicrobial peptide experimental data management method provided in Embodiment 1, this embodiment of the invention also provides an antimicrobial peptide experimental data management system. This system can serve as the core architecture of a cloud platform, implementing the antimicrobial peptide experimental data management method by incorporating a smart chip. Multiple laboratories act as edge data acquisition nodes, and the smart chip is responsible for processing the antimicrobial peptide experimental data from each edge data acquisition node, and performing core functions such as data confidence analysis, environmental impact assessment, and reliable data screening.

[0067] The antimicrobial peptide experimental data management system includes: The data confidence analysis module 11 is used to acquire multiple antimicrobial peptide experimental datasets and experimental condition datasets from multiple laboratories, and to perform data confidence analysis based on multiple experimental condition datasets to obtain multiple data confidence sets. The test accuracy evaluation module 12 is used to select training datasets and test datasets from multiple antimicrobial peptide experimental datasets according to the multiple data confidence sets, integrate and train multiple basic analyzers using multiple training datasets, test multiple test datasets, and obtain multiple test accuracy sets. The credibility verification and correction module 13 is used to verify the data confidence of multiple test precision sets and multiple test datasets, obtain multiple environmental influence credibility, and perform fluctuation anomaly analysis on multiple experimental condition datasets and correct the multiple environmental influence credibility to obtain multiple corrected environmental credibility. The experimental data management module 14 is used to select the optimal analyzer based on the credibility of the multiple calibration environments, and to select a reliable experimental dataset within the antimicrobial peptide experimental dataset for experimental data management.

[0068] Furthermore, the execution steps of the data confidence analysis module 11 include: Multiple antimicrobial peptide experimental datasets and multiple experimental condition datasets were obtained from multiple laboratory experimental records. Each experimental condition dataset included pH, ion concentration, and temperature, and each antimicrobial peptide experimental dataset included antimicrobial peptide characteristics and minimum inhibitory concentration. Data confidence analysis was performed on multiple experimental condition datasets to obtain multiple data confidence sets.

[0069] Furthermore, the execution steps of the data confidence analysis module 11 also include: Standard experimental condition data were obtained by statistical processing based on multiple experimental condition datasets. Calculate the similarity between each experimental condition data point and the standard experimental condition data, and use this as the data confidence score to obtain multiple data confidence score sets.

[0070] Furthermore, the execution steps of the test accuracy evaluation module 12 include: Based on the multiple data confidence sets, antimicrobial peptide experimental data that are greater than or equal to and less than the data confidence threshold are respectively filtered to obtain multiple first antimicrobial peptide experimental datasets and multiple second antimicrobial peptide experimental datasets; Randomly select antimicrobial peptide experimental data from the multiple first antimicrobial peptide experimental datasets and multiple second antimicrobial peptide experimental datasets to obtain multiple training datasets and multiple test datasets; Based on machine learning, multiple basic analyzers with identical structures are constructed using antimicrobial peptide experimental data, including antimicrobial peptide characteristics and minimum inhibitory concentration, as inputs and outputs. The multiple basic analyzers are trained and validated using the multiple training datasets respectively, and the training is completed after convergence.

[0071] Furthermore, the execution steps of the test accuracy evaluation module 12 also include: Using the aforementioned multiple test datasets, multiple basic analyzers were tested to obtain multiple sets of minimum inhibitory concentrations (MICs) for output. Calculate the similarity between multiple test minimum inhibitory concentration sets and the true minimum inhibitory concentrations in multiple test datasets to obtain multiple test accuracy sets.

[0072] Furthermore, the execution steps of the credibility verification and correction module 13 include: Obtain multiple data confidence sets corresponding to multiple test datasets, calculate the similarity with the multiple test precision sets, and calculate the mean to obtain multiple environmental influence confidence levels; Fluctuation anomaly analysis was performed on the multiple experimental condition datasets to obtain multiple experimental condition fluctuation coefficients. The confidence levels of the multiple environmental influences were then corrected to obtain multiple corrected environmental confidence levels.

[0073] Furthermore, the execution steps of the credibility verification and correction module 13 also include: The number of times the difference between adjacent experimental condition data within the multiple experimental condition datasets reaches the condition difference threshold is counted to obtain multiple fluctuation quantities. Based on multiple fluctuation quantities, multiple fluctuation frequencies are calculated and used as fluctuation coefficients for multiple experimental conditions; Based on the fluctuation coefficients of the multiple experimental conditions, the credibility of multiple environmental impacts is corrected and calculated to obtain multiple corrected environmental credibility.

[0074] Furthermore, the execution steps of the experimental data management module 14 include: The analysis accuracy of multiple basic analyzers is obtained, and the analysis precision of multiple basic analyzers is calculated by combining the confidence levels of multiple calibration environments. The basic analyzer with the highest analysis accuracy is selected as the optimal analyzer; The optimal analyzer is used to perform traversal tests on multiple antimicrobial peptide experimental datasets. Antimicrobial peptide experimental datasets with traversal test accuracy greater than the test accuracy threshold are used as reliable experimental datasets for experimental data management.

[0075] Furthermore, the execution steps of the experimental data management module 14 also include: The accuracy rates verified during the training of multiple fundamental analyzers are used as multiple analysis accuracies. The analytical accuracy of multiple basic analyzers is obtained by weighted calculation based on multiple analytical accuracy rates and multiple calibration environment confidence levels.

[0076] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0077] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0078] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0079] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0080] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0081] Although preferred embodiments of the invention have been described, those skilled in the art, once they have learned the basic inventive concept, can make other changes and modifications to these embodiments.

[0082] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.

Claims

1. An antibacterial peptide experimental data management method, characterized by, The method includes: We obtained multiple experimental datasets of antimicrobial peptides and experimental conditions from various laboratories, and conducted data confidence analysis based on these datasets to obtain multiple data confidence sets. Based on the multiple data confidence sets, training datasets and test datasets are selected from multiple antimicrobial peptide experimental datasets. Multiple basic analyzers are trained by integrating multiple training datasets and tested on multiple test datasets to obtain multiple test accuracy sets. The confidence levels of data from multiple test precision sets and multiple test datasets are verified to obtain multiple environmental influence confidence levels. Fluctuation anomaly analysis is performed on multiple experimental condition datasets and the multiple environmental influence confidence levels are corrected to obtain multiple corrected environmental confidence levels. Based on the credibility of the multiple calibration environments, the optimal analyzer is selected, and a reliable experimental dataset is selected within the antimicrobial peptide experimental dataset for experimental data management. Multiple experimental datasets of antimicrobial peptides and experimental conditions from various laboratories were obtained. Data confidence analysis was performed on these datasets to obtain multiple data confidence sets, including: Multiple antimicrobial peptide experimental datasets and multiple experimental condition datasets were obtained from multiple laboratory experimental records. Each experimental condition dataset included pH, ion concentration, and temperature, and each antimicrobial peptide experimental dataset included antimicrobial peptide characteristics and minimum inhibitory concentration. Data confidence analysis was performed on multiple experimental condition datasets to obtain multiple data confidence sets; Based on the multiple data confidence sets, training and test datasets are selected from multiple antimicrobial peptide experimental datasets, and multiple basic analyzers are trained by integrating the multiple training datasets, including: Based on the multiple data confidence sets, antimicrobial peptide experimental data that are greater than or equal to and less than the data confidence threshold are respectively filtered to obtain multiple first antimicrobial peptide experimental datasets and multiple second antimicrobial peptide experimental datasets; Randomly select antimicrobial peptide experimental data from the multiple first antimicrobial peptide experimental datasets and multiple second antimicrobial peptide experimental datasets to obtain multiple training datasets and multiple test datasets; Based on machine learning, multiple basic analyzers with identical structures are constructed using antimicrobial peptide experimental data, including antimicrobial peptide characteristics and minimum inhibitory concentration, as inputs and outputs. The multiple basic analyzers are trained and validated using the multiple training datasets respectively, and the training is completed after convergence.

2. The method of claim 1, wherein the method is characterized by, Data confidence analysis was performed on multiple experimental condition datasets to obtain multiple data confidence sets, including: Standard experimental condition data were obtained by statistical processing based on multiple experimental condition datasets. Calculate the similarity between each experimental condition data point and the standard experimental condition data, and use this as the data confidence score to obtain multiple data confidence score sets.

3. The method for managing experimental data of antimicrobial peptides according to claim 1, characterized in that, Multiple test datasets were tested to obtain multiple test accuracy sets, including: Using the aforementioned multiple test datasets, multiple basic analyzers were tested to obtain multiple sets of minimum inhibitory concentrations (MICs) for output. Calculate the similarity between multiple test minimum inhibitory concentration sets and the true minimum inhibitory concentrations in multiple test datasets to obtain multiple test accuracy sets.

4. The method for managing experimental data of antimicrobial peptides according to claim 1, characterized in that, The confidence levels of data from multiple test precision sets and multiple test datasets are validated to obtain multiple environmental impact confidence sets. Furthermore, fluctuation anomaly analysis is performed on multiple experimental condition datasets, and the multiple environmental impact confidence sets are corrected to obtain multiple corrected environmental confidence sets, including: Obtain multiple data confidence sets corresponding to multiple test datasets, calculate the similarity with the multiple test precision sets, and calculate the mean to obtain multiple environmental influence confidence levels; Fluctuation anomaly analysis was performed on the multiple experimental condition datasets to obtain multiple experimental condition fluctuation coefficients. The confidence levels of the multiple environmental influences were then corrected to obtain multiple corrected environmental confidence levels.

5. The method for managing experimental data of antimicrobial peptides according to claim 4, characterized in that, Fluctuation anomaly analysis was performed on the multiple experimental condition datasets to obtain multiple experimental condition fluctuation coefficients. The confidence levels of the multiple environmental influences were then corrected to obtain multiple corrected environmental confidence levels, including: The number of times the difference between adjacent experimental condition data within the multiple experimental condition datasets reaches the condition difference threshold is counted to obtain multiple fluctuation quantities. Based on multiple fluctuation quantities, multiple fluctuation frequencies are calculated and used as fluctuation coefficients for multiple experimental conditions; Based on the fluctuation coefficients of the multiple experimental conditions, the credibility of multiple environmental impacts is corrected and calculated to obtain multiple corrected environmental credibility.

6. The method for managing experimental data of antimicrobial peptides according to claim 1, characterized in that, Based on the credibility of the multiple calibration environments, the optimal analyzer is selected, and a reliable experimental dataset is selected within the antimicrobial peptide experimental dataset for experimental data management, including: The analysis accuracy of multiple basic analyzers is obtained, and the analysis precision of multiple basic analyzers is calculated by combining the confidence levels of multiple calibration environments. The basic analyzer with the highest analysis accuracy is selected as the optimal analyzer; The optimal analyzer is used to perform traversal tests on multiple antimicrobial peptide experimental datasets. Antimicrobial peptide experimental datasets with traversal test accuracy greater than the test accuracy threshold are used as reliable experimental datasets for experimental data management.

7. The method for managing experimental data of antimicrobial peptides according to claim 6, characterized in that, The analysis accuracy of multiple fundamental analyzers is obtained, and combined with the confidence levels of multiple calibration environments, the analysis precision of multiple fundamental analyzers is calculated, including: The accuracy rates verified during the training of multiple fundamental analyzers are used as multiple analysis accuracies. The analytical accuracy of multiple basic analyzers is obtained by weighted calculation based on multiple analytical accuracy rates and multiple calibration environment confidence levels.

8. A data management system for antimicrobial peptide experiments, characterized in that, For implementing the antimicrobial peptide experimental data management method as described in any one of claims 1 to 7, the system comprises: The data confidence analysis module is used to acquire multiple experimental datasets of antimicrobial peptides and experimental conditions from multiple laboratories, and to perform data confidence analysis based on multiple experimental condition datasets to obtain multiple data confidence sets. The test accuracy evaluation module is used to select training datasets and test datasets from multiple antimicrobial peptide experimental datasets based on the multiple data confidence sets, integrate and train multiple basic analyzers using multiple training datasets, test multiple test datasets, and obtain multiple test accuracy sets. The credibility verification and correction module is used to verify the data confidence of multiple test precision sets and multiple test datasets, obtain multiple environmental influence credibility, and perform fluctuation anomaly analysis on multiple experimental condition datasets and correct the multiple environmental influence credibility to obtain multiple corrected environmental credibility. The experimental data management module is used to select the optimal analyzer based on the credibility of the multiple calibration environments, and to select a reliable experimental dataset within the antimicrobial peptide experimental dataset for experimental data management.

Citation Information

Patent Citations

  • Reliability measurement in data analysis of altered data sets

    CN107851465A