A sample analysis method and system based on a multi-dimensional gas chromatograph data fusion algorithm

Through the data fusion algorithm of multidimensional gas chromatograph, the inconsistency problem in data processing of multidimensional gas chromatography has been solved, the reliability and accuracy of the analysis results have been improved, and its application in multiple fields has been promoted.

CN119783038BActive Publication Date: 2025-10-10河南省许昌生态环境监测中心 +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411990230.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-10-10
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

Traditional single-dimensional gas chromatography is difficult to meet the needs of complex sample analysis, and multidimensional gas chromatography has problems with data format and standard inconsistency in data processing, which affects the accuracy and efficiency of the analysis results.

Method used

A data fusion algorithm based on multidimensional gas chromatograph is used to ensure the consistency and accuracy of data during the fusion process by identifying sample types, statistically analyzing detection proportions, screening fusion algorithms, setting debugging parameters, and remotely controlling debugging.

Benefits of technology

It has improved the reliability and accuracy of the analytical results of multidimensional gas chromatography, and promoted its widespread application in chemistry, environmental science, food safety, and drug development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119783038B_ABST
    Figure CN119783038B_ABST
Patent Text Reader

Abstract

The application discloses a sample analysis method and system based on a multi-dimensional gas chromatograph data fusion algorithm, and belongs to the technical field of sample analysis. The method comprises the following steps: step one, collecting multi-dimensional gas chromatograph information, and setting a data fusion algorithm according to the multi-dimensional gas chromatograph information; step two, identifying a detection sample, detecting the detection sample through a multi-dimensional gas chromatograph, and obtaining corresponding sample detection data; step three, determining the debugging parameters of the data fusion algorithm according to the sample detection data, and debugging the data fusion algorithm according to the debugging parameters; step four, performing fusion processing on the sample detection data through the debugged data fusion algorithm, and obtaining fusion detection data; and step five, obtaining a sample analysis target of a user, processing the fusion detection data according to the sample analysis target, and obtaining a sample analysis result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of sample analysis, and in particular relates to a sample analysis method and system based on a multidimensional gas chromatograph data fusion algorithm. Background Art

[0002] Gas chromatography is a precision technique widely used in chemical analysis, demonstrating its strengths in analyzing complex samples. However, with the advancement of science and technology and the increasing complexity of samples, traditional single-dimensional gas chromatography has become inadequate for analytical purposes. To overcome this challenge, multidimensional gas chromatography has emerged.

[0003] The core of multidimensional gas chromatography lies in the combined use of two or more distinct analytical methods, each with its own advantages, to analyze compounds. Among these, the most well-known multidimensional analytical technique is gas chromatography-mass spectrometry (GC-MS). This technique combines the separation capabilities of gas chromatography with the detection capabilities of mass spectrometry, significantly improving the accuracy and sensitivity of analytical results. However, traditional multidimensional gas chromatography also presents certain challenges in data processing. Because data from different dimensions and detectors may have different formats and standards, data inconsistencies may be encountered during data fusion. Furthermore, as data volumes increase, the complexity of data processing and analysis continues to increase, placing higher demands on data processing algorithms and computing power.

[0004] In order to solve the above problems, the present invention provides a sample analysis method and system based on a multi-dimensional gas chromatograph data fusion algorithm. Summary of the Invention

[0005] In order to solve the problems existing in the above solutions, the present invention provides a sample analysis method and system based on a multi-dimensional gas chromatograph data fusion algorithm.

[0006] The purpose of the present invention can be achieved through the following technical solutions:

[0007] A sample analysis method based on a multidimensional gas chromatograph data fusion algorithm, the method comprising:

[0008] Step 1: Collect multi-dimensional gas chromatograph information, and set a data fusion algorithm according to the multi-dimensional gas chromatograph information;

[0009] Furthermore, the setting method of the data fusion algorithm includes:

[0010] determining a sample detection range of the multidimensional gas chromatograph according to the multidimensional gas chromatograph information;

[0011] Determining potential fusion data based on multidimensional gas chromatograph information and sample detection range; performing a search based on the potential fusion data to obtain a plurality of fusion algorithms to be selected;

[0012] The candidate fusion algorithms are screened according to the sample detection range to determine the data fusion algorithm to be applied.

[0013] Furthermore, the method for screening the fusion algorithms to be selected according to the sample detection range includes:

[0014] Identify the sample types corresponding to the sample detection range, and calculate the detection proportion of the sample types;

[0015] Performing a detection simulation on the sample type by the candidate fusion algorithm to obtain a fusion effect value of the data fusion performed by the candidate fusion algorithm on the sample type;

[0016] The sample type is marked as i, i = 1, 2, ..., n, and n is the number of sample types;

[0017] The obtained fusion effect value is marked as RH i ; Mark the detection proportion of the sample type as δ i ;

[0018] The screening value of the corresponding candidate fusion algorithm is calculated according to the screening formula. The screening formula is:

[0019] ;

[0020] Where: SW is the screening value;

[0021] The candidate fusion algorithm with the largest screening value is selected as the applied data fusion algorithm.

[0022] Furthermore, the method of simulating detection of sample types by using the selected fusion algorithm includes:

[0023] Obtaining fusion materials corresponding to the sample types, and fusing the fusion materials using the selected fusion algorithm to obtain fusion result data;

[0024] Preset fusion effect evaluation indicators and indicator intervals, and set corresponding reference standards and standard discount methods based on the fusion effect evaluation indicators;

[0025] Collecting the fusion result data according to the fusion effect evaluation index to obtain index analysis data corresponding to the fusion effect evaluation index; analyzing the index analysis data according to the reference standard and the standard discount method to obtain the index score corresponding to the fusion effect evaluation index;

[0026] The indicator scores corresponding to each fusion effect evaluation indicator are accumulated to obtain the fusion effect value.

[0027] Step 2: Identify the test sample, test the test sample through a multi-dimensional gas chromatograph, and obtain corresponding sample test data;

[0028] Step 3: determining the debugging parameters of the data fusion algorithm according to the sample detection data, and debugging the data fusion algorithm according to the debugging parameters;

[0029] Furthermore, the method for determining the debugging parameters of the data fusion algorithm includes:

[0030] The platform sets a parameter representative curve, wherein the horizontal axis of the parameter representative curve is the serial number value and the vertical axis is the parameter representative value; each serial number value corresponds to a sample feature set data; each parameter representative value corresponds to a debugging reference parameter;

[0031] Identify the sample detection data of the detection sample, identify the sample feature set data corresponding to the sample detection data, match the corresponding debugging reference parameters for the sample feature set data according to the parameter representative curve, and mark them as debugging parameters.

[0032] Furthermore, the parameter representative curve setting method includes:

[0033] Acquire historical test data of a sample type, determine historical sample test data of the sample type based on the historical test data, and mark the data as sample material data; and determine a historical debugging parameter set corresponding to the sample material data based on the historical test data;

[0034] The platform presets each sample feature, performs feature processing on the sample material data according to the sample features, and obtains sample feature set data corresponding to the sample material data;

[0035] Identifying a historical debugging parameter set corresponding to the sample feature set data, evaluating a fusion effect value corresponding to each historical debugging parameter in the historical debugging parameter set; generating a fusion effect curve based on the fusion effect value and the corresponding historical debugging parameter; and determining a debugging reference parameter for the sample feature set data based on the fusion effect curve;

[0036] Sorting the sample feature set data corresponding to the sample types to obtain a first sequence; setting corresponding debugging reference parameters for the corresponding sample feature set data in the first sequence;

[0037] Setting a corresponding sequence number value for each sample feature set data according to the first sequence;

[0038] Mark the debugging reference parameter with a serial number value of 1 as a benchmark parameter, and set corresponding parameter representative values ​​for each debugging reference parameter according to the benchmark parameter;

[0039] A parameter representative curve is generated according to the serial number value and the parameter representative value, with the horizontal axis being the serial number value and the vertical axis being the parameter representative value.

[0040] Furthermore, the setting method of the first sequence includes:

[0041] Step SA1: Determine the first-ranked sample feature set data according to a preset method to form an initial sequence, and determine the benchmark data based on the initial sequence. The benchmark data is the last-ranked sample feature set data in the initial sequence;

[0042] Step SA2: Calculate the similarity between the benchmark data and the remaining sample feature set data, add the sample feature set data with the highest similarity to the initial sequence to obtain a new initial sequence, and determine new benchmark data based on the new initial sequence;

[0043] Step SA3: Loop step SA2 until there is no remaining sample feature set data, and mark the initial sequence as the first sequence.

[0044] Furthermore, when the debugging result of debugging the data fusion algorithm according to the debugging parameters does not meet the user's requirements, the platform performs remote control debugging.

[0045] Furthermore, the method for remote control debugging by the platform includes:

[0046] When a user has a need for remote control, the user activates the data management authority and sends remote control information to the platform. When the platform receives the remote control information, it collects debugging and analysis data in real time after the data management authority is activated.

[0047] The debugging analysis data is analyzed to determine debugging parameters, and the data fusion algorithm is debugged according to the debugging parameters.

[0048] Step 4: Use the debugged data fusion algorithm to fuse the sample detection data to obtain fused detection data;

[0049] Step 5: Obtain the user's sample analysis target, process the fusion detection data according to the sample analysis target, and obtain the sample analysis results.

[0050] A sample analysis system based on a multi-dimensional gas chromatograph data fusion algorithm includes an equipment analysis module, a detection module, an algorithm debugging module, and a data analysis module;

[0051] The equipment analysis module is used to analyze the multidimensional gas chromatograph and set the applied data fusion algorithm;

[0052] The detection module is used for detecting a detection sample to obtain corresponding sample detection data.

[0053] The algorithm debugging module is used for debugging the data fusion algorithm.

[0054] The data analysis module is used for detection analysis, fusing the sample detection data through the debugged data fusion algorithm to obtain fused detection data, obtaining a sample analysis target of a user, processing the fused detection data according to the sample analysis target to obtain a sample analysis result.

[0055] Compared with the prior art, the beneficial effects of the present application are:

[0056] The data fusion algorithm based on the multi-dimensional gas chromatograph provided by the present application can effectively solve the inconsistency of data formats and standards encountered in the data processing of the traditional multi-dimensional gas chromatography. The algorithm can automatically identify and match the data generated by different dimensions and different detectors, ensuring the consistency and accuracy of the data in the fusion process, thereby greatly improving the reliability and accuracy of the analysis results. Since the bottleneck problem in data processing is solved, the present application greatly promotes the wide application of multi-dimensional gas chromatography-mass spectrometry (GC-MS) and other multi-dimensional analysis technologies in the fields of chemistry, environmental science, food safety, drug research and development, etc. BRIEF DESCRIPTION OF DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of these drawings.

[0058] Figure 1 The method flowchart of the present application. DETAILED DESCRIPTION

[0059] The technical solutions of the present application will be described in detail below in conjunction with the embodiments. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0060] As shown in Figure 1 A sample analysis method based on a multi-dimensional gas chromatograph data fusion algorithm, the method comprising:

[0061] Step 1: Collect multi-dimensional gas chromatograph information and set up corresponding data fusion algorithm according to the multi-dimensional gas chromatograph information.

[0062] The multi-dimensional gas chromatograph information is collected based on an actual multi-dimensional gas chromatograph.

[0063] The data fusion algorithm is used to fuse the sample detection data from different dimensions and different detectors.

[0064] In one embodiment, the data fusion algorithm may be set based on existing methods, such as based on the working experience and professional capabilities of professionals.

[0065] In one embodiment, a method for setting a data fusion algorithm based on multi-dimensional gas chromatograph information includes:

[0066] Determine the sample types that the multidimensional gas chromatograph can detect and analyze based on the multidimensional gas chromatograph information and integrate them into the sample detection range;

[0067] According to the multi-dimensional gas chromatograph information and the sample detection range, various subsequent sample detection data are determined, and the potential fusion data is set, that is, the sample detection data that may need to be fused subsequently; according to the potential fusion data, a search is performed to determine the various existing data fusion algorithms that can meet the potential fusion data fusion processing and can be applied, and mark them as candidate fusion algorithms, that is, the candidate fusion algorithm is an existing data fusion algorithm, or a data fusion algorithm set by the platform, etc., which can be used in conjunction with the multi-dimensional gas chromatograph.

[0068] The fusion algorithms to be selected are screened according to the sample detection range, and the data fusion algorithm to be applied is determined.

[0069] In one embodiment, the fusion algorithms to be selected are screened according to the sample detection range, and the screening can be performed based on existing methods.

[0070] In one embodiment, the method for screening the fusion algorithms to be selected based on the sample detection range includes:

[0071] Identify the sample types corresponding to the sample detection range, and determine the detection proportion of different sample types through historical detection data or prediction, that is, the proportion of the sample type in the subsequent detection of all sample types, and compare them according to the corresponding detection quantity;

[0072] By performing detection simulation on each sample type through the candidate fusion algorithm, the fusion effect value of the data fusion of the candidate fusion algorithm on the sample type is obtained. The historical sample detection data corresponding to the sample type can be directly applied as the fusion material, and the candidate fusion algorithm is applied to perform fusion processing to determine the corresponding fusion result data, and the data with the best effect is selected, that is, the data with the best effect is selected from the data obtained by multiple simulations as the fusion result data; and then the corresponding fusion effect value is evaluated according to the fusion result data;

[0073] The sample type is marked as i, i = 1, 2, ..., n, and n is the number of sample types;

[0074] The obtained fusion effect value is marked as RH i ; Mark the detection proportion of the sample type as δ i ;

[0075] The screening value of the corresponding candidate fusion algorithm is calculated according to the screening formula. The screening formula is:

[0076] ;

[0077] Where: SW is the screening value;

[0078] The candidate fusion algorithm with the largest screening value is selected as the applied data fusion algorithm.

[0079] In one embodiment, the fusion effect value can be based on a percentage system, and various fusion effect evaluation indicators and indicator partition intervals are preset. The indicator partition interval is the fusion effect value range corresponding to the fusion effect evaluation indicator, such as [0, 20]; and the reference standard and standard discount method corresponding to the fusion effect evaluation indicator are set; the reference standard is the requirement that needs to be achieved for the full score of the fusion effect evaluation indicator, that is, to achieve the reference standard, and its indicator is divided into the maximum value between the indicator partitions; the standard discount method is the indicator score that needs to be discounted to different degrees below the reference standard; the fusion effect evaluation indicator, reference standard, standard discount method and indicator partition interval are all preset according to actual evaluation needs;

[0080] Obtain fusion result data obtained by performing detection simulation on corresponding sample types by the selected fusion algorithm;

[0081] Collect fusion result data according to each fusion effect evaluation indicator to obtain the indicator analysis data corresponding to the fusion effect evaluation indicator; analyze the indicator analysis data according to the corresponding reference standard and standard discount method to obtain the indicator score corresponding to the corresponding fusion effect evaluation indicator;

[0082] The index scores corresponding to each fusion effect evaluation index are accumulated to obtain the fusion effect value.

[0083] Exemplary:

[0084] Pre-set fusion effect evaluation index:

[0085] Data integrity: assess whether the fused data contains all the information of the original data without omission.

[0086] Reference standard: data integrity reaches more than 95%.

[0087] Data consistency: assess whether the fused data is consistent between different dimensions and different detectors without contradiction.

[0088] Reference standard: data consistency error is within 5%.

[0089] Data accuracy: assess whether the fused data accurately reflects the true situation of the original data without deviation.

[0090] Reference standard: data accuracy reaches more than 98%.

[0091] Data correlation: assess whether there is reasonable correlation between different dimensions in the fused data, which can reflect the sample characteristics.

[0092] Reference standard: data correlation score is more than 80 points (based on a certain correlation evaluation model).

[0093] Data redundancy: assess whether there is redundant information in the fused data, i.e. repeated or invalid data.

[0094] Reference standard: data redundancy is controlled within 5%.

[0095] Set the index partition and standard discount method corresponding to the evaluation index:

[0096] According to the pre-set fusion effect evaluation index and its reference standard, set the corresponding score range for each index. For example:

[0097] Data integrity: 95%-100% full score (such as 20 points), each decrease by 1% deduction 2 points.

[0098] Data consistency: error within 5% full score (such as 20 points), each increase by 1% error deduction 4 points.

[0099] Data accuracy: 98%-100% full score (such as 20 points), each decrease by 1% deduction 2 points.

[0100] Data correlation: more than 80 points full score (such as 20 points), each decrease by 5 points deduction 4 points.

[0101] Data redundancy: within 5% full score (such as 20 points), each increase by 1% redundancy deduction 4 points.

[0102] Perform matching analysis and accumulate scores:

[0103] Data matching analysis: Compare and analyze the fused data with the preset reference standards to determine the actual score of each evaluation indicator.

[0104] Accumulated score: The actual score of each evaluation indicator is accumulated to obtain the total score of the fusion effect.

[0105] Example of fusion effect value evaluation:

[0106] Assume that after data fusion processing, the following evaluation results are obtained:

[0107] Data integrity: 98% (scored 18 points);

[0108] Data consistency: 3% error (20 points);

[0109] Data accuracy: 99% (18 points);

[0110] Data relevance: 85 points (out of 20 points);

[0111] Data redundancy: 4% (20 points);

[0112] The total score of the fusion effect is: 18+20+18+20+20=96 points.

[0113] In one embodiment, the fusion effect value can also be evaluated based on existing methods, such as establishing an intelligent evaluation model based on intelligent algorithms such as neural networks, and performing intelligent evaluation through the successfully trained intelligent evaluation model to obtain the fusion effect value.

[0114] Step 2: Identify the samples that need to be tested and analyzed, mark them as test samples, test the test samples through a multi-dimensional gas chromatograph, and obtain corresponding sample test data.

[0115] Step 3: Determine the debugging parameters of the data fusion algorithm based on the sample detection data, and debug the data fusion algorithm according to the debugging parameters.

[0116] In one embodiment, a method for determining debugging parameters of a data fusion algorithm includes:

[0117] Obtain the sample detection range of the multidimensional gas chromatograph and identify each sample type based on the sample detection range;

[0118] Obtain historical detection data of each sample category, determine various different historical sample detection data of the sample category according to the historical detection data, and mark as sample material data; and determine the sample material data corresponding to the debugging parameters that have been obtained according to the historical detection data, and integrate into a historical debugging parameter set;

[0119] The platform party presets the sample detection data of the parameter debugging of the data fusion algorithm, such as signal strength, noise level, timestamp, spatial position, data quality index, data distribution characteristics, correlation analysis, etc.; according to the sample characteristics, the sample material data is processed, and the sample feature set data corresponding to the sample material data is obtained, that is, composed of sample feature data corresponding to each sample feature, and the corresponding sample feature data is analyzed by using the existing method;

[0120] The historical debugging parameter set corresponding to the sample feature set data is identified, and the fusion effect value corresponding to each historical debugging parameter in the historical debugging parameter set is evaluated; a fusion effect curve is generated according to the fusion effect value and the corresponding historical debugging parameter, the debugging parameter is used as the horizontal axis, and the fusion effect value is used as the vertical axis; according to the fusion effect curve, the historical debugging parameter corresponding to the best fusion effect of the sample feature set data is determined, which is marked as a debugging reference parameter, which can directly identify the debugging reference parameter according to the fusion effect value, or can determine the highest possible fusion effect value in combination with the prediction data, and then determine the debugging reference parameter;

[0121] The sample feature set data corresponding to the sample category is sorted to obtain a first sequence; the corresponding debugging reference parameter is set for the corresponding sample feature set data in the first sequence;

[0122] According to the first sequence, the corresponding serial number value, that is, the corresponding serial number, is set for each sample feature set data;

[0123] The debugging reference parameter with the serial number value of 1 is marked as a reference parameter, and the corresponding parameter representative value is set for each debugging reference parameter according to the reference parameter;

[0124] A parameter representative curve is generated according to the serial number value and the parameter representative value, the horizontal axis is the serial number value, and the vertical axis is the parameter representative value;

[0125] The sample detection data is identified, the sample feature set data corresponding to the sample detection data is identified, the corresponding debugging reference parameter is matched for the sample feature set data according to the parameter representative curve, and is marked as a debugging parameter; it can be directly matched, or it can be matched after the corresponding position is matched, and the best debugging parameter is determined, according to the curve change, the optional range of the best debugging parameter of the sample feature set data can be greatly reduced, and the debugging times are reduced.

[0126] In one embodiment, the sample feature set data corresponding to the sample category is sorted, which can be based on the existing sorting method, such as evaluation priority, sorting in order from good to bad, sorting in order of difficulty, manual sorting, and other sorting methods.

[0127] In one embodiment, the sample feature set data corresponding to the sample category is sorted, which can be sorted according to the difference of the sample feature set data. The smaller the difference, the closer the sorting. The sample feature set data with the highest weight coefficient can be determined by using optional, difficulty, and probability, and the similarity between the remaining sample feature set and the sample feature set data with the highest weight coefficient is calculated. The sample feature set data with the highest similarity is considered to be the second sorted, and the similarity between the remaining sample feature set and the sample feature set data with the highest weight coefficient is calculated. The specific method is:

[0128] Step SA1: determining the sample feature set data with the highest weight coefficient according to the preset method, forming an initial sequence, and determining the reference data according to the initial sequence. The reference data is the sample feature set data sorted last in the initial sequence.

[0129] Step SA2: calculating the similarity between the reference data and the remaining sample feature set data, supplementing the sample feature set data with the highest similarity to the initial sequence to obtain a new initial sequence, and determining a new reference data according to the new initial sequence.

[0130] Step SA3: repeating step SA2 until there is no remaining sample feature set data, and marking the initial sequence as the first sequence.

[0131] In one embodiment, the reference parameter is used to set the corresponding parameter representative value for each debugging reference parameter. When the parameter representative value of the reference parameter is 1, the above method is used for step-by-step sorting, i.e. calculating the similarity between the remaining debugging reference parameter and the reference parameter, marking the parameter representative value of the debugging reference parameter with the highest similarity as 2, and so on, marking as 3, 4, 5, etc.

[0132] In one embodiment, the reference parameter is used to set the corresponding parameter representative value for each debugging reference parameter, which can be determined based on the existing method.

[0133] In one embodiment, signal strength: In many sensor detection systems, signal strength is a key feature. It reflects the size of the signal energy received by the sensor, and for data fusion algorithms, signal strength can help determine the reliability and weight of the data. In parameter debugging, signal strength can be used to adjust the weight distribution between different data sources in the fusion algorithm.

[0134] Noise level: Noise is the unwanted, random variation in data. It can originate from the sensor itself, environmental interference, or errors in data transmission. The level of noise directly impacts the performance of the data fusion algorithm. Understanding and accounting for noise levels during parameter tuning helps optimize the robustness and accuracy of the algorithm.

[0135] Timestamps: Timestamps are a crucial feature of time series data. They record the time at which the data was collected and help align time series from different data sources during data fusion. The accuracy and consistency of timestamps are crucial to ensuring the correctness and reliability of fusion results. During parameter tuning, you may need to adjust parameters such as the time alignment algorithm and the fusion window size.

[0136] Spatial location: Spatial location is a core feature in spatially distributed data fusion. It describes the geographic location or spatial distribution of data points and is crucial for understanding the spatial relationships between data. During parameter tuning, spatial location information can be used to optimize spatial weight allocation and spatial interpolation methods in the fusion algorithm.

[0137] Data quality metrics: Data quality metrics such as completeness, accuracy, consistency, and timeliness are also important factors influencing the parameter tuning of data fusion algorithms. These metrics reflect the overall quality and availability of the data and are crucial for ensuring the accuracy and reliability of fusion results. During parameter tuning, it may be necessary to adjust fusion algorithm parameters based on data quality metrics to optimize fusion performance.

[0138] Correlation analysis: Correlation analysis is an important method for assessing the correlation between different data sources. By calculating metrics such as the correlation coefficient or mutual information between different features, we can understand the degree of association between them. During the parameter tuning phase, correlation analysis helps identify key features and optimize the weight allocation and fusion strategy in the fusion algorithm.

[0139] Data distribution characteristics: Data distribution characteristics such as mean, variance, skewness, and kurtosis are also important factors influencing the parameter tuning of data fusion algorithms. These characteristics reflect the statistical laws and distribution patterns of the data and are crucial for optimizing parameter settings and fusion strategies in fusion algorithms.

[0140] In one embodiment, the debugging parameters of the data fusion algorithm are determined according to the detected samples, and can be determined based on existing methods, such as establishing a debugging analysis model and performing intelligent analysis through the debugging analysis model.

[0141] Exemplarily, sample detection data from different dimensions and different detectors are collected; these data are preprocessed, including data cleaning, data format unification, data normalization, etc., to ensure data quality and consistency.

[0142] Extract useful features from the preprocessed data. These features should reflect the core information of the sample test data. Use feature selection methods, such as statistical methods and machine learning methods, to screen out features that have a significant impact on model building.

[0143] Based on the complexity of the problem and the characteristics of the data, select an appropriate intelligent algorithm as the basis for debugging the analysis model, such as neural networks, support vector machines, decision trees, etc. Use the training data set to train the model, adjust the model parameters and structure, and enable it to accurately classify, predict, or cluster the sample test data.

[0144] Use the validation dataset to validate the trained model and evaluate its performance, such as accuracy, recall, and F1 score. Optimize the model based on the validation results, including adjusting model parameters, adding features, and improving the algorithm.

[0145] The optimized model is deployed in actual application scenarios to process and analyze sample test data.

[0146] In one embodiment, when the debugging result does not meet the user's requirements, the user can make manual adjustments. However, in some cases, the user lacks the corresponding debugging experience and professional knowledge, resulting in the debugging result not meeting the requirements. Therefore, in this embodiment, remote debugging can also be performed. That is, when the debugging result of debugging the data fusion algorithm according to the debugging parameters does not meet the user's requirements, remote control debugging is performed; the remote control debugging method includes:

[0147] When the user has remote control needs, the data management permission is activated, which allows the platform to directly access and debug the data fusion algorithm; remote control information is sent to the platform, and the platform collects debugging and analysis data in real time after the data management permission is activated, that is, sample type, data fusion algorithm parameters, presentation effect and other related data; the debugging and analysis data is analyzed, the debugging parameters are determined, and the data fusion algorithm is debugged according to the debugging parameters.

[0148] In one embodiment, the debugging analysis data may be analyzed by platform staff to determine debugging parameters.

[0149] In one embodiment, the debugging analysis data is analyzed, and a debugging analysis model can be established based on a neural network, etc., and intelligent analysis can be performed through the debugging analysis model; that is, the debugging analysis model in the above embodiment.

[0150] For example, the platform establishes a remote connection with the user, using remote access tools such as VPN, remote desktop, and SSH to ensure the security and stability of the connection for subsequent debugging.

[0151] After data management permissions are activated, the user's sample test data is transmitted to the platform's server in real time, or the platform's debugging tools and data sets are transmitted to the user's device. This ensures data integrity and accuracy, and avoids data loss or corruption during transmission.

[0152] The platform uses debugging tools to remotely debug the user's equipment, including modifying code, adjusting parameters, viewing logs, etc. It monitors data changes and system status during the debugging process in real time to identify and resolve problems in a timely manner.

[0153] Based on debugging results and user feedback, the model is iteratively optimized to improve its performance and accuracy. The optimized model is then redeployed to the user's device for a new round of testing and verification. The platform can provide users with training and guidance to help them better understand and use data fusion algorithms and debug analysis models. Technical support and consulting services are also provided to resolve any issues or difficulties encountered during use.

[0154] Step 4: Use the debugged data fusion algorithm to fuse the sample detection data to obtain fused detection data;

[0155] Step 5: Obtain the user's sample analysis target, process the fusion detection data according to the sample analysis target, and obtain the sample analysis results.

[0156] For example, an analysis report can be generated based on the user's sample analysis objectives, including information such as sample composition and concentration. The analysis can be performed using existing data summarization and test report generation technologies.

[0157] A sample analysis system based on a multi-dimensional gas chromatograph data fusion algorithm includes an equipment analysis module, a detection module, an algorithm debugging module, and a data analysis module;

[0158] The equipment analysis module is used to analyze the multidimensional gas chromatograph and set the applied data fusion algorithm;

[0159] The detection module is used to detect the test sample and obtain corresponding sample detection data;

[0160] The algorithm debugging module is used to debug the data fusion algorithm;

[0161] The data analysis module is used for detection analysis, fusion processing of sample detection data is performed through the debugged data fusion algorithm, and fusion detection data is obtained; a sample analysis target of a user is acquired, the fusion detection data is processed according to the sample analysis target, and a sample analysis result is obtained.

[0162] The above formulas are calculated by removing the dimension and taking the numerical value, the formula is obtained by collecting a large amount of data to simulate the closest real situation, and the preset parameters and the preset threshold in the formula are set by a person skilled in the art according to the actual situation or obtained by a large amount of data simulation.

[0163] It is apparent for a person skilled in the art that the present application is not limited to the details of the above exemplary embodiments, but can be implemented in other concrete forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be considered as exemplary and non-limiting, the scope of the present application is defined by the appended claims rather than the above description, and all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be considered as limiting the claims involved.

Claims

1. A sample analysis method based on a multidimensional gas chromatograph data fusion algorithm, characterized in that: Methods include: Step 1: Collect multi-dimensional gas chromatograph information, and set a data fusion algorithm according to the multi-dimensional gas chromatograph information; Step 2: Identify the test sample, test the test sample through a multi-dimensional gas chromatograph, and obtain corresponding sample test data; Step 3: determining the debugging parameters of the data fusion algorithm according to the sample detection data, and debugging the data fusion algorithm according to the debugging parameters; Step 4: Use the debugged data fusion algorithm to fuse the sample detection data to obtain fused detection data; Step 5: Obtain the user's sample analysis target, process the fusion detection data according to the sample analysis target, and obtain the sample analysis results; Methods for determining the debugging parameters of the data fusion algorithm include: The platform sets a parameter representative curve, wherein the horizontal axis of the parameter representative curve is the serial number value and the vertical axis is the parameter representative value; each serial number value corresponds to a sample feature set data; each parameter representative value corresponds to a debugging reference parameter; Identifying sample detection data of the detection sample, identifying sample feature set data corresponding to the sample detection data, matching corresponding debugging reference parameters for the sample feature set data according to the parameter representative curve, and marking the parameters as debugging parameters; The parameter representative curve setting methods include: Determine the sample types that the multidimensional gas chromatograph can detect, obtain historical detection data of the sample types, determine historical sample detection data of the sample types based on the historical detection data, and mark them as sample material data; and determine a historical debugging parameter set corresponding to the sample material data based on the historical detection data; The platform presets each sample feature, performs feature processing on the sample material data according to the sample features, and obtains sample feature set data corresponding to the sample material data; Identifying a historical debugging parameter set corresponding to the sample feature set data, evaluating a fusion effect value corresponding to each historical debugging parameter in the historical debugging parameter set; generating a fusion effect curve based on the fusion effect value and the corresponding historical debugging parameter; and determining a debugging reference parameter for the sample feature set data based on the fusion effect curve; Sorting the sample feature set data corresponding to the sample types to obtain a first sequence; setting corresponding debugging reference parameters for the corresponding sample feature set data in the first sequence; Setting a corresponding sequence number value for each sample feature set data according to the first sequence; Mark the debugging reference parameter with a serial number value of 1 as a benchmark parameter, and set corresponding parameter representative values ​​for each debugging reference parameter according to the benchmark parameter; A parameter representative curve is generated according to the serial number value and the parameter representative value, with the horizontal axis being the serial number value and the vertical axis being the parameter representative value.

2. The sample analysis method based on multidimensional gas chromatograph data fusion algorithm according to claim 1, characterized in that: The setting methods of the data fusion algorithm include: determining a sample detection range of the multidimensional gas chromatograph according to the multidimensional gas chromatograph information; Determining potential fusion data based on multidimensional gas chromatograph information and sample detection range; performing a search based on the potential fusion data to obtain a plurality of fusion algorithms to be selected; The candidate fusion algorithms are screened according to the sample detection range to determine the data fusion algorithm to be applied.

3. The sample analysis method based on multidimensional gas chromatograph data fusion algorithm according to claim 2, characterized in that: Methods for screening candidate fusion algorithms based on sample detection range include: Identify the sample types corresponding to the sample detection range, and calculate the detection proportion of the sample types; Performing a detection simulation on the sample type by the candidate fusion algorithm to obtain a fusion effect value of the data fusion performed by the candidate fusion algorithm on the sample type; The sample type is marked as i, i = 1, 2, ..., n, and n is the number of sample types; The obtained fusion effect value is marked as RH i ; Mark the detection proportion of the sample type as δ i ; The screening value of the corresponding candidate fusion algorithm is calculated according to the screening formula. The screening formula is: ; Where: SW is the screening value; The candidate fusion algorithm with the largest screening value is selected as the applied data fusion algorithm.

4. The sample analysis method based on multidimensional gas chromatograph data fusion algorithm according to claim 3, characterized in that: Methods for simulating detection of sample types using candidate fusion algorithms include: Obtaining fusion materials corresponding to the sample types, and fusing the fusion materials using the selected fusion algorithm to obtain fusion result data; Preset fusion effect evaluation indicators and indicator intervals, and set corresponding reference standards and standard discount methods based on the fusion effect evaluation indicators; Collecting the fusion result data according to the fusion effect evaluation index to obtain index analysis data corresponding to the fusion effect evaluation index; analyzing the index analysis data according to the reference standard and the standard discount method to obtain the index score corresponding to the fusion effect evaluation index; The indicator scores corresponding to each fusion effect evaluation indicator are accumulated to obtain the fusion effect value.

5. The sample analysis method based on multidimensional gas chromatograph data fusion algorithm according to claim 1, characterized in that: The first sequence of settings includes: Step SA1: Determine the first-ranked sample feature set data according to a preset method to form an initial sequence, and determine the benchmark data based on the initial sequence. The benchmark data is the last-ranked sample feature set data in the initial sequence; Step SA2: Calculate the similarity between the benchmark data and the remaining sample feature set data, add the sample feature set data with the highest similarity to the initial sequence to obtain a new initial sequence, and determine new benchmark data based on the new initial sequence; Step SA3: Loop step SA2 until there is no remaining sample feature set data, and mark the initial sequence as the first sequence.

6. The sample analysis method based on multidimensional gas chromatograph data fusion algorithm according to claim 1, characterized in that: When the debugging results of the data fusion algorithm according to the debugging parameters do not meet the user's requirements, the platform will perform remote control debugging.

7. The sample analysis method based on multidimensional gas chromatograph data fusion algorithm according to claim 6, characterized in that: Methods for remote control debugging by the platform include: When a user has a need for remote control, the user activates the data management authority and sends remote control information to the platform. When the platform receives the remote control information, it collects debugging and analysis data in real time after the data management authority is activated. The debugging analysis data is analyzed to determine debugging parameters, and the data fusion algorithm is debugged according to the debugging parameters.

8. A sample analysis system based on a multidimensional gas chromatograph data fusion algorithm, characterized in that: Execute a sample analysis method based on a multidimensional gas chromatograph data fusion algorithm as described in any one of claims 1 to 7, comprising an equipment analysis module, a detection module, an algorithm debugging module and a data analysis module; The equipment analysis module is used to analyze the multidimensional gas chromatograph and set the applied data fusion algorithm; The detection module is used to detect the test sample and obtain corresponding sample detection data; The algorithm debugging module is used to debug the data fusion algorithm; The data analysis module is used to perform detection analysis and fuse the sample detection data through the debugged data fusion algorithm to obtain fused detection data; Obtain the user's sample analysis target, process the fusion detection data according to the sample analysis target, and obtain the sample analysis results.

Citation Information

Patent Citations

  • Traditional Chinese medicine preparation quality detection method based on multi-dimensional data analysis

    CN118782269A