A cloud computing analysis system and method based on a high-performance liquid chromatograph

By using a cloud computing analysis system based on high-performance liquid chromatography (HPLC), combined with intelligent clustering analysis and multimodal data fusion technology, the problems of manual dependence and insufficient sensitivity in existing gas detection have been solved, achieving efficient, automated and accurate gas analysis.

CN119438477BActive Publication Date: 2026-04-14JIANGSU NANTONG INTELLIGENT CLOUD COMPUTING EXPERIMENTAL EQUIP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-17
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing chemiluminescence detection and spectroscopic analysis methods rely on manual operation in gas detection, resulting in low detection efficiency and insufficient sensitivity. They are also difficult to automate and support with big data, which affects the accuracy of the detection results.

Method used

Design a cloud computing analysis system based on high performance liquid chromatography, including data acquisition, transmission, storage, analysis and user interface modules. Employ intelligent clustering analysis model and multimodal data fusion technology to achieve automated data analysis and real-time feedback.

Benefits of technology

It improves the speed and accuracy of data analysis, enhances the reliability of analysis results through multimodal data fusion, and enables automated and real-time gas detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119438477B_ABST
    Figure CN119438477B_ABST
Patent Text Reader

Abstract

The application discloses a cloud computing analysis system and method based on a high-performance liquid chromatograph, and belongs to the fields of analytical chemistry and information technology.The system comprises a data acquisition module, a data transmission module, a cloud storage module, a data analysis module, a user interface module and a real-time analysis and feedback module; the data acquisition module is used for acquiring data detected by the high-performance liquid chromatograph; the data transmission module is used for determining a protocol adopted by data transmission; the cloud storage module is used for storing data transmitted to the cloud; the data analysis module is used for training an intelligent clustering analysis model and inputting data for automatic data analysis; the user interface module is used for providing an interface for users to view analysis results; and the real-time analysis and feedback module is used for performing real-time analysis and judgment on data.The application uses a machine learning training model to realize automatic data analysis, and reduces manual intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of analytical chemistry and information technology, specifically to a cloud computing analysis system and method based on high performance liquid chromatography. Background Technology

[0002] High Performance Liquid Chromatography (HPLC) is an advanced analytical instrument used for the separation, identification, and quantitative analysis of components in mixtures.

[0003] Current mainstream methods for chemical gas analysis include chemiluminescence detection and spectroscopic analysis. However, chemiluminescence detection relies on specific reagents and optical equipment, requiring real-time monitoring by personnel and involving significant manual intervention, resulting in low detection efficiency. Spectroscopic analysis, on the other hand, is often affected by gas concentration and background noise during detection, leading to insufficient sensitivity and impacting the accuracy of the results. Both methods are highly dependent on manual operation in gas classification, making automation difficult, and lacking support from large datasets, further limiting detection efficiency and accuracy. Summary of the Invention

[0004] The purpose of this invention is to provide a cloud computing analysis system and method based on high performance liquid chromatography to solve the problems mentioned in the background art.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a cloud computing analysis system based on high performance liquid chromatography, characterized in that: the system includes a data acquisition module, a data transmission module, a cloud storage module, a data analysis module, a user interface module, and a real-time analysis and feedback module;

[0006] The data acquisition module is used to acquire data detected by high-performance liquid chromatography (HPLC) and perform baseline adjustment on the data. The data transmission module is used to determine the protocol used for data transmission and to compress and encrypt the baseline-adjusted data. The cloud storage module is used to store the data transmitted to the cloud and provides data backup and recovery options. The data analysis module is used to decrypt and decompress the data, input it into a trained intelligent clustering analysis model for data analysis, and use artificial intelligence algorithms to optimize the model to improve analysis efficiency and accuracy. The user interface module is used to provide users with an interface to view data analysis results and record logs. The real-time analysis and feedback module is used to perform real-time data analysis, and when the analysis results are returned, it makes a judgment. If the error between the analysis result and the prediction result exceeds the sum of squared errors, the analysis scheme and model parameters are immediately adjusted for a second analysis.

[0007] The baseline adjustment is used to remove noise and make the signal display peak values;

[0008] The output of the data acquisition module is connected to the input of the data transmission module; the output of the data transmission module is connected to the input of the cloud storage module; the output of the cloud storage module is connected to the input of the data analysis module; the output of the data analysis module is connected to the input of the user interface module; and the output of the user interface module is connected to the input of the real-time analysis and feedback module.

[0009] The data acquisition module includes an HPLC instrument interface unit and a data processing unit;

[0010] The HPLC instrument interface unit provides an interface between the data acquisition module and the HPLC instrument, transmitting the data detected by the HPLC instrument to the analysis system; the data processing unit is used for baseline adjustment of the data.

[0011] The output of the HPLC instrument interface unit is connected to the input of the data processing unit; the output of the data processing unit is connected to the input of the data transmission module.

[0012] The data transmission module includes a transmission protocol unit, a data compression unit, and a data encryption unit;

[0013] The transmission protocol unit uses the secure transmission protocol HTTPS to ensure the security and integrity of data upload; the data compression unit is used to compress the collected data; and the data encryption unit is used to encrypt the compressed data before transmission.

[0014] The output of the transmission protocol unit is connected to the input of the data compression unit; the output of the data compression unit is connected to the input of the data encryption unit; and the output of the data encryption unit is connected to the input of the cloud storage module.

[0015] The cloud storage module includes a data storage unit and a backup and recovery unit;

[0016] The data storage unit is used to store the received data in the cloud; the backup and recovery unit is used to add a data backup and recovery mechanism, so that the data can be viewed and downloaded again when the data is abnormal or lost.

[0017] The output of the data storage unit is connected to the input of the backup and recovery unit.

[0018] The data analysis module includes a data decryption unit, a data decompression unit, a data preprocessing unit, an intelligent clustering analysis model unit, a multimodal data fusion unit, a model optimization unit, and an analysis result visualization unit.

[0019] The data decryption unit is used to decrypt encrypted data; the data decompression unit is used to decompress the decrypted data to obtain the transmitted data; the data preprocessing unit is used to perform baseline correction, peak detection, and peak integration on the data; the intelligent clustering analysis model unit is used to transmit data to a trained model for automated analysis; the multimodal data fusion unit is used to fuse high-performance liquid chromatography data with mass spectrometry data, ultraviolet-visible spectroscopy data, and infrared spectroscopy data to improve the accuracy and reliability of the analysis results through multimodal data analysis; the model optimization unit is used to adjust the model parameters based on gas analysis data in the database; and the result visualization unit is used to display the data analysis results in a visual manner.

[0020] The baseline correction is to remove baseline drift from chromatographic data and to quantify and qualitatively identify peaks in the chromatogram.

[0021] The peak detection is the process of identifying peaks from the corrected signal;

[0022] The peak integral is the area of ​​each detected peak.

[0023] The output of the data decryption unit is connected to the input of the data decompression unit; the output of the data decompression unit is connected to the input of the data preprocessing unit; the output of the data preprocessing unit is connected to the input of the intelligent clustering analysis model unit; the output of the intelligent clustering analysis model unit is connected to the input of the multimodal data fusion unit; the output of the multimodal data fusion unit is connected to the input of the model optimization unit; the output of the model optimization unit is connected to the inputs of the analysis result visualization unit and the intelligent clustering analysis model unit, respectively; and the output of the analysis result visualization unit is connected to the input of the user interface module.

[0024] The user interface module includes an authorization management unit, a user interface unit, and a monitoring and logging unit;

[0025] The authorization management unit is used to authenticate users who view the analysis results; the user interface unit is the interface displayed to the user; the monitoring and logging unit is used to record the operations, users, and times of viewing the analysis results.

[0026] The output of the authorization management unit is connected to the input of the user interface unit; the output of the user interface unit is connected to the input of the monitoring and logging unit; and the output of the monitoring and logging unit is connected to the input of the real-time analysis and feedback module.

[0027] A cloud computing analysis method based on high performance liquid chromatography, the method comprising the following steps:

[0028] S1. Use an HPLC instrument to analyze the sample, obtain chromatographic data, perform noise reduction processing on the data, and the detector converts the noise-reduced chromatographic data into digital signals and records them in the local data system.

[0029] S2. After compressing and encrypting the locally stored chromatographic data, upload it to the cloud storage platform using the secure transmission protocol HTTPS to ensure the security and integrity of the data upload.

[0030] S3. After decrypting and decompressing the received data, use a cloud computing platform for automated data preprocessing and data standardization;

[0031] S4. Use big data analytics tools on a cloud computing platform for data modeling and analysis; apply the machine learning algorithm K-means clustering to predict and classify gases from preprocessed chromatographic data, and train an intelligent clustering analysis model.

[0032] S5. Combining multimodal data fusion technology, high-performance liquid chromatography data, mass spectrometry data, ultraviolet-visible spectroscopy data and infrared spectroscopy data are fused to improve the accuracy and reliability of analytical results through multimodal data analysis;

[0033] S6. Add a model optimization mechanism to the intelligent clustering analysis model, and combine the database to adjust the model parameters and optimize the model.

[0034] S7. Provide an interface for users to view analysis results, add an access control mechanism to verify the permissions of users viewing analysis results, and record logs.

[0035] S8. Real-time analysis and feedback: After obtaining access permissions, a judgment is made at the end of the analysis. If the error between the analysis result and the prediction exceeds the sum of squared errors, the scheme and model parameters are immediately adjusted for secondary data analysis.

[0036] In step S3,

[0037] The preprocessing includes baseline correction, peak detection, and peak integration; the data standardization is performed on each feature at the same scale using a standard normal distribution with a mean of 0 and a standard deviation of 1.

[0038] The baseline correction is to remove baseline drift from chromatographic data and to quantify and qualitatively identify peaks in the chromatogram.

[0039] The peak detection is the process of identifying peaks from the corrected signal;

[0040] The peak integral is the area of ​​each detected peak.

[0041] In step S4, the training of the intelligent clustering analysis model specifically includes the following steps:

[0042] S9-1, will Data points , , , Divided into Group , , , This makes the data points similar within the same group, but different between different groups;

[0043] Centroid calculation:

[0044]

[0045] in, For each group The center of mass, ; It is a group The number of data points in the middle; It is a group Data points in;

[0046] S9-2, Initialization: Select Initial centroid of the group , , , ;

[0047] S9-3, Group Assignment: For each data point

[0048] Calculate the distance from each centroid to the nearest centroid and assign it to the group to which the nearest centroid belongs. ;

[0049]

[0050] S9-4, Centroid Update: Calculate the new centroid for each group and update it:

[0051]

[0052] in, For each group The center of mass, ; It is a group The number of data points in the middle; It is a group Data points in;

[0053] S9-5, Iteration: Repeat steps S9-3 and S9-4 until the centroid no longer changes;

[0054] The silhouette coefficient method is used to determine the number of groups in K-means clustering; specifically:

[0055] When the number of data points At that time, the initial number of groups Take 2, 3, and 4 respectively, when the number of data points At that time, the initial number of groups Take 5, 6, and 7 respectively; for each The value is calculated by performing the K-means algorithm. The corresponding profile coefficient :

[0056]

[0057] in Data points The average distance to other data points in the same set. Data points The average distance to the nearest other data point;

[0058]

[0059] in, Data points The group you belong to; It is a group The number of data points in the middle; Data points and data points The Euclidean distance between them;

[0060]

[0061] in, Data points The group you belong to; It is a group The number of data points in the middle; Data points and data points The Euclidean distance between them;

[0062] calculate Corresponding contour coefficient The largest Corresponding This is the number of groups in the K-means clustering.

[0063] In steps S5-S8,

[0064] The multimodal data fusion technology involves fusing four types of data: high-performance liquid chromatography (HPLC), mass spectrometry (MS), ultraviolet-visible spectroscopy (UV-Vis), and infrared spectroscopy (IR). This multimodal data fusion improves the accuracy and reliability of the analytical results.

[0065] The real-time analysis and feedback process involves: obtaining access permissions and making a judgment after the analysis is completed. If the error between the analysis result and the prediction result exceeds the sum of squared errors, the scheme and model parameters are immediately adjusted for a second data analysis. If the error between the second analysis result and the prediction result is less than the sum of squared errors, the second analysis result is selected and entered into the database. If the error between the second analysis result and the prediction result is greater than or equal to the sum of squared errors, but the error between the second analysis result and the first analysis result is less than the sum of squared errors, the prediction result part of the intelligent clustering analysis model is adjusted and logged.

[0066] The squared error and :

[0067]

[0068] in, Represents the sum of squared errors; Data points its group's centroid The Euclidean distance between them ; It is a group Data points in; Data points The quantity.

[0069] Compared with the prior art, the beneficial effects achieved by the present invention are:

[0070] 1. This invention uses cloud computing to automate the analysis of massive and complex data, greatly improving the speed of data analysis;

[0071] 2. This invention adds a real-time data analysis and feedback mechanism, which can analyze the data in the database and then adjust the analysis parameters to obtain more accurate analysis results;

[0072] 3. This invention utilizes multimodal data fusion technology, which compares mass spectrometry data, ultraviolet-visible spectral data, and infrared spectral data, thereby further improving the performance of the analysis system and the reliability of the analysis results, and providing users with a more comprehensive analysis solution. Attached Figure Description

[0073] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0074] Figure 1 This is a schematic flowchart of a cloud computing analysis system and method based on high performance liquid chromatography according to the present invention.

[0075] Figure 2 This is a schematic diagram of the steps of a cloud computing analysis method based on high performance liquid chromatography according to the present invention. Detailed Implementation

[0076] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0077] Please see Figures 1-2 The present invention provides a technical solution: a cloud computing analysis system based on high performance liquid chromatography, characterized in that: the system includes a data acquisition module, a data transmission module, a cloud storage module, a data analysis module, a user interface module, and a real-time analysis and feedback module;

[0078] The data acquisition module is used to acquire data detected by high-performance liquid chromatography (HPLC) and perform baseline adjustment on the data. The data transmission module is used to determine the protocol used for data transmission and to compress and encrypt the baseline-adjusted data. The cloud storage module is used to store the data transmitted to the cloud and provides data backup and recovery options. The data analysis module is used to decrypt and decompress the data, input it into a trained intelligent clustering analysis model for data analysis, and use artificial intelligence algorithms to optimize the model to improve analysis efficiency and accuracy. The user interface module is used to provide users with an interface to view data analysis results and record logs. The real-time analysis and feedback module is used to perform real-time data analysis, and when the analysis results are returned, it makes a judgment. If the error between the analysis result and the prediction result exceeds the sum of squared errors, the analysis scheme and model parameters are immediately adjusted for a second analysis.

[0079] The baseline adjustment is used to remove noise and make the signal display peak values;

[0080] The output of the data acquisition module is connected to the input of the data transmission module; the output of the data transmission module is connected to the input of the cloud storage module; the output of the cloud storage module is connected to the input of the data analysis module; the output of the data analysis module is connected to the input of the user interface module; and the output of the user interface module is connected to the input of the real-time analysis and feedback module.

[0081] The data acquisition module includes an HPLC instrument interface unit and a data processing unit;

[0082] The HPLC instrument interface unit provides an interface between the data acquisition module and the HPLC instrument, transmitting the data detected by the HPLC instrument to the analysis system; the data processing unit is used for baseline adjustment of the data.

[0083] The output of the HPLC instrument interface unit is connected to the input of the data processing unit; the output of the data processing unit is connected to the input of the data transmission module.

[0084] The data transmission module includes a transmission protocol unit, a data compression unit, and a data encryption unit;

[0085] The transmission protocol unit uses the secure transmission protocol HTTPS to ensure the security and integrity of data upload; the data compression unit is used to compress the collected data; and the data encryption unit is used to encrypt the compressed data before transmission.

[0086] The output of the transmission protocol unit is connected to the input of the data compression unit; the output of the data compression unit is connected to the input of the data encryption unit; and the output of the data encryption unit is connected to the input of the cloud storage module.

[0087] The cloud storage module includes a data storage unit and a backup and recovery unit;

[0088] The data storage unit is used to store the received data in the cloud; the backup and recovery unit is used to add a data backup and recovery mechanism, so that the data can be viewed and downloaded again when the data is abnormal or lost.

[0089] The output of the data storage unit is connected to the input of the backup and recovery unit.

[0090] The data analysis module includes a data decryption unit, a data decompression unit, a data preprocessing unit, an intelligent clustering analysis model unit, a multimodal data fusion unit, a model optimization unit, and an analysis result visualization unit.

[0091] The data decryption unit is used to decrypt encrypted data; the data decompression unit is used to decompress the decrypted data to obtain the transmitted data; the data preprocessing unit is used to perform baseline correction, peak detection, and peak integration on the data; the intelligent clustering analysis model unit is used to transmit data to a trained model for automated analysis; the multimodal data fusion unit is used to fuse high-performance liquid chromatography data with mass spectrometry data, ultraviolet-visible spectroscopy data, and infrared spectroscopy data to improve the accuracy and reliability of the analysis results through multimodal data analysis; the model optimization unit is used to adjust the model parameters based on gas analysis data in the database; and the result visualization unit is used to display the data analysis results in a visual manner.

[0092] The baseline correction is to remove baseline drift from chromatographic data and to quantify and qualitatively identify peaks in the chromatogram.

[0093] The peak detection is the process of identifying peaks from the corrected signal;

[0094] The peak integral is the area of ​​each detected peak.

[0095] The output of the data decryption unit is connected to the input of the data decompression unit; the output of the data decompression unit is connected to the input of the data preprocessing unit; the output of the data preprocessing unit is connected to the input of the intelligent clustering analysis model unit; the output of the intelligent clustering analysis model unit is connected to the input of the multimodal data fusion unit; the output of the multimodal data fusion unit is connected to the input of the model optimization unit; the output of the model optimization unit is connected to the inputs of the analysis result visualization unit and the intelligent clustering analysis model unit, respectively; and the output of the analysis result visualization unit is connected to the input of the user interface module.

[0096] The user interface module includes an authorization management unit, a user interface unit, and a monitoring and logging unit;

[0097] The authorization management unit is used to authenticate users who view the analysis results; the user interface unit is the interface displayed to the user; the monitoring and logging unit is used to record the operations, users, and times of viewing the analysis results.

[0098] The output of the authorization management unit is connected to the input of the user interface unit; the output of the user interface unit is connected to the input of the monitoring and logging unit; and the output of the monitoring and logging unit is connected to the input of the real-time analysis and feedback module.

[0099] A cloud computing analysis method based on high performance liquid chromatography, the method comprising the following steps:

[0100] S1. Use an HPLC instrument to analyze the sample, obtain chromatographic data, perform noise reduction processing on the data, and the detector converts the noise-reduced chromatographic data into digital signals and records them in the local data system.

[0101] S2. After compressing and encrypting the locally stored chromatographic data, upload it to the cloud storage platform using the secure transmission protocol HTTPS to ensure the security and integrity of the data upload.

[0102] S3. After decrypting and decompressing the received data, use a cloud computing platform for automated data preprocessing and data standardization;

[0103] S4. Use big data analytics tools on a cloud computing platform for data modeling and analysis; apply the machine learning algorithm K-means clustering to predict and classify gases from preprocessed chromatographic data, and train an intelligent clustering analysis model.

[0104] S5. Combining multimodal data fusion technology, high-performance liquid chromatography data, mass spectrometry data, ultraviolet-visible spectroscopy data and infrared spectroscopy data are fused to improve the accuracy and reliability of analytical results through multimodal data analysis;

[0105] S6. Add a model optimization mechanism to the intelligent clustering analysis model, and combine the database to adjust the model parameters and optimize the model.

[0106] S7. Provide an interface for users to view analysis results, add an access control mechanism to verify the permissions of users viewing analysis results, and record logs.

[0107] S8. Real-time analysis and feedback: After obtaining access permissions, a judgment is made at the end of the analysis. If the error between the analysis result and the prediction exceeds the sum of squared errors, the scheme and model parameters are immediately adjusted for secondary data analysis.

[0108] In step S3,

[0109] The preprocessing includes baseline correction, peak detection, and peak integration; the data standardization is performed on each feature at the same scale using a standard normal distribution with a mean of 0 and a standard deviation of 1.

[0110] The baseline correction is to remove baseline drift from chromatographic data and to quantify and qualitatively identify peaks in the chromatogram.

[0111] The peak detection is the process of identifying peaks from the corrected signal;

[0112] The peak integral is the area of ​​each detected peak.

[0113] In step S4, the training of the intelligent clustering analysis model specifically includes the following steps:

[0114] S9-1, will Data points , , , Divided into Group , , , This makes the data points similar within the same group, but different between different groups;

[0115] Centroid calculation:

[0116]

[0117] in, For each group The center of mass, ; It is a group The number of data points in the middle; It is a group Data points in;

[0118] S9-2, Initialization: Select Initial centroid of the group , , , ;

[0119] S9-3, Group Assignment: For each data point Calculate the distance from each centroid to the nearest centroid and assign it to the group to which the nearest centroid belongs. ;

[0120]

[0121] S9-4, Centroid Update: Calculate the new centroid for each group and update it:

[0122]

[0123] in, For each group The center of mass, ; It is a group The number of data points in the middle; It is a group Data points in;

[0124] S9-5, Iteration: Repeat steps S9-3 and S9-4 until the centroid no longer changes;

[0125] The silhouette coefficient method is used to determine the number of groups in K-means clustering; specifically:

[0126] When the number of data points At that time, the initial number of groups Take 2, 3, and 4 respectively, when the number of data points At that time, the initial number of groups Take 5, 6, and 7 respectively; for each The value is calculated by performing the K-means algorithm. The corresponding profile coefficient :

[0127]

[0128] in Data points The average distance to other data points in the same set. Data points The average distance to the nearest other data point;

[0129]

[0130] in, Data points The group you belong to; It is a group The number of data points in the middle; Data points and data points The Euclidean distance between them;

[0131]

[0132] in, Data points The group you belong to; It is a group The number of data points in the middle; Data points and data points The Euclidean distance between them;

[0133] calculate Corresponding contour coefficient The largest Corresponding This is the number of groups in the K-means clustering.

[0134] In steps S5-S8,

[0135] The multimodal data fusion technology involves fusing four types of data: high-performance liquid chromatography (HPLC), mass spectrometry (MS), ultraviolet-visible spectroscopy (UV-Vis), and infrared spectroscopy (IR). This multimodal data fusion improves the accuracy and reliability of the analytical results.

[0136] The real-time analysis and feedback process involves: obtaining access permissions and making a judgment after the analysis is completed. If the error between the analysis result and the prediction result exceeds the sum of squared errors, the scheme and model parameters are immediately adjusted for a second data analysis. If the error between the second analysis result and the prediction result is less than the sum of squared errors, the second analysis result is selected and entered into the database. If the error between the second analysis result and the prediction result is greater than or equal to the sum of squared errors, but the error between the second analysis result and the first analysis result is less than the sum of squared errors, the prediction result part of the intelligent clustering analysis model is adjusted and logged.

[0137] The squared error and :

[0138]

[0139] in, Represents the sum of squared errors; Data points its group's centroid The Euclidean distance between them ; It is a group Data points in; Data points The quantity.

[0140] In the embodiments of the present invention, the hardware facilities used include: high performance liquid chromatograph (HPLC), computer and data acquisition system, network connection equipment (router, modem), cloud server and cloud storage platform, and data processing and analysis software;

[0141] Materials used: the sample to be analyzed, the mobile phase solvent (water, methanol), and the standard samples used for calibration;

[0142] Test results: Test result label 1, Test result label 2;

[0143] Sample analysis: The sample to be analyzed is injected into the HPLC instrument, and the components are separated by the chromatographic column;

[0144] Data recording: Chromatogram data, including retention time, peak area, and peak height parameters, are recorded through a data acquisition system; the data recording format is {sample number, retention time (min), peak area, peak height}.

[0145] {"1", "5.2", "15000", "300"}

[0146] {"2", "7.8", "12000", "250"}

[0147] {“3”, “5.3”, “15500”, “310”}

[0148] {"4", "9.1", "13000", "270"}

[0149] {"5", "5.1", "14900", "295"}

[0150] {"6", "7.7", "12300", "255"}

[0151] {"7", "9.0", "12800", "265"}

[0152] Based on the contour coefficient method , Take 2, 3 and 4;

[0153] Calculate the profile coefficient , and ;

[0154] , , ;

[0155] The number of K-means clusters is 2.

[0156] The two cluster centers are:

[0157] {" "5.2", "15000", "300"

[0158] {" "7.8", "12000", "250"

[0159]

[0160]

[0161] Sample "1" was assigned to Group;

[0162] The results were obtained by calculating in sequence:

[0163] Sample "1" to The distance is 0, to The distance is 3000.03, and it is assigned to Group;

[0164] Sample "2" to The distance is 3000.03, to The distance is 0, and it is assigned to Group;

[0165] Sample "3" to The distance is:

[0166]

[0167] arrive The distance is:

[0168]

[0169] Assigned to Group;

[0170] Sample "4" to The distance is:

[0171]

[0172] arrive The distance is:

[0173]

[0174] Assigned to Group;

[0175] Sample "5" to The distance is:

[0176]

[0177] arrive The distance is:

[0178]

[0179] Assigned to Group;

[0180] Sample "6" to The distance is:

[0181]

[0182] arrive The distance is:

[0183]

[0184] Assigned to Group;

[0185] Sample "7" to The distance is:

[0186]

[0187] arrive The distance is:

[0188]

[0189] Assigned to Group;

[0190] Mean values ​​of samples "1", "3", and "5":

[0191] Retention time (min):

[0192] Peak area:

[0193] Peak height:

[0194] Samples “2”, “4”, “6”, “7”:

[0195] Retention time (min):

[0196] Peak area:

[0197] Peak height:

[0198] The updated cluster centers are as follows:

[0199] {" "5.2", "15133.33", "301.67"

[0200] {" "8.4", "12525", "260"}

[0201] Recalculate the Euclidean distance and group the data:

[0202] Sample "1" to The distance is:

[0203]

[0204] arrive The distance is:

[0205]

[0206] Assigned to Group;

[0207] Sample "2" to The distance is:

[0208]

[0209] arrive The distance is:

[0210]

[0211] Assigned to Group;

[0212] Sample "3" to The distance is:

[0213]

[0214] arrive The distance is:

[0215]

[0216] Assigned to Group;

[0217] Sample "4" to The distance is:

[0218]

[0219] arrive The distance is:

[0220]

[0221] Assigned to Group;

[0222] Sample "5" to The distance is:

[0223]

[0224] arrive The distance is:

[0225]

[0226] Assigned to Group;

[0227] Sample "6" to The distance is:

[0228]

[0229] arrive The distance is:

[0230]

[0231] Assigned to Group;

[0232] Sample "7" to The distance is:

[0233]

[0234] arrive The distance is:

[0235]

[0236] Assigned to Group;

[0237] Update the calculation for each cluster center:

[0238] {" "5.2", "15133.33", "301.67"

[0239] {" "8.4", "12525", "260"}

[0240] No changes were observed, and the iteration ended. The results were obtained in the following format: data record format: {sample number, retention time (min), peak area, peak height, detection result label}.

[0241] {“1”,“5.2”,“15000”,“300”,“Label 1”}

[0242] {"2", "7.8", "12000", "250", "Label 2"}

[0243] {“3”, “5.3”, “15500”, “310”, “Label 1”}

[0244] {"4", "9.1", "13000", "270", "Label 2"}

[0245] {“5”, “5.1”, “14900”, “295”, “Label 1”}

[0246] {“6”, “7.7”, “12300”, “255”, “Label 2”}

[0247] {"7", "9.0", "12800", "265", "tag 2"}

[0248] Predicted results; data recording format is {sample number, retention time (min), peak area, peak height, detection result label};

[0249] {“1”, “5.3”, “15000”, “290”, “Label 1”}

[0250] {"2", "7.2", "13000", "250", "Label 2"}

[0251] {“3”, “5.3”, “15500”, “300”, “Label 1”}

[0252] {"4", "8.6", "13500", "270", "Label 2"}

[0253] {“5”,“5.6”,“14700”,“285”,“Label 1”}

[0254] {“6”, “7.1”, “13300”, “285”, “Label 2”}

[0255] {"7", "8.5", "12100", "225", "Label 2"}

[0256] Calculate the sum of squared errors:

[0257]

[0258] in, Represents the sum of squared errors; Data points its group's centroid The Euclidean distance between them ; It is a group Data points in; Data points The quantity.

[0259]

[0260] The error between the analysis results and the prediction results is Since the sum of squared errors is less than the sum of squared errors, secondary analysis is not required.

[0261] Analysis results:

[0262] {“1”,“5.2”,“15000”,“300”,“Label 1”}

[0263] {"2", "7.8", "12000", "250", "Label 2"}

[0264] {“3”, “5.3”, “15500”, “310”, “Label 1”}

[0265] {"4", "9.1", "13000", "270", "Label 2"}

[0266] {“5”, “5.1”, “14900”, “295”, “Label 1”}

[0267] {“6”, “7.7”, “12300”, “255”, “Label 2”}

[0268] {"7", "9.0", "12800", "265", "tag 2"}

[0269] Gas analysis results were obtained through multimodal data fusion and data analysis feedback.

[0270] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0271] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A cloud computing analysis system based on high-performance liquid chromatography, characterized in that: The system includes a data acquisition module, a data transmission module, a cloud storage module, a data analysis module, a user interface module, and a real-time analysis and feedback module; The data acquisition module is used to acquire data detected by the high-performance liquid chromatograph and to perform baseline adjustment on the data; the data transmission module is used to determine the protocol used for data transmission and to compress and encrypt the baseline-adjusted data; the cloud storage module is used to store the data transmitted to the cloud and provides data backup and recovery options. The data analysis module is used to decrypt and decompress the data, input it into the trained intelligent clustering analysis model for data analysis, and optimize the model to improve analysis efficiency and accuracy. The user interface module is used to provide users with an interface to view the data analysis results and record logs. The real-time analysis and feedback module is used to perform real-time analysis of the data, and make judgments when the analysis results are returned. If the error between the analysis result and the prediction result exceeds the sum of squared errors, the analysis scheme and model parameters are immediately adjusted for secondary analysis. The baseline adjustment is used to remove noise and make the signal display peak values; The output of the data acquisition module is connected to the input of the data transmission module; the output of the data transmission module is connected to the input of the cloud storage module; the output of the cloud storage module is connected to the input of the data analysis module; the output of the data analysis module is connected to the input of the user interface module; and the output of the user interface module is connected to the input of the real-time analysis and feedback module. The data analysis module includes a data decryption unit, a data decompression unit, a data preprocessing unit, an intelligent clustering analysis model unit, a multimodal data fusion unit, a model optimization unit, and an analysis result visualization unit. The data decryption unit is used to decrypt encrypted data; the data decompression unit is used to decompress the decrypted data; the data preprocessing unit is used to perform baseline correction, peak detection, and peak integration on the data; the intelligent clustering analysis model unit is used to transmit the preprocessed data to the trained model for automated analysis; the multimodal data fusion unit is used to fuse high-performance liquid chromatography data, mass spectrometry data, ultraviolet-visible spectroscopy data, and infrared spectroscopy data to improve the accuracy and reliability of the analysis results through multimodal data analysis; the model optimization unit is used to adjust the model parameters based on the gas analysis data in the database; and the analysis result visualization unit is used to display the data analysis results in a visual manner. The baseline correction is to remove baseline drift from chromatographic data and to quantify and qualitatively identify peaks in the chromatogram. The peak detection is the process of identifying peaks from the corrected signal; The peak integral is the area of ​​each detected peak. The output of the data decryption unit is connected to the input of the data decompression unit; the output of the data decompression unit is connected to the input of the data preprocessing unit; the output of the data preprocessing unit is connected to the input of the intelligent clustering analysis model unit; and the output of the intelligent clustering analysis model unit is connected to the input of the multimodal data fusion unit. The output of the multimodal data fusion unit is connected to the input of the model optimization unit; the output of the model optimization unit is connected to the input of the analysis result visualization unit and the intelligent clustering analysis model unit, respectively; the output of the analysis result visualization unit is connected to the input of the user interface module.

2. The cloud computing analysis system based on high performance liquid chromatography according to claim 1, characterized in that: The data acquisition module includes an HPLC instrument interface unit and a data processing unit; The HPLC instrument interface unit provides an interface between the data acquisition module and the HPLC instrument, transmitting the data detected by the HPLC instrument to the system; the data processing unit performs baseline adjustment on the data. The output of the HPLC instrument interface unit is connected to the input of the data processing unit; the output of the data processing unit is connected to the input of the data transmission module.

3. The cloud computing analysis system based on high-performance liquid chromatography according to claim 1, characterized in that: The data transmission module includes a transmission protocol unit, a data compression unit, and a data encryption unit; The transmission protocol unit uses the secure transmission protocol HTTPS to ensure the security and integrity of data uploads; The data compression unit is used to compress the collected data; the data encryption unit is used to encrypt and transmit the compressed data. The output of the transmission protocol unit is connected to the input of the data compression unit; the output of the data compression unit is connected to the input of the data encryption unit; and the output of the data encryption unit is connected to the input of the cloud storage module.

4. The cloud computing analysis system based on high performance liquid chromatography according to claim 1, characterized in that: The cloud storage module includes a data storage unit and a backup and recovery unit; The data storage unit is used to store the received data in the cloud. The backup and recovery unit is used to add a mechanism for data backup and recovery, so that data can be viewed and downloaded again in case of data anomalies or loss; The output of the data storage unit is connected to the input of the backup and recovery unit.

5. The cloud computing analysis system based on high performance liquid chromatography according to claim 1, characterized in that: The user interface module includes an authorization management unit, a user interface unit, and a monitoring and logging unit; The authorization management unit is used to authenticate users who view the analysis results; the user interface unit is the interface displayed to the user; the monitoring and logging unit is used to record the users and time of viewing the analysis results. The output of the authorization management unit is connected to the input of the user interface unit; the output of the user interface unit is connected to the input of the monitoring and logging unit; and the output of the monitoring and logging unit is connected to the input of the real-time analysis and feedback module.

6. A cloud computing analysis method based on high performance liquid chromatography, characterized in that: The method includes the following steps: S1. Use an HPLC instrument to analyze the sample, obtain chromatographic data, perform noise reduction processing on the data, and the detector converts the noise-reduced chromatographic data into digital signals and records them in the local data system. S2. After compressing and encrypting the locally stored chromatographic data, upload it to the cloud storage platform using the secure transmission protocol HTTPS to ensure the security and integrity of the data upload. S3. After decrypting and decompressing the received data, use a cloud computing platform for automated data preprocessing and data standardization; S4. Use big data analytics tools on a cloud computing platform for data modeling and analysis; apply the K-means clustering machine learning algorithm to predict and classify gases from the preprocessed chromatographic data, and train an intelligent clustering analysis model; specifically including the following steps: S9-1, will Data points , , , Divided into Group , , , , The value was calculated using the profile coefficient method; Centroid calculation: in, For each group The center of mass, ; It is a group The number of data points in the middle; It is a group Data points in; S9-2, Initialization: Select Initial centroid of the group , , , ; S9-3, Group Assignment: For each data point Calculate the distance from each centroid to the nearest centroid and assign it to the group to which the nearest centroid belongs. ; S9-4, Centroid Update: Calculate the new centroid for each group and update it: in, For each group The center of mass, ; It is a group The number of data points in the middle; It is a group Data points in; S9-5, Iteration: Repeat steps S9-3 and S9-4 until the centroid no longer changes; The silhouette coefficient method is used to determine the number of groups in K-means clustering; specifically: When the number of data points At that time, the initial number of groups Take 2, 3, and 4 respectively, when the number of data points At that time, the initial number of groups Take 5, 6, and 7 respectively; for each The value is calculated by performing the K-means algorithm. The corresponding profile coefficient : in Data points The average distance to other data points in the same set. Data points The average distance to the nearest other data point; in, Data points The group you belong to; It is a group The number of data points in the middle; Data points and data points The Euclidean distance between them; in, Data points The group you belong to; It is a group The number of data points in the middle; Data points and data points The Euclidean distance between them; calculate Corresponding contour coefficient The largest Corresponding That is, the number of groups in the K-means clustering; S5. Combining multimodal data fusion technology, high-performance liquid chromatography data, mass spectrometry data, ultraviolet-visible spectroscopy data and infrared spectroscopy data are fused to improve the accuracy and reliability of analytical results through multimodal data analysis; S6. Add a model optimization mechanism to the intelligent clustering analysis model, and combine the database to adjust the model parameters and optimize the model. S7. Provide an interface for users to view analysis results, add an access control mechanism to verify the permissions of users viewing analysis results, and record logs. S8. Real-time analysis and feedback: After obtaining access rights, a judgment is made at the end of the analysis. If the error between the analysis result and the prediction exceeds the sum of squared errors, the scheme and model parameters are immediately adjusted and a second data analysis is performed. In steps S5-S8, the multimodal data fusion technology involves using a high-performance liquid chromatograph to acquire high-performance liquid chromatography data, a mass spectrometer to acquire mass spectrometry data, a UV-Vis spectrometer to acquire UV-Vis spectrometer data, and an infrared spectrometer to acquire infrared spectrometer data, fusing the four types of data, performing multiple data analyses, and improving the accuracy and reliability of the analysis results through multimodal data analysis. The real-time analysis and feedback process involves: obtaining access permissions and making a judgment after the analysis is completed. If the error between the analysis result and the prediction result exceeds the sum of squared errors, the scheme and model parameters are immediately adjusted for a second data analysis. If the error between the second analysis result and the prediction result is less than the sum of squared errors, the second analysis result is selected and entered into the database. If the error between the second analysis result and the prediction result is greater than or equal to the sum of squared errors, and the error between the second analysis result and the first analysis result is less than the sum of squared errors, the prediction result part of the intelligent clustering analysis model is adjusted and logged. The squared error and : in, Represents the sum of squared errors; Data points its group's centroid The Euclidean distance between them ; It is a group Data points in; Data points The quantity.

7. The cloud computing analysis method based on high performance liquid chromatography according to claim 6, characterized in that: In step S3, The preprocessing includes baseline correction, peak detection, and peak integration; the data standardization is performed on each feature at the same scale using a standard normal distribution with a mean of 0 and a standard deviation of 1. The baseline correction is to remove baseline drift from chromatographic data and to quantify and qualitatively identify peaks in the chromatogram. The peak detection is the process of identifying peaks from the corrected signal; The peak integral is the area of ​​each detected peak.

Citation Information

Patent Citations

  • Data analysis system based on gas / liquid chromatogram and mass spectrum platform

    CN109061020A

  • Human body behavior recognition and data acquisition system based on artificial intelligence

    CN118430070A