Gas chromatograph data optimization storage method

Through the method of structured partitioning of gas chromatograph data and dynamically adjusting the indexing accuracy, the problem of resource allocation imbalance in the existing technology is solved, data storage and retrieval efficiency is improved, and data efficient utilization and accuracy are ensured.

CN120234344AActive Publication Date: 2025-07-01ZENITH SHANGHAI AUTO TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510703183.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

In the existing gas chromatograph data optimization storage technology, the indexing accuracy cannot be dynamically regulated based on the criticality of the data segment in the overall structure, resulting in insufficient indexing capabilities of important data, redundant non-critical data resources, reduced compression efficiency and poor retrieval performance.

Method used

The data collected by the gas chromatograph is divided into data segments according to preset boundary recognition rules, the criticality of each segment is evaluated through structural behavior information, and the index accuracy is dynamically regulated based on the key evaluation index to generate differentiated index bitmaps and compression strategies.

Benefits of technology

It realizes accurate identification and optimal resource allocation of different data segments, improves storage resource utilization, retrieval response capabilities and compression efficiency, and ensures data accuracy and completeness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234344A_ABST
    Figure CN120234344A_ABST
Patent Text Reader

Abstract

The invention discloses a gas chromatograph data optimization storage method, and relates to the technical field of gas chromatograph data storage, and the method specifically comprises the following steps: dividing original data collected by a gas chromatograph into a plurality of data segments according to a preset boundary identification rule; performing an initial feature indexing operation on each data segment, collecting structure behavior information corresponding to each data segment in the initial feature indexing operation process, evaluating the key degree of each data segment in the overall data structure based on the collected structure behavior information, and classifying each data segment according to an evaluation result; dynamically regulating and controlling the indexing precision of each data segment according to a classification result; and generating a corresponding indexing bitmap based on the indexing precision of each data segment after dynamic regulation and control. According to the method, the problem that the indexing precision of the gas chromatography data cannot be dynamically regulated and controlled according to the key degree of the structure is solved, and the resource optimization distribution and the compression retrieval efficiency are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gas chromatograph data storage, and particularly relates to a method for optimizing the storage of gas chromatograph data. Background Art

[0002] A gas chromatograph is an instrument commonly used for separating and analyzing chemical substances, and is widely used in fields such as chemistry, pharmaceuticals, and environmental monitoring. After gasifying the sample, it reacts with the stationary phase and the mobile phase to separate each component, and finally outputs signals related to time through a detector. These signals are usually recorded in the form of data points, forming a series of curve data reflecting the concentration changes of chemical substances. Due to the large amount, complexity, and fineness of the data generated by the gas chromatograph, as the types of experimental samples and the number of tests increase, the amount of data will increase exponentially. To effectively utilize this data, traditional storage methods may encounter problems such as insufficient storage space, data redundancy, and slow access speed. Especially when faced with massive data, it is often impossible to ensure real-time processing and efficient storage. Therefore, it is particularly important to optimize the storage of gas chromatograph data. Through optimized storage, redundancy can be reduced, storage efficiency can be improved, data reading and analysis can be accelerated, while ensuring data accuracy. At the same time, it can better meet the requirements of big data processing, improve the overall performance of the instrument and experimental efficiency. The optimized storage method can not only save hard disk space, but also improve the retrieval speed and processing accuracy of data, making it more efficient and accurate for researchers during data analysis, and ultimately improving the quality of research results and the repeatability of the experimental process.

[0003] The existing gas chromatograph data optimization storage technology mainly reduces the redundancy of data storage and improves storage efficiency and data processing speed through multiple links. First, in the data acquisition stage, through the preprocessing of signal data, such as denoising and data compression, the storage space occupied by irrelevant data and noise can be reduced to ensure that only key information is retained. Secondly, in terms of data storage, an efficient compression algorithm is used to compress the original data into smaller storage files to save storage space without losing important experimental information. At the same time, in view of the temporal and structural characteristics of the data, some optimization technologies will perform block storage or index processing on the data, making subsequent data retrieval and reading more efficient. In addition, some advanced technologies will also use cloud storage and distributed storage architecture to store and back up data in a distributed manner, which not only improves data security, but also optimizes data reading speed through load balancing and reduces storage bottlenecks. In the data processing stage, storage optimization technology usually combines data analysis algorithms to effectively identify and classify stored data, ensuring that different types of data can be quickly accessed according to priority and needs, thereby improving data utilization and experimental processing efficiency. Through the joint action of these links, existing technologies can significantly improve the storage efficiency, access speed and processing accuracy of gas chromatograph data, meeting the needs of efficient data analysis and long-term storage.

[0004] The prior art has the following deficiencies: In the process of optimizing the storage of data collected by the gas chromatograph, in order to support subsequent data compression and fast retrieval, it is usually necessary to extract the structural features of each data segment and generate an indexing bitmap before storage. During this operation, if a uniform precision level is used for indexing settings, when the indexing strategy fails to distinguish the roles of different data segments in the overall structure, there will be an imbalance in the allocation of indexing resources. Specifically, when a data set contains both key segments that bear the main information and non-key segments that are only used for auxiliary purposes, the system often uniformly executes the same rules for feature extraction and indexing precision setting for all segments due to the failure to identify the criticality of each segment in the data structure, resulting in an even distribution of indexing resources, the main segment cannot obtain high-precision indexing support, and the auxiliary segment is redundantly indexed. The existing gas chromatograph data optimization storage technology cannot dynamically adjust the indexing accuracy according to the criticality of each data segment in the overall data structure during the data feature indexing operation, resulting in insufficient indexing capabilities for important data and redundant resources for non-critical data, which in turn leads to a series of problems such as decreased compression efficiency, inaccurate decompression and restoration, and poor data retrieval and positioning performance, which have a substantial impact on the data storage optimization goals.

[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not constitute the prior art that is already known to one of ordinary skill in the art. Summary of the Invention

[0006] The object of the present invention is to provide a method for optimizing the storage of gas chromatograph data to solve the problems in the above-mentioned background technology.

[0007] To achieve the above object, the present invention provides the following technical solution: A method for optimizing the storage of gas chromatograph data, specifically including the following steps: Divide the original data collected by the gas chromatograph into several data segments according to a pre-set boundary recognition rule; Perform an initial feature indexing operation on each data segment. During the initial feature indexing operation, collect the structural behavior information corresponding to each data segment, evaluate the key degree of each data segment in the overall data structure based on the collected structural behavior information, and classify each data segment according to the evaluation result; Dynamically adjust the indexing accuracy of each data segment according to the classification result; Generate a corresponding indexing bitmap based on the indexing accuracy of each data segment after dynamic adjustment; Associate the indexing bitmap of each data segment with the corresponding data number, physical location, and indexing accuracy of the data segment, determine the compression strategy based on the indexing accuracy, perform compression processing, and write the compression result, indexing bitmap, and association information into the storage medium.

[0008] Preferably, performing an initial feature indexing operation on each data segment specifically means: extracting the peak position information, peak spacing information, and change inflection point information in each data segment, and encoding the extraction results in the form of structural markers.

[0009] Preferably, during the initial feature indexing operation, collect the structural behavior information corresponding to each data segment, evaluate the key degree of each data segment in the overall data structure based on the collected structural behavior information, and classify each data segment according to the evaluation result, specifically including the following steps: Collect the structural behavior information corresponding to each data segment during the initial feature indexing operation, and perform preprocessing after collection; Extract the structural coupling behavior information and feature-bearing distribution information from the preprocessed structural behavior information, and perform analysis after extraction to generate the structural coupling coefficient and feature-bearing coefficient of each data segment respectively; Based on the generated structural coupling coefficients and feature-bearing coefficients of each data segment, generate the key evaluation index of each data segment through weighted summation; Determine the pre-set key evaluation index threshold interval, and compare it with the generated key evaluation index of each data segment after determination. Evaluate the key degree of each data segment in the overall data structure according to the comparison result, and classify each data segment according to the evaluation result.

[0010] Preferably, the acquisition logic of the structural coupling coefficient of each data segment is as follows: Extract the structural coupling behavior information from the preprocessed structural behavior information, specifically including the average spacing between all main peaks within each data segment, the average difference in main peak spacing between each data segment and adjacent data segments, the signal slope difference between the head and tail inflection points within each data segment, and the local fluctuation amplitude variance of each data segment during the initial feature indexing operation, and respectively label them as , , and , represents the average spacing between all main peaks within the th data segment during the initial feature indexing operation, represents the average difference in main peak spacing between the th data segment and adjacent data segments during the initial feature indexing operation, represents the signal slope difference between the head and tail inflection points within the th data segment during the initial feature indexing operation, represents the local fluctuation amplitude variance of the th data segment during the initial feature indexing operation, , is a positive integer; Calculate the average value of the average spacing between all main peaks within all data segments during the initial feature indexing operation , according to the formula: ; Calculate the structural coupling coefficient of each data segment, and the specific calculation formula is as follows: ; In the formula, is the structural coupling coefficient of the th data segment.

[0011] Preferably, the acquisition logic of the feature bearing coefficient of each data segment is as follows: Extract the feature bearing distribution information from the preprocessed structural behavior information, specifically including the number of unique main peaks identified in each data segment, the total peak area of all main peaks, and the average number of times all main peaks appear in other data segments during the initial feature indexing operation, and respectively label them as , and , represents the number of unique main peaks identified in the th data segment during the initial feature indexing operation, represents the total peak area of all main peaks in the th data segment during the initial feature indexing operation, represents the average number of times all main peaks in the th data segment appear in other data segments during the initial feature indexing operation, , is a positive integer; Calculate the sum of the total peak areas of all main peaks in all data segments during the initial feature indexing operation , according to the formula: ; Calculate the feature bearing coefficient of each data segment. The specific calculation formula is as follows: ; In the formula, is the feature bearing coefficient of the th data segment.

[0012] Preferably, based on the structural coupling coefficient and the feature bearing coefficient of each generated data segment, generate the key evaluation index of each data segment through weighted summation. The specific calculation formula is as follows: ; In the formula, is the key evaluation index of the th data segment, and are the non-zero weight coefficients of the reciprocal and the feature bearing coefficient of each data segment respectively, and . .

[0013] Preferably, determine the preset key evaluation index threshold interval , and after determination, compare it with the key evaluation index of each generated data segment, evaluate the key degree of each data segment in the overall data structure according to the comparison result, and classify each data segment according to the evaluation result. The specific comparison analysis and classification are as follows: If , the key degree of this data segment in the overall data structure is low key degree, and this data segment is classified as a low key data segment; If , the key degree of this data segment in the overall data structure is medium key degree, and this data segment is classified as a medium key data segment; If , the key degree of this data segment in the overall data structure is high, and this data segment is divided into a high-key data segment.

[0014] Preferably, according to the classification result, the indexing accuracy of each data segment is dynamically adjusted, specifically as follows: For the high-key data segment, set the indexing accuracy to the highest level, and index all structural fields, including all indexing contents of peak sites, structural jump fields, and peak group boundaries; For the medium-key data segment, set the indexing accuracy to the intermediate level, and the only fields to be indexed include the peak site index and the structural field number; For the low-key data segment, set the indexing accuracy to the lowest level, and the only fields to be indexed include the start position and end position of the data segment.

[0015] Preferably, an indexing bitmap corresponding to the indexing accuracy of each data segment after dynamic adjustment is generated, specifically as follows: According to the indexing accuracy level of each data segment, select the corresponding field set for constructing the indexing bitmap; when the indexing accuracy is the highest level, the indexing bitmap includes the peak site index field, the peak group boundary field, the structural field number field, and the structural jump position field; when the indexing accuracy is the intermediate level, the indexing bitmap includes the peak site index field and the structural field number field; when the indexing accuracy is the lowest level, the indexing bitmap only includes the data segment start position field and the end position field.

[0016] In the above technical solution, the technical effects and advantages provided by the present invention are as follows: 1. By introducing two types of quantitative indicators, namely the structure coupling coefficient and the feature bearing coefficient, and combining mathematical modeling and weighted summation methods, the present invention constructs a key evaluation index to comprehensively evaluate the structural stability and information contribution degree of each data segment in the overall data structure. Compared with the traditional method of uniformly processing the indexing accuracy of gas chromatograph data, this method realizes the refined determination from structural behavior analysis, content density measurement to comprehensive scoring, can accurately identify the differences between the structural backbone segments and the redundant auxiliary segments, and improves the accuracy and pertinence of data structure recognition.

[0017] 2. By mapping the key evaluation index to multi-level key degree classification and then dynamically adjusting the indexing accuracy of each data segment based on the classification result, the present invention realizes the differential configuration of the indexing strategy. The system no longer processes all data with a fixed template, but automatically determines the indexing field range and accuracy level according to the criticality of the data segment, enabling high-value segments to obtain complete indexing coverage and low-value segments to only record the necessary boundaries. This mechanism significantly improves the utilization rate of storage resources, avoids the waste of high-density indexing resources on redundant data segments, and at the same time ensures the retrieval accuracy and compression and restoration quality of the core segments.

[0018] 3. During the process of performing optimization compression and storage, the present invention structurally associates the indexing bitmap with information such as data numbers, physical locations, and indexing accuracies, and automatically matches the compression strategy template according to different accuracy levels, realizing the compression path control driven by key factors. Finally, the compression results, indexing bitmap, and associated information are uniformly packaged and written into the storage medium, which not only improves the compression efficiency but also enhances the retrieval response ability and structure restoration ability of the system in scenarios of multi-source samples and large-scale data storage. The overall solution takes into account high-performance compression and structure fidelity, and has good engineering practical value and system scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.

[0020] Figure 1 It is a schematic flowchart of a method for optimizing the storage of gas chromatograph data according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] Now, the exemplary embodiments will be described more comprehensively with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these exemplary embodiments are provided so that the present disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art.

[0022] The present invention provides a method for optimizing the storage of gas chromatograph data as shown in Figure 1 the following, which specifically includes the following steps: Dividing the original data collected by the gas chromatograph into several data segments according to the preset boundary recognition rules; To achieve the automatic division of the original data collected by the gas chromatograph, the system can scan and analyze the original chromatographic signal in software, use the signal feature recognition algorithm to identify the points with boundary features, and divide the continuous data stream into several structurally complete data segments. The specific division method can be based on the sliding window algorithm, signal gradient detection algorithm, or change rate threshold judgment method. The system first sets a continuous sampling interval as the sliding window, and analyzes the first derivative or difference value in the signal curve in real time to identify the obvious starting and ending points of the peak, inflection points, flat segments, and mutation segments; when these boundary features are detected, the system regards this point as the starting and ending positions of the data segment division, and automatically divides the continuous data into structurally independent data segments. This division method can be implemented in software at the back end of data collection or in the preprocessing module, and has the capabilities of real-time and automation.

[0023] The pre-set boundary recognition rules refer to a set of judgment criteria for identifying paragraph boundaries that are set in advance based on the structural characteristics of the gas chromatograph signal before performing the data segmentation operation. The rules mainly include the following parameters or thresholds: peak height change threshold, peak area mutation value, signal stability duration, inflection point number determination, baseline fluctuation intensity, signal-to-noise ratio change interval, etc. When the system is running, it will segment the signal according to these rules. For example, when the change rate of N consecutive sampling points is lower than the set threshold and there is an inflection point structure before, it can be judged that the position is a paragraph boundary. Boundary recognition rules can be established through historical data statistical analysis, adaptive learning or manual setting, and can be flexibly configured in the software system parameters for structural division adaptation of different samples or different types of chromatographic data.

[0024] In the process of optimizing the storage of gas chromatography data, dividing the continuous raw data stream into several structurally complete data segments is a prerequisite for achieving structured indexing and differentiated storage regulation. Since gas chromatography data has significant stage-by-stage change characteristics, for example, the signal may stably express the characteristics of a single component in a certain segment, while another segment contains overlapping peaks of multiple components. Therefore, each data segment may be different in terms of information structure, importance, and decompression utilization value. If effective division is not performed and the entire signal is treated as a unified structure, it will be impossible to perform structure recognition, precision control, and compression strategies in a targeted manner, resulting in waste of indexing resources, redundant storage, and low efficiency of subsequent retrieval. By achieving precise division through software, independent evaluation, indexing, and compression paths can be established for each subsequent data segment, realizing structural optimization and resource allocation at the smallest granularity unit level.

[0025] Performing an initial feature indexing operation on each data segment, collecting structural behavior information corresponding to each data segment during the initial feature indexing operation, evaluating the criticality of each data segment in the overall data structure based on the collected structural behavior information, and classifying each data segment according to the evaluation result; In this embodiment, an initial feature indexing operation is performed on each data segment, specifically: peak position information, peak distance information and change inflection point information in each data segment are extracted, and the extraction result is encoded in a structural tagging manner.

[0026] The numerical sequence of each data segment can be analyzed by a signal analysis algorithm to perform one-dimensional curve structure analysis, and extract peak position information, peak spacing information, and inflection point information of changes. First, a first-order difference or local extreme value detection method is used to scan the continuous signal sequence to identify the local maximum positions, which are potential peak points. Then, false detection noise is filtered by a set peak height threshold and minimum peak width to extract effective peak position information. Next, based on the extracted effective peak sequence, the sampling point spacing or time interval between adjacent peaks is calculated to form a peak spacing information sequence, reflecting the signal structure density and change rhythm. Subsequently, by analyzing the change trend of the second derivative of the data segment, the inflection points of the function concavity and convexity are identified, thereby extracting the inflection points of changes in the signal. These operations can be implemented based on a sliding window function, digital filters (such as Savitzky–Golay filtering), and threshold judgment logic, and are completed in a segment-by-segment sliding manner in the algorithm module. All calculations can be efficiently and automatically completed in the software.

[0027] After the key features of each data segment are extracted, the software will convert this information into a structural tag form for encoding processing to support subsequent structure recognition and indexing operations. The encoding process can be completed by constructing a feature vector. First, a unified structural tag template is defined, such as fields like "peak start point number", "peak end point number", "peak spacing level", "inflection point type", "inflection point position", etc. Subsequently, the actually extracted feature values are mapped into this template structure to form standardized structural tag entries. Each entry can be encoded as a specific data segment, such as an integer displacement value, a Boolean status bit, or an index label, and is organized in a structured data table or bitmap structure. During the encoding process, it can be represented in the form of a multi-dimensional array, key-value pair structure, or sparse matrix, and is uniformly numbered or block-divided so that the structural features of different data segments are comparable and retrievable. The encoding result is finally used as the structural basis input for subsequent processing steps such as evaluation, classification, and indexing, and is represented in logical bits or participates in subsequent model processing as an input vector within the system.

[0028] In this embodiment, during the initial feature indexing operation, the structural behavior information corresponding to each data segment is collected, and based on the collected structural behavior information, the key degree of each data segment in the overall data structure is evaluated, and each data segment is classified according to the evaluation result. The specific steps are as follows: During the initial feature indexing operation, the structural behavior information corresponding to each data segment is collected and preprocessed after collection; During the initial feature indexing operation, the acquisition of structural behavior information can be completed by the software through real-time analysis and feature modeling of chromatographic signals in each data segment. The acquisition methods include derivative analysis, threshold detection, and local statistics-based methods. Specifically, the system first traverses in units of data segments, performs first-order derivative calculations within a sliding window based on the numerical sequence of consecutive sampling points, and extracts the signal change rate at each sampling point. At the same time, it identifies local maximum and minimum points, locates the start and end positions of the main peaks, inflection points, and fluctuation boundaries in the signal, and calculates the distance and slope change between every two main peaks. In addition, the volatility and signal density of this segment are quantified by statistically calculating the data variance and peak area within the window. All the above structural behavior characteristics are extracted through an automatic algorithm during the in-segment scanning process and are uniformly organized into a set of structural behavior information containing multi-dimensional numerical values such as main peak distribution, inflection point slope, peak distance, and signal fluctuation, which constitutes the basic input for subsequent structural analysis and criticality assessment.

[0029] The purpose of preprocessing the structural behavior information is to improve the accuracy and stability of subsequent parameter calculations, eliminate invalid or abnormal interference factors, unify the data scale, and construct a standard feature format for mathematical modeling. In software implementation, preprocessing usually includes three aspects: First, perform normalization processing, standardize each item of structural behavior information (such as peak distance, slope, fluctuation value, etc.) according to the maximum value or mean of all data segments to eliminate calculation offsets caused by differences in numerical dimensions between different segments. Second, perform outlier detection and suppression processing, identify abnormal peak distances or slope mutation data through statistical judgments (such as the Z-score method or IQR method), and interpolate and repair or eliminate them to ensure that the parameter results are not misled by extreme points. Third, unify the vector format, encode all numerical features into feature vectors with consistent structures, which is convenient for subsequent use as input variables in the calculation of structural coupling coefficients and feature bearing coefficients. The entire preprocessing process is completely automatically completed by the software logic module without relying on manual intervention, and has high stability and engineering versatility.

[0030] Extract the structural coupling behavior information and feature bearing distribution information from the preprocessed structural behavior information, and perform analysis after extraction to generate the structural coupling coefficient and feature bearing coefficient for each data segment respectively; After the preprocessing of the structural behavior information is completed, the extraction of the structural coupling behavior information and the feature-bearing distribution information can be automatically completed by the software by performing specific feature mapping and numerical screening rules on the standardized data features. When extracting the structural coupling behavior information, the system first calculates the difference in the average main peak spacing between the current data segment and the adjacent data segments based on the main peak position index and inflection point information of each data segment, as well as the difference in the signal slope at the inflection points at the beginning and end of this segment. Then, the local variance value of the internal signal of this segment is calculated within a sliding window. The above three dimensions constitute the structural coupling behavior information of this segment. When extracting the feature-bearing distribution information, the system calls the main peak marking index, counts the number of unique main peaks in this segment that do not appear in other data segments, calculates the total peak area corresponding to all main peaks, and retrieves the average occurrence frequency of each main peak in other segments in the database, which constitutes the feature-bearing distribution information of this segment. The entire extraction process is completed based on the built-in feature index logic and statistical analysis algorithm of the software, and each piece of information is automatically matched and classified according to the data segment number without manual operation.

[0031] Based on the structural coupling coefficients and feature-bearing coefficients of each generated data segment, the key evaluation index of each data segment is generated through weighted summation. Determine the preset key evaluation index threshold interval, and compare it with the key evaluation index of each generated data segment after determination. According to the comparison result, evaluate the key degree of each data segment in the overall data structure, and classify each data segment according to the evaluation result.

[0032] The determination of the key evaluation index threshold interval can be automatically generated by the software through adaptive learning in combination with the historical sample data set during the modeling stage or through empirical statistical analysis, ensuring that the classification boundary is representative and discriminative. The specific implementation method is as follows: The system first collects the structural behavior information and calculates the characteristic parameters of a large number of historical chromatographic data with known labels or clear classification results, and generates a complete key evaluation index distribution sequence based on this; then, clustering analysis, density estimation or distribution function fitting processing is performed on this sequence, such as using Gaussian mixture models, K-means clustering or kernel density estimation, etc., to identify the natural aggregation intervals and boundary transition points in the key index distribution; according to the clustering results or the demarcation points in the distribution curve, two critical values are automatically generated, which are used as the threshold boundaries for low key to medium key and medium key to high key respectively, thus constructing a complete key evaluation index threshold interval. This process can be completely run once by the software during the initialization stage, or can be updated automatically based on new data regularly to ensure that the threshold boundary dynamically matches the data structure characteristics.

[0033] In this embodiment, the acquisition logic of the structural coupling coefficient of each data segment is as follows: Extract the structural coupling behavior information from the preprocessed structural behavior information, specifically including the average spacing between all main peaks within each data segment, the difference in average main peak spacing between each data segment and adjacent data segments, the difference in signal slope between the first and last inflection points within each data segment, and the variance of the local fluctuation amplitude of each data segment, and calibrate them respectively as 、 、 and , represents the average spacing between all main peaks within the -th data segment during the initial feature indexing operation, represents the difference in average main peak spacing between the -th data segment and adjacent data segments during the initial feature indexing operation, represents the difference in signal slope between the first and last inflection points within the -th data segment during the initial feature indexing operation, represents the variance of the local fluctuation amplitude of the -th data segment during the initial feature indexing operation, , is a positive integer; During the initial feature indexing operation, the average spacing between all main peaks within each data segment, the difference in average main peak spacing between adjacent data segments, the difference in signal slope between the first and last inflection points, and the variance of the local fluctuation amplitude can all be automatically extracted and calculated by software algorithms based on continuously sampled data. First, the system scans the chromatographic signal curve within each data segment through the main peak recognition algorithm, identifies the local maximum points that meet the set peak height, peak width, and symmetry conditions, and extracts the position indices of these main peaks; subsequently, calculate the spacing between adjacent main peaks within each data segment and take the average to obtain the average main peak spacing ; To obtain the difference in main peak spacing , the system will simultaneously extract the average main peak spacing of the previous segment and the next segment, perform a difference calculation with the current segment and normalize it to reflect the degree of mutation in the main peak distribution rhythm of the segment; at the same time, the system automatically identifies the first and last inflection points of each data segment based on curve derivative operations, calculates the derivative values of the signals at these two points, and then calculates their slope difference to measure the overall transition strength of the segment structure; for the variance of the local fluctuation amplitude , the system establishes a sliding window within this segment, squares and averages the deviations between each sampling point within the window and its local mean, and performs a global mean aggregation across the entire segment to reflect the signal fluctuations and volatility within this segment. All of the above data can be automatically completed in the software through signal derivative analysis, extreme value detection, and local statistical functions, with a stable and repeatable calculation logic that does not rely on manual intervention.

[0034] The "main peak" refers to the local maximum point with significant intensity characteristics and satisfying morphological constraints identified through feature extraction of the chromatographic signal curve during the initial feature indexing operation, with clear physical and structural meanings. Specifically, the main peak must simultaneously meet the following conditions: First, the peak height must be significantly higher than the local baseline level, usually using the local window mean multiplied by a set multiple as the judgment threshold; Second, the peak width should be within the preset effective range, that is, the horizontal expansion distance of its left half-peak and right half-peak cannot be too narrow (excluding high-frequency noise points) nor too wide (excluding platform drift); Third, the peak shape should exhibit an obvious rising-peak-descending structure, showing a positive-negative change trend in the first derivative, with basic symmetry or can be approximated as a single-peak structure by a Gaussian function; Fourth, this peak must form an independent extreme value within the current data segment and cannot extend across segments or attach to the boundary. The identification of the main peak is completed based on the derivative analysis, extreme value search, and morphological screening rules of the software on continuous sampling signals, and it is the basic feature unit for subsequent extraction of key indicators such as the main peak spacing, main peak area, and feature distribution. The main peak refers to the local maximum point located within each data segment during the initial feature indexing operation, with a peak height greater than a set multiple of the local baseline mean, a peak width within the effective threshold range, and a monotonically rising and then monotonically descending shape.

[0035] Calculate the average of the average distances between all main peaks within all data segments during the initial feature indexing operation , according to the formula: ; Calculate the structural coupling coefficient of each data segment. The specific calculation formula is as follows: ; In the formula, is the structural coupling coefficient of the th data segment.

[0036] Using this calculation method for the structural coupling coefficients of each data segment aims to comprehensively evaluate the dispersion degree of the data segment in terms of structural continuity, transition stability, and internal volatility, and unify and quantify the structural change characteristics in multiple dimensions into a single scoring index. In the formula, represents the relative change amplitude of the main peak distribution rhythm between this data segment and the adjacent data segment. After adding 1 to ensure that the value range is non-zero, it is squared to amplify the impact of rhythm mutations; It is used to reflect the change intensity of the start and end positions of the internal structure of this data segment. The square operation enhances the recognition sensitivity to the structural mutation segment; It represents the microscopic fluctuation intensity within the segment. The direct accumulation is used to reflect the negative impact of local noise or signal instability on structural coupling; it is overall wrapped in the natural logarithm function to compress non-linear growth and prevent high-variability data from biasing the subsequent evaluation coefficients in weighted synthesis, so that the calculation results have stable gradient discrimination ability and physical interpretation consistency. The design of this formula ensures that each structural deviation factor can be accurately captured without causing imbalance amplification to the overall index, reflecting the comprehensive expression of the coupling intensity under multi-source heterogeneous structural perturbations.

[0037] In the overall data structure, for the structural coupling coefficient of the th data segment, the smaller it is, the closer this data segment is to its adjacent data segments in terms of structural rhythm, transition slope, and internal fluctuations, showing stronger structural continuity and coupling stability. It is usually located on the main path of the data structure and has a connecting function of linking the preceding and the following. Therefore, it has more retention value in structural analysis and data compression and can be evaluated as a segment with a higher degree of criticality; conversely, if the value is larger, it indicates that there are significant mutations in the main peak distribution rhythm between this segment and the preceding and following data segments, the changes at the start and end of the structure are drastic, or the internal signal fluctuations are frequent, showing stronger structural discontinuity and marginalization characteristics. It is often a local abnormal segment, a turning segment, or an information redundancy segment, and its importance to the overall structure is relatively low, so its degree of criticality is correspondingly low. Therefore, through the numerical judgment of the structural coupling coefficient, the embedding degree and supporting role of this segment in the structural hierarchy can be effectively reflected, and thus an accurate basis in the structural dimension can be provided for the classification evaluation of the degree of criticality.

[0038] In this embodiment, the acquisition logic of the characteristic bearing coefficients of each data segment is as follows: Extract the characteristic bearing distribution information from the preprocessed structural behavior information, specifically including the number of unique main peaks identified in each data segment during the initial feature indexing operation, the total peak area of all main peaks, and the average number of times all main peaks appear in other data segments, and they are respectively calibrated as 、 and , represents the number of unique main peaks identified in the th data segment during the initial feature indexing operation, represents the total peak area of all main peaks in the th data segment during the initial feature indexing operation, represents the average number of times all main peaks appear in the The average number of occurrences of all the main peaks in a data segment in other data segments , is a positive integer; During the initial feature indexing operation, the number of unique main peaks identified in each data segment, the total peak area of all main peaks, and the average number of occurrences of all main peaks in other data segments can all be automatically extracted and analyzed by software. The specific implementation method is as follows: First, the system detects main peaks in each data segment through a peak recognition algorithm. The main peak needs to meet the conditions that the peak height is higher than a multiple of the local baseline, the peak width is within an effective interval, and the shape has a standard structure of rising - peak top - falling, and shows typical maximum value characteristics in the derivative change. The system unifies the position index, intensity, and peak shape characteristics of all main peaks into the main peak feature library. Subsequently, for the data item of "the number of unique main peaks", the system compares the main peak set in each data segment with the main peak sets of all other data segments one by one. If a certain main peak does not have a similar peak in other data segments (a matching threshold is set according to the position difference, peak shape similarity, and relative peak height error), then this main peak is regarded as a unique main peak, and the system counts its number as the number of unique main peaks in this segment . For the data item of "the total peak area of all main peaks", the system takes the start and end points of each main peak as boundaries and uses the integral method to calculate the curve area under the main peak, and then sums up the areas of all main peaks to obtain the total peak area of this data segment , which is used to reflect the overall signal strength contribution of this segment. And the "average number of occurrences of all main peaks in other data segments" is obtained by counting the frequencies of each main peak in this data segment in all other data segments. The matching criteria include peak position difference, shape coincidence degree, etc. After averaging the frequencies of all main peaks in other segments, it is the average number of occurrences of this segment , which reflects the scarcity degree of the information in this segment. The above three types of data are all implemented in the software based on the main peak matching and feature comparison algorithm, with high automation and no ambiguity, and are the basic data sources for calculating the feature bearing coefficient later

[0039] Calculate the sum of the total peak areas of all main peaks in all data segments during the initial feature indexing operation , according to the formula: ; Calculate the feature bearing coefficient of each data segment. The specific calculation formula is as follows: ; In the formula, is the feature bearing coefficient of the th data segment

[0040] The core purpose of using this formula is to integrate the quantitative performance of a data segment in terms of feature strength, feature scarcity, and structural uniqueness into a single score value to evaluate its value in data compression and information retention. It indicates the ratio of the main peak area of ​​the data segment to the total main peak area, reflecting its contribution to the overall energy or signal intensity; the second item A nonlinear amplification function representing the scarcity of the main peak, The smaller the value, the less likely the feature of this segment is repeated in other segments. The larger the index value, the more effective it is in highlighting the importance of rare segments. The addition of 1 prevents the denominator from being zero and weakens the expansion effect of extreme values. The third item The number of unique main peaks in this section is then enhanced twice and then logarithmically compressed to balance the order of magnitude difference and avoid scoring bias caused by too many main peaks. The combination of the three not only considers the total amount of characteristic signals, but also takes into account their irreplaceability and uniqueness, so that the final The value can fully reflect the comprehensive carrying capacity of the data segment at the content information level, and support accurate judgment of the criticality.

[0041] In the overall data structure, Characteristic load factor of each data segment The larger the value, the more important the data segment is in terms of content information, and thus it should be given a higher priority in the criticality assessment. When the value is large, it means that the data segment not only bears a higher proportion of signal intensity in all segments (reflected by the proportion of the total peak area), but also contains multiple main peak features that are rare or even unique in other data segments (expressed by the main peak scarcity index item), and its structural content has strong exclusivity and irreplaceability (reflected by the enhanced logarithm of the number of unique main peaks). These factors together show that the data segment carries rich, unique and recognizable key information. If it is compressed or ignored, it will seriously affect the complete restoration of the data and the accuracy of subsequent retrieval. Therefore, the data segment with a higher feature load factor is more likely to be classified as "highly critical" in the criticality assessment, and is the object that must be retained and indexed with high precision in the optimized storage strategy.

[0042] In this embodiment, based on the generated structural coupling coefficients of each data segment and characteristic load factor , the key evaluation index of each data segment is generated by weighted summation. The specific calculation formula is as follows: ; In the formula, For the Key evaluation index for each data segment, and The structural coupling coefficients for each data segment reciprocal and the feature bearing coefficients with non - zero weight coefficients, and .

[0043] In the process of generating the key evaluation index for each data segment , the system performs a combined weighted calculation on the structural coupling coefficient and the feature bearing coefficient calculated in the early stage through software logic to achieve a comprehensive evaluation of the structural features and content value. Specifically, when implementing, the system first takes the reciprocal form for , which is used to reflect the positive contribution of the structural coupling strength, that is, the more continuous the structure and the stronger the coupling, the larger this value, indicating a higher structural criticality; then retains in its original form as a direct reflection of the paragraph content bearing capacity, representing the comprehensive value of the paragraph in terms of feature strength, scarcity, and information density. To balance the weight influence of structure and content, the system pre - sets two weight coefficients and , where is used to control the proportion of the structural factor ( ) in the overall evaluation, is used to control the weight of the feature factor ( ), both of which are non - zero real numbers and satisfy the normalization condition . These two weight coefficients can be adjusted according to different application scenarios: if more emphasis is placed on structural continuity, can be set; if more emphasis is placed on feature information density, then is set. Finally, the key evaluation index calculated through is used as a unified quantitative index to measure the critical degree of each data segment, providing an accurate basis for subsequent classification processing.

[0044] In this embodiment, a pre - set key evaluation index threshold interval is determined, and after determination, it is compared with the key evaluation index of each generated data segment. According to the comparison result, the critical degree of each data segment in the overall data structure is evaluated, and each data segment is classified according to the evaluation result. The specific comparison analysis and classification are as follows: If , the critical degree of this data segment in the overall data structure is a low critical degree, and this data segment is classified as a low - critical data segment; This situation indicates that the comprehensive performance of this data segment is poor in terms of both structural coupling and content bearing. It neither shows significant structural continuity nor contains valuable feature information, and is often an edge segment of the structure, a noise interference segment, or a data redundancy segment. Such data segments do not play a core connection or feature expression role in the overall data structure. If high-precision indexing or retaining the complete signal is performed, it will cause waste of storage resources. Therefore, during the storage optimization process, a low-precision indexing or skipping indexing strategy can be adopted for it, only retaining the boundary position, summary information, or pointer index, and maintaining the most basic structure positioning function with the smallest data cost, so as to effectively improve the overall compression ratio and retrieval efficiency.

[0045] If , the criticality of this data segment in the overall data structure is medium criticality, and this data segment is classified as a medium-critical data segment; This situation indicates that it makes a certain contribution in terms of structural coupling or content bearing, but is not sufficient to be recognized as a core structure segment or a feature backbone segment. It usually appears as an auxiliary connection segment, an information transition segment, or a feature-bearing segment with medium confidence. Such data segments may provide context support in specific query, restoration, or partial decoding scenarios, have a certain degree of compressibility but cannot be completely discarded. Therefore, a medium-precision strategy should be adopted during indexing and compression processing, retaining the main structure fields, significant feature loci, or peak group indexes to ensure its usability in retrieval coverage and restoration inference, while avoiding resource redundancy caused by high-intensity coding.

[0046] If , the criticality of this data segment in the overall data structure is high criticality, and this data segment is classified as a high-critical data segment.

[0047] This situation indicates that this data segment shows good performance in terms of structural continuity, and its content features are dense, unique, and rare. It not only has the meaning of backbone connection but also carries high-value feature signals, belonging to the core critical segment in the overall data structure. Such data segments usually determine the accuracy of the compressed structure restoration and the hit rate of the retrieval results, and are the key objects for information retention. Therefore, during the storage optimization process, a high-precision indexing strategy should be implemented for it, extracting and recording all structural features, main peak information, position information, and auxiliary feature bitmaps to ensure its maximum role in subsequent high-precision retrieval and reconstruction, while avoiding any information loss or incorrect deletion from having an adverse impact on the system integrity.

[0048] According to the classification results, dynamically adjust the indexing precision of each data segment; In this embodiment, according to the classification results, dynamically adjust the indexing precision of each data segment, specifically as follows: For high-critical data segments, set the indexing accuracy to the highest level, and index all structural fields, including all indexing content of peak positions, structural jump fields, and peak group boundaries; For data segments classified as high-critical, when the system performs indexing accuracy regulation, it calls the highest indexing template through software, activates all available structural fields, and performs field-by-field traversal indexing. Specifically, it includes: extracting and recording the position indexes of all main peaks in the segment to construct a peak position index field; identifying and marking the start and end boundaries of all peak groups to form a peak group boundary field; analyzing the sequence of structural fields in the data segment (such as chromatographic logical grouping, sample channel labels, etc.) to generate structural field numbers; analyzing the logical jump sequence between structural fields to construct a structural jump position field. All these fields will be accurately recorded as the content of high-precision indexing. The reason for this is that high-critical segments usually undertake the functions of backbone information transmission and important feature expression, and their structural information and feature density determine whether they can be restored with high fidelity after compression. Therefore, it is necessary to cover all aspects and handle without omission during the indexing stage.

[0049] For medium-critical data segments, set the indexing accuracy to the intermediate level, and the only fields to be indexed include the peak position index and the structural field number; For medium-critical data segments, the system uses a medium indexing template to control the indexing accuracy, and limits the indexing range through software by calling a field selector, only enabling the extraction and recording of specific fields. The specific implementation is: indexing the position index field of the main peak to support peak-level retrieval and positioning; at the same time, indexing the structural field number to ensure the basic embedding logic during the structural restoration of the segment. The remaining fields, such as structural jump positions and peak group boundaries, are excluded by the system during this stage to reduce the data processing burden. The basis for implementing this strategy is that although medium-critical segments do not form the information core, they still play an auxiliary role in structural connection or feature complementation. Therefore, retaining their key identification fields helps to restore the integrity of the compressed information, but there is no need to process them as comprehensively as high-critical segments, achieving a balance between processing efficiency and information retention.

[0050] For low-critical data segments, set the indexing accuracy to the lowest level, and the only fields to be indexed include the start position and end position of the data segment.

[0051] For low-critical data segments, the system executes a minimization indexing strategy. By invoking the lowest-precision template, the extraction logic for all structural and feature fields is turned off, and only the boundary identifiers of the data segments are retained. The specific method is as follows: The system sets a boundary marker value at each of the starting sampling point and the ending sampling point of this segment to form an indexing field, which is only used to record the physical position of this segment in the data stream. High-cost fields such as main peaks, structural fields, and feature indexes are no longer extracted. The core reason for this simplified processing method is that low-critical segments often only carry repetitive backgrounds, low-confidence features, or noise segments. Completely retaining their structural features will not significantly improve the compression effect or retrieval performance, but will instead waste computing resources and storage space. Therefore, this strategy achieves extreme data compression and optimization of processing efficiency on the premise of ensuring the integrity of the basic structure.

[0052] Generate a corresponding indexing bitmap based on the indexing precision of each data segment after dynamic regulation; In this embodiment, generating a corresponding indexing bitmap based on the indexing precision of each data segment after dynamic regulation is specifically as follows: According to the indexing precision level of each data segment, select the corresponding field set for constructing the indexing bitmap; when the indexing precision is at the highest level, the indexing bitmap includes a peak point index field, a peak group boundary field, a structural field number field, and a structural jump position field; when the indexing precision is at the intermediate level, the indexing bitmap includes a peak point index field and a structural field number field; when the indexing precision is at the lowest level, the indexing bitmap only includes a data segment start position field and an end position field.

[0053] In the process of generating the corresponding indexing bitmap according to the indexing accuracy of each data segment, the system can preset multiple sets of indexing field configuration templates through software modules, and automatically select the corresponding template for bitmap construction according to the indexing accuracy level of the data segment, so as to achieve the generation of accuracy-driven structured indexing. The specific implementation method is as follows: after the indexing accuracy is determined, the system first identifies the accuracy level of the data segment (such as the highest, intermediate or lowest), and then loads the corresponding field template, which contains the field types and structure definitions to be represented in the indexing bitmap. For data segments with the highest accuracy level, the software will fully enable the peak position index field (used to mark the position index of the main peak in the original signal), the peak group boundary field (used to locate the start and end points of the peak group), the structure field number field (encoding the logical structure units in the data segment), and the structure jump position field (marking the jump structure relationship between segments or fields), and perform bitmap encoding and assembly in the order of field definitions; for data segments with the intermediate accuracy level, only the peak position index field and the structure field number field are loaded for bitmap construction, and the other fields are skipped to control the complexity; for data segments with the lowest accuracy level, the system only extracts the start and end sampling point indexes of the data segment and constructs the bitmap information with the smallest range. By generating the bitmap in this way of mapping the field set by level, the indexing range and density of each data segment can be accurately controlled at the software level, avoiding excessive storage overhead caused by unified processing of redundant fields, while retaining the necessary information of data segments with different critical degrees during the structure retrieval and compression restoration processes, and ensuring the balance between indexing efficiency and data integrity.

[0054] Associate the indexing bitmap of each data segment with the corresponding data number, physical location and indexing accuracy of this data segment, determine the compression strategy based on the indexing accuracy, perform compression processing, and write the compression result, indexing bitmap and associated information into the storage medium.

[0055] In the process of associating the indexing bitmap of each data segment with the corresponding data number, physical location and indexing accuracy, the system can construct a unified data mapping structure through software to organize and index the indexing information of all data segments. The specific implementation method is as follows: after the system completes the generation of the indexing bitmap of each data segment, it immediately obtains the number identifier of this data segment in the sampling sequence, the start and end sampling point positions (i.e., the physical location), and the indexing accuracy level adopted by this segment, and encapsulates these information into a group of associated metadata. Subsequently, the system maps and binds the indexing bitmap with this group of metadata to form a structured data segment index unit, and stores it in the indexing mapping table uniformly. This mapping table can be organized in the form of a hash table or a multi-level index, which is convenient for subsequent rapid retrieval and segment-level positioning and accuracy restoration during decompression. The purpose of doing this is to ensure that each indexing bitmap not only has an independent structural meaning, but also can establish a clear corresponding relationship with the original data structure, providing basic support for data management, reconstruction and verification, and avoiding information misalignment or accuracy confusion.

[0056] After the indexing and meta - information binding are completed, the system automatically matches the corresponding compression strategy template based on the recorded indexing accuracy level, and performs differential compression processing on each data segment. The software system presets several compression strategies, and each strategy controls the compression granularity, redundant data deletion rules, feature retention range, etc. according to the indexing accuracy level. After the compression processing is completed, the system packages the compressed data content, the corresponding indexing bitmap, and their bound numbers, positions, and accuracy information, and writes them into the specified storage medium (such as a database, a file system, or a distributed storage node) in a unified data block format. This structured and hierarchical compression writing method not only realizes the coordination of the information compression rate and data integrity, but also significantly improves the efficiency of subsequent decompression on demand and feature location, which is a key link to optimize the response speed of the storage system and reduce the occupancy of storage resources.

[0057] The above - mentioned formulas are all dimensionless and take their numerical values for calculation. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.

[0058] The above - mentioned embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above - mentioned embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general - purpose computer, a special - purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer - readable storage medium, or transmitted from one computer - readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer - readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or a data center that contains one or more sets of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid - state drive.

[0059] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above - mentioned processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0060] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0061] In several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the above-described embodiments are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0062] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0063] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0064] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A method for optimizing the storage of gas chromatograph data, characterized in that, Specifically, the following steps are included: Divide the original data collected by the gas chromatograph into several data segments according to the pre-set boundary recognition rules; Perform an initial feature indexing operation on each data segment. During the initial feature indexing operation, collect the structural behavior information corresponding to each data segment, evaluate the key degree of each data segment in the overall data structure based on the collected structural behavior information, and classify each data segment according to the evaluation results; Dynamically adjust the indexing accuracy of each data segment according to the classification results; Generate a corresponding indexing bitmap based on the indexing accuracy of each data segment after dynamic adjustment; Associate the indexing bitmap of each data segment with the corresponding data number, physical location, and indexing accuracy of the data segment. Determine the compression strategy based on the indexing accuracy, perform compression processing, and write the compression results, indexing bitmap, and association information into the storage medium; 2. The data optimization storage method of a gas chromatograph according to claim 1, wherein Perform an initial feature indexing operation on each data segment. Specifically: extract the peak position information, peak spacing information, and inflection point information of each data segment, and encode the extraction results in the form of structural tags; 3. A method for optimizing the storage of gas chromatograph data according to claim 2, characterized in that, During the initial feature indexing operation, collect the structural behavior information corresponding to each data segment, evaluate the key degree of each data segment in the overall data structure based on the collected structural behavior information, and classify each data segment according to the evaluation results. Specifically, the following steps are included: Collect the structural behavior information corresponding to each data segment during the initial feature indexing operation, and perform preprocessing after collection; Extract the structural coupling behavior information and feature-bearing distribution information from the preprocessed structural behavior information, and perform analysis after extraction to generate the structural coupling coefficient and feature-bearing coefficient of each data segment respectively; Based on the generated structural coupling coefficient and feature-bearing coefficient of each data segment, generate the key evaluation index of each data segment through weighted summation; Determine the pre-set key evaluation index threshold interval, compare it with the generated key evaluation index of each data segment after determination, evaluate the key degree of each data segment in the overall data structure according to the comparison results, and classify each data segment according to the evaluation results; 4. A method for optimizing the storage of gas chromatograph data according to claim 3, characterized in that, The acquisition logic of the structural coupling coefficient of each data segment is as follows: Extract the structural coupling behavior information from the preprocessed structural behavior information, specifically including the average spacing between all main peaks within each data segment during the initial feature indexing operation, the average difference in the main peak spacing between each data segment and the adjacent data segment, the signal slope difference between the head and tail inflection points within each data segment, and the local fluctuation amplitude variance of each data segment, and label them respectively as , , and , represents the average spacing between all main peaks within the th data segment during the initial feature indexing operation, represents the average difference in the main peak spacing between the th data segment and the adjacent data segment during the initial feature indexing operation, represents the signal slope difference between the head and tail inflection points within the th data segment during the initial feature indexing operation, represents the local fluctuation amplitude variance of the th data segment during the initial feature indexing operation, , is a positive integer; Calculate the average of the average spacings between all the main peaks within all the data segments during the initial feature indexing operation , according to the formula: ; Calculate the structural coupling coefficient of each data segment. The specific calculation formula is as follows: ; In the formula, is the structural coupling coefficient of the th data segment.

5. A method for optimizing the storage of gas chromatograph data according to claim 4, characterized in that, The acquisition logic of the feature-bearing coefficient of each data segment is as follows: Extract the feature-bearing distribution information from the preprocessed structural behavior information, specifically including the number of unique main peaks identified in each data segment during the initial feature indexing operation, the total peak area of all main peaks, and the average number of occurrences of all main peaks in other data segments, and calibrate them respectively as , and , represents the number of unique main peaks identified in the -th data segment during the initial feature indexing operation, represents the total peak area of all main peaks in the -th data segment during the initial feature indexing operation, represents the average number of occurrences of all main peaks in other data segments in the -th data segment during the initial feature indexing operation, , is a positive integer; Calculate the sum of the total peak areas of all main peaks in all data segments during the initial feature indexing operation , according to the formula: ; Calculate the feature-bearing coefficient of each data segment. The specific calculation formula is as follows: ; In the formula, is the characteristic bearing coefficient of the th data segment.

6. The data optimization storage method of a gas chromatograph according to claim 5, characterized in that Structure coupling coefficients based on the generated individual data segments and feature bearing coefficients , generate the key evaluation index of each data segment through weighted summation, and the specific calculation formula is as follows: ; Wherein, is the key evaluation index of the th data segment, and are respectively the structure coupling coefficients of each data segment reciprocal and the characteristic bearing coefficient non-zero weight coefficients, and .

7. A method for optimizing the storage of gas chromatograph data according to claim 6, characterized in that, Determine the pre-set threshold range of key evaluation indices , and after determination, compare with the key evaluation indices of each generated data segment , evaluate the key degree of each data segment in the overall data structure according to the comparison result, and classify each data segment according to the evaluation result. The specific comparison analysis and classification are as follows: If , the criticality level of this data segment in the overall data structure is a low criticality level, and this data segment is divided into a low-critical data segment; If , the criticality level of this data segment in the overall data structure is medium criticality, and this data segment is divided into a medium critical data segment; If , the criticality level of this data segment in the overall data structure is high criticality, and this data segment is divided into a high-critical data segment.

8. A method for optimizing the storage of gas chromatograph data according to claim 7, characterized in that, Dynamically adjust the indexing accuracy of each data segment according to the classification results. Specifically: For high-key data segments, set the indexing accuracy to the highest level, and index all structural fields, including all indexing contents of peak positions, structural jump fields, and peak group boundaries; For medium-key data segments, set the indexing accuracy to the middle level, and the only indexed fields include the peak position index and the structural field number; For low-key data segments, set the indexing accuracy to the lowest level, and the only indexed fields include the start position and end position of the data segment.

9. A method for optimizing the storage of gas chromatograph data according to claim 8, characterized in that, Generate a corresponding indexing bitmap based on the indexing accuracy of each data segment after dynamic regulation. Specifically: according to the indexing accuracy level of each data segment, select the corresponding field set for constructing the indexing bitmap; when the indexing accuracy is the highest level, the indexing bitmap includes a peak position index field, a peak group boundary field, a structure field number field, and a structure jump position field; when the indexing accuracy is the intermediate level, the indexing bitmap includes a peak position index field and a structure field number field; when the indexing accuracy is the lowest level, the indexing bitmap only includes a data segment start position field and an end position field.

Citation Information

Patent Citations

  • Test script generation method and system based on BS architecture

    CN115408258A

  • Gas chromatograph data optimization storage method and system

    CN117785818A

  • Chromatographic peak detection method and device, computer equipment, storage medium and computer program product

    CN119413939A

  • Power grid multi-agent large model safety evaluation index calculation method

    CN120046718A