A method, device and medium for chromatographic overlapping peak resolution based on pattern recognition

By training a chromatographic overlapping peak resolution model using a pattern recognition-based method, the problems of unclear peak shoulder origins and ambiguous tailing fragment boundaries were solved, achieving accurate peak shape reconstruction and peak area calculation.

CN122283033APending Publication Date: 2026-06-26QINGDAO GUOKE QUALITY DETECTION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QINGDAO GUOKE QUALITY DETECTION CO LTD
Filing Date
2026-05-18
Publication Date
2026-06-26

Smart Images

  • Figure CN122283033A_ABST
    Figure CN122283033A_ABST
Patent Text Reader

Abstract

This invention discloses a method, device, and medium for analyzing overlapping chromatographic peaks based on pattern recognition, relating to the field of pattern recognition technology. The method includes: collecting clear single-peak segments, slightly overlapping peak segments, and verified overlapping peak segments to form chromatographic peak segment data; performing single-peak reference normalization on clear single-peak segments to obtain a single-peak reference peak shape set; combining the single-peak reference peak shape set with annotations of shoulder origin relationships and sub-peak structural parameters to form overlapping peak training samples and train a peak shape pattern recognition model; inputting the overlapping peak segments to be analyzed into the peak shape pattern recognition model to obtain shoulder origin relationships and sub-peak structural parameters, classifying rising-side candidate segments, tailing-side candidate segments, and interfering shoulder segments; forming tailing-influence bands and potential sub-peak analysis records, generating chromatographic overlapping peak analysis results. This invention distinguishes shoulder origin relationships, separates tailing segments from true sub-peak segments, and reduces the interference of boundary contamination on potential sub-peak shape reconstruction and potential sub-peak area calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pattern recognition technology, and in particular to a method, apparatus and medium for analyzing overlapping chromatographic peaks based on pattern recognition. Background Technology

[0002] Chromatographic analysis is commonly used in practical testing scenarios such as drug impurity inspection, food additive detection, environmental pollutant monitoring, and chemical process sample analysis. After different components in a sample are separated by a chromatographic column, the detector outputs a response intensity curve based on retention time. Analysts typically judge the component separation based on peak position, retention time window, peak area, peak shoulder morphology, and baseline changes. When the sample composition is complex, adjacent components have similar retention times, the main peak has obvious tailing, or the detection baseline has slight fluctuations, overlapping peak regions with two or more component peaks may easily appear in the chromatogram. For overlapping peak regions, conventional processing methods usually combine baseline preparation, peak segmentation, reference peak comparison, curve fitting, and manual verification to analyze the morphology and area of ​​each component peak.

[0003] However, existing methods still have two limitations: First, the peak shoulder sampling section in overlapping peaks may originate from the rising side of the subsequent peak group, the trailing side of the previous peak group, baseline disturbances, or noise disturbances. If the peak shoulder origin relationship is not stably distinguished based solely on the peak top, valley, and local curve morphology, it is difficult to distinguish the origin relationship. Second, when the trailing region of the previous peak group overlaps with the rising region of the subsequent peak group, the boundaries of the trailing segment, the real sub-peak segment, and the disturbance segment are not easily defined, affecting peak shape reconstruction and peak area calculation. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a pattern recognition-based method for analyzing overlapping chromatographic peaks, which solves the problems of unclear peak shoulder origin relationships and ambiguous boundaries between tailing fragments and true sub-peak fragments in existing methods for analyzing overlapping chromatographic peaks.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a method for analyzing overlapping chromatographic peaks based on pattern recognition. The method includes: collecting clear single-peak segments, slightly overlapping peak segments, and verified overlapping peak segments to form chromatographic peak segment data; performing single-peak reference normalization on the clear single-peak segments to obtain a set of single-peak reference peak shapes; and combining the set of single-peak reference peak shapes, slightly overlapping peak segments, and verified overlapping peak segments to label the shoulder origin relationships and sub-peak structural parameters, forming overlapping peak training samples, and training a peak shape pattern recognition model; inputting the overlapping peak segments to be analyzed into the peak shape pattern recognition model, identifying the shoulder sampling sections in the overlapping peak segments to be analyzed, obtaining the shoulder origin relationships and sub-peak structural parameters of each shoulder sampling section, and then... The sampling segment is divided into rising-side candidate fragments, tailing-side candidate fragments, and interfering peak shoulder fragments. The single-peak reference peak shape type in the sub-peak structure parameters is matched with the single-peak reference peak shape set. The tailing-side morphology of the matching single-peak reference peak shape is extracted and mapped to the overlapping peak segment to be analyzed, forming a tailing influence band. The rising-side candidate fragments, tailing-side candidate fragments, and interfering peak shoulder fragments are segmented according to the tailing influence band to form a potential sub-peak analysis record. Based on the single-peak reference peak shape type and sub-peak structure parameters in the potential sub-peak analysis record, the matching single-peak reference peak shape is selected from the single-peak reference peak shape set, the potential sub-peak shape is reconstructed, and the potential sub-peak peak area is calculated to generate the chromatographic overlapping peak analysis results.

[0007] As a preferred embodiment of the pattern recognition-based chromatographic overlapping peak analysis method of the present invention, the method involves: collecting clear single-peak segments, slightly overlapping peak segments, and verified overlapping peak segments to form chromatographic peak segment data; performing single-peak reference normalization processing on clear single-peak segments to obtain a single-peak reference peak shape set, including: collecting chromatographic peak segments; reading the chromatographic method name, detection batch number, retention time series, and detection response intensity series of the chromatographic peak segments; and recording the peak segment source marker; classifying peak segments that do not overlap with adjacent component peaks as clear single-peak segments; classifying peak segments that form peak shoulder sampling sections with identifiable peak shoulder sources as slightly overlapping peak segments; and classifying peak segments where two or more component peaks overlap and the potential sub-peak attribution can be identified as verified overlapping peak segments; and organizing clear single-peak segments, slightly overlapping peak segments, and verified overlapping peak segments according to the chromatographic method name, detection batch number, retention time series, detection response intensity series, and peak segment source marker. Overlapping peak segments and verified overlapping peak segments are used to form chromatographic peak segment data. For clear single peak segments in the chromatographic peak segment data, baseline straightening, peak apex alignment, peak width normalization, response intensity normalization, and peak shape segmentation are performed to obtain normalized retention time, normalized response intensity, rising-side morphology, peak apex neighborhood, falling-side morphology, and tailing-side morphology. Based on the normalized retention time, normalized response intensity, rising-side morphology, peak apex neighborhood, falling-side morphology, and tailing-side morphology, the single-peak reference peak shape type, normalization start position, normalization end position, left-side normalized half-peak width, and right-side normalized half-peak width are organized. These single-peak reference peak shape types, normalization start position, normalization end position, left-side normalized half-peak width, right-side normalized half-peak width, rising-side morphology, peak apex neighborhood, falling-side morphology, and tailing-side morphology are then associated and organized into a single-peak reference peak shape. All single-peak reference peak shapes are aggregated according to their single-peak reference peak shape type to obtain a single-peak reference peak shape set.

[0008] As a preferred embodiment of the chromatographic overlapping peak analysis method based on pattern recognition described in this invention, the step of combining a single-peak reference peak shape set, slightly overlapping peak segments, and verified overlapping peak segments to label the peak shoulder origin relationship and sub-peak structure parameters, forming overlapping peak training samples, and training the peak shape pattern recognition model includes: retrieving slightly overlapping peak segments and verified overlapping peak segments from chromatographic peak segment data, organizing the overlapping peak segment retention time series, performing overlapping peak segment baseline organization on the slightly overlapping peak segments and verified overlapping peak segments, obtaining the overlapping peak segment baseline organization response intensity, and marking the peak shoulder sampling segment and... The sampling range of the potential sub-peak is defined, and peak segments are formed on both sides of the potential sub-peak within this range. Peak shoulder sampling sections are matched with the rising and trailing peak shapes in the single-peak reference peak shape set to obtain the peak shoulder origin relationship. Peak segment sampling sections on both sides of the potential sub-peak are matched with the single-peak reference peak shapes in the single-peak reference peak shape set to obtain sub-peak structure parameters, including single-peak reference peak shape type, potential sub-peak peak retention time, potential sub-peak peak height, potential sub-peak peak width expansion / contraction, and potential sub-peak trailing direction correction. The overlapping peak segment retention time series and overlapping peak segment baseline adjustment response strength are also analyzed. The peak height, peak shoulder sampling section location, peak shoulder origin relationship, and sub-peak structure parameters are organized into overlapping peak training samples. Peak shape variation data is then organized around the peak shoulder sampling section location in the overlapping peak training samples. This peak shape variation data includes variations on the rising peak side, the falling peak side, peak shoulder variations, valley shallowing variations, and tail extension variations. Using this peak shape variation data as input and the peak shoulder origin relationship and sub-peak structure parameters as supervised annotations, a peak shape pattern recognition model is trained. This model employs a gated recurrent unit structure, including an input layer, a gated recurrent unit temporal feature extraction layer, a peak shoulder origin relationship output layer, and a sub-peak structure parameter layer. The input layer organizes the peak shape change data into a sampling point feature sequence according to the sampling order of the overlapping peak segments and the time series. The gated recurrent unit temporal feature extraction layer extracts the peak shape changes before and after the peak shoulder sampling segment along the sampling point feature sequence. The peak shoulder source relationship output layer outputs the peak shoulder source relationship training output, and the sub-peak structure parameter output layer outputs the sub-peak structure parameter training output. Based on the difference between the peak shoulder source relationship training output, the sub-peak structure parameter training output, and the supervised annotation content, the training difference quantity is obtained. The model parameters of the peak shape pattern recognition model are updated according to the training difference quantity to obtain the trained peak shape pattern recognition model.

[0009] As a preferred embodiment of the pattern recognition-based chromatographic overlapping peak analysis method of the present invention, the step of inputting the overlapping peak segment to be analyzed into the peak shape pattern recognition model and identifying the peak shoulder sampling segment in the overlapping peak segment to be analyzed includes: extracting the overlapping peak region to be analyzed from the chromatogram to be analyzed as the overlapping peak segment to be analyzed, the overlapping peak segment to be analyzed carrying the retention time series and the detection response intensity series of the overlapping peak segment to be analyzed; performing baseline adjustment on the overlapping peak segment to be analyzed to obtain the baseline adjustment response intensity of the overlapping peak segment to be analyzed, and checking the main peaks in the overlapping peak segment to be analyzed in the retention time series and the baseline adjustment response intensity of the overlapping peak segment according to the sampling order from smallest to largest retention time. The peak-side morphology, local bending morphology, and adjacent peak-side morphology are used to identify the sampling segment between the main peak-side morphology and the local bending morphology as the peak shoulder sampling segment to be identified. The peak shape change data to be analyzed is organized around the peak shoulder sampling segment. This peak shape change data includes the time series of the peak segment to be analyzed, the baseline response intensity of the peak segment to be analyzed, the position of the peak shoulder sampling segment to be identified, the changes on the rising peak side, the changes on the falling peak side, the peak shoulder changes, the shallowing of the peak valley, and the tailing extension changes. The peak shape change data to be analyzed is input into the trained peak shape pattern recognition model. The trained peak shape pattern recognition model identifies the peak shoulder sampling segment to be identified, obtaining the peak shoulder origin relationship and sub-peak structure parameters of each peak shoulder sampling segment.

[0010] As a preferred embodiment of the pattern recognition-based chromatographic overlapping peak analysis method of the present invention, the step of dividing the peak shoulder sampling segment into rising-side candidate segments, tailing-side candidate segments, and interfering peak shoulder segments includes: the trained peak shape pattern recognition model outputting four types of source probabilities and sub-peak structure parameters based on the peak shape change data to be analyzed; the four types of source probabilities represent the identification probabilities of the peak shoulder sampling segment to be identified originating from the rising side of a subsequent potential sub-peak, the peak shoulder sampling segment to be identified originating from the tailing side of a previous potential sub-peak, the peak shoulder sampling segment to be identified originating from baseline perturbation, and the peak shoulder sampling segment to be identified originating from noise perturbation; the category with the largest value among the four types of source probabilities is taken as the peak shoulder sampling segment to be identified. The peak-shoulder origin relationship of the sample segment is determined, and the identified peak-shoulder sampling segment is designated as the peak-shoulder sampling segment. When the peak-shoulder origin relationship indicates that the peak-shoulder sampling segment originates from the rising side of the subsequent potential sub-peak, the peak-shoulder sampling segment is recorded as a rising side candidate segment. When the peak-shoulder origin relationship indicates that the peak-shoulder sampling segment originates from the trailing side of the preceding potential sub-peak, the peak-shoulder sampling segment is recorded as a trailing side candidate segment. When the peak-shoulder origin relationship indicates that the peak-shoulder sampling segment originates from baseline disturbance or noise disturbance, the peak-shoulder sampling segment is recorded as an interfering peak-shoulder segment. The rising side candidate segment, trailing side candidate segment, interfering peak-shoulder segment, peak-shoulder origin relationship, and sub-peak structure parameters are compiled into a peak-shoulder identification record to be analyzed.

[0011] As a preferred embodiment of the pattern recognition-based chromatographic overlapping peak analysis method of the present invention, the following steps are included: matching the single-peak reference peak shape type in the sub-peak structure parameters with the single-peak reference peak shape set, extracting the tailing side morphology of the matching single-peak reference peak shape, and mapping the tailing side morphology to the overlapping peak segment to be analyzed to form a tailing influence band. This includes calling up the rising side candidate fragment, tailing side candidate fragment, interfering peak shoulder fragment, peak shoulder source relationship, and sub-peak structure parameters in the identification record of the peak shoulder to be analyzed, and sorting the identification outputs with potential sub-peak peak retention times according to the potential sub-peak peak retention times in the sub-peak structure parameters to form a candidate potential sub-peak sequence; according to the retention time order in the overlapping peak segment to be analyzed, within the candidate potential sub-peak sequence, the potential sub-peak located before the candidate fragment and having a potential sub-peak peak retention time is taken as the previous potential sub-peak, and the potential sub-peak located before the candidate fragment is taken as the previous potential sub-peak. The potential sub-peak with a retention time at the peak of the selected sub-peak is taken as the next potential sub-peak. The single-peak reference peak type of the previous potential sub-peak is matched with the single-peak reference peak type in the single-peak reference peak type set. The single-peak reference peak type with the same type is selected as the matching single-peak reference peak type. The trailing side shape located to the right of the peak is extracted from the matching single-peak reference peak type. The normalized retention time of each sampling point in the trailing side shape is multiplied by the peak width scaling amount of the previous potential sub-peak. The retention time of the peak of the previous potential sub-peak is used as the translation reference to obtain the retention time position of the trailing side shape in the overlapping peak segment to be analyzed. The smaller of the retention time positions after mapping at both ends of the trailing side shape is taken as the starting position of the trailing influence band, and the larger one is taken as the ending position of the trailing influence band. The retention time segment between the starting position and the ending position of the trailing influence band is taken as the trailing influence band.

[0012] As a preferred embodiment of the pattern recognition-based chromatographic overlapping peak resolution method of the present invention, the step of segmenting the rising-side candidate fragments, the tailing-side candidate fragments, and the interfering peak shoulder fragments according to the tailing influence band to form a potential sub-peak resolution record includes: matching the single-peak reference peak shape type of the subsequent potential sub-peak with the single-peak reference peak shape set, retrieving the peak apex neighborhood in the matched single-peak reference peak shape, and mapping the peak apex neighborhood to the overlapping peak segment to be resolved according to the peak apex retention time and peak width scaling of the subsequent potential sub-peak, thereby obtaining the peak apex neighborhood of the subsequent potential sub-peak; when all sampling points of the rising-side candidate fragments are located outside the tailing influence band. When the candidate fragment on the rising side and the peak neighborhood of the next potential sub-peak are adjacent in retention time order and have the same direction of response intensity change, the candidate fragment on the rising side is used as the rising side fragment of the next potential sub-peak, generating the record of the next potential sub-peak; when all sampling points of the candidate fragment on the rising side are located within the trailing influence zone, the candidate fragment on the rising side is merged into the trailing side of the previous potential sub-peak, forming a sampling segment merged into the trailing side of the previous potential sub-peak; when some sampling points of the candidate fragment on the rising side are located within the trailing influence zone and other sampling points are located outside the trailing influence zone, the segmentation is performed with the end position of the trailing influence zone as the segment boundary position, obtaining the record of the next potential sub-peak. The sampling segment is recorded in the sub-peak record, merged into the tail side of the previous potential sub-peak, or excluded from the segment record; when all the sampling points of the tail side candidate segment are located within the tail influence zone, the tail side candidate segment is merged into the tail side of the previous potential sub-peak, forming a sampling segment merged into the tail side of the previous potential sub-peak; when all the sampling points of the tail side candidate segment are located outside the tail influence zone, and the sub-peak structure parameters carry the potential sub-peak peak retention time, potential sub-peak peak height, and potential sub-peak peak width expansion and contraction, the tail side candidate segment is rewritten as the rising peak side candidate segment, and continues to be processed according to the segment processing method of the rising peak side candidate segment; when the tail side candidate segment lacks the potential sub-peak peak retention time .... When considering the retention time, peak height, or peak width scaling of potential sub-peaks, candidate segments on the trailing side are written into the excluded segment record; interfering shoulder segments are written into the excluded segment record. The generated subsequent potential sub-peak record, the sampling segment merged into the trailing side of the previous potential sub-peak, and the excluded segment record are then organized to form a potential sub-peak analysis record. The potential sub-peak analysis record includes the number of potential sub-peaks, the single-peak reference peak shape type of each potential sub-peak, the retention time of the potential sub-peak peak, the peak height of the potential sub-peak, the peak width scaling of the potential sub-peak, the tailing direction correction of the potential sub-peak, the shoulder source relationship, the tailing influence zone, the merged trailing side segment, and the excluded segment record.

[0013] As a preferred embodiment of the pattern recognition-based chromatographic overlapping peak analysis method of the present invention, the step of selecting a matching single-peak reference peak shape from the single-peak reference peak shape set, reconstructing the potential sub-peak shape and calculating the peak area of ​​the potential sub-peak to generate chromatographic overlapping peak analysis results includes retrieving the single-peak reference peak shape type, potential sub-peak retention time, potential sub-peak height, potential sub-peak peak width scaling, potential sub-peak tailing direction correction, peak shoulder origin relationship, tailing influence band, and records of merged and excluded tailing side segments from the potential sub-peak analysis records; according to each potential... In the sub-peak reference peak type, select a single-peak reference peak of the same type from the single-peak reference peak type set as the matching single-peak reference peak type. The matching single-peak reference peak type carries the normalized start position, normalized end position, left half-peak width, right half-peak width, rising peak side shape, peak peak neighborhood, falling peak side shape, and tail side shape. The peak width expansion and contraction of the potential sub-peak is used as the width of the matching single-peak reference peak type within the overlapping peak segment to be analyzed, based on the difference between the retention time position and the retention time of the potential sub-peak peak, and the aforementioned difference and the potential sub-peak peak width expansion and contraction. The ratio between the peak width scaling factors of sub-peaks yields the normalized retention time. Based on the left-right relationship of the normalized retention time relative to the peak position, the left and right halves of the peak width in the matching single-peak reference shape are respectively called. The extent of expansion on both sides of the peak is corrected using the potential sub-peak tailing direction correction, resulting in the contour broadening. The normalized retention time and contour broadening are processed using a natural exponential decay method to obtain the single-peak reference shape response. The single-peak reference shape response is then reconstructed by detecting the response scale based on the potential sub-peak height, yielding the reconstruction of each potential sub-peak. Response intensity; based on the normalized start position, normalized end position, and the incorporated tailing fragment of the matching single-peak reference peak shape, organize the peak area integration interval of each potential sub-peak, and calculate the potential sub-peak peak area of ​​each potential sub-peak by integrating the potential sub-peak height, potential sub-peak width scaling, and the response of the single-peak reference peak shape within the peak area integration interval; organize the reconstructed response intensity, potential sub-peak peak area, peak shoulder origin relationship, tailing influence band, incorporated tailing fragment, and excluded fragment records of each potential sub-peak to generate chromatographic overlapping peak resolution results.

[0014] In a second aspect, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein when the computer program is executed by the processor, it implements any step of the pattern recognition-based chromatographic overlapping peak analysis method as described in the first aspect of the present invention.

[0015] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the pattern recognition-based chromatographic overlapping peak analysis method as described in the first aspect of the present invention.

[0016] The beneficial effects of this invention are as follows: By inputting the peak shape change data to be analyzed into the peak shape pattern recognition model, the source division of the peak shoulder sampling segment among the rising side candidate fragment, the tailing side candidate fragment, and the interfering peak shoulder fragment is realized, reducing the influence of the mixing of peak shoulder source relationships on the chromatographic overlapping peak analysis process; by constraining the segment assignment of candidate fragments by the tailing influence band, the boundary division of the tailing side of the previous potential sub-peak, the rising side of the next potential sub-peak, and the excluded fragment record is realized, so that the tailing fragment and the real sub-peak fragment are separated in the potential sub-peak analysis record, reducing the interference of boundary mixing on the reconstruction of potential sub-peak peak shape and the calculation of potential sub-peak peak area. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of a pattern recognition-based chromatographic overlapping peak analysis method.

[0019] Figure 2 This is a flowchart for training a peak-shaped pattern recognition model.

[0020] Figure 3 This is a flowchart of the peak shoulder identification record to be analyzed.

[0021] Figure 4 A flowchart for generating the resolution records of potential sub-peaks and the resolution results of overlapping chromatographic peaks. Detailed Implementation

[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0025] Reference Figures 1-4This is one embodiment of the present invention, which provides a method for analyzing chromatographic overlapping peaks based on pattern recognition, including the following steps: S1. Collect clear single-peak segments, slightly overlapping peak segments, and verified overlapping peak segments to form chromatographic peak segment data. Perform single-peak reference normalization processing on clear single-peak segments to obtain a single-peak reference peak shape set. Combine the single-peak reference peak shape set, slightly overlapping peak segments, and verified overlapping peak segments to label the peak shoulder origin relationship and sub-peak structure parameters to form overlapping peak training samples and train the peak shape pattern recognition model.

[0026] Furthermore, the chromatographic method name, batch number, file source, sample category, retention time series, and detection response intensity series are read from the detection files, single-component standard sample injection files, mixed sample injection files, and manual review files exported from the chromatography workstation. These are then organized according to peak source markings to form a chromatographic peak source record. Specifically, the detection files are operation log files exported from the chromatography workstation during sample injection and detection, recording retention time sampling points and detector response intensities; the single-component standard sample injection files are chromatographic record files generated after the injection of a single-component standard sample, used to provide the morphology of non-overlapping single peaks; the mixed sample injection files are chromatographic record files generated after the co-injection of multiple components, used to provide the overlapping peaks, peak shoulders, and tailing patterns; and the manual review files are files generated by analysts combining standard sample injection results with chromatographic observations. The peak confirmation document generated from the observation results is used to record the source of the peak shoulder and the attribution of components within overlapping peaks; the retention time series records the time position of the chromatographic sampling points during the chromatographic operation; the detection response intensity series records the detector response intensity at each retention time sampling point, serving as the basis for organizing peak shape change data; the document source includes detection documents, single-component standard sample injection documents, mixed sample injection documents, and manual review documents; the sample category includes single-component standard samples and mixed samples; the peak source markers include clear single peaks, slightly overlapping peaks, and verified overlapping peaks; the document source, sample category, and peak source markers in the chromatographic peak source record together serve as the classification basis for clear single peaks, slightly overlapping peaks, and verified overlapping peaks; the chromatographic method name, detection batch number, retention time series, and detection response intensity series serve as the organizing fields for chromatographic peak data.

[0027] Read the file source, sample category, and peak source marker in the chromatographic peak source record. In the chromatographic peaks of a single-component standard sample injection file, peaks that sequentially exhibit the rising side, peak neighborhood, falling side, and tailing side, and do not overlap with adjacent component peaks, are considered clear single-peak peaks. In the chromatographic peaks of a mixed sample injection file, peaks whose main peak edge has locally bent and formed a peak shoulder sampling area, and whose peak shoulder source can be identified by combining the single-component standard sample injection results or manual verification files, are considered slightly overlapping peaks. In the chromatographic peaks of a mixed sample injection file, overlapping peaks of two or more components, and whose potential sub-peak attribution can be identified by combining the single-component standard sample injection results or manual verification files, are considered verified overlapping peaks.

[0028] Clear single-peak segments refer to non-overlapping chromatographic peak segments used to extract the reference peak shape; slightly overlapping peak segments refer to weakly overlapping chromatographic peak segments used to observe the source of the peak shoulder sampling section; verified overlapping peak segments refer to overlapping chromatographic peak segments used to label the sub-peak attribution and sub-peak structural parameters; peak shoulder sampling sections refer to sampling sections located between the main peak side morphology and local bending morphology within overlapping peak segments, used to label the peak shoulder origin relationship; potential sub-peaks refer to sub-peaks within overlapping peak segments that are obscured by the main peak tail, the rising side of adjacent peaks, or local overlapping areas, but whose peak apex position and peak shape attribution can be identified by combining the injection results of single-component standard samples or manual verification documents.

[0029] Clear single-peak segments, slightly overlapping peak segments, and verified overlapping peak segments are organized according to the chromatographic method name, detection batch number, retention time series, detection response intensity series, and peak segment source label to form chromatographic peak segment data.

[0030] Furthermore, for each clear single peak segment in the chromatographic peak data, single-peak reference normalization is performed. Single-peak reference normalization includes baseline straightening, peak top alignment, peak width normalization, response intensity normalization, and peak shape segmentation recording.

[0031] During baseline preparation, the detection response intensities of the starting and ending sampling points of a clear single-peak segment are connected to form a baseline reference curve. The difference between the detection response intensity of a sampling point at the same retention time and the response intensity of the baseline reference curve is used as the baseline preparation response intensity. During peak alignment, the sampling point with the largest baseline preparation response intensity is taken as the peak sampling point, and the difference between the retention time of each sampling point and the retention time of the peak sampling point is used as the centered retention time. If a clear single-peak segment has a plateau-shaped peak sampling point group, the sampling point located in the middle of the plateau-shaped peak sampling point group is taken as the peak sampling point, and the entire plateau-shaped peak sampling point group is taken as the peak neighborhood. When peak width is normalized, half of the baseline-aligned response intensity at the peak sampling point is taken as the half-peak height response intensity. To the left of the peak sampling point, along the retention time series, find two adjacent sampling points that cross the half-peak height response intensity, and perform linear interpolation using the retention time and baseline-aligned response intensity of the two adjacent sampling points to obtain the left half-peak height position. To the right of the peak sampling point, along the retention time series, find two adjacent sampling points that cross the half-peak height response intensity, and perform linear interpolation using the retention time and baseline-aligned response intensity of the two adjacent sampling points to obtain the right half-peak height position. The left half-peak height position is then... The retention time between the peak sampling point and the peak is used as the original value of the left half-peak width, and the retention time between the peak sampling point and the right half-peak height is used as the original value of the right half-peak width. The ratio of the centralized retention time of each sampling point to the left of the peak sampling point to the original value of the left half-peak width is used as the normalized retention time of the sampling point to the left of the peak, and the ratio of the centralized retention time of each sampling point to the original value of the right half-peak width is used as the normalized retention time of the sampling point to the right of the peak. When normalizing the response intensity, the ratio of the baseline-sorted response intensity of each sampling point to the baseline-sorted response intensity of the peak sampling point is used as the normalized response intensity. By recording retention time and normalized response intensity, the peak shape changes of a single peak segment under a unified time scale and a unified response scale are clearly organized. When recording the peak shape segment, the sampling point group with the normalized retention time located to the left of the peak sampling point and the normalized response intensity approaching the peak neighborhood is recorded as the rising peak side shape. The normalized retention time and normalized response intensity change shape of the peak sampling point and the adjacent sampling points before and after the peak sampling point are recorded as the peak neighborhood. The sampling point group with the normalized retention time located to the right of the peak sampling point and the normalized response intensity moving away from the peak neighborhood is recorded as the falling peak side shape. The sampling point group with the longer side of the original value of the left half peak width and the original value of the right half peak width is recorded as the trailing side shape.

[0032] The single-peak reference peak shape type is obtained by combining and organizing the chromatographic method name, component name in the single-component standard sample, the side where the tailing morphology is located, and the normalized sampling order. The normalized retention time of the starting sampling point of the clear single-peak segment is taken as the normalization start position, and the normalized retention time of the ending sampling point of the clear single-peak segment is taken as the normalization end position. The combined length of the original values ​​of the left half peak width and the original values ​​of the right half peak width is taken as the normalized peak width benchmark of the clear single-peak segment. The ratio between the original value of the left half peak width and the normalized peak width benchmark of the clear single-peak segment is recorded as the left half peak width. The ratio between the original value of the right half peak width and the normalized peak width benchmark of the clear single-peak segment is recorded as the right half peak width.

[0033] The single-peak reference peak shape is obtained by associating and organizing the single-peak reference peak shape type, normalization start position, normalization end position, left half peak width, right half peak width, rising peak side shape, peak neighbor, falling peak side shape, and tail side shape with the clear single-peak segments after baseline processing, peak alignment, peak width normalization, response intensity normalization, and peak shape segmentation recording. All single-peak reference peak shapes are then collected according to their single-peak reference peak shape type to obtain the single-peak reference peak shape set.

[0034] Furthermore, after the single-peak reference peak shape set is formed, slightly overlapping peak segments and verified overlapping peak segments are retrieved from the chromatographic peak segment data. Both slightly overlapping peak segments and verified overlapping peak segments carry retention time series and detection response intensity series. The retention time series in the slightly overlapping peak segments and verified overlapping peak segments are organized into overlapping peak segment retention time series, and the detection response intensity series in the slightly overlapping peak segments and verified overlapping peak segments are organized into overlapping peak segment detection response intensity series.

[0035] In the time series and response intensity series of overlapping peak segments, the morphology of the main peak side, local bending morphology, and adjacent peak side morphology within the overlapping peak segments are examined in ascending order of the time series. The sampling segment located between the main peak side morphology and the local bending morphology is marked as the peak shoulder sampling segment. For verified overlapping peak segments, the sampling range of potential sub-peaks within the overlapping peak segment is marked using the component attribution location recorded in the single-component standard sample injection results or manual verification documents. For slightly overlapping peak segments, the sampling range of potential sub-peaks adjacent to the peak shoulder sampling segment is marked using the peak shoulder source recorded in the single-component standard sample injection results or manual verification documents.

[0036] For slightly overlapping peak segments and verified overlapping peak segments, baseline simplification of the overlapping peak segments is performed. The detection response intensities at the start and end sampling points of the overlapping peak segments are connected to form a baseline reference curve for the overlapping peak segments. The difference between the detection response intensity at the sampling point at the same retention time and the response intensity of the baseline reference curve at the same retention time position is used as the baseline simplification response intensity of the overlapping peak segments. The baseline simplification response intensity of the overlapping peak segments is used to mark the peak apex sampling point, left half-peak height position, right half-peak height position, and potential sub-peak height.

[0037] Within the sampling range of the potential sub-peak, the sampling point with the largest baseline recovery response intensity of the overlapping peak segment is marked as the peak sampling point of the potential sub-peak, and half of the baseline recovery response intensity of the overlapping peak segment at the peak sampling point of the potential sub-peak is taken as the half-peak height response intensity. To the left of the peak sampling point of the potential sub-peak, along the retention time series, two adjacent sampling points that cross the half-peak height response intensity are found forward from the peak sampling point of the potential sub-peak. The left half-peak height position is obtained by linear interpolation using the retention time and the baseline recovery response intensity of the overlapping peak segment of the two adjacent sampling points. To the right of the peak sampling point of the potential sub-peak, along the retention time series... The time series is used to find two adjacent sampling points that cross the half-peak height response intensity from the sampling point at the peak of the potential sub-peak. The response intensity is linearly interpolated by the retention time of the two adjacent sampling points and the baseline of the overlapping peak segment to obtain the position of the right half-peak height. The sampling point group between the left half-peak height position and the sampling point at the peak of the potential sub-peak is organized into the left peak shape segment of the potential sub-peak, and the sampling point group between the sampling point at the peak of the potential sub-peak and the right half-peak height position is organized into the right peak shape segment of the potential sub-peak. The left peak shape segment and the right peak shape segment of the potential sub-peak are merged into the two peak shape segments of the potential sub-peak.

[0038] When labeling the origin relationship of peak shoulders, peak shoulder sampling segments are matched with the rising and trailing peak shapes in the single-peak reference peak shape set. During peak shape matching, peak shoulder sampling segments are normalized and aligned according to the segment endpoints. The direction of change of normalized response intensity, the slope change of adjacent sampling points, and the continuity of response intensity at the segment endpoints are compared between the peak shoulder sampling segments and the rising and trailing peak shapes. The shape with the same direction of change and the closest continuity of response intensity at the segment endpoints is taken as the matching shape. When peak shoulder sampling segments are connected to the rising peak shape of a subsequent potential sub-peak in ascending order of the retained time series, the peak shoulder origin relationship is labeled as the peak shoulder sampling segment originating from the subsequent potential sub-peak. The peak rising side; when the peak shoulder sampling section continues the tail side morphology of the previous potential sub-peak in the order of the retention time series from small to large, the peak shoulder source relationship is marked as the peak shoulder sampling section originates from the tail side of the previous potential sub-peak; when the peak shoulder sampling section extends with the baseline fluctuation in the overlapping peak section and does not connect to the rising side morphology or the tail side morphology, the peak shoulder source relationship is marked as the peak shoulder sampling section originates from the baseline disturbance; when the peak shoulder sampling section only shows a local response intensity transition between adjacent sampling points, and the sampling point group before and after the local response intensity transition does not connect to the rising side morphology, the tail side morphology, or the baseline fluctuation morphology, the peak shoulder source relationship is marked as the peak shoulder sampling section originates from the noise disturbance.

[0039] When labeling the sub-peak structure parameters, the peak shape segments on both sides of the potential sub-peak are matched with the single-peak reference peak shapes in the single-peak reference peak shape set to obtain the single-peak reference peak shape type. During single-peak reference peak shape matching, the peak shape segments on both sides of the potential sub-peak are normalized and aligned according to the sampling points at the peak top of the potential sub-peak. The direction of change of normalized response intensity, the slope change of adjacent sampling points, and the half-peak width ratio on both sides of the peak top are compared between the peak shape segments on both sides of the potential sub-peak and the single-peak reference peak shape. The single-peak reference peak shape with the consistent direction of change and the smallest difference in the half-peak width ratio on both sides of the peak top is selected to obtain the single-peak reference peak shape type. The retention time of the sampling points at the peak top of the potential sub-peak is taken as the retention time of the potential sub-peak peak. The baseline adjustment response intensity of the overlapping peak segments at the sampling points at the peak top of the potential sub-peak is taken as the peak height of the potential sub-peak. The retention time length between the left half-peak height position and the right half-peak height position is taken as the retention time. The original peak width of the potential sub-peak is used as the base value. The left and right half-peak widths of the matching single-peak reference shape are merged to form the normalized peak width of the matching single-peak reference shape. The peak width scaling of the potential sub-peak is obtained by the ratio between the original peak width of the potential sub-peak and the normalized peak width of the matching single-peak reference shape. The peak width scaling of the potential sub-peak represents the expansion width of the potential sub-peak on the retention time scale. The retention time length between the left half-peak height position and the peak top sampling point of the potential sub-peak is used as the left half-peak width of the potential sub-peak, and the retention time length between the peak top sampling point and the right half-peak height position of the potential sub-peak is used as the right half-peak width of the potential sub-peak. The tailing direction correction of the potential sub-peak is obtained by the ratio between the difference between the right half-peak width and the left half-peak width of the potential sub-peak, and the merged length between the right half-peak width and the left half-peak width of the potential sub-peak.

[0040] Slightly overlapping peak segments, verified overlapping peak segments, time series of overlapping peak segments, baseline response intensity of overlapping peak segments, location of peak shoulder sampling sections, peak shoulder origin relationships, and sub-peak structure parameters are organized into the same training record to form overlapping peak training samples. The time series of overlapping peak segments and the baseline response intensity of overlapping peak segments are used to organize peak shape variation data, the location of peak shoulder sampling sections is used to define the annotation location of peak shoulder origin relationships, and the peak shoulder origin relationships and sub-peak structure parameters serve as supervised annotation content during the training of the peak shape pattern recognition model. Following the same organization method as the peak shape variation data to be analyzed, the peak rising side variation, peak falling side variation, peak shoulder variation, valley shallowing variation, and tail extension variation are calculated around the peak shoulder sampling section location in the overlapping peak training samples, and the calculation results are combined with the time series of overlapping peak segments and the baseline response intensity of overlapping peak segments to form peak shape variation data.

[0041] Furthermore, the peak shape change data in the overlapping peak training samples are used as input, and the peak shoulder origin relationship and sub-peak structure parameters in the overlapping peak training samples are used as supervised annotations to train the peak shape pattern recognition model. Among them, the peak shape change data includes the overlapping peak segment retention time series, the overlapping peak segment baseline adjustment response intensity, the peak shoulder sampling segment position, the change on the rising peak side, the change on the falling peak side, the peak shoulder change, the peak valley shallowing change, and the tail extension change; the sub-peak structure parameters include the single peak reference peak shape type, the potential sub-peak peak retention time, the potential sub-peak peak height, the potential sub-peak peak width expansion and contraction, and the potential sub-peak tail direction correction.

[0042] The peak shape pattern recognition model adopts a gated recurrent unit structure, including an input layer, a gated recurrent unit temporal feature extraction layer, a peak shoulder origin relationship output layer, and a sub-peak structure parameter output layer. The input layer organizes the peak shape change data into a sampling point feature sequence according to the sampling order of the overlapping peak segments' time series, and marks the start, end, and intra-segment sampling points of the peak shoulder sampling segment in the sampling point feature sequence based on the peak shoulder sampling segment's location. The gated recurrent unit temporal feature extraction layer extracts the peak-rising side changes, peak-falling side changes, peak shoulder changes, valley-shrinking changes, and tail extension changes before and after the peak shoulder sampling segment within the range defined by the start, end, and intra-segment sampling points and the sampling intervals before and after the peak shoulder sampling segment, along the sampling point feature sequence. The peak shoulder origin relationship output layer outputs the peak shoulder origin relationship training output, and the sub-peak structure parameter output layer outputs the sub-peak structure parameter training output.

[0043] The peak-shoulder origin relationship training output includes four types of output values. These four types represent the peak-shoulder sampling segment originating from the rising side of the next potential sub-peak, the trailing side of the previous potential sub-peak, the baseline perturbation, and the noise perturbation, respectively. The peak-shoulder origin relationships in the supervised annotation content are organized into four-digit category labels, arranged in the following order: rising side of the next potential sub-peak, trailing side of the previous potential sub-peak, baseline perturbation, and noise perturbation. The category position of the peak-shoulder origin relationship is marked with a value of 1, and the other category positions are marked with a value of 0. The natural exponent of each of the four types of output values ​​is taken, and the ratio of the natural exponent of each type of output value to the sum of the natural exponents of the four types of output values ​​is used as the origin probability to obtain the four types of origin probabilities. The negative logarithm of the origin probability at the position of the value 1 in the four-digit category label is taken as the origin relationship training difference. The consistency between the position of the largest source probability among the four types and the position of the value 1 in the four-digit category label is organized into the origin relationship classification error record.

[0044] The single-peak reference peak type in the sub-peak structure parameter training output is used as the classification output content. The single-peak reference peak types in the single-peak reference peak type set are arranged in order of type number to obtain the single-peak reference peak type label position. The position of the single-peak reference peak type in the supervised annotation content is marked with a value of 1, and the positions of other types are marked with a value of 0, thus obtaining the single-peak reference peak type label. The natural exponent of each type output value in the single-peak reference peak type training output is taken, and the ratio of the natural exponent of each type output value to the sum of the natural exponents of all type output values ​​is taken as the single-peak reference peak type probability. The position of the single-peak reference peak type label with the value 1 is taken, and the negative logarithm of the single-peak reference peak type probability at the same position is taken as the peak type training difference. The consistency between the position of the largest single-peak reference peak type probability and the position of the single-peak reference peak type label with the position of the value 1 is compiled into a peak type classification error record.

[0045] The training output of sub-peak structural parameters includes the potential sub-peak retention time, peak height, peak width scaling, and tail direction correction. The absolute difference between the training output and the supervised annotation of the potential sub-peak retention time is recorded as the peak retention time difference; the absolute difference between the training output and the supervised annotation of the potential sub-peak height is recorded as the peak height difference; the absolute difference between the training output and the supervised annotation of the potential sub-peak width scaling is recorded as the peak width scaling difference; and the absolute difference between the training output and the supervised annotation of the potential sub-peak tail direction correction is recorded as the tail direction correction difference. These differences in peak retention time, peak height, peak width scaling, and tail direction correction are then recorded as structural parameter difference records.

[0046] The training variances for source relationship training, peak shape type training, and structural parameter differences are collectively recorded as training variance quantities. Specifically, the source relationship training variance is used to update the model parameters of the peak shoulder source relationship output layer and the gated recurrent unit temporal feature extraction layer; the peak shape type training variance is used to update the type output parameters in the sub-peak structural parameter output layer and the model parameters of the gated recurrent unit temporal feature extraction layer; and the structural parameter variance record is used to update the numerical output parameters in the sub-peak structural parameter output layer and the model parameters of the gated recurrent unit temporal feature extraction layer. The training variance quantity is propagated backward along the connection relationships between the sub-peak structural parameter output layer, the peak shoulder source relationship output layer, and the gated recurrent unit temporal feature extraction layer, and the model parameters are updated according to the training variance quantity.

[0047] During training, the training samples with overlapping peaks are divided into training batches according to the batch number. After each training batch is completed and the model parameters are updated, a set of model parameter records is saved. The training samples with overlapping peaks that have been manually verified are used as verification training records. The verification training records are processed using each set of model parameter records to generate verification source relationship classification error records, verification peak shape type classification error records, and verification structure parameter difference records.

[0048] Arrange the model parameter records in each group from smallest to largest according to the number of inconsistent records in the classification error records based on the source relationship of the verification. If the number of inconsistent records in two or more groups of model parameter records is the same, then arrange them from smallest to largest according to the number of inconsistent records in the classification error records based on the peak shape type. If the number of inconsistent records in two or more groups of model parameter records is the same, then arrange them from smallest to largest according to the total difference in peak retention time of all verification training records in the verification structure parameter difference records. If the total difference in peak retention time is the same, then arrange them from smallest to largest according to the total difference in peak height. If the total difference in peak height is the same, then arrange them from smallest to largest according to the total difference in peak width scaling. If the total difference in peak width scaling is the same, then arrange them from smallest to largest according to the total difference in tail direction correction. Write the model parameter record at the top of the list into the peak shape pattern recognition model to obtain the trained peak shape pattern recognition model.

[0049] After training, the peak shape pattern recognition model receives peak shape change data of the overlapping peak segments to be analyzed, and outputs the peak shoulder source relationship and sub-peak structure parameters of the peak shoulder sampling section in the overlapping peak segments to be analyzed.

[0050] S2. Input the overlapping peak segment to be analyzed into the peak pattern recognition model, identify the peak shoulder sampling segment in the overlapping peak segment to be analyzed, obtain the peak shoulder source relationship and sub-peak structure parameters of each peak shoulder sampling segment, and divide the peak shoulder sampling segment into rising side candidate segment, trailing side candidate segment and interference peak shoulder segment.

[0051] Furthermore, overlapping peak regions to be analyzed are extracted from the chromatogram to be analyzed, and these are designated as overlapping peak segments to be analyzed. Each overlapping peak segment carries a retention time series and a detection response intensity series. The retention time series records the time position of each sampling point within the overlapping peak segment, and the detection response intensity series records the detector response intensity of each sampling point within the overlapping peak segment.

[0052] For overlapping peak segments to be analyzed, baseline simplification is performed. The detection response intensities of the starting and ending sampling points of the overlapping peak segments are connected to form a baseline reference curve for the peak segments to be analyzed. The baseline simplification response intensity of the peak segments to be analyzed is obtained by comparing the detection response intensity at the same retention time sampling point with the response intensity of the baseline reference curve at the same retention time position. The baseline simplification response intensity of the peak segments to be analyzed serves as the basis for identifying the morphology of the main peak side, local bending morphology, adjacent peak side morphology, and tail side morphology.

[0053] Furthermore, in the time series of the peak segment to be analyzed and the baseline of the peak segment to be analyzed, the morphology of the main peak side, the local bending morphology and the adjacent peak side morphology in the overlapping peak segment to be analyzed are checked in the sampling order from smallest to largest retention time; the sampling segment located between the main peak side morphology and the local bending morphology is marked as the peak shoulder sampling segment to be identified.

[0054] The peak shape variation data to be analyzed is organized around the sampling section of the peak shoulder to be identified. The peak shape variation data to be analyzed includes the time series of the peak segment to be analyzed, the baseline response intensity of the peak segment to be analyzed, the location of the sampling section of the peak shoulder to be identified, the changes on the rising side, the changes on the falling side, the changes on the peak shoulder, the changes in the shallowing of the peak valley, and the changes in the tail extension.

[0055] Specifically, the adjacent sampling point group near the side with shorter retention time in the sampling section of the peak shoulder to be identified is designated as the front sampling point group, and the adjacent sampling point group near the side with longer retention time in the sampling section of the peak shoulder to be identified is designated as the rear sampling point group. The change on the rising peak side is obtained by the ratio between the difference in response intensity and the difference in retention time of adjacent sampling points in the front sampling point group; the change on the falling peak side is obtained by the ratio between the difference in response intensity and the difference in retention time of adjacent sampling points in the rear sampling point group; the location where the direction of change of the difference in response intensity of adjacent sampling points in the sampling section of the peak shoulder to be identified turns is obtained by the location where the change in the direction of change of the difference in response intensity of adjacent sampling points in the sampling section of the peak shoulder to be identified turns by the location where the change in peak shoulder to be identified turns by the location where the change in peak shoulder to be identified turns by the difference in response intensity and the difference in retention time of the sampling points on both sides of the location where the change in peak shoulder to be identified turns by ...

[0056] Furthermore, the peak shape change data to be analyzed is input into the trained peak shape pattern recognition model. The input layer of the peak shape pattern recognition model organizes the peak shape change data to be analyzed into a feature sequence of sampling points according to the sampling order of the time series of the peak segments to be analyzed, and marks the start sampling point, end sampling point, and intra-segment sampling point of the peak shoulder sampling segment in the feature sequence of the sampling points to be analyzed according to the position of the peak shoulder sampling segment to be identified; the gated recurrent unit time series feature extraction layer extracts the peak shape change before and after the peak shoulder sampling segment along the feature sequence of the sampling points to be analyzed, within the range defined by the start sampling point, end sampling point, and intra-segment sampling point of the peak shoulder sampling segment to be identified, and within the sampling intervals before and after it; the peak shoulder source relationship output layer outputs four types of source probabilities; the sub-peak structure parameter output layer outputs the sub-peak structure parameters.

[0057] The four source probabilities represent the identification probabilities that the sampled section of the shoulder to be identified originates from the rising side of the next potential sub-peak, the trailing side of the previous potential sub-peak, baseline perturbation, and noise perturbation. The category with the highest value among the four source probabilities is taken as the shoulder source relationship of the sampled section. When two or more source probabilities are the same, the category with the highest probability is selected as the shoulder source relationship, following the order of rising side of the next potential sub-peak, trailing side of the previous potential sub-peak, baseline perturbation, and noise perturbation. The shoulder source relationship of each sampled section is obtained through the recognition output of the peak pattern recognition model.

[0058] Sub-peak structure parameters include single-peak reference peak shape type, potential sub-peak peak retention time, potential sub-peak peak height, potential sub-peak peak width scaling, and potential sub-peak tail direction correction; sub-peak structure parameters are used to form the tail influence band and potential sub-peak analytical records.

[0059] Furthermore, the identified peak shoulder sampling sections are designated as peak shoulder sampling sections, and these sections are divided according to the peak shoulder origin relationship.

[0060] When the peak shoulder originates from the rising side of a subsequent potential sub-peak, the peak shoulder sampling segment is recorded as a candidate segment for the rising side; when the peak shoulder originates from the trailing side of a previous potential sub-peak, the peak shoulder sampling segment is recorded as a candidate segment for the trailing side; when the peak shoulder originates from baseline disturbance or noise disturbance, the peak shoulder sampling segment is recorded as an interfering peak shoulder segment.

[0061] The candidate segments on the rising peak side, the candidate segments on the trailing peak side, the interfering peak shoulder segments, the peak shoulder origin relationships, and the sub-peak structure parameters are organized into a peak shoulder identification record to be analyzed. The candidate segments on the rising peak side, the candidate segments on the trailing peak side, the interfering peak shoulder segments, the peak shoulder origin relationships, and the sub-peak structure parameters in the peak shoulder identification record to be analyzed are used to form the trailing influence band and generate potential sub-peak analysis records.

[0062] S3. Match the single-peak reference peak type in the sub-peak structure parameters with the single-peak reference peak set, extract the trailing side morphology of the matching single-peak reference peak, and map the trailing side morphology to the overlapping peak segment to be analyzed to form a trailing influence band. Perform segment processing on the rising side candidate segment, trailing side candidate segment and interference peak shoulder segment according to the trailing influence band to form a potential sub-peak analysis record.

[0063] Furthermore, the process involves accessing candidate segments from the rising peak side, candidate segments from the trailing peak side, interfering peak shoulder segments, peak shoulder origin relationships, and sub-peak structure parameters from the record to be analyzed for peak shoulder identification. These sub-peak structure parameters include the single-peak reference peak shape type, the retention time of the potential sub-peak peak, the peak height of the potential sub-peak, the scaling of the potential sub-peak peak width, and the correction amount for the trailing direction of the potential sub-peak. Based on the retention time of the potential sub-peak peak carried by each sub-peak structure parameter in the record to be analyzed for peak shoulder identification, the identification outputs with the retention time of the potential sub-peak peak are sorted to form a candidate potential sub-peak sequence.

[0064] Based on the retention time order in the overlapping peak segments to be analyzed, within the candidate potential sub-peak sequence, the potential sub-peak that precedes the candidate segment and has a retention time at its peak is designated as the preceding potential sub-peak, and the potential sub-peak that follows the candidate segment and has a retention time at its peak is designated as the following potential sub-peak. If no preceding potential sub-peak exists before the candidate segment, no tailing influence band is generated for the preceding potential sub-peak. If the candidate segment is a rising-side candidate segment, it is processed according to the segment processing method for rising-side candidate segments. If the candidate segment is a tailing-side candidate segment, it is written into the excluded segment record; or, if the sub-peak structure parameters carry the potential sub-peak peak retention time, potential sub-peak height, and potential sub-peak width scaling, the tailing-side candidate segment is rewritten as a rising-side candidate segment before processing continues. If the candidate segment is an interfering peak shoulder segment, it is written into the excluded segment record.

[0065] Furthermore, the single-peak reference peak shape type of the previous potential sub-peak is matched with the single-peak reference peak shape types in the single-peak reference peak shape set, and the single-peak reference peak shape with the same type is selected as the matching single-peak reference peak shape. The matching single-peak reference peak shape carries the rising peak side shape, peak neighborhood, falling peak side shape, tail side shape, normalization start position, normalization end position, left half-peak width, and right half-peak width.

[0066] When a potential sub-peak precedes a candidate segment, the trailing side shape to the right of the peak in the matching single-peak reference shape is selected as the reference shape for the trailing effect of the previous potential sub-peak on the candidate segment; the potential sub-peak trailing direction correction is used to adjust the extension degree of the trailing side shape to the right of the peak. When the candidate segment precedes the peak retention time of the previous potential sub-peak, and it is necessary to determine the forward extension effect to the left of the peak of the previous potential sub-peak, the trailing side shape to the left of the peak in the matching single-peak reference shape is selected as the reference shape for the forward extension effect to the left of the peak.

[0067] The normalized retention time of each sampling point in the selected trailing side morphology is multiplied by the peak width scaling of the previous potential sub-peak. Using the retention time at the peak apex of the previous potential sub-peak as the translation reference, the retention time position of the trailing side morphology within the overlapping peak segment to be analyzed is obtained. The smaller of the mapped retention time positions at both ends of the trailing side morphology is taken as the starting position of the trailing influence band, and the larger one is taken as the ending position. The retention time interval between the starting and ending positions of the trailing influence band is defined as the trailing influence band.

[0068] Furthermore, the candidate segments on the rising side are segmented according to the trailing influence band. The single-peak reference peak shape type of the subsequent potential sub-peak is matched with the single-peak reference peak shape set. The peak apex neighborhood in the matched single-peak reference peak shape is retrieved, and the peak apex neighborhood is mapped to the overlapping peak segment to be parsed according to the peak retention time and peak width scaling amount of the subsequent potential sub-peak, thus obtaining the peak apex neighborhood of the subsequent potential sub-peak. When all sampling points of the rising side candidate segments are located outside the trailing influence band of the previous potential sub-peak, and the rising side candidate segments can be sequentially connected to the peak apex neighborhood of the subsequent potential sub-peak along the retention time, the rising side candidate segments are used as the rising side segments of the subsequent potential sub-peak, generating the subsequent potential sub-peak record.

[0069] Being able to access the peak neighborhood of the next potential sub-peak in the order of retention time means that the end position of the candidate segment on the rising side is adjacent to the start position of the peak neighborhood of the next potential sub-peak in the order of retention time, and the direction of change of response intensity at the end of the candidate segment on the rising side is consistent with the direction of change of response intensity at the start of the peak neighborhood.

[0070] When all sampling points of the candidate fragments on the rising side are located within the trailing influence band of the previous potential sub-peak, the candidate fragments on the rising side are merged into the trailing side of the previous potential sub-peak, and no record of the next potential sub-peak is generated.

[0071] When some sampling points of a candidate segment on the rising side are located within the trailing influence band of the previous potential sub-peak, and other sampling points are located outside the trailing influence band, the end position of the trailing influence band is taken as the segment boundary position of the candidate segment on the rising side. Sampling points located within the trailing influence band are incorporated into the trailing side of the previous potential sub-peak, while sampling points located outside the trailing influence band continue to participate in segment processing as rising side segments of the next potential sub-peak. If sampling points located outside the trailing influence band can be connected to the peak neighborhood of the next potential sub-peak in the order of retention time, a record of the next potential sub-peak is generated; if sampling points located outside the trailing influence band cannot be connected to the peak neighborhood of the next potential sub-peak, they are written into the excluded segment record.

[0072] Furthermore, the candidate segments on the trailing side are segmented according to the trailing influence band. When all the sampling points of the candidate segments on the trailing side are located within the trailing influence band of the previous potential sub-peak, the candidate segments on the trailing side are merged into the trailing side of the previous potential sub-peak.

[0073] When all sampling points of the trailing-side candidate segment are located outside the trailing influence zone of the previous potential sub-peak, and the sub-peak structure parameters carry the retention time of the potential sub-peak peak, the potential sub-peak height, and the potential sub-peak width expansion and contraction, the trailing-side candidate segment is rewritten as the rising-side candidate segment, and the next potential sub-peak record is generated or written into the excluded segment record according to the segment processing method of the rising-side candidate segment.

[0074] When all sampling points of the candidate segment on the trailing side are located outside the trailing influence zone of the previous potential sub-peak, but the sub-peak structure parameters lack the potential sub-peak retention time, potential sub-peak height, or potential sub-peak width scaling, the candidate segment on the trailing side will be written into the excluded segment record.

[0075] When some sampling points of the trailing side candidate segment are located within the trailing influence band of the previous potential sub-peak, and other sampling points are located outside the trailing influence band, the sampling points located within the trailing influence band are merged into the trailing side of the previous potential sub-peak, and the sampling points located outside the trailing influence band are rewritten as rising peak candidate segments, and the processing continues according to the segment processing method of rising peak candidate segments.

[0076] Furthermore, interfering peak shoulder segments are written into the exclusion segment record. The exclusion segment record includes peak shoulder sampling segments originating from baseline perturbations or noise perturbations, which are not involved in the potential sub-peak shape reconstruction and potential sub-peak area calculation.

[0077] The generated potential sub-peak records, the sampling segments merged into the tail side of the previous potential sub-peak, and the excluded segment records are organized to form potential sub-peak analysis records. The potential sub-peak analysis records include the number of potential sub-peaks, the single-peak reference peak shape type for each potential sub-peak, the retention time of the potential sub-peak peak top, the potential sub-peak peak height, the potential sub-peak peak width scaling, the potential sub-peak tail direction correction, the peak shoulder origin relationship, the tail influence zone, and records of merged tail side segments and excluded segments.

[0078] The potential sub-peak analysis record is used for potential sub-peak shape reconstruction and potential sub-peak area calculation; among them, the single-peak reference shape type is used to select a matching single-peak reference shape from the single-peak reference shape set, the potential sub-peak peak retention time, potential sub-peak peak height, potential sub-peak peak width scaling amount and potential sub-peak tail direction correction amount are used to reconstruct the potential sub-peak shape, and the excluded segment record is used to exclude peak shoulder sampling sections caused by baseline disturbance and noise disturbance.

[0079] S4. Based on the single-peak reference peak shape type and sub-peak structure parameters in the potential sub-peak analysis record, select matching single-peak reference peak shapes from the single-peak reference peak shape set, reconstruct the potential sub-peak peak shape, calculate the potential sub-peak peak area, and generate chromatographic overlapping peak analysis results.

[0080] Furthermore, retrieve the number of potential sub-peaks, the single-peak reference peak type of each potential sub-peak, the retention time of the peak top of the potential sub-peak, the peak height of the potential sub-peak, the peak width expansion and contraction amount of the potential sub-peak, the tail direction correction amount of the potential sub-peak, the source relationship of the peak shoulder, the tail influence zone, and the records of the segments incorporated into the tail side and the segments excluded from the potential sub-peak analysis records.

[0081] Based on the single-peak reference peak shape type of each potential sub-peak, a single-peak reference peak shape of the same type is selected from the single-peak reference peak shape set as the matching single-peak reference peak shape. The matching single-peak reference peak shape carries the normalized start position, normalized end position, left half-peak width, right half-peak width, rising peak side shape, peak apex neighborhood, falling peak side shape, and tail side shape. Peak shoulder sampling segments in the excluded fragment records do not participate in the potential sub-peak shape reconstruction and potential sub-peak peak area calculation; they are incorporated into the tail side segments as the extension of the tail side of the previous potential sub-peak and participate in the reconstruction interval arrangement of the previous potential sub-peak.

[0082] Furthermore, to ensure that the matching single-peak reference peak shape expands according to the retention time and width scaling of the potential sub-peaks, the width scaling of the potential sub-peaks is used as the width of the matching single-peak reference peak shape within the overlapping peak segment to be analyzed, based on the difference between the retention time position and the retention time of the potential sub-peak peak, and the ratio between the difference and the width scaling of the potential sub-peaks, the normalized retention time centered on the peak of the potential sub-peak and measured by the width scaling is obtained, expressed as: ; in, For the first A potential sub-peak during retention time The time of unified retention at the location; Number the potential sub-peaks; For retention period; For the first The retention time of the peak of each potential sub-peak; For the first The peak width scaling of the potential sub-peak represents the peak width scaling of the th potential sub-peak. The expansion width of each potential sub-peak on the retention time scale is used to map the normalized retention time of the single-peak reference peak shape to the actual retention time position of the overlapping peak segment to be resolved.

[0083] Furthermore, the potential sub-peak tail direction correction is split into positive and negative components according to the numerical sign. The positive component retains the right tail extension information, and the negative component retains the left tail extension information, represented as: ; ; in, For the first Positive component of the tailing direction correction of each potential sub-peak; For the first The tailing direction correction of each potential sub-peak; For the first The absolute value of the tailing direction correction for each potential sub-peak; For the first The reverse component of the tailing direction correction of each potential sub-peak.

[0084] When the trailing direction correction is greater than zero, the positive component of the trailing direction correction retains the extension strength of the right trailing side, and the negative component of the trailing direction correction is zero; when the trailing direction correction is less than zero, the negative component of the trailing direction correction retains the extension strength of the left trailing side, and the positive component of the trailing direction correction is zero; when the trailing direction correction is equal to zero, both the positive and negative components of the trailing direction correction are zero.

[0085] Furthermore, based on the left-right relationship between the normalized retention time and the peak position, the left and right half-peak widths of the matched single-peak reference peaks are respectively called, and combined with the positive and negative components of the tailing direction correction, the contour broadening amount of the potential sub-peak at the normalized retention time position is calculated, expressed as: ; in, For the first The profile widening amount of each potential sub-peak under the normalized retention time and tailing direction correction; For the first The left side of the reference peak shape for matching a potential sub-peak is reduced to half the peak width; For the first The right side of the reference peak shape of the single peak matched by the potential sub-peak is returned to half the peak width.

[0086] When the normalized retention time is less than zero, the profile widening amount follows the width reference on the left side of the peak and is superimposed with the left trailing extension; when the normalized retention time is not less than zero, the profile widening amount follows the width reference on the right side of the peak and is superimposed with the right trailing extension.

[0087] Furthermore, based on the distance of the normalized retention time from the peak position and the adjustment of the peak shape expansion degree on both sides of the peak by the contour broadening amount, the natural exponential decay form is used to calculate the single-peak reference peak shape response. The natural exponential decay form, as a parameterized reconstruction method for matching the single-peak reference peak shape, utilizes the left half-peak width, right half-peak width, and tail side morphology in the matched single-peak reference peak shape to continuously expand the rising peak side morphology, peak neighborhood, falling peak side morphology, and tail side morphology. The natural exponential decay form maintains a high response in the peak neighborhood and gradually reduces the responses of the rising and falling peak sides as the normalized retention time moves further away from the peak position, expressed as: ; in, For the first The single-peak reference peak shape response of each potential sub-peak under the normalized retention time and tailing direction correction; It is a natural exponential function.

[0088] Furthermore, based on the peak height of the potential sub-peak in the potential sub-peak analysis record, the response scale of the single-peak reference peak shape response is restored to obtain the reconstructed response intensity of the potential sub-peak at the retention time position, expressed as: ; in, For the first The reconstructed response intensity of each potential sub-peak at the retention time; For the first The peak height of each potential sub-peak.

[0089] Furthermore, the peak area integration interval for each potential sub-peak is organized. When the potential sub-peak is not recorded in the analysis record and is not included in the trailing side segment, the normalized start position of the matching single-peak reference peak shape is taken as the start position of the peak area integration, and the normalized end position of the matching single-peak reference peak shape is taken as the end position of the peak area integration.

[0090] When the potential sub-peak analysis record contains a segment incorporated into the tail side, the normalized start position and normalized end position of the segment are obtained by substituting the start and end retention times of the segment into the normalized retention time calculation method. The earlier position between the normalized start position of the matching single-peak reference peak and the normalized start position of the segment incorporated into the tail side is taken as the start position of the peak area integration, and the later position between the normalized end position of the matching single-peak reference peak and the normalized end position of the segment incorporated into the tail side is taken as the end position of the peak area integration.

[0091] Furthermore, the integral variable is substituted into the contour broadening calculation method as the normalized retention time, and the peak area of ​​each potential sub-peak is calculated using the integral results of peak height, peak width scaling, and single-peak reference peak shape response, expressed as: ; in, For the first The peak area of ​​each potential sub-peak; For the first The starting position of the peak area integral of each potential sub-peak; For the first The endpoint of the peak area integration of each potential sub-peak; For integration, represents the normalized retention time in the unimodal reference peak shape; This refers to the contour widening amount obtained by substituting the integral variable into the contour widening calculation method.

[0092] Furthermore, the reconstruction response intensity, peak area, shoulder origin, tailing effect band, and records of incorporated and excluded segments for each potential sub-peak are compiled to generate chromatographic overlapping peak resolution results. These results include the single-peak reference peak shape type, peak retention time, peak height, peak width scaling, tailing direction correction, peak area, tailing effect band, and excluded segment records for each potential sub-peak.

[0093] This embodiment also provides a computer device applicable to the pattern recognition-based chromatographic overlapping peak analysis method, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the pattern recognition-based chromatographic overlapping peak analysis method proposed in the above embodiment.

[0094] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0095] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the pattern recognition-based chromatographic overlapping peak resolution method proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0096] In summary, this invention: by inputting the peak shape change data to be analyzed into a peak shape pattern recognition model, it achieves source segmentation of the peak shoulder sampling segment among the rising-side candidate fragment, the tailing-side candidate fragment, and the interfering peak shoulder fragment, reducing the impact of mixed peak shoulder source relationships on the chromatographic overlapping peak analysis process; by constraining the segment assignment of candidate fragments by the tailing influence band, it achieves boundary segmentation of the tailing side of the previous potential sub-peak, the rising-side of the next potential sub-peak, and the excluded fragment records, so that the tailing fragment and the real sub-peak fragment are separated in the potential sub-peak analysis record, reducing the interference of boundary mixing on the reconstruction of potential sub-peak peak shape and the calculation of potential sub-peak peak area.

[0097] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for analyzing overlapping chromatographic peaks based on pattern recognition, characterized in that, include: Clear single-peak segments, slightly overlapping peak segments, and verified overlapping peak segments are collected to form chromatographic peak segment data. Single-peak reference normalization is performed on clear single-peak segments to obtain a set of single-peak reference peak shapes. The source relationship of peak shoulders and sub-peak structural parameters are labeled by combining the set of single-peak reference peak shapes, slightly overlapping peak segments, and verified overlapping peak segments to form overlapping peak training samples and train the peak shape pattern recognition model. The overlapping peak segments to be analyzed are input into the peak shape pattern recognition model to identify the peak shoulder sampling segments in the overlapping peak segments to be analyzed, obtain the peak shoulder source relationship and sub-peak structure parameters of each peak shoulder sampling segment, and divide the peak shoulder sampling segments into rising side candidate segments, trailing side candidate segments and interference peak shoulder segments. Match the single-peak reference peak type in the sub-peak structure parameters with the single-peak reference peak set, extract the trailing side morphology of the matching single-peak reference peak, and map the trailing side morphology to the overlapping peak segment to be analyzed to form a trailing influence band. According to the trailing influence band, segment processing is performed on the rising side candidate segment, the trailing side candidate segment, and the interference peak shoulder segment to form a potential sub-peak analysis record. Based on the single-peak reference peak shape type and sub-peak structure parameters in the potential sub-peak analysis record, a matching single-peak reference peak shape is selected from the single-peak reference peak shape set, the potential sub-peak shape is reconstructed, and the potential sub-peak area is calculated to generate the chromatographic overlapping peak analysis results.

2. The chromatographic overlapping peak analysis method based on pattern recognition as described in claim 1, characterized in that, The process involves collecting clear single-peak segments, slightly overlapping peak segments, and verified overlapping peak segments to form chromatographic peak segment data. Clear single-peak segments are subjected to single-peak reference normalization processing to obtain a single-peak reference peak shape set. Chromatographic peak segments are collected, and the chromatographic method name, detection batch number, retention time series, and detection response intensity series of the chromatographic peak segments are read. Peak segment source markers are also recorded. Peak segments that do not overlap with adjacent component peaks are designated as clear single-peak segments; peak segments that form shoulder sampling areas and whose shoulder origins can be identified are designated as slightly overlapping peak segments; and peak segments where two or more component peaks overlap and whose potential sub-peak affiliations can be identified are designated as validated overlapping peak segments. Clear single-peak segments, slightly overlapping peak segments, and validated overlapping peak segments are then organized according to the chromatographic method name, detection batch number, retention time series, detection response intensity series, and peak segment origin markers to form chromatographic peak segment data. Baseline straightening, peak alignment, peak width normalization, response intensity normalization, and peak shape segmentation are performed on clear single peak segments in chromatographic peak data to obtain normalized retention time, normalized response intensity, rising side morphology, peak apex neighborhood, falling side morphology, and tailing side morphology. Based on the normalized retention time, normalized response intensity, peak-rising side morphology, peak-top neighborhood, peak-falling side morphology, and tail-side morphology, the single-peak reference peak shape type, normalization start position, normalization end position, left half-peak width, and right half-peak width are organized. The single-peak reference peak shape type, normalization start position, normalization end position, left half-peak width, right half-peak width, peak-rising side morphology, peak-top neighborhood, peak-falling side morphology, and tail-side morphology are then associated and organized into single-peak reference peak shapes. All single-peak reference peak shapes are then aggregated according to their single-peak reference peak shape type to obtain a single-peak reference peak shape set.

3. The chromatographic overlapping peak analysis method based on pattern recognition as described in claim 1, characterized in that, The process of combining single-peak reference peak shape sets, slightly overlapping peak segments, and verified overlapping peak segments to label the source relationship of peak shoulders and sub-peak structural parameters, forming overlapping peak training samples, and training the peak shape pattern recognition model includes: retrieving slightly overlapping peak segments and verified overlapping peak segments from chromatographic peak segment data, organizing the time series of overlapping peak segments, performing baseline organization on slightly overlapping peak segments and verified overlapping peak segments, obtaining the response intensity of the baseline organization of overlapping peak segments, and marking the peak shoulder sampling area and the sampling range to which the potential sub-peak belongs based on the response intensity of the baseline organization of overlapping peak segments, forming peak shape segments on both sides of the potential sub-peak within the sampling range to which the potential sub-peak belongs; Peak shoulder sampling sections are matched with the rising and trailing peak shapes in the single-peak reference peak shape set to obtain the peak shoulder origin relationship; peak shape segments on both sides of potential sub-peaks are matched with single-peak reference peak shapes in the single-peak reference peak shape set to obtain sub-peak structure parameters, including single-peak reference peak shape type, potential sub-peak peak retention time, potential sub-peak peak height, potential sub-peak peak width expansion / contraction, and potential sub-peak trailing direction correction amount; The overlapping peak segments retain time series, overlapping peak segment baseline processing response intensity, peak shoulder sampling segment position, peak shoulder origin relationship and sub-peak structure parameters are processed into overlapping peak training samples. Peak shape change data are then processed around the peak shoulder sampling segment position in the overlapping peak training samples. The peak shape change data includes rising peak side change, falling peak side change, peak shoulder change, peak valley shallowing change and tail extension change. Using peak shape variation data as input, and peak shoulder origin relationship and sub-peak structure parameters as supervised annotations, a peak shape pattern recognition model is trained. The peak shape pattern recognition model adopts a gated recurrent unit structure, including an input layer, a gated recurrent unit temporal feature extraction layer, a peak shoulder origin relationship output layer, and a sub-peak structure parameter output layer. The input layer organizes the peak shape variation data into a sampling point feature sequence according to the sampling order of the time series of overlapping peak segments. The gated recurrent unit temporal feature extraction layer extracts the peak shape variation before and after the peak shoulder sampling segment along the sampling point feature sequence. The peak shoulder origin relationship output layer outputs the peak shoulder origin relationship training output, and the sub-peak structure parameter output layer outputs the sub-peak structure parameter training output. The training difference is obtained based on the differences between the training output of the peak shoulder origin relationship, the training output of the sub-peak structure parameters, and the supervised annotation content. The model parameters of the peak pattern recognition model are then updated according to the training difference to obtain the trained peak pattern recognition model.

4. The chromatographic overlapping peak analysis method based on pattern recognition as described in claim 1, characterized in that, The step of inputting the overlapping peak segment to be analyzed into the peak shape pattern recognition model and identifying the peak shoulder sampling section in the overlapping peak segment to be analyzed includes: extracting the overlapping peak region to be analyzed from the chromatogram to be analyzed as the overlapping peak segment to be analyzed; the overlapping peak segment to be analyzed carries the retention time series of the peak segment to be analyzed and the detection response intensity series of the peak segment to be analyzed. The baseline of the overlapping peak segment to be analyzed is processed to obtain the response intensity of the baseline processing of the peak segment to be analyzed. In the time series of the peak segment to be analyzed and the response intensity of the baseline processing of the peak segment to be analyzed, the main peak side morphology, local bending morphology and adjacent peak side morphology in the overlapping peak segment to be analyzed are checked in the sampling order from small to large retention time. The sampling segment located between the main peak side morphology and the local bending morphology is marked as the peak shoulder sampling segment to be identified. The peak shape change data to be analyzed is organized around the sampling section of the peak shoulder to be identified. The peak shape change data to be analyzed includes the time series of the peak segment to be analyzed, the response intensity of the baseline of the peak segment to be analyzed, the location of the sampling section of the peak shoulder to be identified, the change on the rising side, the change on the falling side, the change on the peak shoulder, the change in the shallowing of the peak valley, and the change in the tail extension. The peak shape change data to be analyzed is input into the trained peak shape pattern recognition model. The trained peak shape pattern recognition model identifies the peak shoulder sampling segment to be identified, and obtains the peak shoulder source relationship and sub-peak structure parameters of each peak shoulder sampling segment.

5. The chromatographic overlapping peak analysis method based on pattern recognition as described in claim 4, characterized in that, The process of dividing the peak shoulder sampling section into rising peak side candidate segments, trailing peak side candidate segments, and interfering peak shoulder segments includes the trained peak shape pattern recognition model outputting four types of source probabilities and sub-peak structure parameters based on the peak shape change data to be analyzed. The four types of source probabilities represent the identification probabilities that the sampling segment to be identified originates from the rising side of the next potential sub-peak, the tail side of the previous potential sub-peak, the baseline disturbance, and the noise disturbance, respectively. The category with the largest value among the four types of source probabilities is taken as the peak shoulder source relationship of the sampling segment to be identified, and the sampling segment to be identified that has been identified is taken as the peak shoulder sampling segment. When the peak shoulder origination relationship is that the peak shoulder sampling segment originates from the rising side of the subsequent potential sub-peak, the peak shoulder sampling segment is recorded as a candidate segment of the rising side; when the peak shoulder origination relationship is that the peak shoulder sampling segment originates from the trailing side of the preceding potential sub-peak, the peak shoulder sampling segment is recorded as a candidate segment of the trailing side; when the peak shoulder origination relationship is that the peak shoulder sampling segment originates from baseline disturbance or from noise disturbance, the peak shoulder sampling segment is recorded as an interfering peak shoulder segment. The candidate segments on the rising peak side, the candidate segments on the trailing peak side, the interfering peak shoulder segments, the peak shoulder source relationship, and the sub-peak structure parameters are organized into a peak shoulder identification record to be analyzed.

6. The chromatographic overlapping peak analysis method based on pattern recognition as described in claim 5, characterized in that, The process of matching the single-peak reference peak type in the sub-peak structure parameters with the single-peak reference peak set, extracting the trailing side morphology of the matching single-peak reference peak, and mapping the trailing side morphology to the overlapping peak segment to be analyzed to form a trailing influence band includes calling up the rising peak candidate segment, trailing side candidate segment, interference peak shoulder segment, peak shoulder source relationship and sub-peak structure parameters in the peak shoulder identification record to be analyzed, and sorting the identification output with potential sub-peak peak retention time according to the potential sub-peak peak retention time in the sub-peak structure parameters to form a candidate potential sub-peak sequence; According to the retention time order in the overlapping peak segments to be analyzed, in the candidate potential sub-peak sequence, the potential sub-peak that is located before the candidate segment and has the retention time of the potential sub-peak peak is taken as the previous potential sub-peak, and the potential sub-peak that is located after the candidate segment and has the retention time of the potential sub-peak peak is taken as the next potential sub-peak. Match the single-peak reference peak type of the previous potential sub-peak with the single-peak reference peak type in the single-peak reference peak type set, select the single-peak reference peak type with the same type as the matching single-peak reference peak type, and extract the trailing side shape located on the right side of the peak from the matching single-peak reference peak type. Multiply the normalized retention time of each sampling point in the trailing side morphology by the peak width scaling of the previous potential sub-peak, and use the retention time of the peak top of the previous potential sub-peak as the translation reference to obtain the retention time position of the trailing side morphology within the overlapping peak segment to be analyzed. Take the smaller of the retention time positions mapped at both ends of the trailing side morphology as the starting position of the trailing influence band, and the larger of the two as the ending position of the trailing influence band. The retention time segment between the starting position and the ending position of the trailing influence band is taken as the trailing influence band.

7. The chromatographic overlapping peak analysis method based on pattern recognition as described in claim 6, characterized in that, The step of performing segment processing on the rising side candidate segment, the trailing side candidate segment, and the interfering peak shoulder segment according to the trailing influence band to form a potential sub-peak analysis record includes matching the single-peak reference peak shape type of the next potential sub-peak with the single-peak reference peak shape set, retrieving the peak apex neighborhood in the matched single-peak reference peak shape, and mapping the peak apex neighborhood to the overlapping peak segment to be analyzed according to the peak apex retention time and peak width scaling amount of the next potential sub-peak to obtain the peak apex neighborhood of the next potential sub-peak. When all sampling points of the rising-side candidate segment are located outside the trailing influence zone, and the peak neighbor of the rising-side candidate segment and the next potential sub-peak are adjacent in retention time order and have the same direction of response intensity change, the rising-side candidate segment is used as the rising-side segment of the next potential sub-peak, generating the next potential sub-peak record; when all sampling points of the rising-side candidate segment are located within the trailing influence zone, the rising-side candidate segment is merged into the trailing side of the previous potential sub-peak, forming a sampling segment merged into the trailing side of the previous potential sub-peak; when some sampling points of the rising-side candidate segment are located within the trailing influence zone and other sampling points are located outside the trailing influence zone, the end position of the trailing influence zone is used as the segment boundary position for segmentation processing, resulting in the next potential sub-peak record, the sampling segment merged into the trailing side of the previous potential sub-peak, or the excluded segment record; When all sampling points of the candidate segment on the trailing side are located within the trailing influence zone, the candidate segment on the trailing side is merged into the trailing side of the previous potential sub-peak, forming a sampling segment merged into the trailing side of the previous potential sub-peak; when all sampling points of the candidate segment on the trailing side are located outside the trailing influence zone, and the sub-peak structure parameters carry the retention time of the potential sub-peak peak, the peak height of the potential sub-peak, and the expansion / contraction of the potential sub-peak peak width, the candidate segment on the trailing side is rewritten as a candidate segment on the rising peak side, and the processing continues according to the segment processing method of the candidate segment on the rising peak side; when the candidate segment on the trailing side lacks the retention time of the potential sub-peak peak, the peak height of the potential sub-peak, or the expansion / contraction of the potential sub-peak peak width, the candidate segment on the trailing side is written into the excluded segment record; Interference peak shoulder segments are written into the exclusion segment record, and the generated subsequent potential sub-peak record, the sampling segment merged into the tail side of the previous potential sub-peak, and the exclusion segment record are organized to form a potential sub-peak analysis record. The potential sub-peak analysis record includes the number of potential sub-peaks, the single peak reference peak shape type of each potential sub-peak, the retention time of the potential sub-peak peak top, the potential sub-peak peak height, the potential sub-peak peak width expansion and contraction, the potential sub-peak tail direction correction, the peak shoulder source relationship, the tail influence zone, the merged tail side segment, and the exclusion segment record.

8. The chromatographic overlapping peak analysis method based on pattern recognition as described in claim 1 or 7, characterized in that, The process of selecting a matching single-peak reference peak shape from the single-peak reference peak shape set, reconstructing the potential sub-peak shape and calculating the peak area of ​​the potential sub-peak to generate chromatographic overlapping peak analysis results includes retrieving the single-peak reference peak shape type, potential sub-peak retention time at the peak top, potential sub-peak height, potential sub-peak peak width scaling, potential sub-peak tailing direction correction, peak shoulder origin relationship, tailing influence band, and records of included and excluded tailing side fragments from the potential sub-peak analysis record. According to the single-peak reference peak shape type of each potential sub-peak, select the single-peak reference peak shape of the same type from the single-peak reference peak shape set as the matching single-peak reference peak shape. The matching single-peak reference peak shape carries the normalization start position, normalization end position, left half peak width, right half peak width, rising peak side shape, peak apex neighborhood, falling peak side shape and tail side shape. The potential sub-peak width expansion is used as the retention time scale of the matching single peak reference peak shape within the overlapping peak segment to be analyzed. The normalized retention time is obtained by the difference between the retention time position and the retention time of the potential sub-peak peak top, and the ratio between the aforementioned difference and the potential sub-peak width expansion. Based on the left-right relationship between the normalized retention time and the peak position, the left half-peak width and the right half-peak width in the matching single-peak reference peak shape are called respectively, and the expansion degree on both sides of the peak is corrected by the potential sub-peak tailing direction correction amount to obtain the contour broadening amount. The normalized retention time and profile broadening amount are processed using the natural exponential decay method to obtain the single-peak reference peak shape response. The single-peak reference peak shape response is then reconstructed according to the peak height of the potential sub-peaks to obtain the reconstructed response intensity of each potential sub-peak. Based on the normalized start position, normalized end position, and the merged tail side segment of the matching single-peak reference peak shape, the peak area integration interval of each potential sub-peak is sorted out, and the potential sub-peak peak area of ​​each potential sub-peak is calculated by integrating the potential sub-peak height, potential sub-peak width scaling amount, and single-peak reference peak shape response within the peak area integration interval. The reconstruction response intensity, peak area, shoulder origin, tailing influence band, and the records of incorporated and excluded segments of each potential sub-peak are compiled to generate chromatographic overlapping peak analysis results.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the chromatographic overlapping peak analysis method based on pattern recognition as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the chromatographic overlapping peak analysis method based on pattern recognition as described in any one of claims 1 to 8.