Multi-dimensional holographic data analysis method and system for hazardous waste supervision

Through the multi-dimensional holographic data analysis method, combined with fluorescence spectroscopy and mass spectrometry data, efficient traceability detection of various hazardous wastes in the field is achieved, solving the problem of deviation in the traceability results in the existing technology, and improving the accuracy and reliability of traceability.

CN120044012AActive Publication Date: 2025-05-27SHANDONG QINGKONG ECOLOGICAL ENVIRONMENT IND DEV CO LTD

Patent Information

Application Number
CN202510267007.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-09-25
Filing Date
2025-03-07
Publication Date
2025-05-27
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

It is difficult for the prior art to efficiently trace the source of a variety of hazardous wastes in the wild environment, especially in the presence of organic pollutants, heavy metal pollutants and radioactive substances, which are prone to misjudgment, resulting in a large deviation from the actual situation.

Method used

The multi-dimensional holographic data analysis method is used to detect hazardous waste samples through a three-dimensional fluorescence spectrometer and mass spectrometer, obtain fluorescence spectroscopy data and mass spectrometer, and after pretreatment, peak detection, matching and characteristic vector analysis are used, and similarity score calculation is performed in combination with standard samples to trace the source of hazardous wastes.

Benefits of technology

It realizes efficient traceability detection of wild hazardous substances, improves the accuracy and reliability of traceability results, and can reflect the similarity between samples more comprehensively and objectively, avoiding the one-sidedness of single-dimensional evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120044012A_ABST
    Figure CN120044012A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, and particularly relates to a multi-dimensional holographic data analysis method and system for hazardous waste supervision, and the method comprises the steps: obtaining fluorescence spectrum data and mass spectrum data; preprocessing the acquired data, and carrying out peak detection by utilizing a first-order derivative of peak intensity based on a preprocessed fluorescence intensity matrix; determining the weight of each peak point in combination with the fluorescence spectrum of the standard sample to complete the matching of peak pairs, and calculating the peak matching distance and the distribution difference of the feature vectors; calculating a similarity score of the spectrum by integrating the peak matching distance and the distribution difference of the feature vectors and combining a matching ratio and a weight ratio; calculating weighted cosine similarity and Euclidean distance between the processed mass spectrum data and the mass spectrum data of the standard sample, and completing calculation of the similarity score of the mass spectrum between the samples; and tracing the source of the hazardous wastes according to the similarity score of the spectrum and the similarity score of the mass spectrum. And accurate traceability is realized through qualitative and quantitative analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a multi-dimensional holographic data analysis method and system for hazardous waste supervision. Background Art

[0002] With the acceleration of the industrialization process, the generation amount of hazardous waste is increasing day by day, and the new pollutants contained therein pose a serious threat to the ecological environment and human health. In the field of environmental supervision, due to the great difficulty in supervision and environmental sensitivity in the wild areas, they have become high-incidence areas for illegal dumping and leakage of hazardous waste, posing a serious threat to the ecological environment and human health. In this situation, it is crucial to achieve accurate traceability detection of wild hazardous substances.

[0003] In the wild, there may be various hazardous substances such as organic pollutants, heavy metal pollutants, and radioactive substances at the same time. Most of the existing traceability algorithms only simply compare the similarity indexes of one kind of data, such as tracing only based on the component data of chemical substances, and cannot comprehensively utilize various data information. When dealing with the data of hazardous substances in the complex wild environment, misjudgment is likely to occur, resulting in a large deviation between the traceability result and the actual situation.

[0004] How to develop a multi-dimensional holographic data analysis method for hazardous waste supervision to achieve efficient traceability detection of wild hazardous substances and provide a strong guarantee for wild environmental protection and ecological security is an urgent problem to be solved at present. Summary of the Invention

[0005] In order to achieve efficient traceability of wild hazardous substances, the present invention provides a multi-dimensional holographic data analysis method and system for hazardous waste supervision.

[0006] In the first aspect, the technical solution of the present invention provides a multi-dimensional holographic data analysis method for hazardous waste supervision, including: Detecting hazardous waste samples respectively using a three-dimensional fluorescence spectrometer and a mass spectrometer to obtain fluorescence spectral data and mass spectral data; Preprocessing the obtained fluorescence spectral data to obtain a preprocessed fluorescence intensity matrix, and at the same time performing alignment operation and normalization processing on the mass spectral data; Performing peak detection based on the preprocessed fluorescence intensity matrix using the first derivative of the peak intensity; Combining the fluorescence spectra of standard samples to determine the weight of each peak point to complete the matching of peak pairs, and calculating the peak matching distance and the distribution difference of eigenvectors; Combining the peak matching distance and the distribution difference of eigenvectors, and combining the matching ratio and the weight ratio, calculating the similarity score of the spectra; Calculate the weighted cosine similarity and Euclidean distance between the processed mass spectrometry data and the mass spectrometry data of the standard sample to complete the calculation of the similarity score of the mass spectrometry between samples; Trace the source of the hazardous waste based on the similarity score of the spectrum and the similarity score of the mass spectrometry.

[0007] As a further limitation of the technical solution of the present invention, the step of preprocessing the obtained fluorescence spectrum data to obtain the preprocessed fluorescence intensity matrix includes: Extract and filter the dot matrix data of the fluorescence spectrum data to generate a fluorescence intensity matrix with the excitation wavelength as the row index, the emission wavelength as the column index, and the corresponding fluorescence intensity value as the matrix element; Perform Gaussian smoothing processing on the fluorescence intensity matrix after dot matrix data extraction and filtering to obtain the preprocessed fluorescence intensity matrix.

[0008] As a further limitation of the technical solution of the present invention, the steps of aligning and normalizing the mass spectrometry data include: Convert the mass spectrometry data of different time steps into a unified format; Analyze the stability of the mass spectrometry data of each time step, and select the time step with the number of peaks within the set range and the peak intensity fluctuation range less than the set value as the reference time step; Determine the mass-to-charge ratio range window, and preliminarily match the peaks in the mass spectrometry diagrams of the reference time step and other time steps that fall within this window; By comparing the mass-to-charge ratio differences of multiple matching peaks, take the average value as the global offset, and then perform an overall translation on the mass spectrometry data of other time steps to align the mass-to-charge ratios of the matching peaks; For the mass spectrometry data of each time step, arrange the intensities of all peaks in ascending order, take the intensity value at the middle position as the median intensity of this time step, and divide the original intensity of each peak within this time step by the median intensity of this time step to obtain the normalized mass spectrometry data, where if the number of peaks is even, take the average value of the two middle intensity values as the median intensity.

[0009] As a further limitation of the technical solution of the present invention, the step of performing peak detection based on the preprocessed fluorescence intensity matrix using the first derivative of the peak intensity includes: Take the derivative of the data of the preprocessed fluorescence intensity matrix along the excitation wavelength axis and the emission wavelength axis, and use the characteristic that the first derivative is 0, combined with the judgment of the point sets and Euclidean distances in the two axis directions to determine the data points with key features in the fluorescence intensity matrix, that is, the peak points, and generate a set of peak points.

[0010] As a further limitation of the technical solution of the present invention, the steps of determining the weight of each peak point in combination with the fluorescence spectrum of the standard sample to complete the matching of peak pairs, and calculating the peak matching distance and the distribution difference of feature vectors include: Calculate the Euclidean distance between the detected peak points, and merge the peak points within the distance threshold according to the set distance threshold; Determine the weight of each peak point, calculate the distances of all point pairs for the two groups of peak points and sort them, select the point pairs that meet the distance constraint, and complete the matching of peak pairs; For the selected matching peak pairs, calculate the Euclidean distance between the matching peak pairs, then use the linear assignment algorithm to find the set of optimal matching pairs, and accumulate the minimum cost to obtain the peak matching distance; Select a set number of points around each peak, calculate the feature vector, and then calculate the distribution difference of the feature vector through the Euclidean distance.

[0011] As a further limitation of the technical solution of the present invention, in the step of calculating the similarity score of the spectrum by synthesizing the peak matching distance and the distribution difference of the feature vector, and combining the matching ratio and the weight ratio, the formula is as follows:

[0012] In the formula, is the weight of the peak matching distance, is the weight of the distribution difference of the feature vector, is the distribution difference of the feature vector, and the peak matching distance , is the Euclidean distance between the peak points in the corresponding matching pair, and match coefficient is the matching coefficient; Optimal Match represents the set of optimal matching pairs obtained by the linear assignment algorithm; among them, by normalizing the peak matching distance and normalizing the distribution difference of the feature vector, and performing weighted summation according to the set weight coefficient, a preliminary matching coefficient value is obtained, and the preliminary matching coefficient value is multiplied by the matching ratio to obtain match coefficient; the ratio of the successful matching of the peak points with the peak points of the standard sample under the distance and width constraint conditions is defined as the matching ratio.

[0013] As a further limitation of the technical solution of the present invention, the steps of calculating the weighted cosine similarity and the Euclidean distance between the processed mass spectrometry data and the mass spectrometry data of the standard sample to complete the calculation of the similarity score of the mass spectrometry between samples include: Respectively organize the processed mass spectrometry data and the mass spectrometry data of the standard sample into vector forms, and set the weight of each mass-to-charge ratio position. The weight vector is ; the vector of the processed mass spectrometry data , the vector of the mass spectrometry data of the standard sample is ; Weighted cosine similarity ; Euclidean distance ; Similarity score of mass spectrometry ; In the formula, represents the weight at the th mass-to-charge ratio position, is the weight of the weighted cosine similarity, is the weight of the Euclidean distance, is the maximum value of the Euclidean distance.

[0014] As a further limitation of the technical solution of the present invention, the steps of tracing the source of hazardous waste according to the similarity score of the spectrum and the similarity score of the mass spectrometry include: Calculate the weight distribution between the similarity score of the spectrum and the similarity score of the mass spectrometry, obtain the similarity score of the final hazardous waste and the standard sample, and trace the source of the hazardous waste based on the finally obtained similarity score.

[0015] The fluorescence spectrum data is used for qualitative analysis of the tracing result, and the accuracy of the mass spectrometry data is used to quantitatively calculate the sample similarity, so that the intelligent tracing algorithm can not only quickly respond to the tracing requirement, but also further perform more reasonable quantitative analysis on the tracing result to achieve accurate tracing.

[0016] In the second aspect, the technical solution of the present invention also provides a multi-dimensional holographic data analysis system for hazardous waste supervision, including a detection data acquisition module, a preprocessing module, a spectrum data processing module, a mass spectrometry data processing module, and a tracing processing module; The detection data acquisition module is used to detect the hazardous waste sample by using a three-dimensional fluorescence spectrometer and a mass spectrometer respectively to obtain fluorescence spectrum data and mass spectrometry data; The preprocessing module is used to preprocess the obtained fluorescence spectrum data to obtain a preprocessed fluorescence intensity matrix, and at the same time perform alignment operation and normalization processing on the mass spectrometry data; The spectrum data processing module is used to perform peak detection based on the preprocessed fluorescence intensity matrix by using the first derivative of the peak intensity; determine the weight of each peak point in combination with the fluorescence spectrum of the standard sample to complete the matching of peak pairs, and calculate the peak matching distance and the distribution difference of the eigenvectors; combine the peak matching distance and the distribution difference of the eigenvectors, and combine the matching ratio and the weight ratio to calculate the similarity score of the spectrum; The mass spectrometry data processing module is used to calculate the weighted cosine similarity and the Euclidean distance between the processed mass spectrometry data and the mass spectrometry data of the standard sample, and complete the calculation of the similarity score of the mass spectrometry between samples; The traceability processing module is used to trace the source of hazardous waste according to the similarity scores of spectra and mass spectra.

[0017] As a further limitation of the technical solution of the present invention, the preprocessing module includes a spectral data preprocessing unit, which is used to extract and filter dot matrix data from fluorescence spectral data, and generate a fluorescence intensity matrix with the excitation wavelength as the row index, the emission wavelength as the column index, and the corresponding fluorescence intensity value as the matrix element; perform Gaussian smoothing processing on the fluorescence intensity matrix after dot matrix data extraction and filtering to obtain the preprocessed fluorescence intensity matrix.

[0018] As a further limitation of the technical solution of the present invention, the preprocessing module further includes a mass spectrometry data preprocessing unit, which converts the mass spectrometry data of different time steps into a unified format; analyzes the stability of the mass spectrometry data of each time step, and selects the time step with the number of peaks within a set range and the peak intensity fluctuation range less than a set value as the reference time step; determines the mass-to-charge ratio range window, and preliminarily matches the peaks in the reference time step and the mass spectrometry diagrams of other time steps that fall within this window; by comparing the mass-to-charge ratio differences of multiple matching peaks, taking the average value as the global offset, and then performing an overall translation on the mass spectrometry data of other time steps to align the mass-to-charge ratios of the matching peaks; for the mass spectrometry data of each time step, arrange the intensities of all peaks in ascending order, take the intensity value at the middle position as the median intensity of this time step, and divide the original intensity of each peak in this time step by the median intensity of this time step to obtain the normalized mass spectrometry data, where if the number of peaks is even, take the average value of the two middle intensity values as the median intensity.

[0019] As a further limitation of the technical solution of the present invention, the spectral data processing module includes a peak detection unit, which is used to take the derivative of the preprocessed fluorescence intensity matrix data along the excitation wavelength axis and the emission wavelength axis, and use the characteristic that the first derivative is 0, combined with the point sets and Euclidean distance judgment in the two axis directions to determine the data points with key features in the fluorescence intensity matrix, that is, peak points, and generate a peak point set.

[0020] As a further limitation of the technical solution of the present invention, the spectral data processing module further includes a peak pair processing unit, which is used to calculate the Euclidean distance between the detected peak points, merge the peak points within the distance threshold range according to the set distance threshold; determine the weight of each peak point, calculate the distances of all point pairs for two groups of peak points and sort them, select the point pairs that meet the distance constraints to complete the matching of peak pairs; for the selected matching peak pairs, calculate the Euclidean distance between the matching peak pairs, and then use the linear assignment algorithm to find the set of optimal matching pairs and accumulate the minimum cost to obtain the peak matching distance; select a set number of points around each peak, calculate the feature vectors, and then calculate the distribution difference of the feature vectors through the Euclidean distance.

[0021] As a further limitation of the technical solution of the present invention, the spectral data processing module includes a first calculation unit for calculating the similarity score of the spectrum; the calculation formula is as follows:

[0022] In the formula, is the weight of the peak matching distance, is the weight of the distribution difference of the feature vectors, is the distribution difference of the feature vectors, the peak matching distance , is the peak point in the corresponding matching pair the Euclidean distance between them, match coefficient is the matching coefficient; Optimal Match represents the set of optimal matching pairs obtained by the linear assignment algorithm; among them, by normalizing the peak matching distance, normalizing the distribution difference of the feature vectors, and performing weighted summation according to the set weight coefficients, a preliminary matching coefficient value is obtained, and the preliminary matching coefficient value is multiplied by the matching ratio to obtain the match coefficient; the ratio of the successful matching of the peak points with the peak points of the standard sample under the distance and width constraint conditions is defined as the matching ratio.

[0023] As a further limitation of the technical solution of the present invention, the calculation formula for the mass spectrometry data processing module to calculate the similarity score of the mass spectrum is as follows: The processed mass spectrometry data and the mass spectrometry data of the standard sample are respectively arranged in the form of vectors, and the weight of each mass-to-charge ratio position is set, and the weight vector is ; the vector of the processed mass spectrometry data , the vector of the mass spectrometry data of the standard sample is ; Weighted cosine similarity ; Euclidean distance ; Similarity score of the mass spectrum ; In the formula, represents the th weight of the mass-to-charge ratio position, is the weight of the weighted cosine similarity, is the weight of the Euclidean distance, is the maximum value of the Euclidean distance.

[0024] As a further limitation of the technical solution of the present invention, the traceability processing module is specifically used to calculate the weight distribution between the similarity score of the spectrum and the similarity score of the mass spectrum, obtain the final similarity score of the hazardous waste and the standard sample, and trace the source of the hazardous waste based on the finally obtained similarity score.

[0025] The beneficial effects of the technical solution of the present invention are as follows: Through the combination of two methods, multi-dimensional holographic analysis of hazardous waste is realized, providing a rich and comprehensive data basis for subsequent accurate traceability and supervision. This multi-dimensional evaluation method can more comprehensively and objectively reflect the similarity between samples, avoiding the one-sidedness that may be brought by single-dimensional evaluation.

[0026] Preprocess the fluorescence spectrum data to obtain the preprocessed fluorescence intensity matrix. At the same time, preprocess the obtained mass spectrometry data, including alignment operation and normalization processing. The alignment operation aims at the data distribution offset generated by multiple time steps of the mass spectrometry data, enabling the algorithm to utilize the mass spectrometry data of multiple time steps rather than a single time step, improving the stability and accuracy of the algorithm. The normalization processing is to enable the algorithm to more stably examine the similarity between different mass spectrometry samples from the data distribution rather than the data size, enabling the algorithm to better handle the analysis process fluctuations caused by sample concentration.

[0027] Based on the preprocessed fluorescence intensity matrix, use the first derivative of the peak intensity for peak detection, and combine the fluorescence spectrum of the standard sample to determine the weight of each peak point to complete the matching of peak pairs. This method can accurately extract the characteristic peaks in the fluorescence spectrum data and reasonably weight the peaks according to the information of the standard sample, thus more accurately reflecting the characteristics of hazardous waste. At the same time, calculate the peak matching distance and the distribution difference of the eigenvectors, comprehensively considering the position of the peaks and the overall distribution of the eigenvectors, further improving the accuracy and reliability of feature extraction.

[0028] Calculate the similarity scores of the spectrum and the mass spectrometry respectively, and conduct similarity evaluation on the hazardous waste samples and the standard samples from two dimensions of fluorescence spectrum and mass spectrometry, which can accurately find out the possible sources of hazardous waste and provide strong support for the supervision and treatment of hazardous waste. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solution of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0030] Figure 1 It is a schematic flowchart of the method provided by the embodiment of the present invention.

[0031] Figure 2 It is a schematic block diagram of the system provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0032] To make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the specific embodiments. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts belong to the scope of protection of this application.

[0033] As Figure 1 shown, an embodiment of the present invention provides a multi-dimensional holographic data analysis method for hazardous waste supervision, including: S1: Detect hazardous waste samples using a three-dimensional fluorescence spectrometer and a mass spectrometer respectively to obtain fluorescence spectral data and mass spectral data; For liquid hazardous waste samples, if there are suspended solids in the samples, they need to be removed by filtration through a filter membrane to avoid scattering interference on the fluorescence signal. If the sample concentration is too high, it needs to be diluted with ionized water to make the fluorescence intensity within the linear detection range of the instrument.

[0034] For solid hazardous waste samples, first grind them, then weigh a certain amount of the sample, add an organic solvent or buffer solution for extraction to dissolve the target fluorescent substance into the extract. After the extract is filtered or centrifuged, take the supernatant for detection. According to the nature of the sample and the detection purpose, set parameters such as the excitation wavelength range, emission wavelength range, scanning speed, and integration time. For example, for common organic pollutants, the excitation wavelength range can be set to 200 - 400 nm, and the emission wavelength range can be set to 250 - 600 nm. The scanning speed is generally set to about 1200 nm / min, and the integration time is appropriately adjusted according to the fluorescence intensity of the sample, usually 0.1 - 1 s. Inject the processed sample into a quartz cuvette and place it in the sample cell of the three-dimensional fluorescence spectrometer. Start the instrument for scanning, and the instrument will automatically record the fluorescence intensity at different combinations of excitation wavelengths and emission wavelengths, thereby obtaining three-dimensional fluorescence spectral data. After the scanning is completed, save and preliminarily process the data, such as subtracting the fluorescence signal of the blank sample. In this step of testing, it may be impossible to obtain three-dimensional fluorescence spectral data. If this situation occurs, directly analyze and trace the source using the mass spectral data.

[0035] Inject the pre-treated sample into the mass spectrometer. The sample is ionized in the ion source, and the formed ions are separated by mass-to-charge ratio in the mass analyzer. Finally, the detector detects and records the intensity and mass-to-charge ratio information of the ions to obtain mass spectral data.

[0036] S2: Preprocess the obtained fluorescence spectral data to obtain a preprocessed fluorescence intensity matrix, and at the same time perform alignment operation and normalization processing on the mass spectral data; The original three-dimensional fluorescence spectral data may contain a large amount of redundant and irrelevant information, such as data generated from background noise of the instrument itself, scattered light unrelated to the target analyte, etc. Through dot matrix data extraction and filtering, valuable data points for analysis can be selectively selected, and those irrelevant or interfering data can be removed; in addition, the amount of three-dimensional fluorescence spectral data is usually large, and direct processing may face problems such as high computational complexity and low analysis efficiency. Dot matrix data extraction can simplify and reduce the dimension of the data on the premise of retaining key information, and transform complex three-dimensional data into a dot matrix form that is easier to process. Furthermore, the specific processing process in this step includes: S21: Extract and filter the dot matrix data from the fluorescence spectral data to generate a fluorescence intensity matrix with the excitation wavelength as the row index, the emission wavelength as the column index, and the corresponding fluorescence intensity value as the matrix element; perform Gaussian smoothing on the fluorescence intensity matrix after dot matrix data extraction and filtering to obtain the preprocessed fluorescence intensity matrix.

[0037] First, convert the matrix into dot matrix data, which is presented in the form of three-dimensional coordinates, and each point contains excitation wavelength EX, emission wavelength EM, and intensity (i.e., weight) information. To highlight the effective information, filter out the points with lower weights, and perform screening through the formula Set the excitation wavelength as Set the emission wavelength as Set the intensity (weight) as Select the effective data points according to the intensity (weight) threshold. Among them is the maximum weight in the current point set, r is the weight threshold ratio. In the embodiments of the present invention, r is set between 0.2 - 0.5 according to experience. This step can remove the data points that contribute less to the subsequent analysis, and improve the data processing efficiency and accuracy. For example, when processing a large amount of fluorescence spectral data, the interference of redundant information can be reduced.

[0038] Apply Gaussian smoothing to the fluorescence intensity matrix F ( x, y ) to reduce noise. The formula is In the formula, is the Gaussian kernel function, is the smoothing parameter. By adjusting the value to control the smoothing degree, this operation can effectively reduce the noise in the data, avoid the interference of noise on peak detection, and improve the accuracy of peak detection.

[0039] During the fluorescence spectrum measurement, various random noises will inevitably be introduced, such as electronic noise, environmental interference, etc. These noises will cause fluctuations and unevenness in the spectral curve, affecting the accurate identification and analysis of spectral features. Gaussian smoothing can effectively suppress these random noises through weighted average processing of the data, making the spectral curve smoother, highlighting the true spectral signal, and improving the quality and reliability of the data.

[0040] After the dot matrix data extraction and filtering, the data points may be discrete or discontinuous to a certain extent. Gaussian smoothing processing can fill the gaps between data points to a certain extent, making the data more continuous and smooth within a local range, which is beneficial for further analysis of the spectrum, such as peak detection and other operations, and more accurate and stable results can be obtained.

[0041] S22: Convert the mass spectrometry data of different time steps into a unified format to ensure that each data file contains the mass-to-charge ratio and the corresponding peak intensity information; analyze the stability of the mass spectrometry data of each time step, and select the time steps with the number of peaks within the set range and the peak intensity fluctuation range less than the set value as the reference time steps; determine the mass-to-charge ratio range window, and preliminarily match the peaks in the reference time steps and the mass spectrometry diagrams of other time steps that fall within this window; by comparing the mass-to-charge ratio differences of multiple matching peaks, take the average value as the global offset, and then perform an overall translation on the mass spectrometry data of other time steps to align the mass-to-charge ratios of the matching peaks; for the mass spectrometry data of each time step, arrange the intensities of all peaks in ascending order, take the intensity value at the middle position as the median intensity of this time step, and divide the original intensity of each peak in this time step by the median intensity of this time step to obtain the normalized mass spectrometry data, where if the number of peaks is even, take the average value of the two middle intensity values as the median intensity.

[0042] S3: Perform peak detection based on the first derivative of the peak intensity using the preprocessed fluorescence intensity matrix; specifically, it includes: taking the derivative of the data in the preprocessed fluorescence intensity matrix along the excitation wavelength axis and the emission wavelength axis, and using the characteristic that the first derivative is 0, combined with the point sets and Euclidean distance judgment in the two axis directions to determine the data points with key features in the fluorescence intensity matrix, that is, the peak points, and generate a set of peak points.

[0043] Using the first derivative of the peak intensity, take the derivative of the data along the EX axis and the EM axis respectively, and use the characteristic that the first derivative is 0, combined with the point sets and Euclidean distance judgment in the two axis directions to determine the peak points. The formula is as follows: ,

[0044] Along the EX axis direction, according to the formula ,traverse the data to determine the satisfaction of Point set ; Here Traverse different positions on the EX axis, Fix at a certain position, find all coordinate combinations that make the function value 0 and the corresponding F values, and form a set ; Along the EM axis direction, according to the formula , Similarly, Fix, Traverse different positions on the EM axis, and find the point set that satisfies Point set ; Is the spatial coordinate of the midpoint of the set And ; Traverse Set, calculate the Euclidean distance from each point in the set To all points in the set , Select the points with Euclidean distance equal to 0 as the peak points.

[0045] S4: Combine the fluorescence spectra of the standard samples to determine the weights of each peak point to complete the matching of peak pairs, and calculate the peak matching distance and the distribution difference of the eigenvectors; This step specifically includes: S41: Calculate the Euclidean distance between the detected peak points, and merge the peak points within the distance threshold according to the set distance threshold; For two points in the peak set And Calculate their two-dimensional Euclidean distance , Provide data support for subsequent deduplication and merging operations. According to the set distance threshold , Merge the peak points with calculated distance less than the distance threshold to obtain a new peak center, with coordinates , The calculation formula is , Where n Is the number of points that satisfy the distance constraint. This step avoids duplicate counting of peaks caused by data errors or instrument noise, making the peak information more accurate.

[0046] S42: Determine the weight of each peak point, calculate the distances of all point pairs for the two groups of peak points and sort them, select the point pairs that satisfy the distance constraint, and complete the matching of peak pairs; Find the corresponding weight (intensity value) through the detected points , Is the fluorescence intensity matrix after Gaussian smoothing. Represents the finally determined weight of this point after weight assignment, Indicates the current point Search within the neighborhood range of for the maximum value, where the neighborhood refers to the set of points within a specified range centered around the current point; For two sets of peak point sets and , calculate the distances between all point pairs and sort them, and select the point pairs that satisfy the distance constraint. The formula is , to provide matching point pairs for calculating the peak matching distance. Specifically, in the two sets of peak point sets and , calculate the Euclidean distance of all possible point pairs . Select the point pairs with distances less than or equal to the set distance threshold to form Matched Pairs, that is, the set of matching peak pairs.

[0047] S43: For the selected matching peak pairs, calculate the Euclidean distance between the matching peak pairs, then use the linear assignment algorithm to find the set of optimal matching pairs and accumulate the minimum cost to obtain the peak matching distance; For the two sets of peak sets peaks 1 and peaks 2 of the corresponding matching pairs, first calculate the Euclidean distance between them; Then use the linear assignment algorithm to find the set of optimal matching pairs and accumulate the minimum cost to obtain the peak matching distance , is the Euclidean distance between the peak points in the corresponding matching pairs, and Optimal Match represents the set of optimal matching pairs obtained by the linear assignment algorithm.

[0048] In the embodiments of the present invention, the process of obtaining the set of optimal matching pairs by the linear assignment algorithm: For the two sets of peaks peaks 1 and peaks 2 of the corresponding matching pairs, there are m peak points in peaks 1 , and there are n peak points in peaks 2 . In the formula for calculating the Euclidean distance , is the coordinate of the 1 th peak point in peaks , peaks 2 is the coordinate of the th peak point in peaks . The calculated distance values form a , the elements in the matrix .

[0049] For each row of the cost matrix , find the minimum value in that row, and then subtract this minimum value from each element in that row. The purpose of this step is to have at least one zero element in each row. For each column of the matrix after row reduction, find the minimum value in that column, and then subtract this minimum value from each element in that column. Similarly, this will result in at least one zero element in each column.

[0050] Find as many independent zero elements as possible in the matrix (i.e., zero elements located in different rows and different columns). The marking method can be used to achieve this: First, check each zero element in the matrix. Mark the zero element whose row and column have no other marked zero elements as an independent zero element. Repeat the above steps until no more independent zero elements can be marked.

[0051] If the number of independent zero elements is equal to , it means that the optimal matching has been found, and the peak point pairs corresponding to the rows and columns where these independent zero elements are located are the set of optimal matching pairs.

[0052] If the number of independent zero elements is less than , the matrix needs to be adjusted until the number of independent zero elements is equal to . Here, the steps for matrix adjustment include: using the minimum number of straight lines (row lines and column lines) to cover all zero elements in the matrix, finding the minimum value K among the elements not covered by the straight lines, then subtracting this minimum value K from each element not covered by the straight lines, adding this minimum value K to each element covered by two straight lines, and then repeating the steps of checking each zero element in the matrix and subsequent steps.

[0053] In one embodiment, the linear assignment algorithm is specifically implemented using the linear_sum_assignment function in the scipy library. Further, it should be noted that the cost matrix is processed using the linear_sum_assignment function in the scipy library to obtain the row indices and column indices of the optimal matching. According to the obtained row indices and column indices, the corresponding peak coordinates are found from peaks 1 and peaks 2 to form the set of optimal matching pairs.

[0054] S44: Select a set number of points around each peak, calculate the eigenvector, and then calculate the distribution difference of the eigenvectors through the Euclidean distance.

[0055] Select n points around each peak and calculate the feature vector and calculate the distribution difference of the feature vectors using the Euclidean distance .

[0056] Specifically represents the feature vector calculated from the points selected around each peak represents the weights (intensity values) of the n points selected around each peak, calculate the average value of the weights of these n points , and then normalize this average value normalize to obtain the feature vector .

[0057] represents the distribution difference between two feature vectors, which is used to measure the difference degree of the features in the regions around two groups of peaks and are the feature vectors corresponding to two different groups of peaks respectively. Calculate the sum of the squares of the differences of the corresponding elements of the two feature vectors, and then take the square root. The larger the obtained result, the greater the feature difference in the regions around the two groups of peaks; on the contrary, the smaller the difference

[0058] S5: Synthesize the peak matching distance and the distribution difference of the feature vectors, combine the matching ratio and the weight ratio, and calculate the similarity score of the spectrum The formula in the steps of calculating the similarity score of the spectrum is as follows

[0059] In the formula is the weight of the peak matching distance is the weight of the distribution difference of the feature vectors is the distribution difference of the feature vectors; match coefficient is the matching coefficient; among them, by normalizing the peak matching distance and normalizing the distribution difference of the feature vectors, and performing weighted summation according to the set weight coefficients, a preliminary matching coefficient value is obtained, and the preliminary matching coefficient value is multiplied by the matching ratio to obtain matchcoefficient; the ratio of the successfully matched peak points of the detected peak points to the peak points of the standard sample is defined as the matching ratio. Here, the detected peak points refer to the peak points that meet the distance and width constraint conditions during peak detection

[0060] The distance constraint condition refers to the distance between peak points, that is, the interval of the detected different peak points in the coordinate space of the fluorescence intensity matrix. For example, if the distance between two adjacent peak points is less than a certain set threshold, it may be considered that they are multiple pseudo-peaks generated by the same peak due to factors such as noise interference, and only one of them is retained as the true peak

[0061] The width constraint condition is the full width at half maximum (FWHM) of the peak, which generally refers to the width of the abscissa (such as wavelength, etc.) range corresponding to the peak at the position where the peak intensity is half. If the width of the detected "peak" is too narrow, it may be a spike caused by noise rather than the peak of the true fluorescence signal; if the width is too wide, it may be due to the overlap of multiple peaks or signal broadening caused by issues such as instrument resolution, and further analysis or data processing is required to separate the true peak.

[0062] S6: Calculate the weighted cosine similarity and Euclidean distance between the processed mass spectrometry data and the mass spectrometry data of the standard sample to complete the calculation of the similarity score of the mass spectrometry between samples; In this step, the processed mass spectrometry data and the mass spectrometry data of the standard sample are respectively organized into vector forms, and weights are set for each peak position. (Here, weights are assigned to each mass-to-charge ratio position according to the actual situation. The determination of weights can be based on factors such as the importance and stability of the peaks. For example, some characteristic peaks have a higher discrimination ability for samples and can be assigned higher weights.) The weight vector is ; the vector of the processed mass spectrometry data , and the vector of the mass spectrometry data of the standard sample is ; Weighted cosine similarity , and the value range of the weighted cosine similarity is between [-1, 1]. The closer the value is to 1, the more similar the directions of the two vectors are; Euclidean distance , and the larger the value of the Euclidean distance, the greater the difference between the two vectors; Similarity score of mass spectrometry ; In the formula, represents the weight at the i-th mass-to-charge ratio position, is the weight of the weighted cosine similarity, is the weight of the Euclidean distance, is the maximum value of the Euclidean distance, which is used to normalize the Euclidean distance to the [0, 1] interval so that it can be combined with the weighted cosine similarity on the same scale.

[0063] In the embodiments of the present invention, the preprocessing of the acquired mass spectrometry data includes alignment operation and normalization processing. The alignment operation addresses the data distribution offset generated by the data of multiple time steps of the mass spectrometry data, enabling the algorithm to utilize the mass spectrometry data of multiple time steps rather than a single time step, thereby enhancing the stability and accuracy of the algorithm. The normalization processing is to enable the algorithm to more stably examine the similarity between different mass spectrometry samples from the data distribution rather than the data size, enabling the algorithm to better handle the analysis process fluctuations brought about by the sample concentration. After the alignment operation and normalization processing, two feature vectors are obtained, representing two samples respectively. The weighted cosine similarity and Euclidean distance between the two vectors are calculated; the weighted cosine similarity is used to calculate the similarity of the mass spectrometry shape, and the Euclidean distance is used to measure the absolute difference size of the mass spectrometry data. Through this calculation method, the algorithm can consider from multiple aspects such as local shape (weighted) and global distribution to complete the calculation of the mass spectrometry similarity measurement between samples.

[0064] S7: Trace the source of the hazardous waste based on the similarity scores of the spectra and the similarity scores of the mass spectra.

[0065] Specifically, by calculating the weight distribution between the similarity scores of the spectra and the similarity scores of the mass spectra, the final similarity score between the hazardous waste and the standard sample is obtained, and the source of the hazardous waste is traced based on the finally obtained similarity score.

[0066] It should be noted that the relevant data of the standard samples used in the present invention are all data obtained by detecting the hazardous waste of each enterprise within the set detection range through reasonable and compliant methods, and the detected data are pre-stored to facilitate subsequent comparison analysis and traceability of the monitored hazardous waste.

[0067] As Figure 2 shown, the embodiments of the present invention also provide a multi-dimensional holographic data analysis system for hazardous waste supervision, including a detection data acquisition module, a preprocessing module, a spectral data processing module, a mass spectrometry data processing module, and a traceability processing module; The detection data acquisition module is used to detect the hazardous waste samples using a three-dimensional fluorescence spectrometer and a mass spectrometer respectively to obtain fluorescence spectral data and mass spectrometry data; The preprocessing module is used to preprocess the acquired fluorescence spectral data to obtain a preprocessed fluorescence intensity matrix, and at the same time perform alignment operation and normalization processing on the mass spectrometry data; The spectral data processing module is used to perform peak detection based on the preprocessed fluorescence intensity matrix using the first derivative of the peak intensity; determine the weight of each peak point in combination with the fluorescence spectrum of the standard sample to complete the matching of peak pairs, and calculate the peak matching distance and the distribution difference of the feature vectors; combine the peak matching distance and the distribution difference of the feature vectors, and combine the matching ratio and the weight ratio to calculate the similarity score of the spectrum; A mass spectrometry data processing module, which is used to calculate the weighted cosine similarity and Euclidean distance between the processed mass spectrometry data and the mass spectrometry data of the standard sample, and complete the calculation of the similarity score of the mass spectrometry between samples; A traceability processing module, which is used to trace the source of hazardous waste according to the similarity scores of spectra and mass spectrometry.

[0068] In some embodiments, the preprocessing module includes a spectral data preprocessing unit, which is used to extract and filter dot matrix data from fluorescence spectral data, and generate a fluorescence intensity matrix with the excitation wavelength as the row index, the emission wavelength as the column index, and the corresponding fluorescence intensity value as the matrix element; perform Gaussian smoothing processing on the fluorescence intensity matrix after dot matrix data extraction and filtering to obtain the preprocessed fluorescence intensity matrix.

[0069] In some embodiments, the preprocessing module further includes a mass spectrometry data preprocessing unit, which converts the mass spectrometry data of different time steps into a unified format; analyzes the stability of the mass spectrometry data of each time step, and selects the time step with the number of peaks within a set range and the peak intensity fluctuation range less than a set value as the reference time step; determines the mass-to-charge ratio range window, and preliminarily matches the peaks in the reference time step and the mass spectrometry diagrams of other time steps that fall within this window; by comparing the mass-to-charge ratio differences of multiple matching peaks, taking the average value as the global offset, and then performing an overall translation on the mass spectrometry data of other time steps to align the mass-to-charge ratios of the matching peaks; for the mass spectrometry data of each time step, arrange the intensities of all peaks in ascending order, take the intensity value at the middle position as the median intensity of this time step, and divide the original intensity of each peak in this time step by the median intensity of this time step to obtain the normalized mass spectrometry data, where if the number of peaks is even, take the average value of the two middle intensity values as the median intensity.

[0070] In some embodiments, the spectral data processing module includes a peak detection unit, which is used to take the derivative of the preprocessed fluorescence intensity matrix data along the excitation wavelength axis and the emission wavelength axis, and use the characteristic that the first derivative is 0, combined with the judgment of the point sets and Euclidean distances in the two axis directions to determine the data points with key features in the fluorescence intensity matrix, that is, peak points, and generate a peak point set.

[0071] In some embodiments, the spectral data processing module further includes a peak pair processing unit, which is used to calculate the Euclidean distance between the detected peak points, merge the peak points within the range of the distance threshold according to the set distance threshold; determine the weight of each peak point, calculate the distances of all point pairs for two groups of peak points and sort them, select the point pairs that meet the distance constraint to complete the matching of peak pairs; for the selected matching peak pairs, calculate the Euclidean distance between the matching peak pairs, then use the linear assignment algorithm to find the set of optimal matching pairs and accumulate the minimum cost to obtain the peak matching distance; select a set number of points around each peak, calculate the feature vector, and then calculate the distribution difference of the feature vector through the Euclidean distance.

[0072] In some embodiments, the spectral data processing module includes a first calculation unit, which is used to calculate the similarity score of the spectrum; the calculation formula is as follows:

[0073] In the formula, is the weight of the peak matching distance, is the weight of the distribution difference of the feature vector, is the distribution difference of the feature vector, and the peak matching distance , is the Euclidean distance between the peak points in the corresponding matching pair, and match coefficient is the matching coefficient; Optimal Match represents the set of optimal matching pairs obtained through the linear assignment algorithm; among them, by normalizing the peak matching distance and normalizing the distribution difference of the feature vector, and performing weighted summation according to the set weight coefficient, a preliminary matching coefficient value is obtained, and the preliminary matching coefficient value is multiplied by the matching ratio to obtain match coefficient; the ratio of the successful matching of the peak points to the peak points of the standard sample under the distance and width constraint conditions is defined as the matching ratio.

[0074] In some embodiments, the calculation formula for the mass spectrometry data processing module to calculate the similarity score of the mass spectrometry is as follows: The processed mass spectrometry data and the mass spectrometry data of the standard sample are respectively arranged in the form of vectors, and the weight of each mass-to-charge ratio position is set, and the weight vector is ; the vector of the processed mass spectrometry data , and the vector of the mass spectrometry data of the standard sample is ; The weighted cosine similarity ; The Euclidean distance ; The similarity score of the mass spectrometry ; In the formula, represents the weight of the i-th mass-to-charge ratio position, is the weight of the weighted cosine similarity, is the weight of the Euclidean distance, is the maximum value of the Euclidean distance.

[0075] In some embodiments, the traceability processing module is specifically configured to calculate the weight distribution between the similarity score of the spectrum and the similarity score of the mass spectrum, obtain the final similarity score between the hazardous waste and the standard sample, and trace the source of the hazardous waste based on the finally obtained similarity score.

[0076] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multidimensional holographic data analysis method for hazardous waste supervision, characterized in that: include: The hazardous waste samples are tested using a three-dimensional fluorescence spectrometer and a mass spectrometer to obtain fluorescence spectrum data and mass spectrum data; The acquired fluorescence spectrum data is preprocessed to obtain a preprocessed fluorescence intensity matrix, and the mass spectrum data is aligned and normalized; Peak detection is performed using the first-order derivative of the peak intensity based on the preprocessed fluorescence intensity matrix; The weight of each peak point is determined by combining the fluorescence spectrum of the standard sample to complete the peak pair matching, and the peak matching distance and the distribution difference of the characteristic vector are calculated; The similarity score of the spectrum is calculated by combining the peak matching distance and the distribution difference of the feature vector, the matching ratio and the weight ratio; Calculate the weighted cosine similarity and Euclidean distance between the processed mass spectrum data and the mass spectrum data of the standard sample to complete the calculation of the similarity score of the mass spectra between the samples; Trace the source of hazardous waste based on the similarity score of the spectrum and the similarity score of the mass spectrum.

2. The multidimensional holographic data analysis method for hazardous waste supervision according to claim 1 is characterized in that: The step of preprocessing the acquired fluorescence spectrum data to obtain a preprocessed fluorescence intensity matrix includes: Extract and filter the fluorescence spectrum data to generate a fluorescence intensity matrix with the excitation wavelength as the row index, the emission wavelength as the column index, and the corresponding fluorescence intensity value as the matrix element; Gaussian smoothing is performed on the fluorescence intensity matrix after dot matrix data extraction and filtering to obtain a preprocessed fluorescence intensity matrix.

3. The multidimensional holographic data analysis method for hazardous waste supervision according to claim 2 is characterized in that: The steps for aligning and normalizing mass spectrometry data include: Convert mass spectrometry data at different time steps into a unified format; Analyze the stability of mass spectrometry data at each time step, and select the time step where the number of peaks is within the set range and the peak intensity fluctuation range is less than the set value as the reference time step; Determine the mass-to-charge ratio range window, and preliminarily match the peaks falling within the window in the mass spectra of the reference time step and other time steps; By comparing the differences in mass-to-charge ratios of multiple matching peaks, the average value is taken as the global offset, and then the mass spectrum data of other time steps are globally translated to align the mass-to-charge ratios of the matching peaks; For the mass spectrometry data of each time step, the intensities of all peaks are arranged in ascending order, the intensity value in the middle position is taken as the median intensity of the time step, and the original intensity of each peak in the time step is divided by the median intensity of the time step to obtain the normalized mass spectrometry data. If the number of peaks is an even number, the average of the two middle intensity values ​​is taken as the median intensity.

4. The multidimensional holographic data analysis method for hazardous waste supervision according to claim 3 is characterized in that: The steps of performing peak detection based on the preprocessed fluorescence intensity matrix using the first-order derivative of the peak intensity include: The preprocessed fluorescence intensity matrix data is differentiated along the excitation wavelength axis and the emission wavelength axis. The characteristic that the first-order derivative is 0 is used, and the point sets in the two axis directions and the Euclidean distance judgment are combined to determine the data points with key features in the fluorescence intensity matrix, namely the peak points, and generate a peak point set.

5. The multi-dimensional holographic data analysis method for hazardous waste supervision according to claim 4 is characterized in that: The steps of determining the weight of each peak point in combination with the fluorescence spectrum of the standard sample to complete the peak pair matching, and calculating the peak matching distance and the distribution difference of the characteristic vector include: Calculate the Euclidean distance between the detected peak points, and merge the peak points within the distance threshold range according to the set distance threshold; Determine the weight of each peak point, calculate the distance of all point pairs for the two groups of peak points and sort them, select the point pairs that meet the distance constraints, and complete the peak pair matching; For the selected matching peak pairs, the Euclidean distance between the matching peak pairs is calculated, and then the linear allocation algorithm is used to find the optimal matching pair set, and the minimum cost is accumulated to obtain the peak matching distance; A set number of points are selected around each peak, the eigenvectors are calculated, and then the distribution difference of the eigenvectors is calculated using the Euclidean distance.

6. The multi-dimensional holographic data analysis method for hazardous waste supervision according to claim 5 is characterized in that: The formula for calculating the similarity score of the spectrum by combining the peak matching distance and the distribution difference of the feature vector, the matching ratio and the weight ratio is as follows: In the formula, is the weight of the peak matching distance, is the weight of the distribution difference of the eigenvector, is the distribution difference of the feature vector, the peak matching distance , The peak point of the corresponding matching pair The Euclidean distance between them, match coefficient is the matching coefficient; Optimal Match represents the optimal matching pair set obtained by the linear assignment algorithm; Among them, the peak matching distance is normalized, the distribution difference of the feature vector is normalized, and a preliminary matching coefficient value is obtained by weighted summation according to the set weight coefficient. The preliminary matching coefficient value is multiplied by the matching ratio to obtain the match coefficient; the ratio of the detected peak point to the peak point of the standard sample that is successfully matched is defined as the matching ratio.

7. The multi-dimensional holographic data analysis method for hazardous waste supervision according to claim 6 is characterized in that: The steps of calculating the weighted cosine similarity and the Euclidean distance between the processed mass spectrum data and the mass spectrum data of the standard sample and completing the calculation of the similarity score of the mass spectra between the samples include: The processed mass spectrometry data and the mass spectrometry data of the standard sample are organized into vectors, and the weight of each mass-to-charge ratio position is set. The weight vector is ; Vector of processed mass spectrometry data , the vector of mass spectrum data of standard sample is ; Weighted cosine similarity ; Euclidean distance ; Similarity score of mass spectra ; In the formula, Indicates The weight of each mass-to-charge ratio position, is the weight of weighted cosine similarity, is the weight of the Euclidean distance, is the maximum value of the Euclidean distance.

8. The multi-dimensional holographic data analysis method for hazardous waste supervision according to claim 7 is characterized in that: The steps for tracing the source of hazardous waste based on the similarity scores of the spectrum and the similarity scores of the mass spectrum include: The weight distribution between the similarity score of the spectrum and the similarity score of the mass spectrum is calculated to obtain the final similarity score between the hazardous waste and the standard sample, and the source of the hazardous waste is traced based on the final similarity score.

9. A multi-dimensional holographic data analysis system for hazardous waste supervision, characterized in that: It includes a detection data acquisition module, a preprocessing module, a spectrum data processing module, a mass spectrum data processing module and a traceability processing module; A detection data acquisition module is used to detect hazardous waste samples using a three-dimensional fluorescence spectrometer and a mass spectrometer to obtain fluorescence spectrum data and mass spectrum data; A preprocessing module is used to preprocess the acquired fluorescence spectrum data to obtain a preprocessed fluorescence intensity matrix, and to align and normalize the mass spectrum data; The spectrum data processing module is used to perform peak detection based on the first-order derivative of the peak intensity based on the preprocessed fluorescence intensity matrix; determine the weight of each peak point in combination with the fluorescence spectrum of the standard sample to complete the peak pair matching, and calculate the peak matching distance and the distribution difference of the characteristic vector; comprehensively consider the peak matching distance and the distribution difference of the characteristic vector, and combine the matching ratio with the weight ratio to calculate the similarity score of the spectrum; The mass spectrum data processing module is used to calculate the weighted cosine similarity and Euclidean distance between the processed mass spectrum data and the mass spectrum data of the standard sample, and complete the calculation of the similarity score of the mass spectra between the samples; The source tracing processing module is used to trace the source of hazardous waste based on the similarity scores of the spectrum and the similarity scores of the mass spectrum.

10. The multi-dimensional holographic data analysis system for hazardous waste supervision according to claim 9, characterized in that: The formula for calculating the similarity score of the spectrum by the spectral data processing module is as follows: In the formula, is the weight of the peak matching distance, is the weight of the distribution difference of the eigenvector, is the distribution difference of the feature vector, the peak matching distance , The peak point of the corresponding matching pair The Euclidean distance between them, match coefficient is the matching coefficient; Optimal Match represents the optimal matching pair set obtained by the linear assignment algorithm; wherein, the peak matching distance is normalized, the distribution difference of the feature vector is normalized, and a preliminary matching coefficient value is obtained by weighted summation according to the set weight coefficient, and the preliminary matching coefficient value is multiplied by the matching ratio to obtain the match coefficient; the ratio of the detected peak point to the peak point of the standard sample that is successfully matched is defined as the matching ratio.

Citation Information

Patent Citations

  • Method for realizing rapid identification and comparison by utilizing fluorescence spectrum characteristic information

    CN110554013A

  • Tracing instrument and system based on multi-dimensional data

    CN117373556A

  • Method and apparatus for analysing samples of biomolecules using mass spectrometry with data-independent acquisition

    EP4047371A1

  • High Mass Accuracy Filtering for Improved Spectral Matching of High-Resolution Gas Chromatography-Mass Spectrometry Data Against Unit-Resolution Reference Databases

    US20150340216A1

  • Three-dimensional spectral data processing device and processing method

    US20170356889A1

Cited By

  • Mass spectrum gas source analysis method and system based on K-means clustering algorithm

    CN120929865A

  • Mass spectrometry gas source analysis method and system based on k-means clustering algorithm

    CN120929865B