Method and device for identifying a mixture of substances based on infrared spectroscopy and deconvolution
By using infrared spectroscopy and peak matching, the components of mixed substances can be automatically identified, solving the problems of low identification efficiency and large errors in existing technologies, and achieving efficient and accurate identification of mixed substances.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 元始智能科技(南通)有限公司
- Filing Date
- 2023-01-29
- Publication Date
- 2026-07-24
AI Technical Summary
Current technologies for identifying mixed substances rely on human experience, resulting in low efficiency, high manpower and material resources, and large identification errors.
The method employs infrared spectroscopy and peak matching to automatically identify the components of the mixture by acquiring the infrared spectra of the mixture and the sample.
It enables automatic and accurate identification of the components of mixed substances, improves identification efficiency, reduces manpower and material resources consumption, and reduces identification errors.
Smart Images

Figure CN116202980B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of materials testing technology, and in particular to a method and apparatus for identifying mixed substances based on infrared spectroscopy and peak matching. Background Technology
[0002] Substance testing and inspection refers to the process by which testing and inspection organizations, commissioned by regulatory agencies, manufacturers, or product users, use professional technical means and equipment under applicable standards and technical specifications to test and inspect the quality, safety, performance, and environmental aspects of identified samples, and issue test reports to assess whether the samples meet the standards and requirements of regulatory agencies, industry, and users in terms of quality, safety, and performance. Therefore, accurate identification of substances is a crucial step in ensuring product safety and performance.
[0003] In existing technologies, substance identification typically relies on the human experience of engineers. However, the composition of mixed substances is complex, requiring manual switching between multiple detection devices to identify each component individually. This approach is inefficient, consumes significant human and material resources, and the identification results are greatly influenced by human experience, leading to large identification errors. Summary of the Invention
[0004] This invention provides a method and apparatus for identifying mixed substances based on infrared spectroscopy and peak matching, which solves the shortcomings of existing technologies that rely on human experience for substance identification, resulting in low identification efficiency, high consumption of manpower and resources, and large identification errors, and achieves automatic and accurate identification of mixed substances.
[0005] This invention provides a method for identifying mixed substances based on infrared spectroscopy and peak matching, comprising:
[0006] Obtain the infrared spectrum of the mixture to be tested and the infrared spectrum of multiple sample substances;
[0007] Based on the infrared spectrum of the mixture to be tested, obtain the set of characteristic peaks to be tested; based on the infrared spectrum of the multiple sample substances, obtain the set of characteristic peaks of the samples.
[0008] Iterative peak-removal matching is performed on the target feature peaks in the target feature peak set and the sample feature peaks in the sample feature peak set;
[0009] Based on the peak matching results, a set of sample substances that match the mixture to be tested is obtained. Based on the composition of the set of sample substances that match the mixture to be tested, the composition identification result of the mixture to be tested is obtained.
[0010] The peak removal matching includes a matching operation and a peak removal operation. The matching operation is used to perform similarity matching between the test feature peaks in the test feature peak set corresponding to each peak removal matching and the sample feature peaks in the sample feature peak set corresponding to each peak removal matching. The peak removal operation is used to delete the test feature peak with the highest similarity obtained from each peak removal matching from the test feature peak set corresponding to each peak removal matching to obtain the test feature peak set corresponding to the next peak removal matching, and to delete all sample feature peaks of the sample substance corresponding to the sample feature peak with the highest similarity obtained from each peak removal matching from the sample feature peak set corresponding to each peak removal matching to obtain the sample feature peak set corresponding to the next peak removal matching.
[0011] According to the present invention, a method for identifying mixed substances based on infrared spectroscopy and peak matching includes iteratively performing peak matching on the target feature peaks in the target feature peak set and the sample feature peaks in the sample feature peak set, comprising:
[0012] For the current peak removal matching, each test feature peak in the set of test feature peaks corresponding to the current peak removal matching is matched with each sample feature peak in the set of sample feature peaks corresponding to the current peak removal matching.
[0013] The sample substance to which the feature peak of the most similar sample is matched is taken as the sample substance that matches the mixture to be tested in the current peak removal matching.
[0014] The infrared spectrum of the sample substance that matches the mixed substance to be tested in the current peak removal matching is superimposed with the combined spectrum corresponding to the previous peak removal matching to obtain the combined spectrum corresponding to the current peak removal matching.
[0015] Calculate the similarity between the combined spectrum corresponding to the current peak removal matching and the infrared spectrum to be measured, and obtain the spectrum similarity corresponding to the current peak removal matching;
[0016] If the spectral similarity corresponding to the current peak removal match is greater than the spectral similarity corresponding to the previous peak removal match, the test feature peak with the highest similarity is deleted from the test feature peak set corresponding to the current peak removal match, and the test feature peak set corresponding to the next peak removal match is obtained.
[0017] The sample feature peak with the highest similarity is removed from the sample feature peak set corresponding to the current peak removal match, and the sample feature peak set corresponding to the next peak removal match is obtained.
[0018] The peak removal matching step is iteratively performed by matching each test feature peak in the set of test feature peaks corresponding to the next peak removal matching with each sample feature peak in the set of sample feature peaks corresponding to the next peak removal matching, until the peak removal matching termination condition is met.
[0019] According to the present invention, a method for identifying mixed substances based on infrared spectroscopy and peak matching includes, wherein the step of performing similarity matching between each target feature peak in the target feature peak set corresponding to the current peak matching and each sample feature peak in the sample feature peak set corresponding to the current peak matching includes:
[0020] Based on the Jaccard coefficient, a similarity match is performed between each test feature peak in the set of test feature peaks corresponding to the current peak removal match and each sample feature peak in the set of sample feature peaks corresponding to the current peak removal match.
[0021] According to the present invention, a method for identifying mixed substances based on infrared spectroscopy and peak matching is provided, wherein the method involves superimposing the infrared spectrum of the sample substance that matches the mixed substance to be tested in the current peak matching with the combined spectrum corresponding to the previous peak matching to obtain the combined spectrum corresponding to the current peak matching, comprising:
[0022] Linear interpolation is performed on the infrared spectra of the sample substances that match the mixed substance to be tested in the current peak-removal matching and the combined spectra corresponding to the previous peak-removal matching, so that the data format of the infrared spectra of the samples after linear interpolation is consistent with the data format of the combined spectra after linear interpolation.
[0023] The sample infrared spectrum after linear interpolation is linearly superimposed with the combined spectrum after linear interpolation to obtain the combined spectrum corresponding to the current peak matching.
[0024] According to the present invention, a method for identifying mixed substances based on infrared spectroscopy and peak matching is provided. The step of obtaining a set of sample substances matching the mixed substance to be tested based on the peak matching result includes:
[0025] Based on the peak matching results, obtain the sample substance that matches the mixture to be tested in each peak matching;
[0026] The sample substances that match the test mixture in all peak matching processes are summarized to obtain the sample substance set that matches the test mixture.
[0027] According to the present invention, a method for identifying mixed substances based on infrared spectroscopy and peak matching is provided, wherein acquiring the infrared spectrum of the mixed substance to be tested and the infrared spectra of multiple sample substances includes:
[0028] Based on an infrared spectrometer, the original infrared spectrum of the mixture to be tested was acquired;
[0029] The original infrared spectrum of the target sample and the original sample infrared spectrum in the database are preprocessed.
[0030] The preprocessed original infrared spectrum to be tested is used as the infrared spectrum to be tested.
[0031] From the preprocessed original sample infrared spectrum in the database, extract the sample infrared spectrum of each sample substance;
[0032] The preprocessing includes filtering and standard state transformation.
[0033] The filtering process includes low-pass filtering and least-squares-based convolution fitting filtering.
[0034] According to the present invention, a method for identifying mixed substances based on infrared spectroscopy and peak matching is provided, wherein extracting the sample infrared spectrum of each sample substance from the preprocessed original sample infrared spectrum in the database includes:
[0035] Spectral clustering is performed on the preprocessed infrared spectra of the original samples in the database;
[0036] Based on the clustering results, determine the original infrared spectra of all pretreated samples for each sample material;
[0037] One pre-processed original infrared spectrum was randomly selected from all the pre-processed original infrared spectra of each sample substance and used as the sample infrared spectrum of each sample substance.
[0038] The present invention also provides a mixed substance identification device based on infrared spectroscopy and peak matching, comprising:
[0039] The data acquisition module is used to acquire the infrared spectrum of the mixture to be tested and the infrared spectrum of multiple sample substances.
[0040] The data processing module is used to obtain the set of characteristic peaks to be tested based on the infrared spectrum of the mixed substance to be tested, and to obtain the set of characteristic peaks of the samples based on the infrared spectrum of the multiple sample substances.
[0041] The substance matching module is used to iteratively perform peak removal matching on the target feature peaks in the target feature peak set and the sample feature peaks in the sample feature peak set;
[0042] The substance identification module is used to obtain a set of sample substances that match the mixture to be tested based on the peak matching result, and to obtain the component identification result of the mixture to be tested based on the composition of the set of sample substances that match the mixture to be tested.
[0043] The peak removal matching includes a matching operation and a peak removal operation. The matching operation is used to perform similarity matching between the test feature peaks in the test feature peak set corresponding to each peak removal matching and the sample feature peaks in the sample feature peak set corresponding to each peak removal matching. The peak removal operation is used to delete the test feature peak with the highest similarity obtained from each peak removal matching from the test feature peak set corresponding to each peak removal matching to obtain the test feature peak set corresponding to the next peak removal matching, and to delete all sample feature peaks of the sample substance corresponding to the sample feature peak with the highest similarity obtained from each peak removal matching from the sample feature peak set corresponding to each peak removal matching to obtain the sample feature peak set corresponding to the next peak removal matching.
[0044] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the mixed substance identification method based on infrared spectroscopy and peak matching as described above.
[0045] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the mixed substance identification method based on infrared spectroscopy and peak matching as described above.
[0046] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the mixed substance identification method based on infrared spectroscopy and peak matching as described above.
[0047] The present invention provides a method and apparatus for identifying mixed substances based on infrared spectroscopy and peak matching. By acquiring the infrared spectrum of the mixed substance to be tested and the infrared spectra of multiple sample substances, and iteratively performing peak matching on the characteristic peaks of the infrared spectrum of the mixed substance to be tested and the characteristic peaks of the infrared spectra of multiple sample substances, the components of the mixed substance to be tested can be automatically and accurately identified based on the peak matching results. This effectively improves the efficiency of substance identification and solves the problems of low identification efficiency, large consumption of manpower and resources, and large identification errors in the existing technology that relies on human experience for substance identification. Attached Figure Description
[0048] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0049] Figure 1 This is a flowchart illustrating the mixed substance identification method based on infrared spectroscopy and peak matching provided by the present invention.
[0050] Figure 2 This is a schematic diagram of the distribution of the infrared absorption spectrum of octene provided by the present invention;
[0051] Figure 3 This is a schematic diagram showing the distribution of the comparison results between the infrared spectrum of the synthesized sample and the infrared spectrum of the mixture to be tested provided by the present invention.
[0052] Figure 4 This is a schematic diagram showing the distribution of the linear interpolation results provided by the present invention;
[0053] Figure 5 This is a schematic diagram of the distribution of the infrared spectrum of the original sample provided by the present invention;
[0054] Figure 6 This is a schematic diagram showing the distribution of the Savitzky Golay filtering results provided by the present invention;
[0055] Figure 7 This is a schematic diagram showing the distribution of the low-pass filtering results provided by the present invention;
[0056] Figure 8 This is a schematic diagram of the distribution of the infrared spectrum of the preprocessed original sample provided by the present invention;
[0057] Figure 9 This is a schematic diagram illustrating the distribution of the spectral clustering results provided by the present invention;
[0058] Figure 10 This is a schematic diagram of the structure of the mixed substance identification device based on infrared spectroscopy and peak matching provided by the present invention;
[0059] Figure 11 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0060] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0061] Typically, testing and inspection require high technical expertise and involve the integration of multiple disciplines such as chemistry, physics, materials science, electronics, biology, and food science.
[0062] In existing technologies, the identification of mixed substances mainly relies on the human experience of engineers. These engineers require up to six months of professional training, followed by a training assessment and test. Even after passing the assessment, they still need extensive training before being allowed to work. This training cycle is lengthy and consumes significant human and material resources. Secondly, in practice, the software can only provide results for identifying single components. Therefore, for mixed substances, repeated manual operations and manual confirmation and adjustments based on experience are necessary. Furthermore, when there are significant differences in the composition of the substances, multiple detection devices are required. However, the data and software between different devices are not interoperable, requiring manual switching between multiple systems. This results in engineers spending a considerable amount of time on image capture and report writing, failing to effectively improve detection efficiency and increasing the risk of errors.
[0063] Therefore, traditional human experience in identifying mixed substances is not only inaccurate and inefficient, but also wastes a lot of engineers' time.
[0064] To address the aforementioned issues, this embodiment provides a method for identifying mixed substances based on infrared spectroscopy and peak matching. This method iteratively performs peak matching on the infrared spectrum of the mixed substance to be tested and the infrared spectra of multiple sample substances to automatically and accurately identify the original components that make up the mixed substance. Only one server is needed to identify each original component of the mixed substance, effectively solving the problems of data incompatibility between various systems and low identification accuracy caused by manual experience in identifying mixed substances. This improves the reliability of mixed substance identification and can also automatically generate reports, saving a lot of manpower, material resources and identification time, and improving identification efficiency.
[0065] The following is combined Figures 1-9This invention describes a method for identifying mixed substances based on infrared spectroscopy and peak matching. The method can be implemented by an electronic device equipped with mixed substance identification capabilities. This electronic device has a complete spectral identification process internally, enabling the identification of components in unknown mixtures. This electronic device can be a terminal, such as a mobile phone or computer, or a server, such as an edge server or cloud server; this embodiment does not specifically limit its application.
[0066] like Figure 1 The diagram shown is one of the flowcharts of the mixed substance identification method based on infrared spectroscopy and peak matching provided in this embodiment. The method includes the following steps:
[0067] Step 101: Obtain the infrared spectrum of the mixture to be tested and the infrared spectrum of multiple sample substances;
[0068] The mixture to be tested can be a substance that requires component identification in fields such as chemistry, materials, electronics, biology, and food. This mixture contains a variety of different substances.
[0069] The sample substance is a substance pre-stored in a database and labeled with its component category. The sample substance can be a single substance.
[0070] Optionally, since the atoms that make up chemical bonds or functional groups in organic molecules are in a state of constant vibration, and their vibrational frequencies are comparable to those of infrared light, when organic molecules are irradiated with infrared light, the chemical bonds or functional groups in the molecules will undergo vibrational absorption. Different chemical bonds or functional groups absorb at different frequencies and will be located at different positions on the infrared spectrum, thus providing information about the types of chemical bonds or functional groups present in the molecule. Therefore, analyzing the infrared spectrum of a substance can accurately identify its composition.
[0071] Optionally, when it is necessary to identify the components of the mixture to be tested, an infrared spectrometer can be used to acquire the infrared spectrum of the mixture to be tested. The acquired infrared spectrum can be directly used as the infrared spectrum of the mixture to be tested. Alternatively, the acquired infrared spectrum can be preprocessed by filtering and / or state transformation, and then the infrared spectrum of the mixture to be tested can be obtained based on the preprocessing results. This improves the data cleaning work in the early stage, thereby improving the identification efficiency and matching accuracy in the later stage. This embodiment does not specifically limit this.
[0072] Similarly, at least one initial infrared spectrum of multiple sample substances can be randomly extracted from the database as the sample infrared spectrum of multiple sample substances; or the initial infrared spectrum of multiple sample substances randomly extracted from the database can be preprocessed, and then the sample infrared spectrum of multiple sample substances can be obtained based on the preprocessing result. This embodiment does not specifically limit this.
[0073] Step 102: Obtain the set of characteristic peaks to be tested based on the infrared spectrum of the mixed substance to be tested, and obtain the set of characteristic peaks of the samples based on the infrared spectrum of the multiple sample substances.
[0074] Characteristic peaks or characteristic frequencies refer to absorption peaks used to identify the presence of chemical bonds or groups. The infrared spectrum of a compound is an objective reflection of its molecular structure. The absorption peaks in the spectrum correspond to the vibrational modes of a certain chemical bond or group in the molecule, and the vibrational frequencies of the same group always appear in a certain region.
[0075] Optionally, after obtaining the infrared spectrum of the mixture to be tested and the infrared spectra of multiple sample substances, characteristic peaks can be extracted from the infrared spectrum of the mixture to be tested to construct a set of characteristic peaks to be tested.
[0076] Furthermore, characteristic peaks were extracted from the infrared spectra of multiple sample substances to construct a sample characteristic peak set.
[0077] Step 103 iteratively performs peak removal matching on the target feature peaks in the target feature peak set and the sample feature peaks in the sample feature peak set; wherein, the peak removal matching includes a matching operation and a peak removal operation; the matching operation is used to perform similarity matching on the target feature peaks in the target feature peak set and the sample feature peaks in the sample feature peak set corresponding to each peak removal matching; the peak removal operation is used to delete the target feature peak with the highest similarity obtained from each peak removal matching from the target feature peak set corresponding to each peak removal matching to obtain the target feature peak set corresponding to the next peak removal matching, and delete all sample feature peaks of the sample substance corresponding to the sample feature peak with the highest similarity obtained from each peak removal matching from the sample feature peak set corresponding to each peak removal matching to obtain the sample feature peak set corresponding to the next peak removal matching;
[0078] Since different infrared spectra have different characteristic peaks, the characteristic peaks of the basic sample that makes up the mixture will all be reflected in the infrared spectrum of the mixture to be tested. Therefore, by utilizing the principle of peak similarity, peak matching can be performed iteratively to find the original sample (i.e., the sample substance) that is highly similar to the mixture to be tested based on the identification of characteristic peaks, thereby accurately identifying the synthetic components of the mixture to be tested.
[0079] For example, such as Figure 2 The image shows the infrared absorption spectrum of octene, which contains four characteristic peaks: one at wavenumbers of 3080, one at 1640, one at 995, and one at 915. These are all characteristic peaks of octene. Therefore, different substances have different characteristic peaks, but the characteristic peaks of the original sample will be present in the mixture. Based on this pattern, peak matching can be used to find the original components that make up the mixture.
[0080] Optionally, after obtaining the set of characteristic peaks to be tested and the set of sample characteristic peaks, it is necessary to iteratively perform peak matching on the characteristic peaks to be tested in the set of characteristic peaks to be tested and the sample characteristic peaks in the set of sample characteristic peaks. In order to utilize the correlation between characteristic peaks in different infrared spectra, peak matching operations are continuously performed to accurately find the original sample that makes up the sample to be tested.
[0081] Optionally, peak removal matching is iteratively performed on the target feature peaks in the target feature peak set and the sample feature peaks in the sample feature peak set. A matching operation is used to perform similarity matching between the target feature peaks in the target feature peak set and the sample feature peaks in the sample feature peak set corresponding to each peak removal match. The sample substance corresponding to the sample feature peak with the highest similarity in each peak removal match is obtained as the sample substance matched with the target mixture in each peak removal match. A peak removal operation is then performed to delete the target feature peak with the highest similarity in each peak removal match from the target feature peak set corresponding to each peak removal match, updating the target feature peak set. Similarly, all sample feature peaks of the sample substance corresponding to the sample feature peak with the highest similarity in each peak removal match are deleted from the sample feature peak set corresponding to each peak removal match, updating the sample feature peak set. Based on the updated sample feature peak set and the updated target feature peak set, peak removal matching is iteratively performed until the peak removal matching termination condition is met, thus obtaining a set of all sample substances matched with the target mixture.
[0082] Step 104: Based on the peak matching results, obtain a set of sample substances that match the mixture to be tested; based on the composition of the set of sample substances that match the mixture to be tested, obtain the composition identification result of the mixture to be tested.
[0083] Optionally, each peak-removal matching result contains the sample substances that match the mixture to be tested in each peak-removal matching. Based on all the sample substances that match the mixture to be tested in all peak-removal matching, a set of sample substances that match the mixture to be tested can be constructed, which is the original sample that makes up the mixture to be tested.
[0084] Since the composition of each sample in the sample set is known, the composition of each sample can be summarized to obtain the composition identification result of the mixture to be tested.
[0085] The mixed substance identification method based on infrared spectroscopy and peak matching provided in this embodiment acquires the infrared spectrum of the mixed substance to be tested and the infrared spectra of multiple sample substances. Iteratively performs peak matching on the characteristic peaks of the infrared spectrum of the mixed substance to be tested and the characteristic peaks of the infrared spectra of multiple sample substances. Based on the peak matching results, the composition of the mixed substance to be tested can be automatically and accurately identified, which effectively improves the substance identification efficiency and solves the problems of low identification efficiency, large consumption of manpower and material resources, and large identification errors in the existing technology that relies on human experience for substance identification.
[0086] In some embodiments, step 102, which iteratively performs peak matching on the target feature peaks in the target feature peak set and the sample feature peaks in the sample feature peak set, includes:
[0087] For the current peak removal matching, each test feature peak in the set of test feature peaks corresponding to the current peak removal matching is matched with each sample feature peak in the set of sample feature peaks corresponding to the current peak removal matching.
[0088] The sample substance to which the feature peak of the most similar sample is matched is taken as the sample substance that matches the mixture to be tested in the current peak removal matching.
[0089] The infrared spectrum of the sample substance that matches the mixed substance to be tested in the current peak removal matching is superimposed with the combined spectrum corresponding to the previous peak removal matching to obtain the combined spectrum corresponding to the current peak removal matching.
[0090] Calculate the similarity between the combined spectrum corresponding to the current peak removal matching and the infrared spectrum to be measured, and obtain the spectrum similarity corresponding to the current peak removal matching;
[0091] If the spectral similarity corresponding to the current peak removal match is greater than the spectral similarity corresponding to the previous peak removal match, the test feature peak with the highest similarity is deleted from the test feature peak set corresponding to the current peak removal match, and the test feature peak set corresponding to the next peak removal match is obtained.
[0092] The sample feature peak with the highest similarity is removed from the sample feature peak set corresponding to the current peak removal match, and the sample feature peak set corresponding to the next peak removal match is obtained.
[0093] The peak removal matching step is iteratively performed by matching each test feature peak in the set of test feature peaks corresponding to the next peak removal matching with each sample feature peak in the set of sample feature peaks corresponding to the next peak removal matching, until the peak removal matching termination condition is met.
[0094] The termination condition for peak removal matching includes the condition that the spectral similarity corresponding to the current peak removal matching is less than or equal to the spectral similarity corresponding to the previous peak removal matching.
[0095] Optionally, the iterative peak-matching operation in step 102 includes:
[0096] For the current peak removal matching, based on the distribution characteristics of each feature peak, the similarity between each test feature peak in the test feature peak set corresponding to the current peak removal matching and each sample feature peak in the sample feature peak set corresponding to the current peak removal matching is calculated; wherein, the distribution characteristics include, but are not limited to, the trend of change and the location information, which are not specifically limited in this embodiment.
[0097] After obtaining the similarity between each target feature peak and each sample feature peak corresponding to the current peak matching, sort them and obtain the sample feature peak and the target feature peak with the highest similarity in the current matching.
[0098] The sample substance to which the sample with the highest similarity in the current matching belongs is taken as the sample substance that matches the mixture to be tested in the current peak removal matching, that is, the original state of the mixture to be tested.
[0099] Then, the infrared spectrum of the sample substance that matches the mixture to be tested in the current peak removal matching is superimposed with the combined spectrum corresponding to the previous peak removal matching (i.e., the superimposed spectrum of the infrared spectra of the sample substances that match the mixture to be tested in all historical peak removal matchings before the current peak removal matching) to obtain the combined spectrum corresponding to the current peak removal matching.
[0100] If the similarity between the combined spectrum corresponding to the current peak-removal matching and the infrared spectrum to be measured no longer increases or even shows a decreasing trend compared to the similarity between the combined spectrum corresponding to the previous peak-removal matching and the infrared spectrum to be measured, it indicates that the peak-removal matching termination condition has been met, and the peak-removal matching operation is stopped.
[0101] If the similarity between the combined spectrum corresponding to the current peak-removal matching and the infrared spectrum to be measured continues to increase compared to the similarity between the combined spectrum corresponding to the previous peak-removal matching and the infrared spectrum to be measured, the characterization does not meet the termination condition for peak-removal matching, and iterative peak-removal matching needs to continue. At this point, peak removal operations need to be performed on both the set of target feature peaks and the set of sample feature peaks corresponding to the current peak-removal matching. Specifically, the sample feature peak with the highest similarity is deleted from the set of sample feature peaks corresponding to the current peak-removal matching, and the target feature peak with the highest similarity is also deleted from the set of target feature peaks corresponding to the current peak-removal matching. These updates are then used to obtain the set of sample feature peaks and the set of target feature peaks corresponding to the next peak-removal matching.
[0102] The peak removal matching process continues iteratively based on the set of characteristic peaks to be measured and the set of characteristic peaks of the sample corresponding to the next peak removal matching, until the similarity between the resulting combined spectrum and the infrared spectrum to be measured no longer increases. At this point, the peak removal matching stops, and the components of the mixture to be measured are identified based on the peak removal matching results.
[0103] After all peak matching is completed, the infrared spectra of the sample set that matches the target mixture during peak matching can be synthesized to obtain the synthetic infrared spectrum of the synthetic sample of the target mixture. For example... Figure 3 As shown, the characteristic peak positions and characteristic peak variation trends of the synthetic infrared spectrum of the synthetic sample of the mixed substance to be tested are basically consistent with the characteristic peak positions and characteristic peak variation trends of the infrared spectrum of the mixed substance to be tested (i.e., the analytical sample). After multiple verifications, the identification accuracy can reach more than 90%.
[0104] In some embodiments, the step of performing similarity matching between each test feature peak in the set of test feature peaks corresponding to the current peak removal matching and each sample feature peak in the set of sample feature peaks corresponding to the current peak removal matching includes:
[0105] Based on the Jaccard coefficient, a similarity match is performed between each test feature peak in the set of test feature peaks corresponding to the current peak removal match and each sample feature peak in the set of sample feature peaks corresponding to the current peak removal match.
[0106] The Jaccard coefficient, also known as the Jaccard similarity coefficient, is a measure of the degree of similarity between different characteristic peaks.
[0107] The Jaccard coefficient is used to compare the differences and similarities between two characteristic peaks. A higher Jaccard coefficient indicates a greater similarity between the two characteristic peaks. The formula for calculating the Jaccard coefficient is:
[0108]
[0109] Where J(A,B) is the Jaccard coefficient between the target feature peak A and the sample feature peak B; |A∩B| is the wavenumber intersection of the same peak values in the target feature peak A and the sample feature peak B; |A∪B| is the wavenumber union of all peak values in the target feature peak A and the sample feature peak B. |A| and |B| are the wavenumbers of the peak values in the target feature peak A and the sample feature peak B, respectively. It should be noted that when both A and B are empty, J(A,B)=1.
[0110] Optionally, based on the Jaccard coefficient, the similarity between each test feature peak in the set of test feature peaks corresponding to the current peak removal matching and each sample feature peak in the set of sample feature peaks corresponding to the current peak removal matching is calculated. Based on the similarity, the test feature peaks and sample feature peaks are accurately matched, thereby achieving efficient and accurate identification of mixed substances.
[0111] In some embodiments, the step of superimposing the infrared spectrum of the sample substance that matches the test mixture in the current peak-reduction matching with the combined spectrum corresponding to the previous peak-reduction matching to obtain the combined spectrum corresponding to the current peak-reduction matching includes:
[0112] Linear interpolation is performed on the infrared spectra of the sample substances that match the mixed substance to be tested in the current peak-removal matching and the combined spectra corresponding to the previous peak-removal matching, so that the data format of the infrared spectra of the samples after linear interpolation is consistent with the data format of the combined spectra after linear interpolation.
[0113] The sample infrared spectrum after linear interpolation is linearly superimposed with the combined spectrum after linear interpolation to obtain the combined spectrum corresponding to the current peak matching.
[0114] Optionally, because different infrared spectrometers are used to collect infrared spectra of different sample substances, and different infrared spectrometers have different acquisition strategies, the data lengths of the infrared spectra of different sample substances vary. Therefore, when superimposing the infrared spectra of the samples, inaccurate superposition results or even failure to superimpose may occur, which seriously restricts the process of substance identification.
[0115] To address this issue, this embodiment employs linear interpolation to downsample or oversample the infrared spectra of the sample substances matched with the target mixture in the current peak-removal matching and the combined spectra corresponding to the previous peak-removal matching. This ensures that the data format of the sample infrared spectra after linear interpolation is consistent with that of the combined spectra after linear interpolation, i.e., the data length is consistent, thus resolving the problem of inconsistent data overlay formats.
[0116] Interpolation, in particular, involves determining the variation patterns of a known data sequence (i.e., the data sequence composed of infrared spectral signals within each time window of the sample's infrared spectrum), and then estimating the values of points for which no data has yet been recorded, based on these variation patterns. It is primarily used to reasonably compensate for missing data and to amplify or reduce the size of data.
[0117] For example, if the values of some infrared spectral signals in a data sequence composed of infrared spectral signals within a certain time window in the sample infrared spectrum are known, i.e., their coordinates (x0, y0) and (x1, y1) are known, while the values of some infrared spectral signals are missing, i.e., their coordinates (x, y) are unknown, specifically as follows: Figure 4 As shown in the figure. Here, x is the sampling time point of the infrared spectral signal, and y is the corresponding value of the infrared spectral signal. At this time, it is necessary to perform numerical estimation of (x,y) based on linear interpolation to achieve reasonable compensation for missing data, and thus unify the data format of the sample infrared spectra to be superimposed.
[0118] The formula for estimating the coordinates (x, y) of the numerically missing infrared spectral signal based on linear interpolation is as follows:
[0119]
[0120] Since x is the sampling time and its value is known, the value of y can be obtained by solving the above formula.
[0121] Linear interpolation is an interpolation method for one-dimensional data. It estimates the value of the missing infrared spectral signal by taking the two nearest neighboring data points (i.e., infrared spectral signals with known values) from the point to be interpolated (i.e., the missing infrared spectral signal). Specifically, it determines the weighting coefficient based on the distance between the two nearest neighboring data points, and then weights and sums the values of the two nearest neighboring data points according to the weighting coefficient to obtain the value of the missing infrared spectral signal. This ensures that the data format of the infrared spectra of each sample after linear interpolation is uniform, i.e., the data length is consistent, thus solving the problem of inconsistent data overlay formats.
[0122] Then, the infrared spectrum of the sample after linear interpolation is linearly superimposed with the combined spectrum corresponding to the previous peak-matching after linear interpolation, so as to accurately obtain the combined spectrum corresponding to the current peak-matching, thereby achieving efficient and accurate identification of mixed substances.
[0123] In some embodiments, step 104, which involves obtaining a set of sample substances that match the mixture to be tested based on the peak matching result, further includes:
[0124] Based on the peak matching results, obtain the sample substance that matches the mixture to be tested in each peak matching;
[0125] The sample substances that match the test mixture in all peak matching processes are summarized to obtain the sample substance set that matches the test mixture.
[0126] Optionally, during each peak matching process, a sample substance that matches the mixture to be tested in each peak matching process is obtained, and this sample substance is the original sample that makes up the mixture to be tested.
[0127] Therefore, all sample substances that match the mixture to be tested obtained during the peak matching process can be summarized to obtain a set of sample substances that match the mixture to be tested, i.e., the synthetic sample of the mixture to be tested; based on the material composition corresponding to the set of sample substances, the material composition of the mixture to be tested can be accurately obtained.
[0128] In some embodiments, the step of obtaining the infrared spectrum of the mixture to be tested and the infrared spectrum of multiple sample substances in step 101 further includes:
[0129] Based on an infrared spectrometer, the original infrared spectrum of the mixture to be tested was acquired;
[0130] The original infrared spectrum of the target sample and the original sample infrared spectrum in the database are preprocessed.
[0131] The preprocessed original infrared spectrum to be tested is used as the infrared spectrum to be tested.
[0132] From the preprocessed original sample infrared spectrum in the database, extract the sample infrared spectrum of each sample substance;
[0133] The preprocessing includes filtering and standard state transformation.
[0134] The filtering process includes low-pass filtering and least-squares-based convolution fitting filtering.
[0135] The infrared spectrometer can be installed on the electronic device or externally placed on the electronic device and communicate with the electronic device. This embodiment does not specifically limit the location of the infrared spectrometer.
[0136] The original infrared spectra of the sample substances in the database were acquired by an infrared spectrometer and stored after being pre-labeled with component categories;
[0137] Optionally, the infrared spectrometer acquires infrared spectra of the mixture to be tested at a preset sampling frequency to obtain the original infrared spectra of the mixture; and retrieves the original sample infrared spectra from the database.
[0138] Because the raw infrared spectra acquired by the infrared spectrometer (i.e., the raw infrared spectrum of the target and the raw infrared spectrum of the sample) are subject to noise and baseline shift due to environmental and equipment factors, they can significantly affect the accuracy and efficiency of subsequent material matching and identification. Therefore, in the early stage of material matching and identification, it is necessary to use filters for data preprocessing to denoise the spectrum and perform standard state transformation to eliminate the influence of baseline shift, thereby greatly improving the efficiency and accuracy of subsequent material identification.
[0139] Optionally, in order to eliminate the influence of noise on substance identification, filtering can be used to preprocess the original infrared spectrum of the target and the original sample to eliminate noise in the original infrared spectrum of the target and the original sample.
[0140] Filtering processes include, but are not limited to, low-pass filtering and filtering based on local polynomial least squares fitting of least squares convolution fitting filters (Savitzky-Golay filters).
[0141] like Figure 5 As shown, the original infrared spectrum contains a lot of noise, interfering with peak matching. Therefore, the Savitzky-Golay filter is used to solve the noise interference problem. It is widely used for data stream smoothing and noise reduction and is a filtering method based on local polynomial least squares fitting in the time domain. The biggest feature of this filter is that it can remove noise while ensuring that the shape and width of the signal remain unchanged.
[0142] like Figure 6 As shown, the Savitzky-Golay filter is a digital filter that can be applied to a set of data to smooth it, improving data accuracy without altering the signal's trend or width. It is implemented through a convolution process, specifically by fitting a continuous subset of adjacent data points to a low-order polynomial using linear least squares.
[0143] The Savitzky-Golay filtering formula is as follows: (The formula is missing from the provided text.)
[0144]
[0145] Among them, X k,smooth The filtered infrared spectral signal of the k-th target within the time window; [x k-w ,…,x k+w [h] represents the infrared spectral signals of all targets within the time window; i / H is the smoothing coefficient, obtained by fitting the polynomial using the least squares method.
[0146] The Savitzky-Golay filter has the advantage that different window widths can be arbitrarily selected at any position on the same curve to meet different smoothing filtering needs; it is particularly advantageous when processing time-series data, especially for processing sequences at different stages. It also performs well in processing noise samples from aperiodic and nonlinear sources.
[0147] Low-pass filtering weakens or blocks high-frequency signals while retaining low-frequency signals. In spectral analysis, some complex samples are simultaneously denoised using low-pass filtering. As a commonly used filter type, low-pass filtering is a moderate filtering method. Figure 7 As shown, the signal waveform after using a first-order low-pass filter exhibits a significantly stronger trend than the signal waveform before using the first-order low-pass filter, thus achieving effective noise reduction. Compared to the Kalman filter algorithm and moving average filter, its computational cost is moderate, while still yielding a suitable result. The low-pass filter algorithm can effectively address the issue of sensors that are reliable in the long term but have high short-term noise levels, effectively filtering out noise.
[0148] The calculation formula for low-pass filtering is as follows:
[0149]
[0150] Where D0 represents the passband radius, u and v are the frequency and amplitude of the signal at that frequency in the original infrared spectrum, respectively, and D(u,v) is the distance from the original infrared spectrum to the center of the spectrum, calculated as follows:
[0151]
[0152] Where M and N represent the horizontal and vertical coordinates of the spectrum image, and (M / 2, N / 2) is the center of the spectrum.
[0153] Furthermore, due to differences in sampling equipment and environmental factors, spectral shifts and deviations can be eliminated after filtering by using Standard Normal Variation (SNV). SNV primarily eliminates the effects of solid particle size, surface scattering, and optical path variations on NIR (Near Infrared) spectroscopy. The process involves processing a single spectrum, specifically based on the rows of the spectral matrix. The formula for Standard Normal Variation is as follows:
[0154]
[0155] In the formula, X i , SNV x is the result of the standard normal transformation of the i-th infrared spectral signal; i The value is the average of the infrared spectral signal, m is the number of wavelength points, and x is the average value. k For each sample, K = 1, 2, 3, ..., m.
[0156] like Figure 8 As shown, the preprocessed optical spectrum is compared to the original spectrum. Figure 5 The resulting spectrum is smoother, and the use of SNV effectively eliminates the influence of the baseline, greatly improving the efficiency of subsequent peak matching and the accuracy of substance identification.
[0157] In this embodiment, Savitzky-Golay filtering, low-pass filtering, and standard normal transformation are used to preprocess the original infrared spectrum acquired by the infrared spectrometer. This effectively eliminates noise and baseline shift in the original infrared spectrum, greatly improving the efficiency of subsequent peak matching and the accuracy and efficiency of substance identification.
[0158] In some embodiments, the step of extracting the sample infrared spectrum of each sample substance from the preprocessed original sample infrared spectrum in the database in step 101 further includes:
[0159] Spectral clustering is performed on the preprocessed infrared spectra of the original samples in the database;
[0160] Based on the clustering results, determine the original infrared spectra of all pretreated samples for each sample material;
[0161] One pre-processed original infrared spectrum was randomly selected from all the pre-processed original infrared spectra of each sample substance and used as the sample infrared spectrum of each sample substance.
[0162] Optionally, after preprocessing the original infrared spectra of the samples in the database, since there are multiple infrared spectra of the same sample substance in the database, repeated matching is required during the peak removal matching process, increasing time consumption and affecting the substance identification efficiency. In order to reduce redundant calculations in the peak removal matching process, a spectral clustering method is used to aggregate the infrared spectra of all sample substances of the same type, so as to extract the infrared spectra of each sample substance based on the aggregation result, thereby improving the substance identification efficiency.
[0163] Spectral clustering only requires solving the similarity matrix between the infrared spectra of samples, making it very effective for clustering sparse data. This is something traditional clustering algorithms, such as K-means clustering, struggle with. Furthermore, it uses dimensionality reduction, making it better than traditional clustering algorithms when handling high-dimensional data.
[0164] Spectral clustering is an algorithm derived from graph theory and is widely used in clustering. The main idea is to treat all data as points in a space, connected by edges. Edges between points that are far apart have lower weights, while edges between points that are close have higher weights. Then, by slicing the graph formed by all data points, the goal is to minimize the sum of edge weights between different subgraphs and maximize the sum of edge weights within each subgraph, thereby achieving the purpose of clustering.
[0165] Spectral clustering mainly consists of two steps: the first step is graph construction, which involves constructing a network graph from the sampled data; the second step is graph slicing, which involves dividing the graph constructed in the first step into different graphs according to certain slicing criteria. These different subgraphs are the corresponding clustering results.
[0166] In the graph construction process, we first need to obtain the adjacency matrix. Currently, there are three main methods: ∈-nearest neighbor, K-nearest neighbor, and fully connected method; the fully connected method is the most commonly used. Since the weights between all points are greater than 0, it is called the fully connected method. This embodiment uses different kernel functions to define the edge weights. Commonly used kernel functions include polynomial kernel functions, Gaussian kernel functions, and Sigmoid kernel functions. The most commonly used is the Gaussian kernel function, such as the RBF (Radial Basis Function Kernel). In this case, the similarity matrix and the adjacency matrix are the same, which can be represented as follows:
[0167]
[0168] Where, x i x j Given two vector samples, W ij For the reconstructed adjacency matrix, S ij Let σ be a similar matrix. 2 The bandwidth controls the radial range of action.
[0169] Then, calculate the Laplace matrix L, defined as L = DW. D is the degree matrix, a diagonal matrix, and W is the adjacency matrix. 1) The Laplace matrix is a symmetric matrix, which can be deduced from the fact that both D and W are symmetric matrices. 2) Since the Laplace matrix is a symmetric matrix, all its eigenvalues are real numbers. 3) For any vector f:
[0170]
[0171] Where f is any vector and n is the number of vector samples.
[0172] After the graph is constructed, slicing is performed, with Ncut being the most common method. Ncut, in addition to minimizing the loss function, also considers the weights between subgraphs. Since a larger number of subgraph samples does not necessarily mean a larger weight, slicing based on weights is more in line with the clustering optimization objective, resulting in more accurate clustering results. The goal of Ncut is to minimize the sum of the edges connecting each subgraph; the specific calculation formula is as follows:
[0173]
[0174] Among them, A i Let i be the set of points contained in the i-th subgraph. For A i The complement of, that is, excluding subset A i The union of all other subsets except for k, where k is the number of subgraphs in the partition, vol(A) i () is subgraph A i The volume is calculated using the following formula:
[0175] vol(A i )=∑ i∈A d i ;
[0176] Where, d i Let A be the edge weights of the subgraph.
[0177] Using subgraph weights in Ncut Let h represent the indicator vector, defined as follows:
[0178]
[0179] Among them, h ji Let be the indicator vector, where i represents the sample subscript, j represents the subset subscript, represents the indicator of sample i to subset j, and vi is the set of the i-th points.
[0180] Accordingly, the optimization objective can be further characterized as:
[0181]
[0182] Where H is the indicator matrix and Tr() is the trace of the matrix.
[0183] Due to H T H≠I, and H T Therefore, DH = I, and the optimization objective can be further derived as follows:
[0184]
[0185] In summary, the optimization objective can ultimately be simplified to:
[0186] min T∈R Tr(T T D -1 / 2 LD -1 / 2 T);
[0187] The constraint condition is T Y T = I;
[0188] In the spectral clustering process, it is necessary to calculate D. -1 / 2 LD -1 / 2 The K smallest eigenvalues are obtained, the corresponding eigenvectors are calculated and standardized, and finally the eigenma is obtained, which is then used to perform clustering.
[0189] It should be noted that in spectral matching, the peak points of all sample substances can be accurately identified. The top five largest peaks among all peak points of each sample substance are taken as features for spectral clustering. Since the infrared spectra of the same sample substance are identical, the peak points are also the same. Through clustering, the infrared spectra of all samples of the same sample substance can be accurately identified.
[0190] like Figure 9 As shown, this demonstrates the effect of using spectral clustering to cluster similar substances. It can be seen that the substance identification method in this embodiment can accurately identify the infrared spectra of all samples of the same substance.
[0191] In this embodiment, spectral clustering is performed on the preprocessed original sample infrared spectra in the database to aggregate all spectra of the same sample substance together, thereby accurately extracting one sample infrared spectrum for each sample substance, avoiding redundancy, effectively reducing the time required for repeated peak matching of the same sample infrared spectrum, and improving the efficiency of substance identification.
[0192] The mixed substance identification device based on infrared spectroscopy and peak matching provided by the present invention will be described below. The mixed substance identification device based on infrared spectroscopy and peak matching described below can be referred to in correspondence with the mixed substance identification method based on infrared spectroscopy and peak matching described above.
[0193] like Figure 10 This embodiment provides a mixed substance identification device based on infrared spectroscopy and peak matching. The device includes a data acquisition module 1001, a data processing module 1002, a substance matching module 1003, and a substance identification module 1004, wherein:
[0194] The data acquisition module 1001 is used to acquire the infrared spectrum of the mixture to be tested and the infrared spectrum of multiple sample substances;
[0195] The data processing module 1002 is used to obtain the set of characteristic peaks to be tested based on the infrared spectrum of the mixed substance to be tested, and to obtain the set of characteristic peaks of the samples based on the infrared spectrum of the multiple sample substances.
[0196] The substance matching module 1003 is used to iteratively perform peak removal matching on the target feature peaks in the target feature peak set and the sample feature peaks in the sample feature peak set;
[0197] The substance identification module 1004 is used to obtain a set of sample substances that match the mixed substance to be tested based on the peak matching result, and to obtain the component identification result of the mixed substance to be tested based on the composition of the set of sample substances that match the mixed substance to be tested.
[0198] The peak removal matching includes a matching operation and a peak removal operation. The matching operation is used to perform similarity matching between the test feature peaks in the test feature peak set corresponding to each peak removal matching and the sample feature peaks in the sample feature peak set corresponding to each peak removal matching. The peak removal operation is used to delete the test feature peak with the highest similarity obtained from each peak removal matching from the test feature peak set corresponding to each peak removal matching to obtain the test feature peak set corresponding to the next peak removal matching, and to delete all sample feature peaks of the sample substance corresponding to the sample feature peak with the highest similarity obtained from each peak removal matching from the sample feature peak set corresponding to each peak removal matching to obtain the sample feature peak set corresponding to the next peak removal matching.
[0199] The mixed substance identification device based on infrared spectroscopy and peak matching provided in this embodiment acquires the infrared spectrum of the mixed substance to be tested and the infrared spectra of multiple sample substances. Iteratively performs peak matching on the characteristic peaks of the infrared spectrum of the mixed substance to be tested and the characteristic peaks of the infrared spectra of multiple sample substances. Based on the peak matching results, the components of the mixed substance to be tested are automatically and accurately identified, which effectively improves the substance identification efficiency and solves the problems of low identification efficiency, large consumption of manpower and material resources, and large identification errors in the existing technology that relies on human experience for substance identification.
[0200] Figure 11An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 11 As shown, the electronic device may include: a processor 1101, a communication interface 1102, a memory 1103, and a communication bus 1104, wherein the processor 1101, the communication interface 1102, and the memory 1103 communicate with each other through the communication bus 1104. The processor 1101 can call logical instructions in the memory 1103 to execute a mixed substance identification method based on infrared spectroscopy and peak matching. This method includes: acquiring the infrared spectrum of the mixed substance to be tested and the infrared spectra of multiple sample substances; acquiring a set of characteristic peaks to be tested based on the infrared spectrum of the mixed substance to be tested, and acquiring a set of sample characteristic peaks based on the infrared spectra of the multiple sample substances; iteratively performing peak matching on the characteristic peaks to be tested in the set of characteristic peaks to be tested and the sample characteristic peaks in the set of sample characteristic peaks; acquiring a set of sample substances that match the mixed substance to be tested based on the peak matching result; and acquiring the mixed substance to be tested based on the composition of the set of sample substances that match the mixed substance to be tested. The component identification result; wherein, the peak removal matching includes a matching operation and a peak removal operation; the matching operation is used to perform similarity matching on the test feature peaks in the test feature peak set corresponding to each peak removal matching and the sample feature peaks in the sample feature peak set corresponding to each peak removal matching; the peak removal operation is used to delete the test feature peak with the highest similarity obtained from each peak removal matching from the test feature peak set corresponding to each peak removal matching to obtain the test feature peak set corresponding to the next peak removal matching, and delete all sample feature peaks of the sample substance corresponding to the sample feature peak with the highest similarity obtained from each peak removal matching from the sample feature peak set corresponding to each peak removal matching to obtain the sample feature peak set corresponding to the next peak removal matching.
[0201] Furthermore, the logical instructions in the aforementioned memory 1103 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0202] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the mixed substance identification method based on infrared spectroscopy and peak matching provided by the above methods. The method includes: acquiring the infrared spectrum of the mixed substance to be tested and the infrared spectra of multiple sample substances; acquiring a set of characteristic peaks to be tested based on the infrared spectrum of the mixed substance to be tested, and acquiring a set of sample characteristic peaks based on the infrared spectra of the multiple sample substances; iteratively performing peak matching on the characteristic peaks to be tested in the set of characteristic peaks to be tested and the sample characteristic peaks in the set of sample characteristic peaks; and acquiring a set of sample substances that match the mixed substance to be tested based on the peak matching result. The composition of the sample substance set matched with the mixture to be tested is obtained to acquire the composition identification result of the mixture to be tested; wherein, the peak removal matching includes a matching operation and a peak removal operation; the matching operation is used to perform similarity matching on the test feature peak in the test feature peak set corresponding to each peak removal matching and the sample feature peak in the sample feature peak set corresponding to each peak removal matching; the peak removal operation is used to delete the test feature peak with the highest similarity obtained from each peak removal matching from the test feature peak set corresponding to each peak removal matching to obtain the test feature peak set corresponding to the next peak removal matching, and delete all sample feature peaks of the sample substance corresponding to the sample feature peak with the highest similarity obtained from each peak removal matching from the sample feature peak set corresponding to each peak removal matching to obtain the sample feature peak set corresponding to the next peak removal matching.
[0203] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program performs the mixed substance identification method based on infrared spectroscopy and peak-removal matching provided by the methods described above. This method includes: acquiring a test infrared spectrum of the mixed substance to be tested and sample infrared spectra of multiple sample substances; acquiring a set of test characteristic peaks based on the test infrared spectrum of the mixed substance to be tested, and acquiring a set of sample characteristic peaks based on the sample infrared spectra of the multiple sample substances; iteratively performing peak-removal matching on the test characteristic peaks in the test characteristic peak set and the sample characteristic peaks in the sample characteristic peak set; acquiring a set of sample substances matching the mixed substance to be tested based on the peak-removal matching result; and acquiring a set of sample substances matching the mixed substance to be tested based on the peak-removal matching result. The composition of the sample set is determined to obtain the composition identification result of the mixture to be tested. The peak removal matching includes a matching operation and a peak removal operation. The matching operation is used to perform similarity matching between the test feature peaks in the test feature peak set corresponding to each peak removal match and the sample feature peaks in the sample feature peak set corresponding to each peak removal match. The peak removal operation is used to delete the test feature peak with the highest similarity obtained from each peak removal match from the test feature peak set corresponding to each peak removal match, to obtain the test feature peak set corresponding to the next peak removal match, and to delete all sample feature peaks of the sample substance corresponding to the sample feature peak with the highest similarity obtained from each peak removal match from the sample feature peak set corresponding to each peak removal match, to obtain the sample feature peak set corresponding to the next peak removal match.
[0204] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0205] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0206] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for identifying mixed substances based on infrared spectroscopy and peak matching, characterized in that, include: Obtain the infrared spectrum of the mixture to be tested and the infrared spectrum of multiple sample substances; Based on the infrared spectrum of the mixture to be tested, obtain the set of characteristic peaks to be tested; based on the infrared spectrum of the multiple sample substances, obtain the set of characteristic peaks of the samples. Iterative peak removal matching is performed on the target feature peaks in the target feature peak set and the sample feature peaks in the sample feature peak set; Based on the peak matching results, a set of sample substances that match the mixture to be tested is obtained. Based on the composition of the set of sample substances that match the mixture to be tested, the composition identification result of the mixture to be tested is obtained. The peak removal matching includes a matching operation and a peak removal operation. The matching operation is used to perform similarity matching between the target feature peaks in the target feature peak set corresponding to each peak removal matching and the sample feature peaks in the sample feature peak set corresponding to each peak removal matching. The peak removal operation is used to delete the target feature peak with the highest similarity obtained from each peak removal matching from the target feature peak set corresponding to each peak removal matching, to obtain the target feature peak set corresponding to the next peak removal matching, and to remove all sample feature peaks of the sample substance corresponding to the sample with the highest similarity obtained from each peak removal matching from the sample feature peak set corresponding to each peak removal matching. The peak set is deleted from the current peak set to obtain the sample peak set corresponding to the next peak removal matching; the peak removal matching termination condition includes the condition that the spectral similarity corresponding to the current peak removal matching is less than or equal to the spectral similarity corresponding to the previous peak removal matching; the spectral similarity corresponding to the current peak removal matching is obtained by calculating the similarity between the combined spectrum corresponding to the current peak removal matching and the infrared spectrum to be tested; the combined spectrum is obtained by superimposing the infrared spectrum of the sample substance that matches the mixed substance to be tested in the current peak removal matching with the combined spectrum corresponding to the previous peak removal matching.
2. The method for identifying mixed substances based on infrared spectroscopy and peak matching according to claim 1, characterized in that, The iterative peak-removal matching of the target feature peaks in the target feature peak set and the sample feature peaks in the sample feature peak set includes: For the current peak removal matching, each test feature peak in the set of test feature peaks corresponding to the current peak removal matching is matched with each sample feature peak in the set of sample feature peaks corresponding to the current peak removal matching. The sample substance to which the feature peak of the most similar sample is matched is taken as the sample substance that matches the mixture to be tested in the current peak removal matching. The infrared spectrum of the sample substance that matches the mixed substance to be tested in the current peak removal matching is superimposed with the combined spectrum corresponding to the previous peak removal matching to obtain the combined spectrum corresponding to the current peak removal matching. Calculate the similarity between the combined spectrum corresponding to the current peak removal matching and the infrared spectrum to be measured, and obtain the spectrum similarity corresponding to the current peak removal matching; If the spectral similarity corresponding to the current peak removal match is greater than the spectral similarity corresponding to the previous peak removal match, the test feature peak with the highest similarity is deleted from the test feature peak set corresponding to the current peak removal match, and the test feature peak set corresponding to the next peak removal match is obtained. The sample feature peak with the highest similarity is removed from the sample feature peak set corresponding to the current peak removal match, and the sample feature peak set corresponding to the next peak removal match is obtained. The peak removal matching step is iteratively performed by matching each test feature peak in the set of test feature peaks corresponding to the next peak removal matching with each sample feature peak in the set of sample feature peaks corresponding to the next peak removal matching, until the peak removal matching termination condition is met.
3. The method for identifying mixed substances based on infrared spectroscopy and peak matching according to claim 2, characterized in that, The step of performing similarity matching between each test feature peak in the set of test feature peaks corresponding to the current peak removal matching and each sample feature peak in the set of sample feature peaks corresponding to the current peak removal matching includes: Based on the Jaccard coefficient, a similarity match is performed between each test feature peak in the set of test feature peaks corresponding to the current peak removal match and each sample feature peak in the set of sample feature peaks corresponding to the current peak removal match.
4. The method for identifying mixed substances based on infrared spectroscopy and peak matching according to claim 2, characterized in that, The step of superimposing the infrared spectrum of the sample substance that matches the mixture to be tested in the current peak-removal matching with the combined spectrum corresponding to the previous peak-removal matching to obtain the combined spectrum corresponding to the current peak-removal matching includes: Linear interpolation is performed on the infrared spectra of the sample substances that match the mixed substance to be tested in the current peak-removal matching and the combined spectra corresponding to the previous peak-removal matching, so that the data format of the infrared spectra of the samples after linear interpolation is consistent with the data format of the combined spectra after linear interpolation. The sample infrared spectrum after linear interpolation is linearly superimposed with the combined spectrum after linear interpolation to obtain the combined spectrum corresponding to the current peak matching.
5. The method for identifying mixed substances based on infrared spectroscopy and peak matching according to claim 2, characterized in that, The step of obtaining a set of sample substances that match the mixture to be tested based on the peak matching results includes: Based on the peak matching results, obtain the sample substance that matches the mixture to be tested in each peak matching; The sample substances that match the test mixture in all peak matching processes are summarized to obtain the sample substance set that matches the test mixture.
6. The method for identifying mixed substances based on infrared spectroscopy and peak matching according to any one of claims 1-5, characterized in that, The acquisition of the infrared spectrum of the mixture to be tested and the infrared spectra of multiple sample substances includes: Based on an infrared spectrometer, the original infrared spectrum of the mixture to be tested was acquired; The original infrared spectrum of the target sample and the original sample infrared spectrum in the database are preprocessed. The preprocessed original infrared spectrum to be tested is used as the infrared spectrum to be tested. From the preprocessed original sample infrared spectrum in the database, extract the sample infrared spectrum of each sample substance; The preprocessing includes filtering and standard state transformation. The filtering process includes low-pass filtering and least-squares-based convolution fitting filtering.
7. The method for identifying mixed substances based on infrared spectroscopy and peak matching according to claim 6, characterized in that, Extracting the infrared spectrum of each sample substance from the preprocessed original sample infrared spectrum in the database includes: Spectral clustering is performed on the preprocessed infrared spectra of the original samples in the database; Based on the clustering results, determine the original infrared spectra of all pretreated samples for each sample material; One pre-processed original infrared spectrum was randomly selected from all the pre-processed original infrared spectra of each sample substance and used as the sample infrared spectrum of each sample substance.
8. A mixed substance identification device based on infrared spectroscopy and peak matching, characterized in that, include: The data acquisition module is used to acquire the infrared spectrum of the mixture to be tested and the infrared spectrum of multiple sample substances. The data processing module is used to obtain the set of characteristic peaks to be tested based on the infrared spectrum of the mixed substance to be tested, and to obtain the set of characteristic peaks of the samples based on the infrared spectrum of the multiple sample substances. The substance matching module is used to iteratively perform peak removal matching on the target feature peaks in the target feature peak set and the sample feature peaks in the sample feature peak set; The substance identification module is used to obtain a set of sample substances that match the mixture to be tested based on the peak matching result, and to obtain the component identification result of the mixture to be tested based on the composition of the set of sample substances that match the mixture to be tested. The peak removal matching includes a matching operation and a peak removal operation. The matching operation is used to perform similarity matching between the target feature peaks in the target feature peak set corresponding to each peak removal matching and the sample feature peaks in the sample feature peak set corresponding to each peak removal matching. The peak removal operation is used to delete the target feature peak with the highest similarity obtained from each peak removal matching from the target feature peak set corresponding to each peak removal matching, to obtain the target feature peak set corresponding to the next peak removal matching, and to remove all sample feature peaks of the sample substance corresponding to the sample with the highest similarity obtained from each peak removal matching from the sample feature peak set corresponding to each peak removal matching. The peak set is deleted from the current peak set to obtain the sample peak set corresponding to the next peak removal matching; the peak removal matching termination condition includes the condition that the spectral similarity corresponding to the current peak removal matching is less than or equal to the spectral similarity corresponding to the previous peak removal matching; the spectral similarity corresponding to the current peak removal matching is obtained by calculating the similarity between the combined spectrum corresponding to the current peak removal matching and the infrared spectrum to be tested; the combined spectrum is obtained by superimposing the infrared spectrum of the sample substance that matches the mixed substance to be tested in the current peak removal matching with the combined spectrum corresponding to the previous peak removal matching.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the mixed substance identification method based on infrared spectroscopy and peak matching as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the mixed substance identification method based on infrared spectroscopy and peak matching as described in any one of claims 1 to 7.