A multispectral information fusion water pollutant identification method based on improved D-S evidence theory

By using an improved DS evidence theory for decision-level fusion, the accuracy problem of multispectral information fusion in water quality analysis is solved, achieving more comprehensive acquisition of pollutant characteristic information and higher identification accuracy, especially in complex water environments.

CN117132862BActive Publication Date: 2025-11-04HANGZHOU COMPASS STAR TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311183942.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-14
Publication Date
2025-11-04
Estimated Expiration
2043-09-14

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively utilize the complementarity of multispectral information in water quality analysis, leading to fusion failures or inaccurate prediction results. In particular, they have failed to effectively address the differences and conflicting evidence among different spectral data types in decision-level fusion.

Method used

An improved Dempster-Shafer (DS) evidence theory is used for decision-level fusion. Through feature extraction and weighted preprocessing, combined with the correlation between absorption spectrum and three-dimensional fluorescence spectrum, a multispectral information fusion model is established. Principal component analysis, wavelet transform and multi-class support vector machine are used for feature extraction and model fusion.

Benefits of technology

It achieves more comprehensive acquisition of pollutant characteristic information, improves the reliability and robustness of the model, and can effectively identify a wide variety of pollutants with a large concentration range, avoiding counterintuitive results in the classic DS evidence theory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117132862B_ABST
    Figure CN117132862B_ABST
Patent Text Reader

Abstract

The application discloses a multispectral information fusion water pollutant identification method based on an improved D-S evidence theory, and specifically comprises the following steps: collecting absorption spectrum and three-dimensional fluorescence spectrum of samples; establishing a pollutant identification model, establishing an identification model according to the spectrum of known pollutant samples, including single-spectrum model establishment and a decision-level fusion process; identifying to-be-detected samples, inputting the absorption spectrum and three-dimensional fluorescence spectrum of the to-be-detected samples into the established pollutant identification model for identification, and obtaining the class label of the possible pollutants contained in the samples. According to the correlation and complementarity of the absorption spectrum and the three-dimensional fluorescence spectrum, the application establishes a fusion model of the two kinds of spectrum, adopts a decision-level fusion method based on the improved D-S evidence theory, fuses the prediction results of each single-spectrum model according to specific fusion rules, avoids the influence of the difference between different spectral data types, obtains more comprehensive pollutant characteristic information, and makes the model have higher reliability and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of water pollutant identification method, and particularly relates to a multispectral information fusion water pollutant identification method based on improved D-S evidence theory. BACKGROUND

[0002] Organic pollution can deteriorate the water environment, destroy the water ecological system, and seriously threaten the health of humans and animals. Effective water quality detection has important practical significance for the management of urban water resources and the identification of pollution sources. Compared with traditional chemical methods for detecting organic pollutants, spectral methods have the advantages of being fast, non-destructive, and free of secondary pollution, and are therefore widely used in water quality analysis. Among them, the effectiveness of absorption spectroscopy and three-dimensional fluorescence spectroscopy in water pollution analysis has been confirmed by a large number of studies, but absorption spectroscopy is easy to obtain but has low sensitivity and many overlapping peaks; three-dimensional fluorescence spectroscopy has the advantages of high sensitivity and good selectivity, but is affected by factors such as self-absorption, internal filtration, and water scattering. The two types of spectra can provide characteristic information of substances from different angles, each with advantages and disadvantages, and are complementary to each other.

[0003] Multi-source spectral information fusion relies on the complementarity and collaboration between different types of spectra, can more comprehensively and deeply mine characteristic information, and thus obtain more reliable and accurate analysis results than single sensors. According to the different data processing methods, multi-spectral information fusion can be divided into three categories: data level fusion, feature level fusion, and decision level fusion. The current research methods for water quality analysis mainly focus on data level and feature level fusion, directly fusing different types of spectra without considering their effectiveness and matching degree; but the decision layer method fuses the prediction results of each sub-model according to a specific fusion rule, which is more suitable for multi-spectral fusion, and the main reasons are as follows:

[0004] (1) The data dimensions, ranges, and physical meanings of various spectra are different. In terms of data dimension, absorption spectroscopy is a one-dimensional vector, while three-dimensional fluorescence spectroscopy is a two-dimensional matrix, and cannot be directly fused; in terms of data range, the data ranges of the two types of spectra may differ by several orders of magnitude, and if they are directly fused, fusion failure or dominance of one type of spectrum may occur, making the fusion meaningless, so normalization preprocessing is needed before fusion; in terms of physical meaning, absorption spectroscopy represents the degree of light absorption by a substance, and three-dimensional fluorescence spectroscopy represents the fluorescence intensity emitted by a substance after excitation, and directly fusing values with different physical meanings may be unreasonable.

[0005] (2) For the identification of a variety of organic pollutants, different spectra have different performances. The absorption spectra of certain pollutants may overlap, while their fluorescence spectra are separated; on the contrary, the absorption peaks of some pollutants can be easily distinguished, but their fluorescence is too weak to be detected or too similar to be distinguished; in the decision-level fusion, decision rules can be made according to the spectral characteristics of each pollutant, and the weighting factor can be adjusted to ensure the accuracy and stability of the prediction results.

[0006] Dempster-Shafer (D-S) evidence theory is one of the decision-level methods of multi-sensor information fusion, which was first proposed by Dempster and popularized by Shafer; the theory analyzes uncertainty or unknown problems by defining concepts such as belief functions and combining various types of evidence in mathematical reasoning. Currently, D-S evidence theory has been widely used in machine fault diagnosis, medical diagnosis, target tracking and multi-attribute decision-making and other fields. With the development of D-S evidence theory, scholars have found that in the classical D-S evidence theory, the fusion of highly conflicting evidence may lead to counterintuitive fusion results. Therefore, many scholars have analyzed the reasons for the inconsistency between conflicting evidence and intuitive fusion, and proposed improvement methods.

[0007] Based on this, the present application adopts a multi-spectral information fusion method to analyze water organic pollution, and establishes a decision-level fusion model of absorption spectrum and three-dimensional fluorescence spectrum according to the correlation and complementarity of absorption spectrum and three-dimensional fluorescence spectrum. SUMMARY

[0008] In order to make up for the shortcomings of the prior art, the present application aims to provide a multi-spectral information fusion water pollutant identification method based on improved D-S evidence theory, so as to obtain more comprehensive pollutant characteristic information and make the model more reliable and robust.

[0009] A multi-spectral information fusion water pollutant identification method based on improved D-S evidence theory, the specific steps comprising:

[0010] S1: Collecting the absorption spectrum and three-dimensional fluorescence spectrum of the sample;

[0011] S2: Establishing a pollutant identification model, establishing an identification model according to the spectrum of known pollutant samples, including single-spectrum model establishment and decision-level fusion process;

[0012] S3: Identifying the sample to be tested, inputting the absorption spectrum and three-dimensional fluorescence spectrum of the sample to be tested into the pollutant identification model established in step S2 for identification, and obtaining the class label of the pollutant possibly contained in the sample.

[0013] Further, the wavelength range of the absorption spectrum collected in step S1 is 200-780 nm. The excitation wavelength range of the three-dimensional fluorescence spectrum is 220-600 nm, and the emission wavelength range is 230-700 nm.

[0014] Further, the specific steps of the single-spectrum model establishment in step S2 are as follows:

[0015] Feature extraction is performed on the spectrum by principal component analysis and wavelet transform;

[0016] Feature extraction is performed on the absorption spectrum and the one-dimensional corrected three-dimensional fluorescence spectrum by principal component analysis, and the cumulative contribution threshold of the feature subset PCs is 95%;

[0017] Wavelet transform is performed on the absorption spectrum image, and the wavelet coefficients of the approximation component are taken as the feature subset. Two-dimensional wavelet transform is performed on the three-dimensional fluorescence spectrum image, and the wavelet coefficients of the approximation component are encoded and corrected to form the feature subset;

[0018] The spectrum features obtained in the above steps are input into a multi-classification support vector machine to establish four single-spectrum recognition models: Abs-Model1, Abs-Model2, Fluo-Model1, and Fluo-Model2.

[0019] Further, the decision-level fusion in step S2 is to fuse the recognition results of the four single-spectrum models. The MSVM output of the four single-spectrum models is the score corresponding to the class label of each sample. The probability converted from the score is taken as the BPA0 of the fusion model. The class labels include single-substance labels, mixed-substance labels, and the empty set. Before applying the Dempster combination rule for fusion, improved D-S evidence theory is used for weighted preprocessing to obtain the BPA function m".

[0020] Further, improved D-S evidence theory is used for fusion, and the specific steps are as follows:

[0021] R1. Construct a recognition framework: In the Dempster-Shafer evidence theory, Θ of the recognition framework FOD is a set of N mutually exclusive and independent elements, and the power set 2 Θ is defined as the set of all subsets of FOD containing 2 M elements, where is the empty set. If A is contained in 2 Θ , it is called a proposition,

[0022] Θ={θ1,θ2,...,θ i ,...θ N},i=1,2,...,N,

[0023]

[0024] If a function m:2 Θ →[0, 1] satisfies:

[0025]

[0026] where m is called basic probability assignment function (BPA) or mass function, m(A) is the reliability of proposition A;

[0027] R2. The belief function Bel(A) and the plausibility function Pl(A) are the upper and lower bounds of the belief interval of proposition A, respectively, and [Bel(A), Pl(A)] is the belief function used to describe the uncertainty, Bel:2 Θ →[0, 1] is the sum of BPA of all subsets of A, which describes the possibility that the evidence believes A is true, i.e. the degree of full support for A, and Pl:2 Θ →[0, 1] is the sum of BPA of the intersection of A and B, which represents the possibility that the evidence does not suspect A to be false, i.e. the degree of A not being denied;

[0028]

[0029]

[0030] R3. The recognition results of single spectral models are fused by combining the weighted conflicting evidence fusion method and the Dempster combination rule, the fusion results of two independent evidence bodies are obtained by using the Dempster combination rule, assuming that m1 and m2 are two BPA in the same Θ, the new BPA after the fusion of m1 and m2 is denoted as m1⊕m2, and the Dempster combination rule is defined as follows:

[0031]

[0032]

[0033] where K is the conflict coefficient, which is the conflict measure between m1 and m2, and satisfies K<1.

[0034] Specifically, the specific steps of improving the weighted conflicting evidence fusion of D-S evidence theory are as follows:

[0035] T1. The recognition results of each single spectral model are converted from scores to probabilities, and these probabilities form a set of BPA functions m0 on the recognition framework Θ;

[0036] T2. According to the availability label, the m0 corresponding to the single spectral model with the “False” label is reset to zero, and the reset BPA function is denoted as m;

[0037] T3. Reassign the belief of multi-subset propositions to single-subset propositions using the probability transformation function to get a new BPA function m':

[0038]

[0039] BEL =∑Bel(θ i );

[0040] T4. Use the Hellinger distance to represent the conflict degree between BPA functions m' i (i = 1, 2, …, n) and m' j (j = 1, 2, …, n), and construct the conflict matrix D:

[0041]

[0042]

[0043] T5. Calculate the average distance of BPA functions Then normalize to get the credibility of BPA function m' i

[0044]

[0045]

[0046]

[0047] T6. Use the entropy E x (m i ) to represent the information amount of BPA functions, and then normalize to get the information amount of BPA function m' i

[0048]

[0049]

[0050]

[0051] T7. Calculate the credibility Cred i of BPA functions based on support and information amount, and then normalize to get the weight factor w i of BPA function m' i :

[0052]

[0053] ​​

[0054] T8. The initial BPA function m' is weighted to obtain the final BPA function m'':

[0055]

[0056] Compared with the prior art, the present application has the following advantages:

[0057] (1) According to the correlation and complementarity of the absorption spectrum and the three-dimensional fluorescence spectrum, a fusion model of the two kinds of spectra is established, a decision-level fusion method is adopted, the prediction results of each single spectrum model are fused according to specific fusion rules, the influence of the difference between different spectrum data types is avoided, more comprehensive pollutant characteristic information is obtained, and the model has higher reliability and robustness;

[0058] (2) The spectrum availability label is combined to make a pre-judgment before fusion, so that the identification of pollutants with a large variety and a large concentration range is realized;

[0059] (3) The improved D-S evidence theory can effectively measure the conflict between evidences, so that the counterintuitive results that may occur when the classical D-S evidence theory is used are avoided. BRIEF DESCRIPTION OF DRAWINGS

[0060] Figure 1 is a total flow chart of the method of the present application;

[0061] Figure 2 is a schematic diagram of the multispectral fusion process;

[0062] Figure 3 is a flow chart of the weighted conflict evidence fusion method;

[0063] Figure 4 is an absorption spectrum diagram of five organic pollutants;

[0064] Figure 5 is a three-dimensional fluorescence spectrum of five organic pollutants: (a) aniline, (b) fluorescent whitening agent, (c) salicylic acid, (d) malachite green and benzoic acid;

[0065] Figure 6 is the absorption spectrum (a) and the three-dimensional fluorescence spectrum (b) of river water. DETAILED DESCRIPTION

[0066] In order for those skilled in the art to understand the technical solutions of the present application, the present application will be further described below with reference to the drawings.

[0067] As Figures 1-3 shown, a multispectral information fusion water pollutant identification method based on an improved D-S evidence theory includes the following specific steps:

[0068] S1: Collecting the absorption spectrum and three-dimensional fluorescence spectrum of the sample;

[0069] The wavelength range of the absorption spectrum is 200-780 nm; the excitation wavelength range of the three-dimensional fluorescence spectrum is 220-600 nm, and the emission wavelength range is 230-700 nm.

[0070] Specifically, the absorption spectrum collection system includes a sample cell, a light source, and a light receiver. The light source is a pulse xenon lamp with a wavelength range of 200-1100 nm. The light receiver is a spectrometer. The sample collected in the cuvette is placed in the sample cell. The composite light emitted by the light source is collimated by a collimating lens and then irradiates the sample in the sample cell. The fluorescence emitted by the sample is condensed by a convex lens and finally received by the spectrometer. The spectral data is displayed and stored on the computer.

[0071] The three-dimensional fluorescence spectrum collection system includes a light source, a monochromator, a sample cell, and a light receiver. The light source is a 150W xenon lamp. The monochromator has a 1200mm slit grating that produces single excitation light at 300nm. The entrance and exit slit widths are both 3mm. The light receiver is a spectrometer that is perpendicular to the excitation light produced by the monochromator to obtain the emission spectrum corresponding to each excitation wavelength. The composite light emitted by the light source is separated by the monochromator, and the light with the specified wavelength obtained is irradiated into the sample cell by a lens. The fluorescence emitted by the sample is condensed by a convex lens and finally received by the spectrometer. The spectral data is displayed and stored on the computer.

[0072] S2: Establishing a pollutant identification model. The identification model is established based on the spectra of known pollutant samples, including single spectrum model establishment and decision-level fusion process.

[0073] (1) The specific steps of single spectrum model establishment are as follows:

[0074] Principal component analysis and wavelet transform are used for feature extraction of the spectrum. Principal component analysis is used to extract features from the absorption spectrum and the one-dimensional corrected three-dimensional fluorescence spectrum. The cumulative contribution threshold of the feature subset PCs is 95%. Wavelet transform is performed on the absorption spectrum image, and the wavelet coefficients of the approximation component are taken as the feature subset. Two-dimensional wavelet transform is performed on the three-dimensional fluorescence spectrum image, and the wavelet coefficients of the approximation component are encoded and corrected to form a feature subset. The obtained spectral features are input into a multi-classification support vector machine to establish four single spectrum identification models: Abs-Model1, Abs-Model2, Fluo-Model1, and Fluo-Model2.

[0075] (2) Decision level fusion is to fuse the recognition results of four single spectral models. The MSVM outputs of four single spectral models are the scores corresponding to the class labels of each sample. The probabilities converted from the scores are taken as the BPA0 of the fusion model. The class labels include single substance label, mixed substance label and empty set. Before applying the Dempster combination rule for fusion, the improved D-S evidence theory is used for weighted preprocessing to obtain the BPA function m".

[0076] (3) Wherein, the improved D-S evidence theory is used for fusion, and the specific steps are as follows:

[0077] R1. Construct the recognition framework: in the Dempster-Shafer evidence theory, Θ of the recognition framework FOD is a set of N mutually exclusive and independent elements, and the power set 2 Θ is defined as the set of all subsets of FOD containing 2 M elements, wherein is an empty set, if A is contained in 2 Θ , A is called a proposition,

[0078] Θ={θ1,θ2,...,θ i ,...θ N},i=1,2,...,N,

[0079]

[0080] If the function m: 2 Θ →[0,1] satisfies:

[0081]

[0082] Wherein, m is called a basic probability assignment function (BPA) or a mass function, and m(A) is the reliability of the proposition A;

[0083] R2. The belief function Bel(A) and the likelihood function Pl(A) are respectively the upper limit and the lower limit of the confidence interval of the proposition A, and [Bel(A), Pl(A)] is the belief function used to describe the uncertainty. Bel: 2 Θ →[0,1] is the sum of the BPA of all subsets of A, which describes the possibility that the evidence believes A to be true, that is, the degree of full support for A. Pl: 2 Θ →[0,1] is the sum of the BPA of the intersection of A and B, which represents the possibility that the evidence does not doubt A to be false, that is, the degree that A is not denied;

[0084]

[0085]

[0086] R3. The recognition results of the single spectral models are fused by combining the weighted conflict evidence fusion method and the Dempster combination rule. The fusion results of two independent evidence bodies are obtained by using the Dempster combination rule. Assuming that m1 and m2 are two BPA in the same Θ, the new BPA after the fusion of m1 and m2 is denoted as m1⊕m2. The Dempster combination rule is defined as follows:

[0087]

[0088]

[0089] wherein K is a conflict coefficient, which is a conflict measure between m1 and m2, and satisfies K<1.

[0090] (4) The specific steps of improving the weighted conflict evidence fusion of the D-S evidence theory are as follows:

[0091] T1. The recognition results of each single spectral model are converted from scores to probabilities. These probabilities form a group of BPA functions m0 on the recognition framework Θ.

[0092] T2. According to the availability label, the m0 corresponding to the single spectral model with the “False” label is reset to zero. The reset BPA function is denoted as m.

[0093] T3. The confidence of the multi-subset proposition is revalued to the single-subset proposition by using the probability conversion function, and a new BPA function m′ is obtained.

[0094]

[0095] BEL=∑Bel(θ i );

[0096] T4. The Hellinger distance is used to represent the conflict degree between the BPA functions m′ i (i=1,2,…,n) and m′ j (j=1,2,…,n), and a conflict matrix D is constructed.

[0097]

[0098]

[0099] T5. The average distance of the BPA functions is calculated and then normalized to obtain the credibility of the BPA function m′ i .

[0100]

[0101]

[0102]

[0103] T6. Use the entropy E x (m i ) to represent the information amount of the BPA function, and then normalize to obtain the information amount of the BPA function m′ i

[0104]

[0105]

[0106]

[0107] T7. Calculate the credibility Cred of the BPA function based on the support degree and the information amount i , and then normalize to obtain the weight factor w i of the BPA function m′ i :

[0108]

[0109]

[0110] T8. Weight the initial BPA function m′ to obtain the final BPA function m″:

[0111]

[0112] S3: Identify the sample to be tested, input the absorption spectrum and three-dimensional fluorescence spectrum of the sample to be tested into the pollutant identification model established in step S2 for identification, and obtain the class label of the pollutant possibly contained in the sample.

[0113] From the above, the application provides a multispectral information fusion water pollutant identification method based on improved D-S evidence theory, including spectrum collection, pollutant identification model establishment and identification of pollutants in samples, identification of one or more pollutants in water, and the experimental method and result analysis include:

[0114] (I) Establish a multispectral fusion model for identifying organic pollutants

[0115] 1. Experimental method

[0116] ​Step 1: In this example, five organic pollutants, aniline, fluorescent whitening agent, salicylic acid, malachite green, and benzoic acid, were used as experimental objects. Deionized water containing dissolved pollutants was provided by a Millipore-q water purification system (Millipore, Billerica, MA, USA). Water samples containing organic pollutants were prepared, and a total of 90 simulated pollution samples were prepared, with 18 samples for each substance. The concentration ranges were as follows: aniline 0.002-0.01 mg / L; fluorescent whitening agent and salicylic acid, 0.1-5 mg / L; malachite green, 2-40 mg / L; and benzoic acid, 1-10 mg / L.

[0117] Step 2: The absorption spectra and three-dimensional fluorescence spectra of the above samples were collected. The absorption spectra of the five organic pollutants are shown in Figure 4 , and the three-dimensional fluorescence spectra are shown in Figure 5 . The concentration ranges described in Step 1 ensure that at least one spectrum is within the measurable range for each pollutant.

[0118] Step 3: The samples were divided into a training set and a validation set in a ratio of 2:1, i.e., 60 samples in the training set and 30 samples in the validation set.

[0119] First, the spectral data of the training set were input into the four single-spectrum models to obtain four sets of recognition results of the single-spectrum models, i.e., the scores of each sample corresponding to each category. Then, the scores output by each single-spectrum model were converted into probabilities of BPAs, and the fusion weights of the four single-spectrum models corresponding to each pollutant were calculated based on the improved D-S evidence theory-based weighted conflict evidence fusion method. Finally, the spectral data of the validation set samples were input into the above fusion recognition model, and the fusion probabilities of each sample belonging to each category were calculated in combination with the availability labels. The category label corresponding to the maximum probability was regarded as the prediction result of the model for the to-be-tested sample.

[0120] 2. Results analysis

[0121] The number of correctly predicted samples in the validation set and the prediction accuracy of each model are shown in Table 1. The prediction accuracy of the two fusion models is between the lowest accuracy and the highest accuracy of the single-spectrum models, because the fusion models integrate the prediction results of the four single-spectrum models. The results of multiple repeated experiments show that the prediction effect of the fusion models is generally more accurate and stable than that of the single-spectrum models, which is consistent with the theoretical analysis.

[0122] In the identification of the above five pollutants, the model made false identifications for some samples of aniline, malachite green and benzoic acid. In order to further analyze, Table 2 lists the probabilities of the four single-spectrum models and the two fusion models for the validation set samples of the three pollutants. These samples are numbered 1-6 and 19-30. In samples 19-30, the fusion model based on the classical D-S evidence theory made several false judgments, the main reason being the Zadeh paradox. Taking sample 19 as an example, among the four single-spectrum models, two models considered that the sample was very likely to be malachite green (M1: 0.4136, M2: 0.4837), one model considered that the sample might be malachite green (M4: 0.2496), and only one model considered that the sample was definitely not malachite green (M3: 0.00). The result of the M3 model is in great conflict with the results of the other three single-spectrum models. The sample is indeed malachite green, but the fusion model based on the classical D-S evidence theory predicts the class to be salicylic acid, which is a false judgment. The proposed fusion model based on the improved D-S evidence theory can effectively avoid the Zadeh paradox, thereby generating the correct predicted class.

[0123] The samples 19-30 in Table 2 are malachite green and benzoic acid, which have no fluorescence peaks (such as Figure 5 (d)) The proposed improved fusion model uses the availability label to handle this situation. In pre-judgment, for samples without obvious fluorescence peaks, the availability label of their fluorescence spectrum is set to "False". In calculating the fusion result, it is determined whether to use the result of the single-spectrum model corresponding to the spectrum as the input of the fusion model according to the availability label. For samples 19-30 in Table 2, the improved fusion model only fuses the results of M1 and M2, and discards the results of M3 and M4. Since it has obvious absorption peak characteristics, its identification accuracy is still 90%. The experimental results show that the process of adding the availability label sometimes reduces the utilization rate of single-spectrum models, but expands the identification range of pollutant classes and concentrations.

[0124] For samples 1-6 in Table 2, both fusion models give the correct class, and the probability of the correct class obtained based on the classical D-S evidence theory is large. This is because the algorithms of the two fusion models are different, and when determining the predicted class of a sample according to its probability, only the probabilities of the five classes obtained by the same model need to be compared.

[0125] Table 1 (identification results of four single-spectrum models and two multi-spectrum fusion models for five organic pollutant samples. For each model, the number of samples that are correctly classified and the prediction accuracy of the validation set are given respectively):

[0126]

[0127] A: aniline, F: fluorescent whitening agent, S: salicylic acid, M: malachite green, B: benzoic acid.

[0128] Table 2 (Predicted probabilities of five organic pollutants by single-spectrum model and multi-spectrum fusion model, bold indicates that the predicted result corresponds to the correct category, box indicates incorrect fusion prediction results of the correct category, italic indicates incorrect categories predicted by the fusion model, and parentheses indicate data that do not participate in fusion based on the improved D-S evidence theory):

[0129]

[0130]

[0131] (B) Analysis of simulated pollution samples of river water

[0132] 1. Only one pollutant in the sample

[0133] Simulated pollution samples were prepared by adding pollutants to river water. The river water used in the experiment was collected from a river near the school. Five organic pollutants were dissolved in the river water to prepare a total of 75 simulated pollution samples, with 15 samples for each pollutant. The concentration ranges were as follows: aniline 0.002-0.01 mg / L; fluorescent whitening agent and salicylic acid, 0.1-5 mg / L, malachite green, 2-40 mg / L; benzoic acid, 1-10 mg / L; these simulated samples constituted the test set of the identification model. The experimental results are shown in Table 3.

[0134] The results show that the fusion model based on the improved D-S fusion theory still has relatively high prediction accuracy, only slightly lower than ABS-Model2. Among the five pollutants identified in this experiment, two do not have fluorescence spectra, so the recognition accuracy of Fluo-Model1 and Fluo-Model2 is relatively low. The five pollutants all have absorption spectra and distinct characteristics, so the absorption single-spectrum model has good recognition results in this experiment. However, in the identification of other pollutants, the performance of the four single-spectrum models may not be stable. For example, in some cases, the absorption spectra of pollutants can significantly overlap, while the fluorescence spectra are distinguishable, in which case the absorption single-spectrum model may not perform better than the fluorescence single-spectrum model. The fusion model proposed in this paper fuses the results of the four single-spectrum models, thus ensuring relatively high and stable recognition accuracy.

[0135] Compared with the results in Table 1, these models have slightly poorer prediction results for simulated river water pollution samples than for pollution samples dissolved in deionized water. For example, Figure 6The absorption spectrum and three-dimensional fluorescence spectrum of the river are shown. The river water may contain other substances such as inorganic suspended matter, amino acids, chlorophyll, lignin, humus, etc. The absorption and fluorescence reactions can be generated. Although the characteristic peaks of these substances are relatively weak, they also cover part of the spectral characteristic information of the pollution sample, affecting the identification result, especially for samples with low concentrations of pollutants.

[0136] Table 3 (the identification results of five organic pollutant river pollution simulation samples by four single-spectrum models and two multi-spectrum fusion models. For each model, the number of correctly classified samples and the prediction accuracy of the validation set are given respectively):

[0137]

[0138] 2、Sample with multiple pollutants

[0139] Two or three pollutants are added to the river water to simulate contaminated samples. In this paper, fluorescent whitening agent, salicylic acid and malachite green are selected as the research objects. The concentration and number of the above three pollutant mixed samples are shown in Table 4. These mixed samples are input into the identification model established above for prediction, and the results are shown in Table 5. For the four mixed samples in this experiment, among the four single-spectrum models, the prediction accuracy of Fluo-Model2 is the highest, followed by Abs-Model2, while the prediction accuracy of Abs-Model1 and Fluo-Model1 is relatively low. The prediction accuracy of the two fusion models is between the best and the worst of the four single-spectrum models. The performance of the fusion model based on improved D-S evidence theory is better than that of the classical D-S model, only second to Fluo-Model2. Compared with the results in Table 3, due to the mutual interference of the spectral characteristics of different pollutants, the prediction accuracy decreases with the increase of the number of pollutants. In summary, the multi-spectrum fusion method proposed in this study can effectively identify mixed samples of multiple pollutants. Compared with single-spectrum models and fusion models based on classical D-S, the fusion model has higher identification accuracy and better robustness.

[0140] Table 4 (concentration and number of mixed samples of multiple pollutants in river water):

[0141]

[0142] Table 5 (identification of three organic pollutant simulation river pollution samples by four single-spectrum models and two multi-spectrum fusion models. For each model, the number of correctly classified samples and the prediction accuracy of the test set are given respectively):

[0143]

[0144] Wherein, FS: fluorescent whitening agent + salicylic acid; FM: fluorescent whitening agent + malachite green; SM: salicylic acid + malachite green; FSM: fluorescent whitening agent + salicylic acid + malachite green.

[0145] The application adopts the D-S evidence theory with improved fusion rules as a decision-level fusion method to obtain more comprehensive pollutant characteristic information, and calculates the weight according to the unique attribute of each sub-model. In order to verify the reliability of the established identification model, the water body organic matter identification model of the application improves the analysis effect of water body organic pollutants by adding one or more organic pollutants to the river water to simulate river pollution, and provides a reference basis for water quality monitoring and pollution source analysis.

[0146] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A multispectral information fusion water pollutant identification method based on improved D-S evidence theory, characterized in that, Specific steps of the method include: S1: collecting absorption spectrum and three-dimensional fluorescence spectrum of the sample; S2: establishing a pollutant identification model, establishing an identification model according to the spectrum of a known pollutant sample, including single-spectrum model establishment and decision-level fusion process; The specific steps of single-spectrum model establishment are: Feature extraction is performed on the spectrum by using principal component analysis and wavelet transform; Feature extraction is performed on the absorption spectrum and the one-dimensionally corrected three-dimensional fluorescence spectrum by using principal component analysis, and the cumulative contribution threshold of the feature subset PCs is 95%; Wavelet transform is performed on the absorption spectrum image, and the wavelet coefficients of the approximation component are taken as the feature subset; two-dimensional wavelet transform is performed on the three-dimensional fluorescence spectrum image, and the wavelet coefficients of the approximation component are encoded and corrected to form the feature subset; The spectrum features obtained in the above steps are input into a multi-classification support vector machine to establish four single-spectrum identification models: Abs-Model1, Abs-Model2, Fluo-Model1 and Fluo-Model2; The decision-level fusion is to fuse the identification results of the four single-spectrum models. The MSVM output of the four single-spectrum models is a score corresponding to the class label of each sample. The score is converted into a probability as the BPA0 of the fusion model. The class labels include a single-substance label, a mixed-substance label and an empty set. Before applying the Dempster combination rule for fusion, an improved D-S evidence theory is used for weighted preprocessing. The identification results of the single-spectrum models are fused by combining the weighted conflict evidence fusion method and the Dempster combination rule to obtain the BPA function m". S3: identifying the to-be-tested sample, inputting the absorption spectrum and three-dimensional fluorescence spectrum of the to-be-tested sample into the pollutant identification model established in step S2 for identification to obtain the class label of the pollutant contained in the sample.

2. The multispectral information fusion water pollutant identification method based on improved D-S evidence theory according to claim 1, characterized in that, In step S1, the wavelength range of the absorption spectrum is 200-780 nm; the excitation wavelength range of the three-dimensional fluorescence spectrum is 220-600 nm, and the emission wavelength range is 230-700 nm.

3. The multispectral information fusion water pollutant identification method based on improved D-S evidence theory according to claim 1, characterized in that, The specific steps of the improved D-S evidence theory are as follows: R1. Constructing the recognition framework: In Dempster-Shafer evidence theory, Θ of the recognition framework FOD is a set of N mutually exclusive and independent elements, power set 2 Θ Θ = {A | A C 2 M Θ, |A| = 1, 2 Θ is empty, if A is contained in 2 Θ , it is called a proposition, Θ = {θ1, θ2,..., θ i ,...θ N}, i = 1, 2,..., N, If a function m: 2 Θ → [0, 1] satisfies: Wherein, m is called BPA function or mass function, m(A) is the reliability of proposition A; R2. The confidence function Bel(A) and the likelihood function Pl(A) are the upper and lower bounds of the confidence interval for proposition A, respectively. [Bel(A), Pl(A)] is the belief function used to describe the uncertainty. Bel: 2 Θ →[0, 1] is the sum of the BPAs of all subsets of A, which describes the probability that the evidence holds that A is true, i.e., the degree of support for A. P1: 2 Θ →[0,1] is the sum of BPA of the intersection of A and B, which represents the probability that the evidence does not raise doubt that A is false, that is, the degree to which A is not denied; R3. The fusion result of two independent evidences is obtained by using the Dempster combination rule. Assuming that m1 and m2 are two BPA in the same Θ, the new BPA after the fusion of m1 and m2 is denoted as m1⊕m2, and the Dempster combination rule is defined as follows: Wherein, K is the conflict coefficient, which is the conflict measure between m1 and m2, and satisfies K<1.

4. The multispectral information fusion water pollutant identification method based on improved D-S evidence theory according to claim 3, characterized in that, The specific steps of the weighted conflict evidence fusion of the improved D-S evidence theory are as follows: T1. The identification results of each single-spectrum model are converted from scores to probabilities, and these probabilities form a group of BPA functions m0 on the identification framework Θ; T2. According to the availability label, the m0 corresponding to the single-spectrum model with the "False" label is reset to zero, and the reset BPA function is denoted as m. T3. Reassign the belief of the multi-hypothesis proposition to the single-hypothesis proposition using the probability transformation function to obtain a new BPA function m': T4. The Hellinger distance is used to represent the BPA function m' i (i = 1, 2,..., n) and m' j (j = 1, 2,..., n) between the conflict matrix D is constructed: T5. Calculate the average distance of the BPA functions Then normalization is performed to obtain the reliability of the BPA function m' i ​ T6. The information content of the BPA function is denoted by E x (m i ) and is normalized to obtain the information content of the BPA function m'i T7. Calculate the credibility Cred of BPA function based on support and information amount i , and then normalize to get BPA function m' i weight factor w i : T8. Weight the initial BPA function m' to obtain the final BPA function m":