Marine temperature data quality control information fusion method and system based on D-S evidence theory

Through the multi-source heterogeneous marine observation data quality control information fusion method based on D-S evidence theory, the results of multiple quality control solutions are combined, and the problems of high misjudgment rate, large resource consumption and low universality in marine observation data quality control are solved, and more accurate and reliable data quality control is achieved.

CN120144570APending Publication Date: 2025-06-13INST OF OCEANOLOGY - CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510181970.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing quality control methods for marine observation data have problems such as high misjudgment rate, high resource consumption and low universality. Different quality control methods have inconsistent marking results for the same data, and lack the internationally recognized "best" solution.

Method used

The multi-source heterogeneous marine observation data quality control information fusion method is adopted based on D-S evidence theory. Through the fusion of results of multiple quality control solutions, an expert system, AI quality control model and automated quality control scheme are used to build an identification framework and basic confidence allocation, and evidence synthesis is carried out to obtain the final quality control results.

Benefits of technology

The integration of multiple quality control solutions results has been achieved, the technical shortcomings of a single quality control solution and the uncertainty of quality control results have been overcome, the accuracy and reliability of data quality control have been improved, and it is suitable for binary and multivariate quality control solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144570A_ABST
    Figure CN120144570A_ABST
Patent Text Reader

Abstract

The invention discloses an ocean temperature data quality control information fusion method and system based on a D-S evidence theory. The data processing module is mainly used for carrying out data cleaning on received multi-source heterogeneous original observation data, extracting key data and serializing the key data into metadata; 2) a quality control module: introducing a plurality of types of quality control schemes, and respectively performing quality control on the cleaned marine observation metadata to obtain a plurality of quality control mark vectors; and 3) a result decision-making module, which is used for performing decision-making information fusion on the plurality of quality control mark vectors by utilizing a D-S evidence theory based on prior knowledge of an expert system, and reconstructing to obtain a final quality control result. The method can integrate results of multiple quality control schemes, get rid of uncertainty in original schemes, and reserve the advantages of the schemes; meanwhile, the method is not only suitable for a current'binary 'quality control scheme for evaluating'good / bad' of ocean observation data, but also suitable for a'multivariate 'data quality control scheme based on data quality grading and scoring, and has very high expansibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing, and specifically relates to a method and system for obtaining more accurate and reliable quality control results by using D-S evidence theory to fuse data labeling information obtained from multi-source heterogeneous ocean observation data through multiple quality control schemes. Background Art

[0002] The correctness of ocean observation data directly affects the accuracy of the description of the basic characteristics of the ocean, the analysis of laws, and management decisions. Distorted or incorrect ocean environmental data will seriously affect the reliability of research results and predictions.

[0003] Quality control can be classified into automated quality control technology and manual inspection quality control technology (also known as "expert quality control technology") from a technical perspective. Automated quality control generally involves writing computer programs to automatically check and mark the observed physical parameters of each profile, such as CODC-QC, ICDC-QC, etc. The expert quality control technology (also known as "manual inspection quality control technology") is to conduct further visual review of the data processed by the automated quality control technology based on the past experience of oceanography experts. Therefore, expert quality control usually has high quality control accuracy, but low universality, consumes a large amount of human, material and financial resources, and requires a relatively high time cost, while the automated quality control technology is more efficient. In addition, considering that the essence of data quality control is to identify and classify ocean observation data ("good", "bad"), machine learning has inherent advantages in solving such problems. Many attempts have been made internationally in the quality control of ocean temperature data based on artificial intelligence. A variety of machine learning technologies such as conditional random field, support vector machine, decision tree, multi-layer perceptron (MLP), recurrent neural network (RNN) or convolutional neural network (CNN) have been continuously applied to explore the quality control technology of ocean temperature and salinity data. Although many data quality control technologies have been accumulated, there is no internationally recognized "best" solution, and each quality control method has its own advantages and disadvantages in various aspects. The IQuOD program tests international mainstream quality control systems including WOD-QC, GTSPP-QC, Argo-QC, EN4-QC, CoTeDe-QC, CSIRO-QC, AOML-QC, ICDC-QC, IQuOD-QC, etc. The results show that AOML-QC, ICDC-QC, and CoTeDe-QC have the highest true positive rate (TPR) of removing false data and are relatively strict quality control schemes, but still up to 15% of the wrong data are not detected; at the same time, there is also a phenomenon of too high false positive rate (FPR), where correct data are mislabeled. The FPR of WOD-QC is the lowest, but the TPR is also relatively low, and more than 40% of the data errors are not effectively identified. Moreover, for the same observation profile or the same observation site, the data labeling results of different quality control methods may also be different. Obviously, none of these many quality control technologies can simultaneously take into account the technical advantages of other quality control technologies.

[0004] D-S evidence theory is a mathematical method proposed by scholars A.P. Dempster and G.Shafer at Harvard University in the United States in the 1960s. It uses evidence and combination to deal with uncertainty reasoning problems. It is a method that analyzes, processes, and fuses multi-source heterogeneous data information by imitating the process of the human brain dealing with complex problems to obtain the final decision. Multi-source data fusion technology can integrate incomplete information collected from multiple different data sources, and perform corresponding processing and fusion, so that the advantages of different data complement each other, and finally obtain a data result with decision-making significance, thereby weakening the uncertain components in the data source, helping users obtain effective fusion judgments and accurate comprehensive measurements, and thus making reasonable judgments and decisions more easily. This theory has been widely applied in fields such as target recognition, automation, situation assessment, and earth science. Compared with other data fusion methods, D-S evidence theory can effectively reason and analyze incomplete information and uncertain information, and can make multi-source data fusion decisions more effectively and quickly. Summary of the Invention

[0005] The purpose of the present invention is to provide a method and system for fusing ocean temperature data quality control information based on D-S evidence theory to overcome the above defects.

[0006] The technical solution adopted by the present invention to achieve the above purpose is:

[0007] A method for fusing ocean temperature data quality control information based on D-S evidence theory includes the following steps:

[0008] 1) Obtain the original multi-source heterogeneous ocean temperature observation profile data to be quality controlled, and preprocess it to obtain metadata;

[0009] 2) Use multiple quality control schemes to perform quality control on the metadata;

[0010] 3) Perform information fusion on multiple quality control result vectors to obtain the final quality control result.

[0011] The specific content of step 1) is:

[0012] Extract key data from the multi-source heterogeneous ocean original observation profile data, serialize and package the key data into metadata, and display and interact with it through a visualization system. Among them, the key data includes: observation time, location, device type, and observed physical parameters.

[0013] The visualization system has the function of displaying and interacting with quality control marker bits, enabling experts to achieve expert quality control through this function.

[0014] Step 2) includes the following steps:

[0015] 2.1) Introduce an expert system, construct an AI-based quality control model for ocean observation data, and train it;

[0016] 2.2) Use the well-trained x AI quality control models to perform AI quality control on the metadata, and obtain the quality control flag vector P for each ocean observation profile i1 , where i1 ∈ [1, x];

[0017] 2.3) Introduce y experts in the expert system, and carry out expert quality control on the metadata through data interaction on the data visualization system. Obtain the quality control flag vector P for each ocean observation profile i2 , where i2 ∈ [1, y];

[0018] 2.4) Introduce z automated quality control schemes, input the metadata into the automated quality control schemes, and obtain the quality control flag vector P for each ocean observation profile i3 , where i3 ∈ [1, z].

[0019] The said step 3) includes the following steps:

[0020] 3.1) Construct an identification framework

[0021] For the in-situ observation data at a certain site, there are only a finite number M of mutually exclusive hypothesis cases Θ in the identification space. The definition method of the identification framework is as follows:

[0022] Θ = {A 1 , A 2 , ……, A M}

[0023] For the "binary" quality control scheme, there are 2 mutually exclusive hypothesis cases for the quality control flag of the observation site. The legal observation data is recorded as hypothesis A 1 , that is, the quality control flag bit is 0, and the illegal observation data is recorded as hypothesis A 2 , that is, the quality control flag bit is 1. Therefore, the defined identification framework FoD is Θ = {A 1 , A 2}; For the "multiple" data quality control scheme based on data quality grading and scoring, the defined identification framework FoD is Θ = {A 1 , A 2 , …, A n}, where n is the number of mutually exclusive hypotheses; Denote the total number of all quality control schemes as N, then N = x + y + z;

[0024] 3.2) Construct the basic belief assignment S i

[0025] S i = HIST(P i) / SUM[HIST(P i )]

[0026] Among them, P i represents the quality control flag vector for the quality control scheme i of a certain ocean observation profile. HIST represents solving the frequency histogram of the elements in the vector, and SUM represents summing the elements in the vector;

[0027] 3.3) Construct the basic belief assignment coefficient

[0028] Take the quality control flag vector of each quality control scheme as an evidence. According to the experience of domain experts, give the event credibility vector of each evidence as the basic belief assignment coefficient, denoted as T i , where i ∈ [1, N], and the evidence is the quality control result;

[0029] 3.4) Calculate the confidence of the hypothesis and construct the evidence vector H i

[0030] H i = S i ⊙ T i

[0031] Among them, the operator ⊙ represents the dot product of the elements in the vector;

[0032] 3.5) Use Dempster's combination rule for evidence combination

[0033] Use the fusion rule to fuse two evidences, and then take the fused result as the new basic belief vector to be fused, and fuse it with the third evidence again. Repeat the operation of fusing two evidences until all evidences are fused into one evidence as the combined evidence;

[0034] 3.6) Use the obtained combined evidence to reconstruct the annotation result of ocean observation data

[0035] P i = ROUND(H i )

[0036] Among them, H i represents the final combined evidence of the i-th ocean observation profile, and P i represents the final quality control result of the ocean observation profile. ROUND represents rounding the floating point number.

[0037] The specific fusion rule is as follows:

[0038] For evidences H m and H n , the way to combine evidence H k in the q-th hypothesis is:

[0039]

[0040] Among them, k represents the evidence synthesis conflict factor:

[0041]

[0042] Among them, M represents the number of hypotheses in the recognition framework, H m,j and H n,j respectively represent the confidence measures of hypotheses A in evidences m and n j .

[0043] An information fusion system for marine temperature data quality control based on D-S evidence theory includes:

[0044] A data preprocessing module, configured to obtain original multi-source heterogeneous marine observation profile data to be quality controlled, and preprocess it to obtain metadata;

[0045] A data quality control module, configured to perform quality control on the metadata using multiple quality control schemes;

[0046] A result decision module, configured to perform data fusion on multiple quality control result vectors to obtain a final quality control result.

[0047] An information fusion device for marine temperature data quality control based on D-S evidence theory includes a memory and a processor; the memory is used to store a computer program; the processor is used to implement the information fusion method for marine temperature data quality control based on D-S evidence theory when executing the computer program.

[0048] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the information fusion method for marine temperature data quality control based on D-S evidence theory is implemented.

[0049] The present invention has the following beneficial effects and advantages:

[0050] 1. The present invention can integrate the results of multiple quality control schemes, which is conducive to the "complementary disadvantages and fusion of advantages" of quality control schemes based on different technical routes, overcomes the technical shortcomings of a single quality control scheme and the uncertainty of quality control results, retains the advantages of these scheme results, and obtains a more accurate fusion quality control result.

[0051] 2. The present invention has strong scalability. It is applicable to both the current "binary" quality control scheme for evaluating the "good / bad" of marine observation data and the "multiple" data quality control scheme based on data quality scoring and grading, which is conducive to the "refined" research on marine observation data quality control technologies and methods. Description of the Drawings

[0052] Figure 1 Schematic diagram of method flow

[0053] Figure 2 Schematic diagram of system module

[0054] Figure 3 Schematic diagram of the information fusion decision-making system based on D-S theory Specific implementation manner

[0055] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0056] As Figure 1 and Figure 2 shown, an information fusion method for ocean temperature data quality control based on D-S evidence theory includes the following steps:

[0057] S1. Data preprocessing module, mainly normalizing multi-source heterogeneous ocean observation data into metadata:

[0058] The S1-1 module receives the original multi-source heterogeneous ocean observation profile data that needs quality control; where the multi-source means diverse data sources, which can come from on-site real-time observations or from general data sets, such as the World Ocean Database WOD; and enters S1-2;

[0059] S1-2 extracts key data from the multi-source heterogeneous ocean original observation profile data received from S1-1, including but not limited to: observation time, location, equipment type, observed physical parameters, etc.; and enters S1-3;

[0060] After serializing and packaging the key data into metadata in S1-3, it enters S1-4;

[0061] S1-4 uses a data visualization system to visually display and interact with the metadata obtained from S1-3; where the visualization system needs to have the display and interaction functions of quality control flag bits;

[0062] S2. Data quality control module, mainly using various quality control schemes to perform quality control on the metadata:

[0063] S2-1 introduces an expert system and uses the professional background knowledge of domain experts to carry out the training of an AI-based ocean observation data quality control model, and enters S2-2;

[0064] S2-2 performs AI quality control on the metadata obtained in S1-3 using the x AI quality control models trained in S2-1, and obtains a quality control flag bit vector P i1 , where, i1 ∈ [1, x], where; and enters S2-3;

[0065] Introduce y experts into the expert system in S2-3, and conduct expert quality control on the metadata obtained in S1-3 through data interaction on the data visualization system for each quality control marker vector P of the ocean observation profile i2 , where i2 ∈ [1, y], and then proceed to S2-4;

[0066] Introduce z automated quality control schemes in S2-4, such as CODC-QC and ICDC-QC. Input the metadata obtained in S1-3 into the automated quality control schemes for each quality control marker vector P of the ocean observation profile i3 , where i3 ∈ [1, z], and then proceed to S3;

[0067] S3. Result decision module, such as Figure 3 As shown, fuse the quality control result vectors obtained from multiple quality control schemes in S2 to obtain the final quality control result. The specific steps are as follows:

[0068] S3-1 Construct the Frame of Discernment (FoD):

[0069] For the in-situ observation data at a certain site, there are only a finite number of M mutually exclusive hypothesis cases in the recognition space. The definition method of the Frame of Discernment is as follows:

[0070] Θ = {A 1 , A 2 , ……, A M}

[0071] Most current quality control schemes have 2 mutually exclusive hypothesis cases for the quality control markers of the observation sites. The legal observation data is denoted as hypothesis A 1 (the quality control marker bit is 0), and the illegal observation data is denoted as hypothesis A 2 (the quality control marker bit is 1). Therefore, the Frame of Discernment FoD is defined as Θ = {A 1 , A 2}. Denote the total number of all quality control schemes as N, and N = x + y + z; then proceed to S3-2;

[0072] S3-2 Construct the basic belief assignment S i :

[0073] Construct the basic belief assignment S under the Frame of Discernment Θ by using the quality control marker vector of the specified quality control scheme i :

[0074] S i = HIST(P i ) / SUM[HIST(P i )]

[0075] Among them, P i represents the quality control marker vector for the quality control scheme i of a certain ocean observation profile. HIST represents the frequency histogram of the elements in the solution vector, and SUM represents the sum of the elements in the vector. Proceed to S3-3;

[0076] S3-3 Construct the basic confidence assignment coefficient:

[0077] Take the quality control marker vector of each quality control scheme as an evidence. According to the experience of domain experts, give the event credibility vector of each evidence as the basic belief assignment coefficient, denoted as T i , where i ∈ [1, N].

[0078] Proceed to S3-4;

[0079] S3-4 Calculate the confidence of the hypothesis and construct the evidence vector H i :

[0080] Use the basic confidence assignment coefficient and the basic confidence assignment to calculate the confidence of the hypothesis and construct the evidence:

[0081] H i = S i ⊙ T i

[0082] where the operator ⊙ represents the dot product of the elements in the vector. Proceed to S3-5.

[0083] For the current ocean observation data quality control scheme using binary hypothesis, the constructed evidence vector is shown in the table:

[0084]

[0085]

[0086] S3-5 Use the Dempster combination rule to combine the evidence:

[0087] First, perform the fusion of two evidences. Among them, the fusion rule is carried out in the following way:

[0088] For the evidences H m and H n , the way to fuse the evidence H k for the q-th hypothesis is:

[0089]

[0090] where k represents the evidence combination conflict factor, and the calculation method is:

[0091]

[0092] where M represents the number of hypotheses in the recognition framework, and H m,j and H n,j represent the confidence measures of the evidence for m and n hypotheses A j respectively.

[0093] Then, the fused result is regarded as the new basic confidence vector to be fused and fused with the third piece of evidence again;

[0094] Repeat the operation of fusing two pieces of evidence until the synthetic evidence is finally obtained. Proceed to S3-6.

[0095] S3-6 uses the obtained synthetic evidence to reconstruct the annotation result of ocean observation data:

[0096] P i = ROUND(H i )

[0097] where H i represents the final fused evidence of the i-th ocean observation profile, and P i represents the final quality control result of the ocean observation profile.

Claims

1. The ocean temperature data quality control information fusion method based on DS evidence theory is characterized by: The following steps are involved: 1) Obtain the original multi-source heterogeneous ocean temperature observation profile data to be quality controlled, and preprocess it to obtain metadata; 2) Use multiple quality control schemes to control metadata quality; 3) Fuse the information of multiple quality control result vectors to obtain the final quality control result.

2. The ocean temperature data quality control information fusion method based on DS evidence theory according to claim 1 is characterized in that: The step 1) is specifically as follows: Key data are extracted from multi-source heterogeneous ocean raw observation profile data, serialized and packaged into metadata, and displayed and interacted through a visualization system, wherein the key data include: observation time, location, equipment type, and observed physical parameters.

3. The ocean temperature data quality control information fusion method based on DS evidence theory according to claim 2 is characterized in that: The visualization system has the function of displaying and interacting with quality control mark positions, so that experts can realize expert quality control through this function.

4. The ocean temperature data quality control information fusion method based on DS evidence theory according to claim 1 is characterized in that: The step 2) comprises the following steps: 2.1) Introduce an expert system, build an AI-based ocean observation data quality control model, and train it; 2.2) Use the trained x AI quality control models to perform AI quality control on the metadata, and obtain the quality control mark bit vector P for each ocean observation profile i1 , where i1∈[1,x]; 2.3) Introduce y experts in the expert system, and conduct expert quality control on the metadata through data interaction on the data visualization system. For each ocean observation profile, a quality control mark vector P is obtained. i2 , where i2∈[1,y]; 2.4) Introduce z automated quality control schemes, input metadata into the automated quality control scheme, and obtain a quality control marker vector P for each ocean observation profile i3 , where i3∈[1,z].

5. The ocean temperature data quality control information fusion method based on DS evidence theory according to claim 1 is characterized in that: The step 3) comprises the following steps: 3.1) Building a recognition framework For the in-situ observation data of a certain site, there are only a limited number of M mutually exclusive hypotheses Θ in the recognition space, and the recognition framework is defined as follows: Θ={A1,A2,……,A M } For the "binary" quality control scheme, there are two mutually exclusive assumptions for the quality control marks of the observation sites. The legal observation data is recorded as assumption A1, that is, the quality control mark bit is 0, and the illegal observation data is recorded as assumption A2, that is, the quality control mark bit is 1. Therefore, the identification framework FoD is defined as Θ = {A1, A2}; for the "multivariate" data quality control scheme based on data quality grading and scoring, the identification framework FoD is defined as Θ = {A1, A2, ..., A n }, where n is the number of mutually exclusive hypotheses; let the total number of all quality control schemes be N, then N = x + y + z; 3.2) Constructing the basic confidence distribution S i S i =HIST(P i ) / SUM[HIST(P i )] Among them, P i represents the quality control mark vector for the quality control scheme i of a certain ocean observation profile, HIST represents the frequency histogram of the elements in the solution vector, and SUM represents the sum of the elements in the vector; 3.3) Constructing basic confidence allocation coefficients The quality control mark vector of each quality control scheme is taken as an evidence. Based on the experience of domain experts, the event credibility vector of each evidence is given as the basic credibility distribution coefficient, denoted as T i , where i∈[1,N], the evidence is the quality control result; 3.4) Calculate the confidence of the hypothesis and construct the evidence vector H i H i =S i ⊙T i Among them, the operator ⊙ represents the dot product of the elements in the vector; 3.5) Evidence synthesis using Dempster synthesis rule Use the fusion rule to fuse the two pieces of evidence, then use the fusion result as the new basic confidence vector to be fused, and fuse it with the third piece of evidence again, repeat the two-piece evidence fusion operation until all the evidence is fused into one piece of evidence as the composite evidence; 3.6) Using the synthesized evidence, reconstruct the annotation results of ocean observation data P i =ROUND(H i ) Among them, H i represents the final fusion evidence of the i-th ocean observation profile, P i It indicates the final quality control result of the ocean observation profile. ROUND indicates the rounding of floating point numbers.

6. The ocean temperature data quality control information fusion method based on DS evidence theory according to claim 5 is characterized in that: The fusion rules are specifically: For evidence H m and H n , integrating evidence H k The way of q-th hypothesis is: Among them, k represents the evidence synthesis conflict factor: Where M represents the number of hypotheses in the recognition framework, and H m,j and H n,j Represents evidence m and n hypotheses A respectively j Confidence measure of .

7. Ocean temperature data quality control information fusion system based on DS evidence theory, characterized by: include: The data preprocessing module is used to obtain the original multi-source heterogeneous ocean observation profile data to be quality controlled, and preprocess it to obtain metadata; Data quality control module, used to perform quality control on metadata using multiple quality control schemes; The result decision module is used to fuse the data of multiple quality control result vectors to obtain the final quality control result.

8. Ocean temperature data quality control information fusion device based on DS evidence theory, characterized in that: It comprises a memory and a processor; the memory is used to store a computer program; the processor is used to implement the ocean temperature data quality control information fusion method based on DS evidence theory as described in any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the ocean temperature data quality control information fusion method based on DS evidence theory as described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Profile data set construction method and system based on multi-source heterogeneous ocean observation

    CN120872936A