Improved d-s evidence theory fusion method and system

By using an improved DS evidence theory fusion method, fuzzy naive Bayes and FCM algorithms are used to generate BPA, and Dempster synthesis rules and Pignistic transformation are used to solve the problems of evidence conflict and high computational complexity in multi-source heterogeneous data fusion, thus achieving more efficient and accurate data fusion.

CN119848766BActive Publication Date: 2026-01-20GUIZHOU POWER GRID CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411911477.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2026-01-20
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

In the process of multi-source heterogeneous data fusion, existing technologies are unable to generate robust basic probability assignments, resulting in insufficient ability to handle evidence conflicts, high computational complexity, and low classification accuracy.

Method used

An improved DS evidence theory fusion method is adopted. By dividing multidimensional data into independent information sources according to attributes, fuzzy naive Bayes method and FCM algorithm are used to construct single and composite hypothesis BPAs. Dempster synthesis rule is used to fuse BPAs of different attributes, and the comprehensive BPA is converted into Pignistic probability to determine the predicted category of test sample.

Benefits of technology

It effectively addresses conflicting evidence, improves the reliability and accuracy of fusion results, reduces computational complexity, and is suitable for efficient integration and processing of complex data scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119848766B_ABST
    Figure CN119848766B_ABST
Patent Text Reader

Abstract

The application discloses an improved D-S evidence theory fusion method and system, relates to the field of data fusion and information processing, and comprises the following steps: dividing multi-dimensional data into independent information sources according to attributes to generate basic data for calculating BPA; using a fuzzy naive Bayes method and an FCM algorithm to construct BPA of single and compound hypotheses; fusing BPA of different attributes through a synthesis rule to obtain comprehensive BPA; and converting the comprehensive BPA into Pignistic probability to determine the prediction category of a test sample according to the maximum value. Through defining an uncertain area and constructing BPA of single and compound hypotheses, the application effectively deals with evidence conflict, improves the reliability of fusion results, uses the fuzzy naive Bayes method and the FCM algorithm to generate BPA, optimizes evidence weight distribution through weight adjustment to reduce complexity, and through probability conversion and the maximum value decision method, the classification result is more accurate and stable.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of data fusion and information processing, in particular to an improved D-S evidence theory fusion method and system. BACKGROUND

[0002] Data preprocessing is an essential step in all data fusion, however, in the literature of data fusion, there is generally not enough attention to the work of this stage. The result of data preprocessing as the data source of data fusion, its quality will directly affect the result of data fusion, a good preprocessing result not only can make the result of fusion more accurate, but also can improve the fusion speed.

[0003] In practical application, due to the multi-source heterogeneity of data, the information collected directly from each data source will exist some problems in different degrees, such as the integrity, uniqueness, authority, consistency of data, the non-uniformity of data dimension, irrelevant information (noise), field redundancy or multiple index values, etc. These problems will lead to high cost of subsequent operation, inaccurate decision, etc., so data preprocessing is an essential link of data fusion. Data preprocessing has various methods of classification, according to the work content, it can be divided into: data cleaning, data integration and transformation, data reduction, etc., these technologies provide an important guarantee for the subsequent data fusion operation, and also improve the performance of fusion. SUMMARY

[0004] In view of the above problems, the present application is proposed.

[0005] Therefore, the technical problem solved by the present application is: how to generate robust basic probability assignment and efficiently handle evidence conflict in the fusion process of multi-source heterogeneous data, realize the accuracy and reliability of classification decision.

[0006] To solve the above technical problems, the present application provides the following technical scheme: an improved D-S evidence theory fusion method, comprising the following steps,

[0007] Divide the multi-dimensional data into independent information sources according to the attributes, generate the basic data for calculating BPA;

[0008] Use fuzzy naive Bayes method and FCM algorithm to construct the BPA of single and compound hypothesis;

[0009] Fuse the BPA of different attributes by Dempster combination rule to obtain the comprehensive BPA;

[0010] Convert the comprehensive BPA into Pignistic probability, and determine the predicted category of the test sample according to the maximum value.

[0011] As a preferred scheme of the improved D-S evidence theory fusion method, the method comprises the following steps:

[0012] For a given multi-dimensional data set, each single attribute is an independent information source, and the BPA of each attribute is calculated and synthesized by using the D-S evidence theory to obtain a reliable comprehensive decision.

[0013] As a preferred scheme of the improved D-S evidence theory fusion method, the method comprises the following steps:

[0014] The BPA function is generated by taking the set C as a recognition framework, and the expression is as follows:

[0015] Θ=C={C1,C2,…,C n}

[0016] The focal elements of the power set 2Θ of the recognition framework are as follows:

[0017] Ω={{C1},…,{C N},{C i ,C j},…,{C N-1 ,C N}}

[0018] Wherein, the compound element {C i ,C j}(i≠j) is an uncertain hypothesis in the D-S evidence theory;

[0019] The ROU is used to represent the compound hypothesis {C i ,C j}, and the uncertain data is divided;

[0020] For each attribute, N Gaussian distributions and C N2 ROU functions are obtained as the models of single hypotheses and compound hypotheses, respectively.

[0021] As a preferred scheme of the improved D-S evidence theory fusion method, the method comprises the following steps:

[0022] When the fuzzy membership value is calculated by using the fuzzy naive Bayes method and the FCM algorithm to distribute the basic probability to each focal element, the fuzzy membership value The membership degree of each attribute belonging to different categories, given an input sample, the membership value calculated for attribute x is:

[0023]

[0024] For the composite hypothesis {C i , C j}, the variance of the membership degree of each fuzzy partition after classification is calculated, and the expression is:

[0025]

[0026] Where M is the expectation;

[0027] The membership matrix is:

[0028]

[0029] Set a variance value D(u) as a threshold, and the average value of the variance of the membership of each row of the membership matrix U in the partition is taken as the value of D(u);

[0030] When D(u i )<D(u), the current sample has the properties of two class labels at the same time.

[0031] As a preferred scheme of the improved D-S evidence theory fusion method, wherein: the BPA constructed by the fuzzy naive Bayes method and the FCM algorithm for single and composite hypotheses further includes,

[0032] The object in the uncertain region belongs to both C i and C j , then the quality function related to the composite hypothesis is assigned using the fuzzy AND operator, and the basic probability assignment function of each hypothesis calculated by the fuzzy naive Bayes method is:

[0033]

[0034] Where μ{Ci} represents the membership value belonging to Ci;

[0035] According to the FCM algorithm, the Euclidean distance between the input sample and the class centroid vector is used to determine the basic probability of the discriminant class, and ROU is used as the composite hypothesis to define the intersection point V{C i , C j} as the class centroid of the composite hypothesis, which is a point with the smallest AND value calculated by the distribution of two different classes C i , C j , and the expression is:

[0036]

[0037] The method for defining the discriminant BPA function is to use the exponential function of the sample and the centroid distance of the class, and the expression is:

[0038]

[0039] Different evidences are collected and integrated using a weighted adjustment framework;

[0040] The generated class BPA generated by the fuzzy naive Bayes method is: And the discriminant class BPA based on distance is: The integration expression is:

[0041]

[0042] Wherein, 0≤α,β≥1 are the adjustment parameters for adaptively determining the importance of two classes of evidence;

[0043] The optimal adjustment parameters are found by using grid search to minimize the training error.

[0044] As a preferred scheme of the improved D-S evidence theory fusion method, wherein: the BPA generated by different attributes is fused by the Dempster combination rule to obtain a comprehensive BPA, which includes,

[0045] The BPA generated by different attributes is combined by the Dempster combination rule to obtain a comprehensive BPA;

[0046] The Dempster combination rule includes single-class combination and multi-class combination;

[0047] The single-class combination includes setting a proposition A, The proposition A is combined with two mass functions m1 and m2 on Θ on the same identification framework, and the combination rule expression is:

[0048]

[0049] Wherein, the symbol Indicates the orthogonal sum;

[0050] If K=1, there is a conflict between the two evidences, so there is no orthogonal sum; if K≠1, the orthogonal sum of the BPA of the two evidences forms a new distribution function, and if K-1=0, m1 and m2 are contradictory, and there is no joint basic probability distribution function;

[0051] The multi-class combination includes that when a plurality of evidence sources of the identification framework Θ need to be processed, a plurality of mass functions m1 and m2, The method for obtaining the orthogonal sum of the plurality of basic probability assignment functions as a basic belief function when m is a plurality of basic probability assignment functions is a combination rule expression:

[0052]

[0053] If two combination objects are not completely conflicting, the orthogonal sum of any two functions is established, regardless of the combination order, and the combination result of the plurality of evidences is certain, and the combination mode of the plurality of evidences is derived from the combination formula of two evidences.

[0054] As a preferred scheme of the improved D-S evidence theory fusion method, the method comprises the following steps of:

[0055] The Pignistic conversion is used for decision making, and the hypothesis class with the maximum Pignistic probability is selected as the prediction class of the sample in the pseudo test data.

[0056] The Pignistic conversion is a mass function conversion into a Pignistic probability function, m(A) is a basic probability assignment function defined on the identification framework Θ, and the Pignistic probability function on the identification framework Θ is Bet P m : Θ→[0, 1], and the expression is:

[0057]

[0058] Wherein, B is a time B.

[0059] Another object of the present application is to provide an improved D-S evidence theory fusion system, which can solve the problems of insufficient conflict processing capability, high computational complexity and low classification accuracy of the existing D-S evidence theory by attribute division of multi-dimensional data, BPA generation of single and compound hypotheses, and evidence fusion and Pignistic probability decision based on the Dempster rule.

[0060] To solve the above technical problems, the present application provides the following technical scheme: an improved D-S evidence theory fusion system, comprising: a data preprocessing module, a BPA generation module, an evidence fusion module and a decision module.

[0061] The data preprocessing module is used for completing the division of multi-dimensional data, regarding each attribute as an independent information source, and generating BPA to provide basic data.

[0062] The BPA generation module is a BPA for generating single hypothesis and compound hypothesis by using a fuzzy naive Bayes method and an FCM algorithm, defining an uncertain region, allocating compound hypothesis probability, constructing a Gaussian distribution model by calculating membership and variance, and determining quality functions of each category and compound category.

[0063] The evidence fusion module is a module for integrating BPA of different attributes by using a Dempster combination rule, generating comprehensive BPA, processing single category and multiple category orthogonal and calculation, and solving evidence conflict problems.

[0064] The decision module is a module for converting comprehensive BPA into Pignistic probability distribution of each hypothesis by Pignistic conversion, selecting a predicted category of the test sample according to a maximum Pignistic probability value, converting a mass function into a Pignistic probability function, and directly outputting a decision result.

[0065] A computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the steps of the improved D-S evidence theory fusion method when executing the computer program.

[0066] A computer readable storage medium stores a computer program, and the computer program implements the steps of the improved D-S evidence theory fusion method when executed by a processor.

[0067] The present application has the following beneficial effects: the present application defines an uncertain region (ROU) and constructs BPA of single and compound hypothesis, effectively deals with evidence conflict, improves reliability of fusion results, generates BPA by using a fuzzy naive Bayes method and an FCM algorithm, optimizes evidence weight distribution by weight adjustment, greatly reduces complexity, makes classification results more accurate and stable by Pignistic probability conversion and maximum value decision method, is suitable for complex data scenarios by multi-dimensional attribute division and independent construction of models, and realizes efficient integration and processing of multi-source data. BRIEF DESCRIPTION OF DRAWINGS

[0068] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. Among them:

[0069] Figure 1 The overall flowchart of the improved D-S evidence theory fusion method provided for the first embodiment of the present application;

[0070] Figure 2 An improved D-S evidence theory fusion method for a second embodiment of the present application provides an uncertain area schematic diagram. DETAILED DESCRIPTION

[0071] In order to make the above objectives, characteristics and advantages of the present application more apparent, obvious and easy to understand, the specific embodiments of the present application are described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should fall within the scope of protection of the present application.

[0072] Embodiment 1, reference Figure 1 For an embodiment of the present application, an improved D-S evidence theory fusion method is provided, comprising:

[0073] Divide the multi-dimensional data into independent information sources according to attributes, and generate basic data for calculating BPA;

[0074] Construct BPA of single and compound hypotheses by using fuzzy naive Bayes method and FCM algorithm;

[0075] Fuse BPA of different attributes by Dempster combination rule to obtain comprehensive BPA;

[0076] Convert the comprehensive BPA into Pignistic probability, and determine the predicted category of the test sample according to the maximum value.

[0077] For a given multi-dimensional data set, each single attribute is an independent information source, and BPA of each attribute is calculated respectively, and then D-S evidence theory is used for combination to obtain reliable comprehensive decision. After attribute division, the test data set with p attributes is divided and converted into p independent models in the present application.

[0078] The BPA function is generated by taking the set C as a recognition framework, and the expression is:

[0079] Θ=C={C1,C2,…,C n}

[0080] The focal element of the power set 2Θ of the recognition framework is expressed as:

[0081] Ω={{C1},…,{C N},{C1,C2},…,{C i ,C j},…,{C N-1 ,C N}}

[0082] Wherein, the compound element {Ci , C j}(i≠j) is the uncertainty hypothesis in D-S evidence theory;

[0083] ROU is used to represent the compound hypothesis {C i , C j} to divide the uncertainty data;

[0084] To understand the compound elements in the recognition framework more intuitively, the following uses intuitive belief assignment to model the uncertainty of the hypothesis. For each class, a Gaussian distribution is used to model, as shown in Figure 2 , represents the membership degree of the k-th attribute belonging to class C i or C j . The left dark green and right blue areas represent the Gaussian distribution of C i and C j , respectively, and the overlapping area in the middle is the uncertainty region (ROU, Region of Uncertainty), so the samples falling in the ROU are difficult to identify because they have a large degree of properties of two different classes at the same time, so the recognition task of this part of the sample may produce classification errors. Therefore, it is obvious that we need to use ROU to represent the compound hypothesis {C i , C j} to divide the uncertainty data. In this way, for each attribute, N Gaussian distributions and C N2 ROU functions can be obtained as the models of single hypothesis and compound hypothesis, respectively.

[0085] When calculating the basic probability assigned to each focus element using the fuzzy naive Bayes method and the FCM algorithm, the fuzzy membership value represents the degree of each attribute belonging to different classes, and given an input sample, for attribute x, the calculated membership value is:

[0086]

[0087] For the compound hypothesis {C i , C j}, the variance of the membership under each fuzzy partition is calculated after classification, and the expression is:

[0088]

[0089] where M is the expectation;

[0090] The membership matrix is:

[0091]

[0092] A variance value D(u) is set as the threshold, the average of the variance of the membership of each row of the membership matrix U is taken as the value of D(u) when partitioning;

[0093] When D(u i )<D(u), the current sample has the property of two class labels at the same time.

[0094] Since the object in the uncertain region can belong to both C i and C j , a fuzzy AND operator is used to assign the mass function related to the compound hypothesis, and the basic probability assignment function of each hypothesis calculated by the fuzzy naive Bayes method is:

[0095]

[0096] Where μ{Ci} represents the membership value belonging to Ci.

[0097] Similarly, without proper normalization, formulas 5 and 6 can not produce valid BPA. For the ∧ operation, any T-Norm can be used, and the minimum value is chosen as the T-Norm in the present application.

[0098] Next, according to the FCM algorithm, the Euclidean distance between the input sample and the class centroid vector is used to determine the discriminant class basic probability, and ROU is used as the compound hypothesis to define the intersection point V{C i , C j} as the class centroid of the compound hypothesis, which is the point with the minimum AND value calculated by the distribution of two different classes C i , C j , and the expression is:

[0099]

[0100] The method of defining the discriminant BPA function is to use the exponential function of the distance between the sample and the class centroid, and the expression is:

[0101]

[0102] In order to make the framework more flexible and better in practical application, a weighted adjustment framework is proposed to collect and integrate different evidences. The generated class BPA produced by the fuzzy naive Bayes method: and the discriminant class BPA based on distance: The integration expression is:

[0103]

[0104] where 0≤α,β≥1 are adaptive adjustment parameters for the importance of two types of evidence; this weighting adjustment mechanism can enable us to find appropriate weights for different evidence sources from training data, and grid search is used to minimize training error to find the optimal adjustment parameters, which are not described in detail here. In addition, it is still necessary to emphasize that m xg and m xd Not the final BPA itself.

[0105] Overall BPA of feature x: m x The definition of ({·}) is:

[0106]

[0107] where K is a normalization factor, and for each attribute there is an optimal set (α,β) corresponding to it:

[0108]

[0109] The BPA generated by different attributes is combined to obtain the comprehensive BPA using the Dempster combination rule;

[0110] Further to be explained:

[0111] To obtain a comprehensive decision, a method is needed to calculate the comprehensive influence of multiple evidences on each hypothesis in the recognition framework, and to obtain the comprehensive trust degree of the hypothesis under the action of multiple evidences. BPA is the basis of belief function and likelihood function. In practical applications, there are often different data sources for the same hypothesis or problem, so two or more different BPA are obtained. Therefore, in order to better utilize the likelihood function and the belief function to measure the credibility of propositions in the future, we need a combination rule to combine BPA from different data sources.

[0112] Dempster's evidence combination rule is also known as evidence combination formula. Sharer made further research on the basis of Dempster and extended the combination rule of evidence theory to more general application scenarios, and proposed

[0113] Dempster's combination rule. In evidence theory, if we can get the belief function or BPA on the same recognition framework Θ when several evidences are not completely conflicting, we can use Dempster's combination rule to calculate a new belief function, which is generated by the above several evidences.

[0114] The Dempster combination rule includes single category combination and multiple category combination;

[0115] Single category synthesis includes, set a proposition A, Proposition A for the same identification framework on the Θ of two quality function m1, m2, synthesis rule expression is:

[0116]

[0117] Where, symbol Indicates the orthogonal sum;

[0118] If K = 1, there is a conflict between the two evidence, so there is no orthogonal sum; if K ≠ 1, the orthogonal sum of the BPA of the two evidence forms a new distribution function, if K-1 = 0, m1, m2 contradiction, no joint basic probability distribution function;

[0119] Multiple category synthesis includes, when the need to deal with multiple evidence sources on the identification framework Θ of finite mass function m1, m2, mn, the method of the orthogonal sum of multiple basic probability distribution functions into a basic trust function, synthesis rule expression is:

[0120]

[0121] From the above formula, if the two synthesis objects are not completely conflicting, the orthogonal sum of any two functions is computable and valid, and the mathematical properties of the synthesis rule satisfy the associative law and commutative rate, so the synthesis result of multiple evidence is certain regardless of the synthesis order, and the synthesis method of multiple evidence can be derived from the synthesis formula of two evidence, refer to Figure 2 .

[0122] Further to be explained is:

[0123] According to the working process of D-S evidence theory introduced above, how to make decisions after completing the combination of evidence, which needs to be analyzed in practical application, depending on the application field. In theory, there are generally the following categories:

[0124] (1) based on the decision of trust function.

[0125] According to the quality function obtained by evidence synthesis, the trust function Bel is calculated, and in general case, the decision result is obtained according to the function value, taking the largest one of the trust degree.

[0126] If we want to further reduce the range of decision set, we can use the "minimum point" principle: the belief function of set A is Bel(A), if the set B is obtained by removing an element bi from A, and |Bel(B)-Bel(A)|<ε, then the element bi can be removed, where ε is a small enough constant. Repeat this operation until no element can be removed.

[0127] (II) Decision based on basic belief satisfies:

[0128]

[0129]

[0130] If: then A1 is the fusion decision result, where ε1, ε2 are pre-set thresholds.

[0131] (III) Decision based on minimum risk

[0132] The identification framework Θ = {θ1, θ2, …, θ n}, the decision set A = {a1, a2, …, a p} in the state θ l , if the selected decision is a i , the risk function is r(a i , x l ), i = 1, 2, …, p, 1 = 1, 2, …, q, if the evidence E produces the corresponding mass function m(A1, …A n ) on Θ, let:

[0133]

[0134] If there exists a k ∈A such that a k = arg min{R(a1), …, R(a p )}, then a k is the optimal decision.

[0135] (IV) Method of class probability function

[0136] The class probability function method is a quantitative discrimination method, first the definition of class probability function is:

[0137]

[0138] After D-S evidence fusion, the class probability function is used as the point estimate of probability P(A), and then the maximum a posteriori probability or minimum Bayesian cost criterion is used to obtain the decision.

[0139] Further to be explained is:

[0140] D-S evidence theory has a strong theoretical basis, and has high processing capacity in the face of uncertainty, fuzziness and incomplete information, can continuously narrow the scope of assumptions through cumulative evidence; when fusing evidence data, it can not need the density of conditional probability and prior probability; it can describe the support of evidence through evidence interval, and has good processing capacity for uncertainty caused by randomness and fuzziness.

[0141] In D-S evidence theory, the number of focal elements in the identification framework Θ will affect the computational complexity of fusion calculation, because the number of focal elements is related to the operation number of evidence synthesis, which will cause the exponential growth trend of computational complexity, which has great limitations in applications with large-scale data. There are usually two ways to solve this problem: reducing the number of focal elements using experience or known information before evidence synthesis, or combining simpler assumptions in the identification framework first and then performing the following operations; For some specific data structures, targeted algorithms can be constructed to quickly process, such as tree structure, approximate calculation, or modification of DS synthesis rules.

[0142] Because of the definition of evidence theory, the evidence synthesis rule cannot be used when there is a large conflict between the evidence involved in the synthesis, otherwise unreliable conclusions will be drawn, which greatly limits the further widespread use of evidence theory. To solve this problem, many scholars have carried out research and proposed some solutions, the main idea is to analyze the reasons for the conflict and modify it according to the application background. Among them, the reasons for the conflict of evidence are as follows:

[0143] (1) The target identification framework is incomplete. Since the generation of target identification depends on people's overall understanding of assumptions, but people's cognitive range is limited, there may be cases such as ignoring a certain assumption or being unable to know the emergence of a new thing, which will result in the lack of some assumptions in the target identification framework, so the description of the target by multi-source data will deviate from the actual situation, resulting in conflict and unable to make accurate decisions.

[0144] (2) There is unreliability in sensor data. Sensors have uncontrollable environment, and when faced with some human or other interference, false information is inevitable, so conflicts will arise between evidence. In addition to external environment, internal factors of sensors such as their own precision, performance, etc. may also get inaccurate evidence information and cause evidence conflict.

[0145] (3) The unreasonable construction of basic probability assignment function. Since the fusion results of D-S evidence theory are constructed based on BPA, and the establishment of BPA is very complex, if the construction method of BPA is not appropriate, the quality function obtained cannot truly reflect the trust of evidence on the hypothesis, and may also cause conflict and affect the fusion result.

[0146] The probability assignment problem in D-S evidence theory refers to the establishment of BPA, which is a very important process of the theory. Although the establishment method of BPA is different in different application scenarios, it must complete the mapping of m: 2 Θ →[0, 1] through the function. The construction of basic probability assignment function has the following two types:

[0147] (1) Neural network. Neural network has high operation speed, strong associative ability, adaptability, fault tolerance and self-organization ability, and can well utilize prior knowledge through training and learning. After training, the data range can be controlled in [0, 1] when processing data, which well meets the basic probability assignment function. But the use of neural network needs certain prior knowledge and certain training time.

[0148] (2) Mathematical calculation. The construction of probability assignment function can also be constructed by converting data between [0, 1] through linear and nonlinear functions in mathematics. The selection of mathematical function should be determined according to the specific application, such as using linear method to correlate image features with image database to be analyzed in image data processing scene, and constructing BPA through the correlation coefficient between them. This method is simple and fast in operation. Nonlinear function has the possibility of losing source data information in the process of processing data, and the obtained data information must also be normalized.

[0149] Using Pignistic conversion to make decision, the hypothesis class with the maximum Pignistic probability is selected as the predicted class of the sample in the pseudo test data;

[0150] It should be further pointed out that:

[0151] Pignistic conversion is to convert mass function into Pignistic probability function, m(A) is a basic probability assignment function defined on the identification framework Θ, and the Pignistic probability function on the identification framework Θ is BetP m : Θ→[0, 1], the expression is:

[0152]

[0153] Among them, B is time B.

[0154] The methods described above quantify the evidence from each information source and construct basic probability assignment functions for single and composite hypotheses, respectively. In our method, ROU is used to define composite hypotheses, and a weighted adjustment architecture is employed to assign probabilities to single and composite hypotheses to account for the characteristics of different sources. In practical applications, a training mechanism can be used to find appropriate weighting coefficients (α, β) and apply them to different evidence categories.

[0155] Example 2, an embodiment of the present invention, provides a system for an improved DS evidence theory fusion method, comprising: a data preprocessing module, a BPA generation module, an evidence fusion module, and a decision module;

[0156] The data preprocessing module completes the division of multidimensional data, treating each attribute as an independent information source, and generates BPA to provide basic data.

[0157] The BPA generation module uses the fuzzy naive Bayes method and FCM algorithm to generate BPAs for single and composite hypotheses, defines the probability of composite hypotheses in uncertain regions, constructs a Gaussian distribution model by calculating membership degree and variance, and determines the quality function of each category and composite category.

[0158] The evidence fusion module uses Dempster's synthesis rules to integrate BPAs of different attributes, generate a comprehensive BPA, handle orthogonal sum calculations for single and multiple categories, and resolve evidence conflict issues.

[0159] The decision module transforms the comprehensive BPA using a Pignistic transformation to calculate the Pignistic probability distribution of each hypothesis. It then selects the prediction category of the test sample based on the maximum Pignistic probability value, converts the mass function into a Pignistic probability function, and outputs the decision result intuitively.

[0160] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0161] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0162] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0163] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0164] Example 3: In this example, to verify the beneficial effects of the present invention, scientific demonstration is conducted through economic benefit calculations and simulation experiments. This example compares the existing conventional methods with the method of this example.

[0165] The performance of the proposed weighted fuzzy DS evidence theory is compared with that of four fusion algorithms: two data fusion methods based on DS evidence theory: k-Nearest Neighbor DS theory (KNN-DST, K-Nearest Neighbor Classification Rule Based on Dempster-Shafer Theory) and a Normal Distribution-Based Classifier (NDBC); and two ordinary fusion methods: Naive Bayes (NBC). Bayes Classifier and NearestMean Classifier (NMC).

[0166] The datasets used were still the three real-world datasets from the UCI Machine Learning Library: the Iris dataset, the Wine Quality dataset, and the Forest Type Mapping dataset. Basic information about these UCI datasets is shown in Table 1. For each comparison model, five-fold cross-validation was applied in the experiments: 80% (four-fold) of the dataset was randomly selected to build the training dataset, while the remaining 20% ​​(one-fold) was used as the test data. This process was repeated five times, and the average of the five runs was used for comparison.

[0167] Table 1 Basic Information of Experimental Dataset

[0168] Dataset Number of samples Number of classes Number of attributes Iris plant 150 3 4 Quality of red paint 178 3 13 Forest type 326 4 27

[0169] The algorithm proposed in this paper was tested on several datasets, and the average classification accuracy was 94.46% ± 3.94%, which is higher than other schemes and lower than other schemes. The specific data are shown in Table 2.

[0170] Table 2 Comparison of Classification Accuracy

[0171]

[0172] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. An improved D-S evidence theory fusion method, characterized in that, The method comprises the following steps: Divide the multi-dimensional data into independent information sources according to attributes, and generate basic data for calculating BPA; Construct BPA of single and compound hypotheses by using fuzzy naive Bayes method and FCM algorithm; Fuse BPA of different attributes by Dempster combination rule to obtain comprehensive BPA; Convert the comprehensive BPA into Pignistic probability, and determine the predicted category of the test sample according to the maximum value; The step of dividing the multi-dimensional data into independent information sources according to attributes and generating basic data for calculating BPA comprises the following steps: For a given multi-dimensional data set, each single attribute is an independent information source, and BPA of each attribute is calculated, and then D-S evidence theory is used for combination to obtain reliable comprehensive decision; after the attribute division, the test data set with p attributes is divided and converted into the p independent models in the application; The step of generating basic data for calculating BPA comprises the following steps: The BPA function is generated by taking set C as a recognition framework, and the expression is as follows: , Identify the framing idempotency set The focal element is represented as: , wherein the composite element is an uncertain hypothesis in the D-S evidence theory; Using ROU to represent composite hypotheses Partitioning of uncertainty data; For each attribute, N Gaussian distributions and ROU functions are obtained as models for the single and composite hypotheses, respectively.

2. The improved D-S evidence theory fusion method of claim 1, wherein: The step of constructing BPA of single and compound hypotheses by using fuzzy naive Bayes method and FCM algorithm comprises the following steps: Fuzzy membership values are calculated using a fuzzy naive Bayes method and FCM algorithm represent the degree to which each attribute belongs to different classes, given For a given input sample, the membership value calculated for attribute x is as follows: , For the composite hypothesis The variance of the membership of each fuzzy partition is calculated after classification, expressed as: , Wherein, M is expectation; The membership matrix is as follows: , Setting a variance value As a threshold, the average of the variance of the membership of each row of the membership matrix at the time of division As the value of the membership matrix As the value of the membership matrix When then the current sample has the property of having both class labels simultaneously.

3. The improved D-S evidence theory fusion method of claim 2, wherein: The step of constructing BPA of single and compound hypotheses by using fuzzy naive Bayes method and FCM algorithm further comprises the following steps: Objects in the uncertain region belong to both class and class, the fuzzy AND operator is used to assign the quality function related to the compound hypothesis, and the basic probability assignment function of each hypothesis calculated by the fuzzy naive Bayes method is: , , wherein, represents a membership value belonging to the class. According to the FCM algorithm, the Euclidean distance between the input sample and the class centroid vector is used to determine the basic probability of the discriminant class , and the intersection point is defined as the ROU under the compound hypothesis The class centroid of the compound hypothesis is the point with the minimum AND value calculated by the distribution of two different classes , , and the expression is , The method for defining the discriminant type BPA function is to use the exponential function of the distance between the sample and the category centroid, and the expression is as follows: , , Different evidences are collected and integrated by using a weighted adjustment framework; the generated class by the fuzzy naive Bayes method and the distance-based discriminative class The integrated expression is: , wherein , is an adaptive determination of the two types of evidence importance adjustment parameters; The optimal adjustment parameter is found by using grid search to minimize the training error.

4. The improved D-S evidence theory fusion method of claim 3, wherein: The step of fusing BPA of different attributes by Dempster combination rule to obtain comprehensive BPA comprises the following steps: The BPA generated by different attributes is combined by using Dempster combination rule to obtain comprehensive BPA; The Dempster combination rule is single-class combination and multi-class combination; Single category synthesis includes, setting a proposition A, , proposition A for the same identification framework on Two quality functions on , , the synthesis rule expression is: , wherein the symbol denotes the orthogonal sum; If K = 1, then the two pieces of evidence are in conflict, so there is no orthogonal sum; if K ≠ 1, then the orthogonal sum of the BPA of the two pieces of evidence forms a new distribution function, and if K−1=0, then , there is no joint basic probability assignment function. Diverse class synthesis includes, when a recognition framework is required to be processed Summing multiple evidence sources of limited mass functions , , , functions into one basic probability assignment function The method of the belief function is as follows: , If two combination objects are not completely conflicting, the orthogonal sum of any two functions is established, regardless of the combination order, and the combination result of multiple evidences is certain, and the combination mode of multiple evidences is derived from the combination formula of two evidences.

5. The improved D-S evidence theory fusion method of claim 4, wherein: The step of converting the comprehensive BPA into Pignistic probability and determining the predicted category of the test sample according to the maximum value comprises the following steps: The Pignistic conversion is used to make a decision, and the hypothesis category with the maximum Pignistic probability is selected as the predicted category of the sample in the test data; The Pignistic transformation is a transformation of a mass function into a Pignistic probability function, , is a basic probability assignment function defined on a recognition framework The Pignistic probability function on the recognition framework is with expression , Wherein, B is time B.

6. A system employing the improved D-S evidence theory fusion method according to any one of claims 1 to 5, characterized in that: The method comprises a data preprocessing module, a BPA generation module, an evidence fusion module and a decision module. The data preprocessing module divides the multi-dimensional data, takes each attribute as an independent information source, and generates BPA to provide basic data; The BPA generation module generates BPA of single and compound hypotheses by using fuzzy naive Bayes method and FCM algorithm, defines the probability of the compound hypothesis in the uncertain region, constructs a Gaussian distribution model by calculating the membership and variance, and determines the quality function of each category and compound category; The evidence fusion module is to integrate BPA of different attributes using Dempster combination rule to generate a comprehensive BPA, process single category and multi-category orthogonal and calculation, and solve the evidence conflict problem. The decision module is to convert the comprehensive BPA into Pignistic probability distribution of each hypothesis through Pignistic conversion, select the prediction category of the test sample according to the maximum Pignistic probability value, convert the mass function into the Pignistic probability function, and directly output the decision result. 7.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-6 when the computer program is executed by the processor. The processor executes the computer program to realize the steps of the improved D-S evidence theory fusion method in any one of claims 1 to 5.

8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the steps of the improved D-S evidence theory fusion method in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Evidence conflict measurement system and method based on relevancy function

    CN105989228A

  • Weighted fuzzy D-S (Dempster-Shafer) evidence theory framework

    CN108763793A