Random arrangement set data construction method and system based on two-layer reliability structure
By constructing a random permutation set data construction method based on a two-layer confidence structure, and utilizing mean and standard deviation error analysis and propensity weighting, the problem of poor adaptability of existing technologies under non-Gaussian distributed data is solved. This enables flexible modeling and robust representation of ordered uncertain information, thereby improving classification performance.
Patent Information
- Application Number
- CN202510956305.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-10-21
AI Technical Summary
Existing technologies have poor adaptability when dealing with real-world data that is not Gaussian distributed or whose distribution is unknown. As a result, the generated random permutations cannot accurately reflect the order uncertainty of the real data, which limits their application effectiveness in complex scenarios.
A random permutation set data construction method based on a two-layer credibility structure is adopted. By constructing the basic probability assignment (BPA) on the first-layer credibility structure, the attribute matching degree is calculated using the mean and standard deviation error analysis, and a random permutation set (RPS) is generated on the second-layer credibility structure. The permutation quality function (PMF) is generated in combination with the propensity weighting to achieve robust and flexible modeling of ordered uncertain information.
It significantly improves the algorithm's generalization ability and robustness in real-world scenarios, enabling it to adapt to different data characteristics, enhancing classification robustness and interpretability in complex decision-making scenarios, and solving the problem of poor adaptability of existing methods to complex data distributions.
Smart Images

Figure CN120822181A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information fusion, and in particular relates to a method and system for constructing random permutation set data based on a two-layer credibility structure. Background Art
[0002] In the field of information fusion and uncertainty modeling, how to effectively represent and process uncertain information in complex environments has always been a research hotspot. With the increasing complexity of data, traditional methods face challenges in characterizing uncertain information containing order relationships, and more refined theoretical tools are urgently needed. Random Permutation Sets (RPS), as an emerging uncertainty representation theory, provides new ideas for modeling order-structured uncertain information by introducing permutation spaces and order structures. In recent years, the successful application of RPS in fields such as situation assessment, intelligent decision-making, and pattern recognition has demonstrated its potential in handling complex uncertainty problems and promoted the development of information fusion technology towards higher dimensions and greater refinement.
[0003] To address the problem of constructing random permutation sets, some researchers have proposed a RPS generation method (RPSGM) based on a Gaussian discriminant model and weight analysis. This method assumes that the data follows a Gaussian distribution, classifies the samples using a Gaussian discriminant model, and generates a permutation quality function (PMF) in conjunction with a weight assignment strategy to construct random permutation sets. RPSGM transforms data distribution characteristics into ordered propositions in the permutation space through parametric modeling, providing a preliminary solution for representing uncertain information in the permutation space.
[0004] However, the RPSGM method has a significant limitation: its heavy reliance on the Gaussian distribution assumption. This parametric modeling approach has poor adaptability and generalization capabilities when applied to real-world data with unknown or non-Gaussian distributions. As a result, the generated random permutation sets may not accurately reflect the order uncertainty of the real data, limiting the method's application in complex scenarios. Summary of the Invention
[0005] Purpose of the invention: The purpose of the present invention is to provide a method for constructing random permutation set data based on a two-layer credibility structure that can solve the problems of poor adaptability and strong dependence of the existing technology when processing actual data with non-Gaussian distribution or unknown distribution, and realize robust and flexible modeling of ordered uncertain information; secondly, to provide a random permutation set data construction system based on a two-layer credibility structure.
[0006] Technical solution: The random permutation set data construction method of the present invention comprises the following steps:
[0007] (1) Construct the basic probability assignment (BPA) on the first-level belief structure: calculate the attribute matching degree based on the mean and standard deviation error analysis of the test sample and the training sample set; generate BPA by normalizing the matching degree to represent the initial belief distribution of the sample for each category.
[0008] (2) Generate a random permutation set (RPS) on the second-level reliability structure: calculate the propensity based on the distance between the sample and the mean of each category to characterize the order information; combine BPA and propensity, weightedly generate the permutation quality function (PMF), and construct RPS.
[0009] Efficient data classification and uncertainty modeling are achieved through a two-layer confidence structure: the first layer uses the BPA generated by the statistical characteristics of the sample (mean and standard deviation) to quantify the initial classification confidence, effectively capturing the matching relationship between data and categories; the second layer introduces order information through distance tendency, and combines the RPS constructed by BPA weighting to enhance the interpretability of classification decisions. It not only retains the characteristics of the initial confidence distribution, but also integrates the tendency information of the sample through the random permutation set, thereby achieving a more robust classification expression in an uncertain environment.
[0010] Preferably, the calculation of attribute matching based on the mean and standard deviation error analysis of the test sample and the training sample set includes:
[0011] (1.1) Construct a power set structure training sample set, according to the power set form of the identification framework Θ2 Θ ={θ1,θ2,...,θ K ,{θ1,θ2},...,{θ1,θ2,...,θ K}}Construct the corresponding power set structure training sample set:
[0012]
[0013] The training sample set L i Divided into single-category training set and multi-category training set: When the training sample set L i When it contains only samples of a single category, it constitutes a single-category training set; when the training sample set L i When it contains samples from two or more categories, it constitutes a multi-category training set;
[0014] (1.2) For the training sample set L i , the sample size is X i Indicates that for a given test sample X0, calculate its j-th attribute on the training sample set L i Matching degree:
[0015] Calculate the training sample set L i The original mean and original standard deviation on the j-th attribute:
[0016]
[0017] in Represents the training sample set L i The value of the lth training sample on the j attribute;
[0018] Add the test sample X0 to the training sample set L i , forming a new sample set L' i , calculate the new sample set L' i New mean and new standard deviation on the jth attribute:
[0019]
[0020] According to the test sample X0, join the training sample set L i The change in the mean and standard deviation before and after is calculated to calculate the effect of X0 on the training sample set L on attribute j. i Matching degree:
[0021]
[0022] By constructing a power set structured training sample set and analyzing the changes in statistical characteristics before and after the addition of test samples, fine-grained attribute matching calculation is achieved: first, the training set structure is dynamically divided based on the power set form of the identification framework to adapt to single-class or multi-class scenarios; second, by comparing the offset of the mean and standard deviation before and after the introduction of test samples, the degree of matching at the attribute level is quantified, which not only reflects the compatibility of the test sample and the training set in statistical distribution, but also captures the impact of data perturbations on the overall distribution through error changes, providing a statistically significant difference measurement basis for subsequent BPA generation.
[0023] Preferably, generating the BPA by normalizing the matching degree includes:
[0024] For the test sample X0, on the jth attribute, all training sample sets L i The calculated matching degree is normalized to obtain the basic probability assignment m j =[m ij |i=1,2,...,2 K -1],m ij =m j (A i ) means X0 belongs to category A in attribute j i ∈PS(Θ) belief value, m j (A i ) is calculated as:
[0025]
[0026] By normalizing the multi-attribute matching degree into a standardized belief distribution, the BPA is effectively constructed: normalization is performed based on the relative proportion of the matching degree of each attribute, so that the matching degree on different attributes is converted into a comparable probability distribution, thereby objectively characterizing the initial credibility distribution of the test sample for each category, retaining the original statistical difference information and eliminating the dimensionality effect, providing a standardized input that conforms to the axioms of probability theory for subsequent uncertainty reasoning.
[0027] Preferably, the RPS construction of the second layer credibility structure includes:
[0028] The tendency is calculated based on the distance between the test sample and the mean of each category. The smaller the distance, the stronger the tendency.
[0029] For ordered propositions, the composite tendency score is calculated by combining the tendency scores of the categories it contains;
[0030] BPA is weighted and fused with the comprehensive propensity score to generate PMF, which is the RPS reliability value.
[0031] In the second-level reliability structure, the ordinal enhancement of classification decision is achieved by fusion of distance-driven propensity calculation and BPA: first, the propensity is quantified by the distance between the sample and the category mean, which strengthens the priority of adjacent categories; second, hierarchical ordinal information is introduced through the comprehensive propensity calculation of ordered propositions; finally, BPA is weightedly fused with propensity to generate PMF, so that RPS not only retains the initial reliability distribution, but also embeds the spatial correlation characteristics of the samples, thereby taking into account both statistical matching and ordered correlation between categories in uncertainty reasoning, and improving the rationality and interpretability of the classification results.
[0032] Preferably, the tendency is calculated as follows:
[0033] The test sample X0 is related to the category θ on the jth attribute h The smaller the mean distance is, the higher the tendency is. The calculation formula is:
[0034]
[0035] The above formula considers the sample's tendency to a single category, so the range of h is [1, K]; for the BPA value m of the test sample X0 ij ,A i Represents the proposition or focal element of the distribution; in the RPS theoretical framework, based on A i Contains different orders of elements, and obtains A i Corresponding ordered propositions where g∈[1,|A i |! ].
[0036] The inclination of samples towards each category is quantified through the inverse distance relationship, which makes the RPS construction have clear spatial correlation: the inclination is dynamically generated based on the mean distance at the attribute level to ensure that adjacent categories obtain higher weights; at the same time, combined with the proposition structure of BPA, the single-category inclination is extended to the comprehensive evaluation of ordered propositions, which not only retains the geometric intuitiveness of distance measurement, but also is compatible with multi-category association scenarios through the hierarchical processing of ordered propositions, providing a fusion basis for the generation of PMF with both local sensitivity and global orderliness, effectively enhancing the spatial consistency of classification decisions.
[0037] Preferably, the comprehensive tendency of the ordered proposition is calculated as follows:
[0038] For ordered propositions Based on the tendency of the test sample X0 to the categories it contains, we can get its tendency to the ordered proposition The comprehensive tendency of
[0039]
[0040] Among them, g(u) represents an ordered proposition The index of the u-th element in the identification frame Θ.
[0041] Through the calculation of weighted comprehensive propensity of ordered propositions, dynamic quantification of multi-category association relationships is achieved: based on the geometric mean calculation of single-category propensity, the monotonicity of the original distance metric is retained, and the influence of key categories in the ordered structure is strengthened through index weighting. The comprehensive propensity can not only reflect the spatial distribution relationship between the test samples and each category, but also effectively characterize the synergistic effect of multi-category combinations, providing a hierarchical order information fusion mechanism for RPS construction and enhancing the rationality of decision-making in complex classification scenarios.
[0042] Preferably, the PMF is calculated as follows:
[0043] Combined BPA value m ij And the comprehensive tendency of the test sample to the ordered proposition, weighted to obtain the ordered proposition The confidence value of the random permutation set is:
[0044]
[0045] in, Indicates that X0 is an ordered proposition in attribute j likelihood or degree of belief.
[0046] Through the weighted fusion of BPA and comprehensive propensity, an organic combination of statistical matching and spatial order information is achieved: the initial belief distribution (BPA) and the propensity assessment of ordered propositions (comprehensive propensity) are dynamically integrated into PMF, so that the final generated RPS not only retains the statistical credibility of the original data matching, but also incorporates the order relationship characteristics of the samples in the category space, thereby taking into account both the accuracy of probability distribution and the logic of category association in uncertainty reasoning, significantly improving the classification robustness and interpretability in complex decision-making scenarios.
[0047] In a second aspect, the random permutation set data construction system of the present invention includes:
[0048] The first trust structure building module is used to build the basic probability allocation BPA on the first layer of trust structure, including:
[0049] The matching degree calculation unit calculates the attribute matching degree based on the mean and standard deviation error analysis of the test sample and the training sample set;
[0050] The BPA generation unit generates BPA by normalizing the matching degree to represent the initial belief distribution of samples for each category.
[0051] The second trust structure building module is used to generate a random permutation set RPS on the second trust structure, including:
[0052] The tendency calculation unit calculates the tendency based on the distance between the sample and the mean of each category to characterize the sequence information;
[0053] The PMF generation unit combines BPA and propensity to perform weighted calculation to generate the ranking quality function PMF;
[0054] RPS construction unit, constructs random permutation set RPS based on PMF.
[0055] In a third aspect, the present invention further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and execute the method for constructing random permutation set data based on a two-layer credibility structure.
[0056] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for constructing random permutation set data based on a two-layer credibility structure.
[0057] Beneficial effects: Compared with the existing technology, the present invention has the following significant advantages: 1. By constructing BPA based on non-parametric error analysis of mean and standard deviation, it gets rid of the dependence on prior assumptions such as Gaussian distribution, effectively solves the problem of poor adaptability of existing technology to complex data distribution, and combines with the propensity weighting mechanism to achieve flexible modeling of ordered uncertain information, significantly improving the generalization ability and robustness of the algorithm in real scenarios; 2. Innovatively construct a two-layer credibility structure that integrates statistical features (BPA layer) and propensity analysis (RPS layer), and for the first time combines error analysis with distance-driven order relationship modeling. This framework not only quantifies the initial belief distribution, but also strengthens it through propensity weighting. 1. The proposed method can obtain the preference relationship of ordered propositions, fill the research gap of RPS generation method, and greatly improve the model's ability to express complex and uncertain information; 2. Based on non-parametric design, BPA and RPS can be dynamically generated without relying on data distribution assumptions, and can adapt to different data characteristics. This feature enables the method to exhibit stable performance in a variety of application scenarios, overcoming the performance degradation problem of traditional parametric methods caused by distribution mismatch; 3. By interpreting ordered information as "tendency" and combining it with data-driven modeling, it not only enriches the theoretical basis of RPS, but also provides a computationally clear and structurally clear implementation framework for actual systems, with good interpretability and practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 Schematic diagram of the method flow of the present invention;
[0059] Figure 2 Flowchart of the classification algorithm based on basic probability distribution and random permutation set of the present invention;
[0060] Figure 3-4 Schematic diagram of the classification accuracy of the present invention and the existing method. DETAILED DESCRIPTION
[0061] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0062] like Figure 1 As shown, from the perspective of interpreting order information as propensity, this paper proposes a method for constructing random permutation set data based on a two-layer reliability structure. The complete logical process is divided into two parts: first, a basic probability assignment (BPA) is constructed on the first-layer reliability structure based on error analysis of mean and standard deviation; then, through distance-based analysis, the propensity information is characterized, and the BPA is integrated with the propensity information to generate a random permutation set (RPS) on the second-layer reliability structure. The detailed process is as follows:
[0063] (1) Generate basic probability distribution on the first layer of credibility structure
[0064] S1: Construct a new training sample set based on the power set structure. According to the identification framework Θ power set form: 2 Θ ={θ1,θ2,...,θ K ,{θ1,θ2},...,{θ1,θ2,...,θ K}}, construct the corresponding power set structure training sample set:
[0065]
[0066] Training set L i There are two types, namely single-category training set and multi-category training set. i Contains only one category of samples, then L i is a single-category training set. If L i If the sample contains two or more categories, it is called L i is a multi-category training set. For example, L i Represents the training samples composed of category θ1, which belongs to the single-category training set. 2Θ-1 Represents the categories θ1, θ2, ..., θ K The training set composed of all categories belongs to the multi-category training set.
[0067] S2: For L i The number of samples is X i For the test sample X0, calculate the j-th attribute pair L of the sample i Matching degree of the dataset:
[0068] S2-1: Calculate the mean and standard deviation of all samples in the dataset corresponding to the jth attribute:
[0069]
[0070] in T i The value of the lth training sample in the training set on the j attribute.
[0071] S2-2: Test sample X0 is added to L i Data set, get a new sample set L' i Similarly, calculate the new sample set L' i The corresponding mean and standard deviation on the jth attribute:
[0072]
[0073] S2-3: Add the test sample X0 to the dataset L iThe standard deviation and mean changes before and after are used to obtain the matching degree of X0 on the j attribute of the data set:
[0074]
[0075] Therefore, according to formula (2-1)-formula (2-5), the matching degree of the test sample X0 on attribute j to all training sample sets can be obtained.
[0076] S3: Based on the normalized matching degree, the basic probability assignment m generated by the test sample X0 on the jth attribute is obtained j =[m ij |i=1,2,...,2 K -1],m ij =m j (A i ) means X0 belongs to category A in attribute j i ∈PS(Θ)’s belief value.
[0077] m j (A i ) is calculated as:
[0078]
[0079] (2) Generate a random permutation set on the second-level trust structure
[0080] The second part of the present invention is to generate the permutation quality function PMF on the second-level credibility structure, which is closely related to the basic probability distribution obtained on the first-level credibility structure. Sequence information, as a sequence of symbols, can further represent the qualitative information of the sample's tendency towards different categories. Based on this explanation, the present invention proposes a tendency analysis method based on the distance between the sample and each category. When each category is represented by its mean as a statistical feature, the smaller the distance between the sample and a certain category, the stronger its tendency to belong to that category. Based on this idea, this paper integrates the analysis of the mean distance between the sample and each category into the BPA distribution framework, thereby deriving the RPS distribution on the second-level credibility structure. The detailed steps are as follows:
[0081] S4: Calculate the value of X0 for different categories θ on the jth attribute h Propensity:
[0082]
[0083] Note that what is considered here is the tendency of the sample to a single category, so the range of h is [1, K]. For the BPA value m of sample X0 ij ,A i Represents the proposition or focal element of the distribution. In the RPS theoretical framework, based on A iDifferent orders of the included elements (i.e. categories) can be obtained i Corresponding ordered propositions where g∈[1,|A i |! ]. For example, in the identification framework with cardinality 2, Θ={θ1,θ2}, A3={θ1,θ2}. Then the ordered proposition corresponding to A3 is
[0084]
[0085] S5: For ordered propositions Based on the tendency of X0 to the categories it contains, we can get its tendency to order propositions The comprehensive tendency of
[0086]
[0087] Where g(u) represents an ordered proposition The index of the u-th element in the identification frame Θ. For example, In this case, g(1)=2, g(2)=1.
[0088] S6: Combined BPA value m ij And the comprehensive tendency of the sample to the ordered proposition, weighted to obtain the ordered proposition The confidence value of the random permutation set is:
[0089]
[0090] Indicates that X0 is an ordered proposition in attribute j likelihood or degree of belief.
[0091] Based on similar inventive concepts, this example provides a random permutation set classification method based on a two-layer credibility structure, combined with Figure 1 and Figure 2 include:
[0092] S1: For any dataset X, use the five-fold cross validation method to divide it into training set and test set;
[0093] S2: Construct single-category training sets and multi-category training sets;
[0094] S3: Calculate the mean and variance of all single-category training sets and multi-category training sets on different attributes;
[0095] S4: For the sample X0 in the test set, add it to the corresponding training set and calculate the new mean and variance;
[0096] S5: Calculate the matching degree of the test sample X0 to different classes on different attributes based on the changes in mean and variance before and after the test sample X0 is added to the training set;
[0097] S6: Normalize the matching degree to obtain the basic probability assignment distribution m0 of the test sample X0 on different attributes;
[0098] S7: Calculate the tendency of the test sample X0 to a single category on different attributes;
[0099] S8: Calculate the comprehensive tendency of the test sample X0 to the ordered propositions on different attributes;
[0100] S9: Calculate the distribution of random permutation sets generated by the test sample X0 on different attributes;
[0101] S10: For each test sample, the permutation quality functions generated on all attributes are combined using the combination rule;
[0102] S11: using the probability conversion rule, convert the fused permutation quality function into a probability distribution;
[0103] S12: Determine the category of the test sample according to the principle of maximum probability.
[0104] Figure 2 It is a framework diagram of the classification algorithm constructed based on the present invention and two comparative classification algorithms (the existing BPA method and the parameterized RPS method). Figure 3 and Figure 4 The classification accuracy performance of the classification algorithm based on the present invention, the existing BPA classification method, and the parameterized RPSGM method on eight public machine learning datasets is shown. It can be seen that the average accuracy of the classification algorithm based on the present invention on all eight datasets is higher than that of the existing BPA classification method and the parameterized RPSGM method, demonstrating that the present invention has greater adaptability and robustness when processing data with different structural features, and can steadily improve the classification performance of ordered uncertain information processing.
[0105] Based on similar inventive concepts, an embodiment of the present invention further provides a random permutation set data construction system corresponding to the random permutation set data construction method, comprising:
[0106] The first trust structure building module is used to build the basic probability allocation BPA on the first layer of trust structure, including:
[0107] The matching degree calculation unit calculates the attribute matching degree based on the mean and standard deviation error analysis of the test sample and the training sample set;
[0108] The BPA generation unit generates BPA by normalizing the matching degree to represent the initial belief distribution of samples for each category.
[0109] The second trust structure building module is used to generate a random permutation set RPS on the second trust structure, including:
[0110] The tendency calculation unit calculates the tendency based on the distance between the sample and the mean of each category to characterize the sequence information;
[0111] The PMF generation unit combines BPA and propensity to perform weighted calculation to generate the ranking quality function PMF;
[0112] RPS construction unit, constructs random permutation set RPS based on PMF.
[0113] The invention also discloses an electronic device.
[0114] Specifically, the electronic device can be a computer device such as a desktop computer, a laptop computer, a PDA, and a cloud server. The computer device may include, but is not limited to, a processor and a memory. The processor and the memory may be connected via a bus or other means. The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, graphics processing units (GPU), embedded neural network processors (NPU) or other dedicated deep learning coprocessors, discrete gate or transistor logic devices, discrete hardware components and other chips, or a combination of the above-mentioned chips.
[0115] As a non-transient computer-readable storage medium, the memory can be used to store non-transient software programs, non-transient computer executable programs and modules. The processor executes various functional applications and data processing of the processor by running the non-transient software programs, instructions and modules stored in the memory. The memory may include a program storage area and a data storage area, wherein the program storage area may store a control unit, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.
[0116] The invention also discloses a computer-readable storage medium.
[0117] Specifically, the computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method in the above method implementation is implemented.
[0118] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments of the present invention can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments of the method. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD). The storage medium can also include a combination of the above-mentioned types of memory.
Claims
1. A method for constructing random permutation set data based on a two-layer credibility structure, characterized in that: The following steps are involved: (1) Constructing the basic probability assignment (BPA) on the first layer of the credibility structure: Calculating the attribute matching degree based on the mean and standard deviation error analysis of the test sample and the training sample set; Normalized matching degree generates BPA, which represents the initial belief distribution of samples for each category; (2) Generate a random permutation set (RPS) on the second-level reliability structure: calculate the propensity based on the distance between the sample and the mean of each category to characterize the order information; combine BPA and propensity, weightedly generate the permutation quality function (PMF), and construct RPS.
2. The random permutation set data construction method according to claim 1, characterized in that: The calculation of attribute matching based on the mean and standard deviation error analysis of the test sample and the training sample set includes: (1.1) Construct a power set structure training sample set, according to the power set form of the identification framework Θ2 Θ ={θ1,θ2,...,θ K ,{θ1,θ2},...,{θ1,θ2,...,θ K }}Construct the corresponding power set structure training sample set: The training sample set L i Divided into single-category training set and multi-category training set: When the training sample set L i When it contains only samples of a single category, it constitutes a single-category training set; when the training sample set L i When it contains samples from two or more categories, it constitutes a multi-category training set; (1.2) For the training sample set L i , the sample size is X i Indicates that for a given test sample X0, calculate its j-th attribute on the training sample set L i Matching degree: Calculate the training sample set L i The original mean and original standard deviation on the j-th attribute: in Represents the training sample set L i The value of the lth training sample on the j attribute; Add the test sample X0 to the training sample set L i , forming a new sample set L' i , calculate the new sample set L' i New mean and new standard deviation on the jth attribute: According to the test sample X0, join the training sample set L i The change in the mean and standard deviation before and after is calculated to calculate the effect of X0 on the training sample set L on attribute j. i Matching degree:
3. The random permutation set data construction method according to claim 2, characterized in that: Generating the BPA by normalizing the matching degree includes: For the test sample X0, on the jth attribute, all training sample sets L i The calculated matching degree is normalized to obtain the basic probability assignment m j =[m ij |i=1,2,...,2 K -1],m ij =m j (A i ) means X0 belongs to category A in attribute j i ∈PS(Θ) belief value, m j (A i ) is calculated as:
4. The random permutation set data construction method according to claim 1, characterized in that: The RPS construction of the second-layer trust structure includes: The tendency is calculated based on the distance between the test sample and the mean of each category. The smaller the distance, the stronger the tendency. For ordered propositions, the composite tendency score is calculated by combining the tendency scores of the categories it contains; BPA is weighted and fused with the comprehensive propensity score to generate PMF, which is the RPS reliability value.
5. The random permutation set data construction method according to claim 4, characterized in that: The calculation method of the propensity is: The test sample X0 is related to the category θ on the jth attribute h The smaller the mean distance is, the higher the tendency is. The calculation formula is: The above formula considers the sample's tendency to a single category, so the range of h is [1, K]; for the BPA value m of the test sample X0 ij ,A i Represents the proposition or focal element of the distribution; in the RPS theoretical framework, based on A i Contains different orders of elements, and obtains A i Corresponding ordered propositions where g∈[1,|A i |! ].
6. The random permutation set data construction method according to claim 4, characterized in that: The calculation method of the comprehensive tendency of the ordered proposition is: For ordered propositions Based on the tendency of the test sample X0 to the categories it contains, we can get its tendency to the ordered proposition The comprehensive tendency of Among them, g(u) represents an ordered proposition The index of the u-th element in the identification frame Θ.
7. The random permutation set data construction method according to claim 4, characterized in that: The PMF is calculated as follows: Combined BPA value m ij And the comprehensive tendency of the test sample to the ordered proposition, weighted to obtain the ordered proposition The confidence value of the random permutation set is: in, Indicates that X0 is an ordered proposition in attribute j likelihood or degree of belief.
8. A random permutation set data construction system based on a two-layer credibility structure, characterized in that: include: The first trust structure building module is used to build the basic probability allocation BPA on the first layer of trust structure, including: The matching degree calculation unit calculates the attribute matching degree based on the mean and standard deviation error analysis of the test sample and the training sample set; The BPA generation unit generates BPA by normalizing the matching degree to represent the initial belief distribution of samples for each category. The second trust structure building module is used to generate a random permutation set RPS on the second trust structure, including: The tendency calculation unit calculates the tendency based on the distance between the sample and the mean of each category to characterize the sequence information; The PMF generation unit combines BPA and propensity to perform weighted calculation to generate the ranking quality function PMF; RPS construction unit, constructs random permutation set RPS based on PMF.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing random permutation set data based on a two-layer credibility structure according to any one of claims 1 to 7 is implemented.
10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein: When the processor executes the program, the random permutation set data construction method based on a two-layer credibility structure according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
RPS grouping prediction method and device based on depth characterization and probability modeling
CN121435398A