Advantage relation rough set classification rule acquisition method, system, device and medium

By constructing and coordinating the decision class approximation set of the distributed preference information system, the problem of low data processing efficiency in the existing technology is solved, and efficient classification rule acquisition under the complete advantage relationship is achieved, ensuring the rationality and accuracy of global decision-making.

CN114548233BActive Publication Date: 2025-11-18WUYI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210096851.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-25
Publication Date
2025-11-18
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

Existing rough set methods based on dominance relations cannot effectively address the incompatibility problem in multi-attribute preference decision-making information systems when processing big data, and cannot directly process ordered distributed preference information systems under complete dominance relations, resulting in low data processing efficiency.

Method used

By constructing a distributed preference information system, coordinating the decision class approximation sets of various preference sub-information systems, updating decision class approximation sets with inconsistencies, and merging them to generate the decision class approximation set of the distributed preference information system, the consistency of advantage relationships and data processing efficiency are ensured.

Benefits of technology

It effectively solves the incompatibility problem of multi-attribute preference decision information systems, improves data processing efficiency, ensures the rationality and accuracy of global classification rules, and avoids unnecessary repetitive calculations in global knowledge acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114548233B_ABST
    Figure CN114548233B_ABST
Patent Text Reader

Abstract

The application provides an advantage relation rough set classification rule acquisition method, system, device and medium, which comprises the following steps: constructing a corresponding distributed preference information system according to the acquired distributed data set and the preset preference attribute; determining a plurality of decision classes of the distributed preference information system through the advantage relation rough set method according to the classification labels of the distributed data set; acquiring the decision class approximation set of each preference sub-information system; obtaining the decision class approximation set of the distributed preference information system according to the decision class approximation set of each preference sub-information system; and obtaining the classification rule, so that the required global knowledge is synthesized based on the existing local knowledge on the premise that the advantage relation consistency is not affected by the data size, unnecessary repeated calculation of global knowledge acquisition is avoided, the data processing efficiency is effectively improved, the inconsistency problem of the preference multi-attribute decision information system is simply and effectively solved, and the rationality and effectiveness of the global classification rule acquisition are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of big data and machine learning, and particularly relates to an advantage relation rough set classification rule acquisition method, system, computer device and storage medium based on a distributed preference information system. BACKGROUND

[0002] The advantage relation rough set method is a rough set model capable of directly processing preference attributes (indicators) proposed by Italian scholars Salvatore Greco et al. in order to solve multi-indicator decision-making problems and multi-indicator sorting problems. The method is characterized in that an advantage relation is used to replace the equivalence relation in the classical rough set model, and is widely studied in the fields of indicator selection, dynamic ordered data mining and hybrid data mining of multi-indicator sorting and multi-indicator decision-making problems.

[0003] The existing advantage relation rough set method can be applied to indicator selection, dynamic ordered data mining and hybrid data mining, but the known rough set method for big data is limited to analyzing and processing all data, such as a neighborhood rough set fast attribute reduction method based on Hadoop and a constraint relation rough set rule acquisition method based on MapReduce. Although the parallel operation framework improves the data calculation efficiency to a certain extent, the rules are mined from data, rather than synthesizing the required global result by using the existing local result. In addition, the prior art solves the inconsistency problem of the preference multi-attribute decision-making information system by adding certain constraint conditions, which cannot be directly used for processing the preference ordered distributed information system, and the classification rule given by weakening the advantage relation is only applicable to the processing of the preference ordered distributed information system under the weakened potential relation, and cannot effectively solve the inconsistency problem of the preference multi-attribute decision-making information system under the complete advantage relation. SUMMARY

[0004] The purpose of the present application is to provide an advantage relation rough set classification rule acquisition method, which detects and updates the existing classification rules of each preference sub-information system of the distributed preference information system, generates the classification rules of the entire distributed preference information system by merging, effectively improves the data processing efficiency, simply and effectively solves the inconsistency problem of the preference multi-attribute decision-making information system, and further ensures the rationality and effectiveness of the global classification rule acquisition.

[0005] In order to achieve the above-mentioned purpose, it is necessary to provide an advantage relation rough set classification rule acquisition method, system, computer device and storage medium in view of the above-mentioned technical problems.

[0006] In a first aspect, an embodiment of the present application provides an advantage relation rough set classification rule acquisition method, the method comprising the following steps:

[0007] acquiring a distributed data set; the distributed data set comprising a plurality of sub data sets to be analyzed; the sub data set to be analyzed comprising data to be analyzed and a corresponding classification label;

[0008] constructing a corresponding distributed preference information system according to the distributed data set and a preset preference attribute; the distributed preference information system comprising a plurality of preference sub information systems; the preference sub information system corresponding to the sub data set to be analyzed one by one;

[0009] determining a plurality of decision classes of the distributed preference information system by an advantage relation rough set method according to the classification label;

[0010] acquiring a decision class approximation set of each preference sub information system; the decision class approximation set comprising an approximation set of each decision class; the approximation set comprising an upper approximation set and a lower approximation set;

[0011] obtaining a decision class approximation set of the distributed preference information system according to the decision class approximation set of each preference sub information system;

[0012] obtaining a classification rule according to the decision class approximation set of the distributed preference information system.

[0013] Further, the step of acquiring the decision class approximation set of each preference sub information system comprises:

[0014] obtaining the decision class approximation set of each preference sub information system by the advantage relation rough set method.

[0015] Further, the step of obtaining the decision class approximation set of the distributed preference information system according to the decision class approximation set of each preference sub information system comprises:

[0016] judging whether there is inconsistency between the approximation sets of the same decision class in the decision class approximation sets of any two preference sub information systems, and updating the decision class approximation set with inconsistency if there is;

[0017] merging the decision class approximation sets of all preference sub information systems according to the decision class to obtain the decision class approximation set of the distributed preference information system.

[0018] Further, the step of updating the decision class approximation set with inconsistency comprises:

[0019] determining an inconsistent object set of the approximation set with inconsistency in the decision class approximation set;

[0020] According to the set of inconsistent objects, the approximation set of the corresponding decision class in the set of decision class approximations is updated.

[0021] Further, the set of inconsistent objects includes an upper approximation set of inconsistent objects and a lower approximation set of inconsistent objects.

[0022] The upper approximation set of inconsistent objects is represented as:

[0023]

[0024] wherein X ki and X kj respectively represent the corresponding kth decision class in the ith and jth preference sub-information system, 1≤i,j≤m, 1≤k≤n. and respectively represent the upper approximation set of the kth decision class of the ith and jth preference sub-information system; U j represents the domain of the jth preference sub-information system. represents the set of inconsistent objects in U j with .

[0025] The lower approximation set of inconsistent objects is represented as:

[0026]

[0027] wherein X ki and X kj respectively represent the corresponding kth decision class in the ith and jth preference sub-information system, 1≤i,j≤m, 1≤k≤n. P (X ki ) and P (X kj ) respectively represent the lower approximation set of the kth decision class of the ith and jth preference sub-information system. P (X ki ) j represents P (X kj ) with P (X ki ).

[0028] Further, the step of updating the approximation set of the corresponding decision class in the set of decision class approximations according to the set of inconsistent objects includes:

[0029] According to the union set of the upper approximation set of inconsistent objects and the upper approximation set of the corresponding decision class, the upper approximation set of the corresponding decision class is updated according to the following formula:

[0030]

[0031] wherein m represents the number of the preference sub-information systems; X ki represents the corresponding kth decision class in the ith preference sub-information system, 1≤i,j≤m, 1≤k≤n; U j represents the universe of the jth preference sub-information system; represents the universe of the jth preference sub-information system; j represents the set of inconsistent objects of U and represent the original upper approximation set and the updated upper approximation set of the kth decision class of the ith preference sub-information system, respectively;

[0032] According to the union set of the lower approximation inconsistent object set and the lower approximation set of the corresponding decision class approximation set, the lower approximation set of the corresponding decision class is updated according to the following formula:

[0033]

[0034] wherein m represents the number of the preference sub-information systems; X ki represents the corresponding kth decision class in the ith preference sub-information system, 1≤i,j≤m, 1≤k≤n; P (X ki ) j represents P (X kj ) and P (X ki ) represents the set of inconsistent objects of U P (X ki ) and P new (X ki ) represent the original lower approximation set and the updated lower approximation set of the kth decision class of the ith preference sub-information system, respectively.

[0035] Further, the step of merging the decision class approximation sets of all the preference sub-information systems according to the decision classes to obtain the decision class approximation set of the distributed preference information system comprises:

[0036] merging the upper approximation sets of the same decision class in the decision class approximation sets of all the preference sub-information systems to obtain the upper approximation set of the corresponding decision class in the decision class approximation set of the distributed preference information system;

[0037] merging the lower approximation sets of the same decision class in the decision class approximation sets of all the preference sub-information systems to obtain the lower approximation set of the corresponding decision class in the decision class approximation set of the distributed preference information system.

[0038] In the second aspect, an embodiment of the present application provides an advantage relation rough set classification rule acquisition system, which comprises: ​​

[0039] The data acquisition module is used to acquire a distributed dataset; the distributed dataset includes multiple sub-datasets to be analyzed; the sub-datasets to be analyzed include data to be analyzed and corresponding classification labels.

[0040] The system construction module is used to construct a corresponding distributed preference information system based on the distributed dataset and preset preference attributes; the distributed preference information system includes multiple preference sub-information systems; each preference sub-information system corresponds one-to-one with the sub-dataset to be analyzed.

[0041] The category determination module is used to determine multiple decision classes of the distributed preference information system based on the classification labels using the advantage relation rough set method.

[0042] The first approximation module is used to obtain the decision class approximation set of each preference sub-information system; the decision class approximation set includes the approximation set of each decision class; the approximation set includes the upper approximation set and the lower approximation set;

[0043] The second approximation module is used to obtain the decision class approximation set of the distributed preference information system based on the decision class approximation set of each preference sub-information system;

[0044] The rule generation module is used to obtain classification rules based on the decision class approximation set of the distributed preference information system.

[0045] Thirdly, embodiments of the present invention also provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.

[0046] Fourthly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described method.

[0047] This application provides a method, system, computer device, and storage medium for obtaining rough set classification rules for dominant relationships. The method constructs a corresponding distributed preference information system based on an acquired distributed dataset and preset preference attributes. Then, based on the classification labels of the distributed dataset, it determines multiple decision classes of the distributed preference information system using the rough set method of dominant relationships. It obtains approximate sets of decision classes for each preference sub-information system, performs coordination analysis on these approximate sets, updates the corresponding approximate sets based on the coordination analysis results, and merges the approximate sets of decision classes according to decision categories to obtain the approximate set of decision classes for the distributed preference information system, thereby obtaining the classification rules. Compared with existing technologies, this method for obtaining rough set classification rules for dominant relationships achieves the synthesis of required global knowledge based on existing local knowledge while ensuring the consistency of dominant relationships is not affected by data scale. This not only avoids unnecessary repetitive calculations in global knowledge acquisition, effectively improving data processing efficiency, but also simply and effectively solves the incompatibility problem of multi-attribute preference decision information systems, thus ensuring the rationality and effectiveness of the obtained global classification rules. Attached Figure Description

[0048] Figure 1 This is a schematic diagram illustrating an application scenario of the method for obtaining rough set classification rules based on dominant relationships in this invention.

[0049] Figure 2 This is a flowchart illustrating the process architecture for obtaining rough set classification rules for dominant relationships in an embodiment of the present invention.

[0050] Figure 3 This is a flowchart illustrating the method for obtaining rough set classification rules based on dominant relationships in an embodiment of the present invention.

[0051] Figure 4 This is a schematic diagram of the simulation results in an embodiment of the present invention;

[0052] Figure 5 This is a schematic diagram of the structure of the rough set classification rule acquisition system for dominant relations in an embodiment of the present invention;

[0053] Figure 6 This is an internal structural diagram of the computer device in an embodiment of the present invention. Detailed Implementation

[0054] To make the objectives, technical solutions, and beneficial effects of this application clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Obviously, the embodiments described below are only part of the embodiments of the present invention and are used to illustrate the present invention, but are not intended to limit the scope of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0055] The method for obtaining rough set classification rules of dominance relations provided by this invention can be applied to, for example... Figure 1 The terminal or server shown. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be a standalone server or a server cluster consisting of multiple servers. The server can actually use the collected data to be analyzed, and adopt the rough set classification rule acquisition method for dominant relations provided by this invention according to... Figure 2 The flowchart shown completes the task of obtaining classification rules for the distributed preference information system, and applies the final classification rules to other learning tasks on the server, or transmits them to the terminal for end users to receive and use.

[0056] In one embodiment, such as Figure 3 As shown, a method for obtaining rough set classification rules for dominance relations is provided, including the following steps:

[0057] S11. Obtain a distributed dataset; the distributed dataset includes multiple subsets to be analyzed; the subsets to be analyzed include data to be analyzed and corresponding classification labels; wherein, a distributed dataset can be understood as a dataset that, in practical applications, can be split into multiple subsets of data to be analyzed based on a certain attribute in the data to be analyzed for separate research. For example, the collected data of all business operations of a certain online cross-border retailer, including sample attributes such as invoice number, invoice date, product code, product description, transaction quantity, transaction unit price, customer ID, and country, can be used as a distributed dataset. It can be divided into several subsets of data to be analyzed according to the country attribute. In addition to the original attributes mentioned above, each subset of data to be analyzed also includes product classification labels obtained by evaluating all products based on some ordered attributes (product description, transaction quantity, and transaction unit price, etc.) as evaluation indicators, such as popular products, favorite products, and unfavorite products.

[0058] S12. Based on the distributed dataset and preset preference attributes, construct a corresponding distributed preference information system; the distributed preference information system includes multiple preference sub-information systems; each preference sub-information system corresponds one-to-one with the dataset to be analyzed; wherein, preference attributes refer to attributes used in actual applications to distinguish different data objects and compare the merits of different objects, such as invoice number, invoice date, product code, product description, transaction quantity, transaction unit price, customer ID, and country in commodity business data, all or some of these indicators can be considered as preference attributes used by the preference information system for data processing, which are determined according to the actual application scenario and application requirements, and are not specifically limited here.

[0059] A distributed preference information system can be understood as an information system integrated from several preference sub-information systems. It can be used to handle both global and local decisions. In this embodiment, the distributed preference information system built based on a distributed dataset can be represented as follows: Where U is the universe of discourse; A = C∪D, A is the attribute set, C is the condition attribute set, and D is the decision attribute set; V is the range of the attribute set; f = U × A → V is the information function that f(x,a) ∈ V for any a ∈ A and x ∈ U. Correspondingly, S i =(U i (A,V,f) represents the i-th preference sub-information system, where U i Let be the universe of discourse of the i-th preference sub-information system. Assume... For decision analysis, the set of preference attributes (indicators) that are of interest are the corresponding This represents an advantage relation based on a set of preference attributes P.

[0060] S13. Based on the classification labels, multiple decision classes of the distributed preference information system are determined using the dominance relation rough set method. As described above, each piece of data to be analyzed in the distributed dataset is labeled with a corresponding classification label. In practical applications, based on data classification requirements, multiple decision classes (concepts) are determined using the dominance relation rough set method based on the upward or downward association of each preset classification. The number of decision classes is related to the number of classification labels. For example, if the classification labels corresponding to product data are: popular products, favorite products, and disliked products, then the decision classes (concepts) can be determined based on the dominance relation rough set method: popular products (popular products are associated upwards), at least favorite products (favorite products are associated upwards), at most favorite products (favorite products are associated downwards), and disliked products (disliked products are associated downwards), etc. The decision classes (concepts) determined in this embodiment are also applicable to each preference sub-information system. Each decision class of the distributed preference information system can be understood as a combination of the same decision class in each of its preference sub-information systems. It should be noted that the above method for determining the corresponding decision classes and their number based on classification labels is only an exemplary description and does not specifically limit the scope of protection.

[0061] S14. Obtain the decision class approximation set of each preference sub-information system; the decision class approximation set includes the approximation set of each decision class; the approximation set includes the upper approximation set and the lower approximation set; wherein, the decision class approximation set of each preference sub-information system can be the known local knowledge information obtained based on the analysis of each sub-dataset to be analyzed, or it can be the upper approximation set and the lower approximation set of each decision class directly calculated based on each preference sub-information system by the traditional advantage relation rough set method.

[0062] S15. Based on the decision class approximation sets of each preference sub-information system, obtain the decision class approximation set of the distributed preference information system. The decision class approximation set of the distributed preference information system can be understood as global knowledge obtained from the local knowledge of each preference sub-information system. In principle, it can be obtained through analysis and processing of distributed datasets. However, given that the decision class approximation sets corresponding to each preference sub-information system are known, obtaining the global knowledge based on the original data would inevitably lead to unnecessary repetitive calculations and time consumption. To further improve the efficiency of obtaining the decision class approximation set of the distributed preference information system, this embodiment preferably merges the known decision class approximation sets corresponding to each preference sub-information system to generate the corresponding decision class approximation set of the distributed preference information system. Specifically, the step of obtaining the decision class approximation set of the distributed preference information system based on the decision class approximation sets of each preference sub-information system includes:

[0063] Determine whether there is inconsistency between approximate sets of the same decision class in the approximate sets of any two preference sub-information systems. If so, update the approximate set of the decision class with inconsistency. Inconsistency can be understood as a conflict in the object partitioning between the upper or lower approximate sets of the same decision class in different preference sub-information systems, such as in... That is, in the i-th preference sub-information system S i The x-th indicator is better than the j-th preference sub-information system S. j The y-index belongs to the lower approximation set of the k-th decision class (concept) of the j-th preference sub-information system, while x does not belong to the lower approximation set of the k-th decision class (concept) of the i-th preference sub-information system. This indicates a conflict between the lower approximation sets of the k-th decision class (concept) of the two preference sub-information systems. Similarly, there may be inconsistencies between the upper approximation sets of the same decision class in different preference sub-information systems. However, directly ignoring these inconsistencies—that is, merging the upper and lower approximation sets of the same decision class in different preference sub-information systems without processing to obtain the corresponding upper and lower approximation sets of the distributed preference information system—is clearly unreasonable and will directly reduce the accuracy of the decision class approximation sets of the distributed preference sub-information system. To reduce computational load and improve data processing efficiency while effectively ensuring the rationality and effectiveness of the decision class approximation sets of the distributed preference sub-information system, this embodiment preferably performs a consistency check on the obtained decision class approximation sets of each preference sub-information system and updates the decision class approximation sets with inconsistencies. Specifically, the step of updating the inconsistent decision class approximation set includes:

[0064] Determine the set of incompatible objects of the approximate set containing inconsistencies within the decision class approximation set; wherein the set of incompatible objects includes the upper approximate incompatible object set and the lower approximate incompatible object set;

[0065] The above approximate set of incompatible objects is represented as:

[0066]

[0067] Among them, X ki and X kj Let i and j represent the k-th decision class corresponding to the i-th and j-th preference sub-information systems, respectively, where 1≤i,j≤m, and 1≤k≤n; and U represents the upper approximation set of the k-th decision class of the i-th and j-th preference sub-information systems, respectively; j Denotes the domain of discourse of the j-th preference subsystem; U j Zhongyu A set of incompatible objects;

[0068] The approximate set of incompatible objects is represented as follows:

[0069]

[0070] Among them, X ki and X kj Let i and j represent the k-th decision class corresponding to the i-th and j-th preference sub-information systems, respectively, where 1≤i,j≤m, and 1≤k≤n; P (X ki )and P (X kj ) represent the lower approximation sets of the k-th decision class of the i-th and j-th preference sub-information systems, respectively; P (X ki ) j express P (X kj ) in and P (X ki The set of incompatible objects;

[0071] Similarly, U i Zhongyu The set of incompatible objects can be represented as:

[0072] P (X kj ) i express P (X i ) in and P (X j The set of incompatible objects can be represented as:

[0073]

[0074] After obtaining the set of incompatible objects of the approximate set with inconsistency in each decision class approximate set according to the above method steps, the corresponding upper approximate set and lower approximate set can be recalculated according to the following method.

[0075] Based on the set of incompatible objects, update the approximate set of the corresponding decision class within the approximate set of decision classes. Specific steps include:

[0076] Based on the union of the set of incompatible objects and the set of approximations of the corresponding decision class, update the set of approximations of the corresponding decision class according to the following formula:

[0077]

[0078] Where m represents the number of preference sub-information systems; X kiU represents the k-th decision class corresponding to the i-th preference sub-information system, where 1 ≤ i, j ≤ m, 1 ≤ k ≤ n; j Denotes the domain of discourse of the j-th preference subsystem; U j Zhongyu A set of incompatible objects; and Let represent the original upper approximation set and the updated upper approximation set of the k-th decision class of the i-th preference sub-information system, respectively;

[0079] Based on the union of the lower approximation set of the incompatible objects and the lower approximation set of the corresponding decision class, update the lower approximation set of the corresponding decision class according to the following formula:

[0080]

[0081] Where m represents the number of preference sub-information systems; X ki Let represent the k-th decision class of the i-th preference sub-information system, where 1≤i,j≤m,1≤k≤n; P (X ki ) j express P (X kj ) in and P (X ki The set of incompatible objects; P (X ki )and P new (X ki Let represent the original lower approximation set and the updated lower approximation set of the k-th decision class of the i-th preference sub-information system, respectively.

[0082] The above steps detail the detection of inconsistencies among approximate sets of the same decision class (concept) in each preference sub-information system, and how to adjust approximate sets with inconsistencies to ensure the validity and rationality of the approximate sets of decision classes in the subsequently merged distributed preference information system. It should be noted that if there are no inconsistencies among the approximate sets of decision classes in each preference sub-information system, the steps of updating the approximate sets of decision classes with inconsistencies can be skipped, and the merging process of approximate sets of the same decision class (concept) in each preference sub-information system can be directly performed in the following steps.

[0083] The decision class approximation sets of all preference sub-information systems are merged according to decision class to obtain the decision class approximation set of the distributed preference information system. The step of merging the decision class approximation sets of all preference sub-information systems according to decision class to obtain the decision class approximation set of the distributed preference information system includes:

[0084] The upper approximation sets of the same decision class in the decision class approximation sets of all preference sub-information systems are merged to obtain the upper approximation sets of the corresponding decision classes within the decision class approximation sets of the distributed preference information system. The specific merging expression is as follows:

[0085] 1) There is no inconsistency among the approximate sets of the original decision classes of each preference sub-information system in the expression of the upper approximate sets:

[0086]

[0087] 2) The expression for the updated upper approximation set of each preference sub-information system due to inconsistencies:

[0088]

[0089] Among them, X ki and X k Let i and k represent the i-th preference sub-information system and the k-th decision class of the distributed preference information system, respectively, where 1≤i≤m and 1≤k≤n; Let represent the upper approximation set of the k-th decision class in a distributed preference information system; and Let represent the original upper approximation set and the updated upper approximation set of the k-th decision class of the i-th preference sub-information system, respectively; the lower approximation sets of the same decision class in the decision class approximation sets of all preference sub-information systems are merged to obtain the lower approximation set of the corresponding decision class within the decision class approximation set of the distributed preference information system. The specific merging expression is as follows:

[0090] 1) The lower approximate set union expression for which there is no inconsistency among the original decision class approximate sets of each preference sub-information system:

[0091]

[0092] 2) The expression for the updated upper approximation set of each preference sub-information system due to inconsistencies:

[0093]

[0094] Among them, X ki and X k Let i and k represent the i-th preference sub-information system and the k-th decision class of the distributed preference information system, respectively, where 1≤i≤m and 1≤k≤n; P (X k ) represents the upper approximation set of the k-th decision class in a distributed preference information system; P (X ki )and P new (X kiLet represent the original lower approximation set and the updated lower approximation set of the k-th decision class of the i-th preference sub-information system, respectively.

[0095] S16. Based on the decision class approximation set of the distributed preference information system, classification rules are obtained. The division of the upper and lower approximation sets of each decision class within the decision class approximation set of the distributed preference information system can be understood as the classification rules for the distributed dataset based on the distributed preference information system. Based on the obtained classification rules, corresponding decision analysis can be performed according to actual application requirements.

[0096] This application embodiment evaluates the coordination of decision class approximation sets of various preference sub-information systems, updates them based on the incoordination object sets of each incoordination decision class approximation set, and then merges them to obtain the decision class approximation set of the corresponding distributed preference information system. This method not only simply and effectively realizes the generation of global knowledge based on the merging of existing local knowledge, avoiding a large amount of unnecessary repetitive calculations and improving data processing efficiency, but also ensures the integrity of the advantage relationships between data in the process of merging to generate global knowledge. It effectively solves the incompatibility problem of multi-attribute preference decision information systems, improves the accuracy of global knowledge acquisition, and thus provides an effective guarantee for the rationality of global decision-making.

[0097] To verify the technical effectiveness of the rough set classification rule acquisition method of the dominant relation of the present invention, this embodiment uses online cross-border retail business data (the size of the distributed dataset is about 46M) for simulation verification. The data sample attributes of the corresponding distributed dataset include invoice number, invoice date, product code, product description, transaction quantity, transaction unit price, customer ID and country, etc. The distributed dataset is divided into multiple sub-datasets to be analyzed according to the country attribute. For each sub-dataset to be analyzed, three ordered attributes, namely product description, transaction quantity and transaction unit price, are selected as evaluation indicators to evaluate the traded products, and the corresponding classification labels (local evaluation) are obtained: popular products, favorite products and unfavorite products. Based on the classification labels of the distributed dataset, four decision classes (concepts) were established using the rough set method of dominance relations: Popular Products (popular products are linked upwards), at least belong to Favorite Products (favorite products are linked upwards), at most belong to Favorite Products (favorite products are linked downwards), and Non-Favorite Products (non-favorite products are linked downwards). For each of these four decision classes (concepts), upper and lower approximation sets were calculated based on each subset of the dataset to be analyzed. These results were then used as existing results. The method of this invention was then applied, employing a multi-core parallel processing method combined with a Huffman tree generation mechanism in a multi-layer synthesis simulation experiment to synthesize the upper and lower approximation sets of these four decision classes (concepts) based on the distributed dataset (the entire business data). The results are as follows: Figure 4The simulation results show that as the number of cores used increases, the runtime decreases; however, as the number of cores increases, the decreasing trend of runtime slows down (because the approximate set synthesis process is divided into multiple levels. Initially, the available worker threads can be used to complete local synthesis. As it approaches global synthesis, the amount of data to be synthesized increases, the number of idle worker threads increases, and the decreasing trend of runtime slows down). This effectively verifies that the method of this invention can obtain global decisions based on the merging of local decisions. While providing assistance for both global and local decisions, it not only ensures the integrity of the advantage relationships between data and the coordination of decisions, but also effectively utilizes existing results to reduce unnecessary computation, reducing computation time and improving data processing efficiency when parallel computing rough set approximations of advantage relationships.

[0098] It should be noted that although the steps in the flowchart above are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order requirement for the execution of these steps, and they can be executed in other orders.

[0099] In one embodiment, such as Figure 5 As shown, a system for obtaining rough set classification rules for dominance relations is provided, the system comprising:

[0100] Data acquisition module 1 is used to acquire a distributed dataset; the distributed dataset includes multiple sub-datasets to be analyzed; the sub-datasets to be analyzed include data to be analyzed and corresponding classification labels;

[0101] System construction module 2 is used to construct a corresponding distributed preference information system based on the distributed dataset and preset preference attributes; the distributed preference information system includes multiple preference sub-information systems; each preference sub-information system corresponds one-to-one with the sub-dataset to be analyzed;

[0102] Category determination module 3 is used to determine multiple decision classes of the distributed preference information system based on the classification labels using the advantage relation rough set method;

[0103] The first approximation module 4 is used to obtain the decision class approximation set of each preference sub-information system; the decision class approximation set includes the approximation set of each decision class; the approximation set includes the upper approximation set and the lower approximation set;

[0104] The second approximation module 5 is used to obtain the decision class approximation set of the distributed preference information system based on the decision class approximation set of each preference sub-information system;

[0105] Rule generation module 6 is used to obtain classification rules based on the decision class approximation set of the distributed preference information system.

[0106] Specific limitations regarding the system for obtaining rough set classification rules for dominance relations can be found in the limitations of the method for obtaining rough set classification rules for dominance relations described above, and will not be repeated here. Each module in the aforementioned system for obtaining rough set classification rules for dominance relations can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0107] Figure 6 An internal structural diagram of a computer device is shown in one embodiment. This computer device may specifically be a terminal or a server. Figure 6 As shown, the computer device includes a processor, memory, network interface, display, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When executed by the processor, the computer program implements a method for obtaining rough set classification rules based on dominance relations. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.

[0108] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computing devices may include more or fewer components than those shown in the figure, or combine certain components, or have the same component arrangement.

[0109] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described above.

[0110] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0111] In summary, the present invention provides a method, system, computer device, and storage medium for obtaining rough set classification rules based on dominant relationships. The method constructs a corresponding distributed preference information system based on an acquired distributed dataset and preset preference attributes. Then, based on the classification labels of the distributed dataset, it determines multiple decision classes of the distributed preference information system using the rough set method of dominant relationships. Next, it obtains approximate sets of decision classes for each preference sub-information system, performs coordination analysis on these approximate sets, updates the corresponding approximate sets based on the coordination analysis results, and merges the approximate sets of decision classes from each sub-information system according to decision categories to obtain the approximate set of decision classes for the distributed preference information system, thereby obtaining classification rules. This technical solution achieves the goal of synthesizing the required global knowledge based on existing local knowledge while ensuring the consistency of dominant relationships is not affected by data scale. It not only avoids unnecessary repetitive calculations in obtaining global knowledge and effectively improves data processing efficiency, but also simply and effectively solves the incompatibility problem of multi-attribute preference decision information systems, thus ensuring the rationality and effectiveness of obtaining global classification rules.

[0112] The various embodiments in this specification are described in a progressive manner. For directly identical or similar parts of the embodiments, refer to each other. Each embodiment focuses on its differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.

[0113] The embodiments described above are merely preferred embodiments of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various improvements and substitutions without departing from the technical principles of this invention, and these improvements and substitutions should also be considered within the scope of protection of this application. Therefore, the scope of protection of this patent application should be determined by the scope of the claims.

Claims

1. A method for obtaining rough set classification rules of dominance relations, characterized in that, The method includes the following steps: Obtain a distributed dataset; the distributed dataset includes multiple sub-datasets to be analyzed; each sub-dataset to be analyzed includes data to be analyzed and corresponding category labels; the data to be analyzed includes invoice number, invoice date, product code, product description, transaction quantity, transaction unit price, customer ID, and country attribute; the category labels include popular products, favorite products, and unfavorite products; Based on the distributed dataset and preset preference attributes, a corresponding distributed preference information system is constructed; the distributed preference information system includes multiple preference sub-information systems; each preference sub-information system corresponds one-to-one with the dataset to be analyzed; the preset preference attributes are used to distinguish different data objects and compare the merits of different objects; the distributed preference information system is used to process global and local decisions; Based on the classification labels, multiple decision classes of the distributed preference information system are determined using the dominance relation rough set method; the decision classes include enthusiastic products, products that are at least liked products, products that are at most liked products, and products that are neither liked nor liked products. Obtain the decision class approximation set for each preference sub-information system; the decision class approximation set includes the approximation set for each decision class; the approximation set includes the upper approximation set and the lower approximation set; Based on the decision class approximation sets of each preference sub-information system, the decision class approximation set of the distributed preference information system is obtained; Based on the decision class approximation set of the distributed preference information system, classification rules are obtained; The step of obtaining the decision class approximation set of the distributed preference information system based on the decision class approximation sets of each preference sub-information system includes: determining whether there is inconsistency between the approximation sets of the same decision class in the decision class approximation sets of any two preference sub-information systems; if so, updating the decision class approximation set with inconsistency, including: determining the inconsistency object set of the approximation sets with inconsistency within the decision class approximation set; and updating the approximation set of the corresponding decision class within the decision class approximation set based on the inconsistency object set. The decision class approximation sets of all preference sub-information systems are merged according to decision class to obtain the decision class approximation set of the distributed preference information system.

2. The method for obtaining rough set classification rules of dominance relations as described in claim 1, characterized in that, The steps for obtaining the approximate set of decision classes for each preference sub-information system include: The rough set method of the advantage relation is used to calculate the approximate set of decision classes for each preference sub-information system.

3. The method for obtaining rough set classification rules based on dominance relations as described in claim 1, characterized in that, The set of incompatible objects includes an upper approximate set of incompatible objects and a lower approximate set of incompatible objects; The above approximate set of incompatible objects is represented as: Among them, X ki and X kj Let i and j represent the k-th decision class corresponding to the i-th and j-th preference sub-information systems, respectively, 1≤i,j≤m, 1≤k≤n, and m represents the number of preference sub-information systems; and U represents the upper approximation set of the k-th decision class of the i-th and j-th preference sub-information systems, respectively; j Denotes the domain of discourse of the j-th preference subsystem; U j Zhongyu A set of incompatible objects; The approximate set of incompatible objects is represented as follows: Among them, X ki and X kj P(X) represents the k-th decision class corresponding to the i-th and j-th preference sub-information systems, respectively, 1≤i,j≤m, 1≤k≤n, where m represents the number of preference sub-information systems; ki ) and P(X kj ) represent the lower approximation sets of the k-th decision class of the i-th and j-th preference sub-information systems, respectively; P (X ki ) j express P (X kj ) in and P (X ki The set of incompatible objects.

4. The method for obtaining rough set classification rules based on dominance relations as described in claim 3, characterized in that, The step of updating the approximation set of the corresponding decision class within the approximation set of the decision class based on the set of incompatible objects includes: Based on the union of the set of incompatible objects and the set of approximations of the corresponding decision class, update the set of approximations of the corresponding decision class according to the following formula: Where m represents the number of preference sub-information systems; X ki U represents the k-th decision class corresponding to the i-th preference sub-information system, where 1 ≤ i, j ≤ m, 1 ≤ k ≤ n; j Denotes the domain of discourse of the j-th preference subsystem; U j Zhongyu A set of incompatible objects; and Let represent the original upper approximation set and the updated upper approximation set of the k-th decision class of the i-th preference sub-information system, respectively; Based on the union of the lower approximation set of the incompatible objects and the lower approximation set of the corresponding decision class, update the lower approximation set of the corresponding decision class according to the following formula: Where m represents the number of preference sub-information systems; X ki P(X) represents the k-th decision class in the i-th preference sub-information system, where 1 ≤ i, j ≤ m, 1 ≤ k ≤ n; ki ) j P(X) represents kj ) in P(X ki The set of incompatible objects of P(X); ki ) and P new (X ki Let represent the original lower approximation set and the updated lower approximation set of the k-th decision class of the i-th preference sub-information system, respectively.

5. The method for obtaining rough set classification rules of dominance relations as described in claim 1, characterized in that, The step of merging the decision class approximation sets of all preference sub-information systems according to decision class to obtain the decision class approximation set of the distributed preference information system includes: The upper approximation sets of the same decision class in the decision class approximation sets of all preference sub-information systems are merged to obtain the upper approximation set of the corresponding decision class within the decision class approximation set of the distributed preference information system. By merging the lower approximation sets of the same decision class in the decision class approximation sets of all preference sub-information systems, we obtain the lower approximation set of the corresponding decision class within the decision class approximation set of the distributed preference information system.

6. A system for obtaining rough set classification rules based on dominance relations, characterized in that, The system, employing the method for obtaining rough set classification rules of dominance relations as described in claim 1, comprises: The data acquisition module is used to acquire a distributed dataset; the distributed dataset includes multiple sub-datasets to be analyzed; the sub-datasets to be analyzed include data to be analyzed and corresponding category labels; the data to be analyzed includes invoice number, invoice date, product code, product description, transaction quantity, transaction unit price, customer ID, and country attribute; the category labels include popular products, favorite products, and unfavorite products; The system construction module is used to construct a corresponding distributed preference information system based on the distributed dataset and preset preference attributes; the distributed preference information system includes multiple preference sub-information systems; each preference sub-information system corresponds one-to-one with the sub-dataset to be analyzed; the preset preference attributes are used to distinguish different data objects and compare the merits of different objects; the distributed preference information system is used to process global and local decisions; The category determination module is used to determine multiple decision classes of the distributed preference information system based on the classification labels using the dominance relation rough set method; the decision classes include enthusiastic products, at least favorite products, at most favorite products, and non-favorite products. The first approximation module is used to obtain the decision class approximation set of each preference sub-information system; the decision class approximation set includes the approximation set of each decision class; the approximation set includes the upper approximation set and the lower approximation set; The second approximation module is used to obtain the decision class approximation set of the distributed preference information system based on the decision class approximation set of each preference sub-information system; The rule generation module is used to obtain classification rules based on the decision class approximation set of the distributed preference information system.

7. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Data-driven variable precision dominance rough set threshold obtaining method under conflict relationship

    CN103955598A

  • Interactive method to reduce the amount of tradeoff information required from decision makers in multi-attribute decision making under uncertainty

    US20140279801A1