A method and system for processing engineering archives based on association rule mining

By using the weighted occurrence frequency threshold and support threshold filter option set in engineering archive data processing, strong association rules are generated, which solves the problems of low efficiency of sparse data processing and poor rule quality in the prior art, and improves the application efficiency of association rules.

CN119337331BActive Publication Date: 2025-05-09CHINA WATER RESOURCES BEIFANG INVESTIGATION DESIGN & RES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411846221.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-09
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

The existing engineering archive processing methods based on association rule mining are inefficient in the face of highly sparse data, and the quality of association rule is reduced, resulting in low practical application efficiency.

Method used

By obtaining the engineering archive data set, extracting features and combining them into item sets, analyzing and generating candidate sets, filtering out frequent item sets using weighted occurrence frequency thresholds and support thresholds, and generating strong association rules to improve application efficiency.

Benefits of technology

The practical application efficiency of association rules is improved, and the generated strong association rules are more universal, and can more scientifically remove invalid features and accurately refine frequent item sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119337331B_ABST
    Figure CN119337331B_ABST
Patent Text Reader

Abstract

The present invention discloses an engineering archive processing method and system based on association rule mining, and relates to the technical field of archive processing. The present invention aims to improve the actual application efficiency of existing association rules, that is, based on the existing association rules, a secondary analysis is performed on a formulated item set in a data-driven manner, wherein in the process of the secondary analysis, firstly, all features contained in the item set are screened for quality in a weighted frequency threshold and the number of occurrences, and then a secondary analysis screening is performed on the candidate item set in a support threshold, so as to screen out a higher-quality item set with a more frequent occurrence, and then based on the frequent item set, a strong association rule with more universal significance than the existing association rule can be generated, so as to effectively improve the actual application efficiency of the association rule.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of archive processing, and in particular to an engineering archive processing method and system based on association rule mining. Background Art

[0002] As the scale and complexity of engineering projects grow, the amount of related engineering archive data is also expanding rapidly. This includes not only structured information such as budgets and schedules, but also a large number of unstructured or semi-structured documents, such as design drawings, construction logs, etc. Therefore, it is necessary to process them based on existing association rule mining algorithms in order to improve the processing efficiency of engineering archive data; however, sparsity is a common feature in engineering archive data, which means that many projects only involve a small number of attributes or items in the total feature set. For example, in a large engineering project, only a few specific types of materials are used, while most other materials will not appear in the data records of the project. This sparsity problem poses a significant challenge to traditional association rule mining algorithms. Therefore, based on the sparsity problem, the existing engineering archive processing methods based on association rule mining still have the following defects:

[0003] First, it is easy to lead to low computational efficiency. Specifically, traditional association rule mining algorithms usually assume that all items share most of the features. Therefore, when faced with highly sparse data, these algorithms will waste a lot of computing resources to process those item set combinations that actually rarely appear. Second, it is easy to lead to a decrease in the quality of association rules. Specifically, due to sparsity, the mined association rules may not have universal significance, thereby reducing the actual application efficiency of the association rules.

[0004] Therefore, the prior art urgently needs a technical solution for an engineering archive processing method and system based on association rule mining. Summary of the invention

[0005] In order to solve the above technical problems, the present invention provides a method for processing engineering archives based on association rule mining, which specifically includes the following steps:

[0006] Step S1, obtaining an engineering archive data set containing at least two projects and performing preprocessing;

[0007] Step S2, extracting features from the preprocessed engineering archive data set, and combining the features extracted from each project into an item set to form at least two item sets;

[0008] Step S3, analyzing all item sets, and generating at least two candidate item sets based on the analysis results;

[0009] Step S3a, analyzing the occurrence frequencies of all features in each item set, and obtaining a weighted occurrence frequency threshold of the features in each item set according to the analysis results;

[0010] Step S3a1, counting the occurrence frequencies of all features in each item set to obtain the occurrence frequency of each feature in each item set;

[0011] Step S3a2: The occurrence frequency of each feature in each item set is synthesized and averaged to obtain the average occurrence frequency of each feature in each item set, and the standard deviation of the occurrence frequency of each feature in each item set is obtained based on the average occurrence frequency;

[0012] Step S3a3, assigning a weight to each feature in each item set according to the frequency of occurrence of each feature in each item set;

[0013] Step S3a4, obtaining a weighted occurrence frequency threshold of the feature in each item set according to the average occurrence frequency of each feature in each item set, the occurrence frequency standard deviation of each feature in each item set, and the weight of each feature in each item set;

[0014] The calculation formula for obtaining the weighted frequency threshold of the feature in each item set is:

[0015] ;

[0016] in, represents the weighted frequency threshold of the feature in the i-th item set; represents the average frequency of occurrence of the jth feature in the i-th item set; Represents the standard deviation of the occurrence frequency of the jth feature in the i-th item set; Represents the weight of the jth feature in the i-th item set; represents adjustment parameters; Represents the total number of features in the i-th item set;

[0017] Step S3b, generating at least two candidate item sets according to the weighted occurrence frequency threshold of the features in each item set;

[0018] Step S3b1, normalizing the weighted occurrence frequency threshold of the feature in each item set and the occurrence frequency of each feature in each item set;

[0019] Step S3b2: using the weighted occurrence frequency threshold of the features in each item set after normalization to determine the occurrence frequency of each feature in each item set, if the occurrence frequency of the feature is greater than or equal to the weighted occurrence frequency threshold, then retain the feature; if the occurrence frequency of the feature is less than the weighted occurrence frequency threshold, then remove the feature;

[0020] Step S3b3, based on the features retained in each item set, perform aggregation and generate a candidate item set according to the aggregation result;

[0021] Step S4: Perform support analysis on each candidate item set, and select candidate item sets whose support meets the support threshold as frequent item sets based on the analysis results;

[0022] Step S41, counting the total number of items in the engineering archive data set, and further counting the number of occurrences of each candidate item set in the total number of items in the engineering archive data set;

[0023] Step S42, obtaining the support of each candidate item set according to the total number of items in the engineering archive data set and the number of occurrences of each candidate item set in the total items in the engineering archive data set;

[0024] The calculation formula for obtaining the support of each candidate item set is:

[0025] ;

[0026] in, Represents the support of the kth candidate item set; represents the number of occurrences of the kth candidate item set in the total items of the engineering archive dataset; Represents the total number of projects in the engineering archive dataset; Represents the kth candidate item set.

[0027] Step S43: Count the support of each candidate item set, take the average, set the average as the support threshold, and use the support threshold to determine the support of each candidate item set;

[0028] Step S44: if the support of the candidate item set is greater than or equal to the support threshold, then the candidate item set is retained and used as a frequent item set; if the support of the candidate item set is less than the support threshold, then the candidate item set is removed;

[0029] Step S5: Generate at least two strong association rules based on each frequent item set, and use the strong association rules to process the engineering archive data set.

[0030] A system for processing engineering archives based on association rule mining, which executes the above-mentioned method for processing engineering archives based on association rule mining, includes the following modules:

[0031] Data acquisition and preprocessing module: used to obtain and preprocess engineering archive data sets containing at least two projects;

[0032] Itemset formulation module: connected to the data acquisition and preprocessing module, used to extract features from the preprocessed engineering archive data set, and combine the features extracted from each project into an itemset to form at least two itemsets;

[0033] Preliminary analysis module: connected to the item set formulation module, used to analyze all item sets and generate at least two candidate item sets based on the analysis results;

[0034] Secondary analysis module: connected to the preliminary analysis module, used to perform support analysis on each candidate item set, and select candidate item sets whose support meets the support threshold as frequent item sets according to the analysis results;

[0035] Strong association rule application module: connected to the secondary analysis module, used to generate at least two strong association rules based on each frequent item set, and use the strong association rules to process the engineering archive data set.

[0036] The embodiments of the present invention have the following technical effects:

[0037] The present invention aims to improve the practical application efficiency of existing association rules, that is, on the basis of existing association rules, a secondary analysis is performed on the formulated item sets in a data-driven manner, so as to screen out better quality and more frequently occurring item sets, and then based on the frequent item sets, a strong association rule with more universal significance than the existing association rules can be generated, so as to effectively improve the practical application efficiency of the association rules; wherein in the secondary analysis process, firstly, a high-quality screening is performed on all features contained in the item set, especially the number of occurrences of each feature is judged by a weighted frequency threshold, and those that meet the judgment criteria are retained, and a candidate item set is formed after being aggregated, and then a secondary analysis screening is performed on the candidate item set to extract a candidate item set with more frequent occurrences, that is, a candidate item set that meets the support threshold, and as the final frequent item set, this secondary screening process can not only highly scientifically remove the so-called invalid features, but also can further accurately extract the frequent item sets, so that the generation of the frequent item sets can meet the high scientificity, accuracy and higher quality, so that the strong association rules generated by them can have more universal significance, so as to effectively improve the practical application efficiency of the association rules. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0039] Figure 1 It is a flow chart of a method for processing engineering archives based on association rule mining provided by an embodiment of the present invention;

[0040] Figure 2 It is a framework diagram of an engineering archive processing system based on association rule mining provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0041] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the scope of protection of the present invention.

[0042] Embodiment 1: Figure 1 As shown, the present invention provides a method for processing engineering archives based on association rule mining, comprising the following steps:

[0043] Step S1, obtaining an engineering archive data set containing at least two projects and performing preprocessing.

[0044] In the engineering archive data set, the items obtained may include but are not limited to the following: 1. Basic information of the engineering project: project name, project number, project location, project type such as building, bridge, road, etc., project scale, project budget, project start date and end date, etc.; 2. Design documents: design drawings, design instructions, technical specifications, etc.; 3. Construction records: construction log, construction progress report, quality inspection report, safety inspection report, etc.; 4. Procurement information: material list, supplier information, procurement contract, procurement price, etc.; 5. Change records: design change notice, engineering change order, change approval document, etc.; 6. Acceptance report: completion acceptance report, quality acceptance report, safety acceptance report, etc.; 7. Maintenance records: repair records, maintenance records, equipment replacement records, etc.; Environmental impact assessment: environmental impact report, implementation of environmental protection measures, etc.

[0045] Preprocessing is a very important step in the data mining process, which ensures the quality of subsequent analysis. The following are the specific preprocessing steps for the above engineering archive dataset:

[0046] 1. Data cleaning: 1.1. Remove duplicate data: Check and delete duplicate records in the data set to avoid bias in the analysis results; 1.2. Missing value processing: For data with missing values, you can choose to fill in, such as using the mean, median or mode, or directly delete records with missing values; 1.3. Outlier detection and processing: Identify and process outliers. Statistical methods such as box plots and Z-scores can be used to detect outliers and decide whether to delete or correct them.

[0047] 2. Data conversion: 2.1. Format unification: unify data from different sources into the same format, such as unifying the date format into YYYY-MM-DD and the currency unit into the same currency; 2.2. Numerical processing: For non-numeric data (such as text data), perform encoding conversion, such as using One-Hot Encoding or Label Encoding to convert it into numerical form.

[0048] 3. Data standardization / normalization: 3.1. Standardization: Convert the data to a standard normal distribution, that is, the mean is 0 and the standard deviation is 1, which helps subsequent algorithms work better; 3.2. Normalization: Scale the data to a specific range, usually [0,1] or [-1,1], which helps prevent certain features from dominating subsequent analysis results because of their larger numerical range.

[0049] For example, suppose there is an archive dataset containing multiple engineering projects, including basic project information, design documents, construction records, etc. The following is an example of a specific preprocessing process:

[0050] 1. Data cleaning: Check and delete duplicate project records; for missing key fields such as project name and project number, consider deleting the record; for missing secondary fields such as a specific design parameter, use the mean or median to fill in; use the Z-score method to detect and handle outliers, such as when the project budget deviates significantly from the normal range;

[0051] 2. Data conversion: unify all date fields into YYYY-MM-DD format; one-hot encode project types such as "building" and "bridge";

[0052] 3. Data standardization / normalization: Standardize numerical fields such as project budget and project scale to make them conform to the standard normal distribution.

[0053] Step S2: extracting features from the preprocessed engineering archive data set, and combining the features extracted from each project into an item set to form at least two item sets.

[0054] It is necessary to extract features from the preprocessed engineering archive dataset and combine these features into item sets. The following is a detailed implementation example showing how to perform feature extraction and item set construction:

[0055] 1. Perform data preparation, that is, prepare the data that has been preprocessed in step S1 to obtain a clean data set with a unified format, which contains information of multiple engineering projects;

[0056] 2. Perform feature extraction;

[0057] 2.1. Select features. According to business needs and analysis purposes, select the following features for extraction: Project type: indicates the type of project (such as building, bridge, road); Project scale: indicates the size or complexity of the project; Project budget: indicates the financial budget of the project; Project location: indicates the geographical location of the project; Design documents: indicates the number or type of design-related documents; Construction records: indicates the number or type of records during the construction process; Procurement information: indicates the number or type of procurement-related records; Change records: indicates the number or type of project changes; Acceptance report: indicates the number or type of acceptance-related records; Maintenance records: indicates the number or type of maintenance-related records; Environmental impact assessment: indicates the number or type of environmental impact assessment-related records; 2.2. Extract features. For each project, extract the above features from the engineering archive data set. For example, for a certain project, the extracted features are as follows: Project type: building; Project scale: large; Project budget: 50 million; Project location: Beijing; Design documents: yes; Construction records: yes; Procurement information: yes; Change records: no; Acceptance report: yes; Maintenance records: no; Environmental impact assessment: yes;

[0058] 3. Construct item sets;

[0059] That is, the features extracted from each item are combined into an itemset, and each itemset represents all the features of a project;

[0060] 3.1. Item set construction. Suppose there are three projects, named Project A, Project B and Project C. The features of each project are combined into an item set as shown below: Item set of Project A: {building, large, 50 million, Beijing, with design documents, with construction records, with procurement information, without change records, with acceptance report, without maintenance records, with environmental impact assessment}; Item set of Project B: {bridge, medium, 30 million, Shanghai, with design documents, with construction records, with procurement information, with change records, with acceptance report, with maintenance records, with environmental impact assessment}; Item set of Project C: {road, small, 10 million, Guangzhou, with design documents, with construction records, with procurement information, without change records, with acceptance report, without maintenance records, without environmental impact assessment}.

[0061] 4. Form at least two item sets;

[0062] Through the above steps, an item set has been constructed for each project. Now, we need to ensure that at least two item sets are formed: Item Set 1: {Building, Large, 50 million, Beijing, with design documents, construction records, procurement information, no change records, acceptance report, no maintenance records, environmental impact assessment}; Item Set 2: {Bridge, Medium, 30 million, Shanghai, with design documents, construction records, procurement information, change records, acceptance report, maintenance records, environmental impact assessment}; Item Set 3: {Road, Small, 10 million, Guangzhou, with design documents, construction records, procurement information, no change records, acceptance report, no maintenance records, no environmental impact assessment};

[0063] 5. Output results;

[0064] Finally, three item sets are obtained, each of which represents a feature set of an item. These item sets will be used for subsequent association rule mining;

[0065] Through the above steps, features are extracted from the preprocessed engineering archive data set, and the features extracted from each project are combined into an item set, forming at least two item sets, thus preparing for the next step of association rule mining.

[0066] Step S3, analyzing all item sets, and generating at least two candidate item sets based on the analysis results;

[0067] Step S3a, analyzing the occurrence frequencies of all features in each item set, and obtaining a weighted occurrence frequency threshold of the features in each item set according to the analysis results;

[0068] Step S3a1, counting the occurrence frequencies of all features in each item set to obtain the occurrence frequency of each feature in each item set;

[0069] Step S3a2: The occurrence frequency of each feature in each item set is synthesized and averaged to obtain the average occurrence frequency of each feature in each item set, and the standard deviation of the occurrence frequency of each feature in each item set is obtained based on the average occurrence frequency;

[0070] Step S3a3, assigning a weight to each feature in each item set according to the frequency of occurrence of each feature in each item set;

[0071] Step S3a4, obtaining a weighted occurrence frequency threshold of the feature in each item set according to the average occurrence frequency of each feature in each item set, the occurrence frequency standard deviation of each feature in each item set, and the weight of each feature in each item set;

[0072] The calculation formula for obtaining the weighted frequency threshold of the feature in each item set is:

[0073] ;

[0074] in, represents the weighted frequency threshold of the feature in the i-th item set; represents the average frequency of occurrence of the jth feature in the i-th item set; Represents the standard deviation of the occurrence frequency of the jth feature in the i-th item set; Represents the weight of the jth feature in the i-th item set; represents adjustment parameters; Represents the total number of features in the i-th item set;

[0075] Step S3b, generating at least two candidate item sets according to the weighted occurrence frequency threshold of the features in each item set;

[0076] Step S3b1, normalizing the weighted occurrence frequency threshold of the feature in each item set and the occurrence frequency of each feature in each item set;

[0077] It is worth mentioning that the weighted occurrence frequency threshold of the feature in each item set and the occurrence frequency of each feature in each item set need to be normalized. The main purpose of normalization is to ensure that the numerical range between the weighted occurrence frequency threshold and the occurrence frequency is consistent, so as to be comparable and avoid calculation and analysis errors.

[0078] Step S3b2: Use the weighted occurrence frequency threshold of the features in each item set after normalization to determine the occurrence frequency of each feature in each item set. If the occurrence frequency of the feature is greater than or equal to the weighted occurrence frequency threshold, retain the feature; if the occurrence frequency of the feature is less than the weighted occurrence frequency threshold, remove the feature.

[0079] Step S3b3: Based on the features retained in each item set, the items are aggregated and a candidate item set is generated according to the aggregated results.

[0080] Step S4: Perform support analysis on each candidate item set, and select candidate item sets whose support meets the support threshold as frequent item sets based on the analysis results.

[0081] Step S41: Count the total number of items in the engineering archive data set, and further count the number of occurrences of each candidate item set in the total number of items in the engineering archive data set.

[0082] Step S42: Obtain the support of each candidate item set according to the total number of items in the engineering archive data set and the number of occurrences of each candidate item set in the total items in the engineering archive data set.

[0083] The calculation formula for obtaining the support of each candidate item set is:

[0084] ;

[0085] in, Represents the support of the kth candidate item set; represents the number of occurrences of the kth candidate item set in the total items of the engineering archive dataset; Represents the total number of projects in the engineering archive dataset; Represents the kth candidate item set.

[0086] Step S43: Count the support of each candidate item set, take the average, set the average as the support threshold, and use the support threshold to determine the support of each candidate item set.

[0087] Step S44: if the support of the candidate item set is greater than or equal to the support threshold, then the candidate item set is retained and used as a frequent item set; if the support of the candidate item set is less than the support threshold, then the candidate item set is eliminated.

[0088] Step S5: Generate at least two strong association rules based on each frequent item set, and use the strong association rules to process the engineering archive data set.

[0089] It is necessary to generate at least two strong association rules based on each frequent item set, and use these strong association rules to process the engineering archive data set. The following is a detailed implementation example showing how to generate strong association rules and how to use these rules for data analysis and project management optimization:

[0090] 1. Data preparation;

[0091] After completing the above steps, the following frequent item sets are obtained:

[0092] Frequent item set 1: {building, large, 50 million, Beijing, with design documents, with construction records, with procurement information, without change records, with acceptance report, without maintenance records, with environmental impact assessment}; Frequent item set 2: {bridge, medium, 30 million, Shanghai, with design documents, with construction records, with procurement information, with change records, with acceptance report, with maintenance records, with environmental impact assessment}; Frequent item set 3: {road, small, 10 million, Guangzhou, with design documents, with construction records, with procurement information, without change records, with acceptance report, without maintenance records, without environmental impact assessment}.

[0093] 2. Generate strong association rules;

[0094] 2.1. In order to generate strong association rules, it is necessary to calculate the support and confidence of each rule. The support represents the frequency of the rule appearing in the engineering archive data set, while the confidence represents the probability of the conclusion occurring when the premise conditions are met.

[0095] 2.2. Generate association rules:

[0096] For each frequent item set, multiple candidate rules can be generated, and their support and confidence can be calculated. Then, the rules that meet the minimum support threshold and the minimum confidence threshold are selected as strong association rules. Taking frequent item set 1 as an example, candidate rules are generated and support and confidence are calculated:

[0097] Rule 1: {building, large, 50 million, Beijing, with design documents, with construction records, with procurement information, without change records, with acceptance reports, without maintenance records} → {with environmental impact assessment};

[0098] Support: supp({building, large, 50 million, Beijing, with design documents, with construction records, with procurement information, without change records, with acceptance reports, without maintenance records, with environmental impact assessment}) / supp({building, large, 50 million, Beijing, with design documents, with construction records, with procurement information, without change records, with acceptance reports, without maintenance records, with environmental impact assessment}) / N;

[0099] Confidence: supp({building, large, 50 million, Beijing, with design documents, construction records, procurement information, no change records, acceptance report, no maintenance records, environmental impact assessment}) / supp({building, large, 50 million, Beijing, with design documents, construction records, procurement information, no change records, acceptance report, no maintenance records})supp({building, large, 50 million, Beijing, with design documents, construction records, procurement information, no change records, acceptance report, no maintenance records, environmental impact assessment}) / supp({building, large, 50 million, Beijing, with design documents, construction records, procurement information, no change records, acceptance report, no maintenance records});

[0100] Rule 2: {building, large, 50 million, Beijing, with design documents, with construction records, with procurement information, without change records, with acceptance reports} → {without maintenance records, with environmental impact assessment};

[0101] Support: supp({building, large, 50 million, Beijing, with design documents, with construction records, with procurement information, without change records, with acceptance reports, without maintenance records, with environmental impact assessment}) / supp({building, large, 50 million, Beijing, with design documents, with construction records, with procurement information, without change records, with acceptance reports, without maintenance records, with environmental impact assessment}) / N;

[0102] Confidence: supp({construction, large, 50 million, Beijing, with design documents, construction records, procurement information, no change records, acceptance report, no maintenance records, environmental impact assessment}) / supp({construction, large, 50 million, Beijing, with design documents, construction records, procurement information, no change records, acceptance report})supp({construction, large, 50 million, Beijing, with design documents, construction records, procurement information, no change records, acceptance report, no maintenance records, environmental impact assessment});

[0103] Assuming that after the above calculation, both Rule 1 and Rule 2 meet the minimum support threshold, such as 0.5, and the minimum confidence threshold, such as 0.8, then these two rules are strong association rules;

[0104] It is worth noting that the support involved here does not represent the same support as the support involved in step S4. The support here is based on the support of the rule, while the support involved in step S4 is based on the support of the candidate item set, so special attention should be paid.

[0105] 3. Use strong association rules to process engineering archive data sets;

[0106] 3.1. Identify potential patterns and trends;

[0107] Through strong association rules, potential patterns and trends in engineering archive data sets can be identified. For example, Rule 1 shows that large-scale construction projects in Beijing, if they have complete construction records, design documents, etc., usually have environmental impact assessments; Rule 2 shows that large-scale construction projects in Beijing, if they have complete construction records, design documents, etc., usually do not have maintenance records, but have environmental impact assessments;

[0108] 3.2. Project management and optimization;

[0109] Based on the patterns and trends identified above, the following actions can be taken to optimize project management and decision making:

[0110] 3.21 Resource Allocation

[0111] Under Rule 1, an environmental impact assessment team can be arranged in advance for large-scale construction projects to ensure that the project's compliance and environmental protection requirements are met.

[0112] According to Rule 2, the reliance on maintenance records can be reduced, as such projects usually do not require maintenance records, thus saving related resources;

[0113] 3.22 Risk Management:

[0114] For rules with high confidence, they can be used as risk warning indicators. For example, if a project lacks necessary design documents or construction records, timely remediation can be carried out according to the rule prompts to avoid subsequent problems;

[0115] 3.23 Quality Control:

[0116] Through strong association rules, more refined quality control processes can be developed. For example, for large-scale construction projects, the review of design documents and construction records can be strengthened to ensure project quality.

[0117] 3.24. Cost Control:

[0118] By analyzing frequent item sets and strong association rules, we can find out which feature combinations will lead to higher budget requirements, so as to reasonably control costs in the project planning stage;

[0119] 3.25 Decision support:

[0120] In the project decision-making process, strong association rules can be referenced to predict the possible outcomes of the project. For example, if a project meets the conditions of Rule 1, it can be predicted that the project will require an environmental impact assessment, so that preparations can be made in advance.

[0121] Through the above steps, at least two strong association rules are generated based on each frequent item set, and these strong association rules are used to process the engineering archive data set. Specifically, potential patterns and trends are identified, and project management and optimization are carried out accordingly, including resource allocation, risk management, quality control, cost control and decision support. This will enable a better understanding and management of engineering projects and improve the overall efficiency and quality of the projects.

[0122] Embodiment 2: Figure 2 As shown, the present invention also proposes a system for processing engineering archives based on association rule mining, which executes the above-mentioned method for processing engineering archives based on association rule mining, and includes the following modules:

[0123] Data acquisition and preprocessing module: used to obtain and preprocess engineering archive data sets containing at least two projects;

[0124] Itemset formulation module: connected to the data acquisition and preprocessing module, used to extract features from the preprocessed engineering archive data set, and combine the features extracted from each project into an itemset to form at least two itemsets;

[0125] Preliminary analysis module: connected to the item set formulation module, used to analyze all item sets and generate at least two candidate item sets based on the analysis results;

[0126] Secondary analysis module: connected to the preliminary analysis module, used to perform support analysis on each candidate item set, and select candidate item sets whose support meets the support threshold as frequent item sets according to the analysis results;

[0127] Strong association rule application module: connected to the secondary analysis module, used to generate at least two strong association rules based on each frequent item set, and use the strong association rules to process the engineering archive data set.

[0128] It should be noted that the terms used in the present invention are only for describing specific embodiments, rather than limiting the scope of the present application. As shown in the present specification, unless the context clearly indicates an exception, the words "one", "a", "a kind of" and / or "the" do not specifically refer to the singular, but may also include the plural. The terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of more restrictions, the elements defined by the sentence "include one..." do not exclude the presence of other identical elements in the process, method or device including the elements.

[0129] It should also be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. Unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", etc. should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be a connection between the two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0130] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. A method for processing engineering archives based on association rule mining, characterized in that: The following steps are involved: Step S1, obtaining an engineering archive data set containing at least two projects and performing preprocessing; Step S2, extracting features from the preprocessed engineering archive data set, and combining the features extracted from each project into an item set to form at least two item sets; Step S3, analyzing all item sets, and generating at least two candidate item sets based on the analysis results; Step S3a, analyzing the occurrence frequencies of all features in each item set, and obtaining a weighted occurrence frequency threshold of the features in each item set according to the analysis results; Step S3a1, counting the occurrence frequencies of all features in each item set to obtain the occurrence frequency of each feature in each item set; Step S3a2: The occurrence frequency of each feature in each item set is synthesized and averaged to obtain the average occurrence frequency of each feature in each item set, and the standard deviation of the occurrence frequency of each feature in each item set is obtained based on the average occurrence frequency; Step S3a3, assigning a weight to each feature in each item set according to the frequency of occurrence of each feature in each item set; Step S3a4, obtaining a weighted occurrence frequency threshold of the feature in each item set according to the average occurrence frequency of each feature in each item set, the occurrence frequency standard deviation of each feature in each item set, and the weight of each feature in each item set; Step S3b, generating at least two candidate item sets according to the weighted occurrence frequency threshold of the features in each item set; Step S4: Perform support analysis on each candidate item set, and select candidate item sets whose support meets the support threshold as frequent item sets based on the analysis results; Step S5: Generate at least two strong association rules based on each frequent item set, and use the strong association rules to process the engineering archive data set.

2. The method for processing engineering archives based on association rule mining according to claim 1 is characterized in that: The step of generating at least two candidate item sets according to the weighted occurrence frequency threshold of the features in each item set includes: Step S3b1, normalizing the weighted occurrence frequency threshold of the feature in each item set and the occurrence frequency of each feature in each item set; Step S3b2: using the weighted occurrence frequency threshold of the features in each item set after normalization to determine the occurrence frequency of each feature in each item set, if the occurrence frequency of the feature is greater than or equal to the weighted occurrence frequency threshold, then retain the feature; if the occurrence frequency of the feature is less than the weighted occurrence frequency threshold, then remove the feature; Step S3b3: Based on the features retained in each item set, the items are aggregated and a candidate item set is generated according to the aggregated results.

3. The method for processing engineering archives based on association rule mining according to claim 2 is characterized in that: The calculation formula for obtaining the weighted occurrence frequency threshold of the feature in each item set is: ; in, represents the weighted frequency threshold of the feature in the i-th item set; represents the average occurrence frequency of the jth feature in the i-th item set; Represents the standard deviation of the occurrence frequency of the jth feature in the i-th item set; Represents the weight of the jth feature in the i-th item set; represents adjustment parameters; Represents the total number of features in the i-th itemset.

4. The method for processing engineering archives based on association rule mining according to claim 1 is characterized in that: The support analysis is performed on each candidate item set, and the candidate item sets whose support meets the support threshold are selected as frequent item sets according to the analysis results, including: Step S41, counting the total number of items in the engineering archive data set, and further counting the number of occurrences of each candidate item set in the total number of items in the engineering archive data set; Step S42, obtaining the support of each candidate item set according to the total number of items in the engineering archive data set and the number of occurrences of each candidate item set in the total items in the engineering archive data set; Step S43: Count the support of each candidate item set, take the average, set the average as the support threshold, and use the support threshold to determine the support of each candidate item set; Step S44: if the support of the candidate item set is greater than or equal to the support threshold, then the candidate item set is retained and used as a frequent item set; if the support of the candidate item set is less than the support threshold, then the candidate item set is eliminated.

5. The method for processing engineering archives based on association rule mining according to claim 4 is characterized in that: The calculation formula for obtaining the support of each candidate item set is: ; in, Represents the support of the kth candidate item set; represents the number of occurrences of the kth candidate item set in the total items of the engineering archive dataset; Represents the total number of projects in the engineering archive dataset; Represents the kth candidate item set.

6. A system for processing engineering archives based on association rule mining, used to execute a method for processing engineering archives based on association rule mining according to any one of claims 1 to 5, characterized in that: The system comprises: Data acquisition and preprocessing module: used to obtain and preprocess engineering archive data sets containing at least two projects; Itemset formulation module: connected to the data acquisition and preprocessing module, used to extract features from the preprocessed engineering archive data set, and combine the features extracted from each project into an itemset to form at least two itemsets; Preliminary analysis module: connected to the item set formulation module, used to analyze all item sets and generate at least two candidate item sets based on the analysis results; Secondary analysis module: connected to the preliminary analysis module, used to perform support analysis on each candidate item set, and select candidate item sets whose support meets the support threshold as frequent item sets according to the analysis results; Strong association rule application module: connected to the secondary analysis module, used to generate at least two strong association rules based on each frequent item set, and use the strong association rules to process the engineering archive data set.

Citation Information

Patent Citations

  • Data processing method, data processing device, equipment and storage medium

    CN116069993A

  • System and method for association itemset mining

    US20040220901A1