A Hidden Danger Warning Method for Construction Engineering Quality Based on Data Mining Technology

By segmenting the initial item set of building data and optimizing the support degree calculation, the problem of inaccurate correlation analysis in the existing technology is solved, the accuracy and comprehensiveness of early warning of hidden dangers for construction projects is achieved, and the efficiency and accuracy of data mining are improved.

CN118484776BActive Publication Date: 2025-07-08CIXI CHENGZHENG CONSTR ENG TESTING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410731816.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-06
Publication Date
2025-07-08
Estimated Expiration
2044-06-06

AI Technical Summary

Technical Problem

The existing correlation analysis methods ignore subtle changes in items in the item set and the indirect impact between item sets, resulting in inaccurate correlation and frequent item sets of important building data, resulting in incomplete mining results and inaccurate warnings for hidden dangers in construction projects.

Method used

By subdividing the initial item set of building data, calculating the initial support and correlation degree, optimizing the calculation method of support and correlation degree, processing the association relationship in special cases, building frequent item sets of building data, and generating association rules using data mining algorithms.

Benefits of technology

It improves the accuracy of correlation analysis and the comprehensiveness of mining results, ensures the accuracy of early warning of hidden dangers for construction projects, and improves the efficiency and accuracy of data mining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118484776B_ABST
    Figure CN118484776B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data mining, and proposes a method for warning potential hazards in construction project quality based on data mining technology, including: establishing an initial item set of construction data; determining the support degrees of two initial item sets of construction data, and determining the frequent item sets of construction data according to the values of the support degrees; in special cases, defining a candidate item set of construction data, obtaining a first special correlation degree, and assigning a value to the correlation degree according to the first special correlation degree; determining the frequent item sets of construction data, repeating the process of obtaining the frequent item sets of construction data until no new frequent item sets of construction data are generated, and using a data mining algorithm to determine the mining result of the association rules of construction-related data based on all the obtained frequent item sets of construction data, so as to realize the warning of potential hazards in construction project quality. The purpose of the present invention is to solve the problem that association analysis is prone to ignoring important information, resulting in incomplete mining results and inaccurate warning of potential hazards in construction project quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data mining, and particularly relates to a method for warning potential quality hazards in construction projects based on data mining technology. Background Art

[0002] In the context of big data, good management effects can be achieved through the application of big data mining. Therefore, the dependence of construction project management on big data mining is increasing. Big data mining can analyze construction project information, discover the associations between various types of information, and then obtain causal associations, temporal associations, and simple associations. Construction project managers can discover potential relationship network vulnerabilities in project management and quality hazards in construction projects based on these associations, thereby improving the quality of construction projects.

[0003] Generally, the Apriori algorithm is used for association analysis. In the Apriori algorithm, the occurrence frequency is used as the support degree of the item set. The support degree directly affects the quantity and quality of the frequent item sets of construction data finally mined, as well as the quality of the association rules mined. However, in the process of using the Apriori algorithm for association analysis, the subtle changes in the items within the item set and the indirect influence between the item sets are ignored. Therefore, the degree of association obtained between the item sets is not accurate, the potential association relationships in the dataset are also inaccurate, important frequent item sets of construction data are easily ignored, and further the mining results are incomplete and the reliability of the association rules is low. Therefore, the existing association analysis is prone to ignoring important information, resulting in incomplete mining results and inaccurate warning of potential quality hazards in construction projects. Summary of the Invention

[0004] The present invention provides a method for warning potential quality hazards in construction projects based on data mining technology to solve the problems that association analysis is prone to ignoring important information, resulting in incomplete mining results and inaccurate warning of potential quality hazards in construction projects. The specific technical solutions adopted are as follows:

[0005] An embodiment of the present invention provides a method for warning potential quality hazards in construction projects based on data mining technology. The method includes the following steps:

[0006] Collect construction-related data and convert it into numerical data to establish an initial item set of construction data;

[0007] Calculate the initial support degree of the initial item set of construction data and the occurrence frequency of binary construction data items. According to the differences between the items within the initial item set of construction data, obtain the association value of two initial item sets of construction data. Combine the occurrence frequency of the items within the initial item set of construction data to obtain the association degree of two initial item sets of construction data. Determine the support degree of two initial item sets of construction data based on the occurrence frequency and association degree of binary construction data items, and construct a frequent item set of construction data;

[0008] When there are two item sets constructed from three initial building data item sets that are not frequent building data item sets, and the item sets constructed from every other two initial building data item sets are frequent building data item sets, the initial building data item sets are used as candidate building data item sets, and the following process for obtaining frequent building data item sets is adopted to obtain frequent building data item sets:

[0009] According to the correlation degree between two different initial building data item sets in the candidate building data item sets and the absolute value of the items of the ternary items of the initial building data item sets, the updated support degree of the candidate building data item sets is obtained. According to the correlation degree between two different initial building data item sets and the updated support degree of the candidate building data item sets, the first special correlation degree is obtained, and the correlation degree of the initial building data item sets is assigned according to the first special correlation degree; According to the ternary items of the candidate building data item sets and the correlation degree between the candidate building data item sets, the final support degree of the candidate building data item sets is obtained, and the frequent building data item sets are determined according to the final support degree;

[0010] Define the candidate building data item sets again, and obtain the frequent building data item sets again according to the obtaining process of the process for obtaining frequent building data item sets. Among them, the number of candidate building data item sets is updated, and the number of candidate building data item sets is the number of the original candidate building data item sets plus 1, until no new frequent building data item sets are generated;

[0011] Use a data mining algorithm to determine the mining result of the association rules of construction-related data based on all the obtained frequent building data item sets, and the mining result is used for early warning of potential quality hazards in construction projects.

[0012] Furthermore, the specific method for obtaining the occurrence frequency of the binary building data items is as follows:

[0013] Arbitrarily select one item from each of the two initial building data item sets for combination to obtain binary building data items, obtain all the binary building data items corresponding to all different initial building data item sets, and calculate the occurrence frequency of each binary building data item among all the binary building data items.

[0014] Furthermore, the specific method for obtaining the correlation value between two initial building data item sets according to the differences between the items within the initial building data item sets includes:

[0015] Sort the occurrence frequencies of all the binary building data items determined by the two initial building data item sets in descending order according to the occurrence frequency of the binary building data items to obtain a binary building data item sequence;

[0016] Denote the absolute value of the difference between the occurrence frequencies of two adjacent ones in the binary building data item sequence as the first occurrence frequency absolute value of the two adjacent occurrence frequencies.

[0017] Denote the latter building data initial item set among the two building data initial item sets as the post-building data initial item set, and denote the absolute value of the difference between the data corresponding to the two adjacent occurrence frequencies in the post-building data initial item set as the second occurrence frequency absolute value.

[0018] Denote the product of the first occurrence frequency absolute value and the second occurrence frequency absolute value as the first occurrence frequency product, and there is a negative correlation between the association value of the two building data initial item sets and the first occurrence frequency product.

[0019] Furthermore, the method for obtaining the association degree of the two building data initial item sets by combining the occurrence frequencies of the items in the building data initial item set specifically includes:

[0020] Set any one of the two building data initial item sets as the reference building data initial item set, the reference building data initial item set;

[0021] Arrange the items included in the reference building data initial item set in ascending order to obtain the first sequence; arrange the occurrence frequencies of each item in the first sequence in the order of the items in the first sequence among all the construction-related data converted into numerical data to obtain the second sequence, and perform curve fitting on the first sequence and the second sequence respectively to obtain the first curve of the previous building data initial item set.

[0022] Denote the correlation coefficient between the first curves of the two building data initial item sets as the first correlation coefficient.

[0023] The association degrees of the two building data initial item sets are positively correlated with the first correlation coefficient and the association value of the two building data initial item sets respectively.

[0024] Furthermore, the method for determining the support degrees of the two building data initial item sets and constructing the building data frequent item set according to the occurrence frequencies and association degrees of the binary building data items specifically includes:

[0025] The support degrees of the two building data initial item sets are positively correlated with the occurrence frequencies and association degrees of all the binary building data items determined by the two building data initial item sets.

[0026] Set a support degree threshold. When the support degree of the two building data initial item sets is greater than or equal to the support degree threshold, denote the item set formed by combining the two building data initial item sets as the building data frequent item set.

[0027] Further, the specific method for including the association degree between two different initial building data item sets in the candidate building data item set and the absolute value of the item of the three-item of the initial building data item set is as follows:

[0028] Denote the initial building data item sets A, B, and C as candidate building data item sets;

[0029] Combine any item in the initial building data item set A, any item in the initial building data item set B, and any item in the initial building data item set C to obtain three-items, obtain all three-items of the three initial building data item sets in all special cases, calculate the occurrence frequency of each three-item in all three-items, and sort the occurrence frequencies of all three-items determined by the initial building data item sets A, B, and C in descending order of the occurrence frequency of the three-items to obtain a three-item sequence;

[0030] Take any three-item in the three-item sequence as the three-item to be analyzed, calculate the absolute value of the difference between the corresponding items of the three-item to be analyzed and the next three-item in the initial building data item sets C, A, and B respectively, and denote them as the first absolute value, the second absolute value, and the third absolute value in sequence; Denote the ratio of the second absolute value to the third absolute value as the first ratio; Denote the absolute value of the difference between the first absolute value and the first ratio as the absolute value of the item of the three-item to be analyzed.

[0031] Further, the specific method for obtaining the updated support degree of the candidate building data item set according to the association degree between two different initial building data item sets in the candidate building data item set and the absolute value of the item of the three-item of the initial building data item set is as follows:

[0032] Denote the association degree between the initial building data item set A and the initial building data item set B as the first association degree, and denote the association degree between the initial building data item set B and the initial building data item set C as the second association degree. The updated support degree of the candidate building data item set is positively correlated with the first association degree, the second association degree, and the absolute value of the item of the three-item determined by the initial building data item set respectively.

[0033] Further, the specific method for obtaining the final support degree of the candidate building data item set according to the three-item of the candidate building data item set and the association degree between the candidate building data item sets is as follows:

[0034] Denote the sum of the occurrence frequencies of all three-items included in the three-item sequence as the second sum value, and denote the sum of the association degrees between all candidate building data item sets as the third sum value. The final support degree of the candidate building data item set is positively correlated with the second sum value and the third sum value respectively.

[0035] Further, the specific method for determining the frequent building data item set according to the final support degree is as follows:

[0036] When the final support of the candidate building data item set is greater than or equal to the support threshold, the item set composed of all candidate building data item sets is recorded as the building data frequent item set.

[0037] Furthermore, the mining result of the association rules of the construction-related data is determined based on all the obtained building data frequent item sets by using a data mining algorithm. The specific method includes:

[0038] Using a data mining algorithm, generate association rules based on all the obtained building data frequent item sets and obtain the confidence of the candidate association rules;

[0039] Set an interference threshold. When the normalized value of the confidence is greater than or equal to the interference threshold, determine the candidate association rule as an actual association rule; otherwise, determine the candidate association rule as an interference association rule;

[0040] Determine all the actual association rules as the mining result of the association rules of the construction-related data.

[0041] The beneficial effects of the present invention are:

[0042] In the present invention, according to the existing data mining algorithm, the occurrence frequency of items is used as the support of the item set. However, the occurrence frequency only reflects the association degree between two item sets from the co-occurrence frequency of the two item sets, which may ignore the subtle changes of items within the item set and the indirect influence between item sets, resulting in inaccurate association degrees between the obtained item sets and inaccurate potential association relationships in the data set, and it is easy to ignore important building data frequent item sets, thus leading to the problems of incomplete mining results and low reliability of association rules. First, the construction-related data in the construction process of a building project is divided into building data initial item sets, and each building data initial item set is analyzed separately. The support of two building data initial item sets is determined according to the difference in the occurrence frequency of the binary building data items corresponding to different building data initial item sets. According to the numerical relationship between the support and the support threshold, the building data frequent item set is determined;

[0043] Then, according to the problem that the actual effect constructed by updating the support based on the difference in the association relationship between different building data initial item sets is poor, candidate building data item sets are defined according to the building data initial item sets. The first special association degree is obtained according to the ternary items corresponding to the candidate building data item sets and the association degree between different building data initial item sets in the candidate building data item sets. The association degree is assigned according to the first special association degree to complete the processing of special cases;

[0044] Finally, according to the three-item terms of the candidate building data item set and the correlation degree between the candidate building data item sets, obtain the final support degree of the candidate building data item sets. Determine the frequent item sets of building data based on the final support degree, and repeat the process of obtaining the frequent item sets of building data until no new frequent item sets of building data are generated. Use the data mining algorithm to determine the mining result of the association rules of the construction-related data based on all the obtained frequent item sets of building data, realize the early warning of potential quality hazards in construction projects based on data mining technology, and solve the problem that the association analysis is likely to ignore important information, resulting in incomplete mining results and inaccurate early warning of potential quality hazards in construction projects. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0046] Figure 1 It is a schematic flow chart of a method for early warning of potential quality hazards in construction projects based on data mining technology provided by an embodiment of the present invention;

[0047] Figure 2 It is a flow chart for obtaining the correlation value. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0049] Please refer to Figure 1 , which shows a flow chart of a method for early warning of potential quality hazards in construction projects based on data mining technology provided by an embodiment of the present invention. The method includes the following steps:

[0050] Step S001: Collect construction-related data and convert it into numerical data to establish an initial item set of building data.

[0051] Obtain the construction-related data during the construction process of the construction project. The construction-related data includes the structural data, construction progress data, material data, safety management data, and quality inspection data of the construction project.

[0052] Among them, the structural data includes the quality inspection data of the structural materials of the building, such as the data of concrete, steel bars, masonry and other materials, and the structural monitoring data includes the real-time monitoring data such as the concrete pouring temperature and the wall thickness; the construction progress data includes the planned progress data and the actual progress data; the material data includes the data such as the qualification or reputation of the material supplier and the usage data of the materials, and the usage data of the materials includes the procurement data, the usage amount and the consumption situation of the materials, etc.; the safety management data is the safety inspection record data of the construction site; the quality inspection data is the engineering quality inspection report data. Implementers can also select construction-related data according to specific actual applications, and this embodiment does not make special restrictions.

[0053] Since the construction-related data includes non-numerical data, in order to ensure the reliability of the analysis, it is necessary to convert the non-numerical data included in the construction-related data into numerical data. In this embodiment, binary coding is selected to convert non-numerical data into numerical data. Implementers can also adopt other methods to convert non-numerical data into numerical data. This embodiment does not make restrictions. The specific process of converting non-numerical data into numerical data using binary coding is a well-known technology and will not be elaborated here.

[0054] The set composed of data of the same type is denoted as the initial item set of building data. For example, the material supplier reputation data includes excellent reputation, good reputation, passing reputation, and poor reputation. The set composed of the corresponding numerical data of these data is the initial item set of building data.

[0055] Thus far, the initial item set of building data is obtained.

[0056] Step S002: Calculate the initial support degree of the initial item set of building data and the occurrence frequency of binary building data items, determine the support degrees of two initial item sets of building data, and determine the frequent item sets of building data according to the values of the support degrees.

[0057] The APriori algorithm is an association rule mining algorithm in data mining algorithms, which is used to discover frequent item sets of building data in a large-scale dataset, and then generate association rules. The association rules reveal the association relationships between items in the dataset. The basic idea of the APriori algorithm is: find all frequent item sets of building data, and the frequency of these frequent item sets of building data appears at least as high as the predefined minimum support degree. Then, generate strong association rules from the frequent item sets of building data, and these rules must meet the minimum support degree and the minimum confidence. Among them, the items in the dataset are the data in the dataset.

[0058] Since the Apriori algorithm usually uses the occurrence frequency of items as the support of item sets, the support directly affects the quantity and quality of the frequent item sets of building data finally mined, as well as the quality of the mined association rules. However, the occurrence frequency only reflects the correlation degree between two item sets from the co-occurrence frequency of the two item sets, and may ignore the subtle changes of items within the item sets and the indirect influence between item sets. For example, there is an indirect influence between the quality inspection data of the structural material concrete of a building and the concrete pouring temperature. When the concrete pouring temperature is too high or too low, the quality of the concrete will be affected, resulting in a difference between the quality inspection data of the concrete and the standard value. Therefore, the correlation degree obtained between item sets is not accurate, and the potential association relationships in the data set are also inaccurate, easily ignoring important frequent item sets of building data, thus leading to incomplete mining results and relatively low reliability of association rules.

[0059] Statistically analyze the occurrence frequency of each data in all construction-related data converted into numerical data in each initial item set of building data, and record the sum of the occurrence frequencies of all data in the initial item set of building data as the initial support of the initial item set of building data.

[0060] In the process of using the Apriori algorithm for association analysis, the amount of construction-related data is large. Directly performing the combination of items and the calculation and judgment of support will result in a large number of useless calculations and interferences, affecting the efficiency and accuracy of the calculation. Therefore, in this embodiment, the initial item set of building data is obtained, and then the initial support of the initial item set of building data is statistically analyzed, and the combination of the initial item set of building data and the calculation and judgment of support are performed to reduce the calculations that are meaningless for data mining and improve the efficiency and accuracy of data mining.

[0061] For example, the material supplier credit data includes excellent credit, good credit, passing credit, and poor credit. Directly performing the combination of items and the calculation and judgment of support will result in calculating the support between the item sets composed of excellent credit and good credit, but such calculations are meaningless and will affect the efficiency and accuracy of data mining. In this embodiment, the combination of the initial item set of building data and the calculation and judgment of support are performed, which will avoid the above meaningless calculations and improve the efficiency and accuracy of data mining.

[0062] This embodiment takes the initial item sets A and B of building data as examples for analysis.

[0063] Combine any item in the initial building data item set A with any item in the initial building data item set B to obtain binary building data items. Obtain all binary building data items corresponding to all different initial building data item sets. Calculate the occurrence frequency of each binary building data item among all binary building data items, and sort the occurrence frequencies of all binary building data items determined by the initial building data item sets A and B in descending order of the occurrence frequency of the binary building data items to obtain a binary building data item sequence.

[0064] Denote the absolute value of the difference between two adjacent occurrence frequencies in the binary building data item sequence as the first absolute occurrence frequency, and denote the absolute value of the difference between the corresponding data of the two adjacent occurrence frequencies corresponding to the first absolute occurrence frequency in the initial building data item set B as the second absolute occurrence frequency.

[0065] Determine the association value between the initial building data item sets A and B according to the first absolute occurrence frequency and the second absolute occurrence frequency. The flowchart for obtaining the association value is as Figure 2 shown. Denote the product of the first absolute occurrence frequency and the second absolute occurrence frequency as the first occurrence frequency product, and there is a negative correlation between the association value of the initial building data item sets A and B and the first occurrence frequency product.

[0066] It can be understood that the negative correlation relationship in this application refers to the relationship between an independent variable and a dependent variable. The negative correlation relationship means that the independent variable decreases (increases) as the dependent variable increases (decreases), which can be an inverse ratio relationship, a subtraction relationship, etc.

[0067] Preferably, in an embodiment of this application, denote the sum of all first occurrence frequency products determined by the binary building data item sequence as the first frequency sum, and denote the power with the natural constant as the base and the opposite of the first frequency sum as the exponent as the association value between the initial building data item sets A and B.

[0068] When the difference between adjacent occurrence frequencies in the binary building data item sequence determined by the initial building data item sets A and B is larger and the difference between the corresponding data of adjacent occurrence frequencies in the binary building data item sequence in the initial building data item set B is larger, when the items in the initial building data item set A do not change, the possibility of the items in the initial building data item set B changing is greater, and the difference between the occurrence frequencies of different binary building data items is greater. At this time, the association value between the initial building data item sets A and B is smaller.

[0069] Arrange the items included in the initial building data item set A in ascending order to obtain the first sequence; arrange the occurrence frequencies of each item in the first sequence in the construction-related data all converted into numerical data in the order of the items in the first sequence to obtain the second sequence, and perform curve fitting on the first sequence and the second sequence respectively to obtain the first curve of the initial building data item set A. According to the acquisition method of the first curve of the initial building data item set A, obtain the first curve of the initial building data item set B. Denote the correlation coefficient between the first curves of the initial building data item set A and the initial building data item set B as the first correlation coefficient.

[0070] In some embodiments of the present application, the Pearson correlation coefficient is used as the correlation coefficient between the first curves of the initial building data item set A and the initial building data item set B. As other implementation manners, implementers can also adopt other calculation methods of correlation coefficients in the prior art. In this embodiment, the Pearson correlation coefficient is used to represent the correlation between the initial building data item sets A and B.

[0071] Determine the association degree between the initial building data item sets A and B according to the first correlation coefficient and the association value. The association degree between the initial building data item sets A and B is positively correlated with the first correlation coefficient and the association value between the initial building data item sets A and B.

[0072] It can be understood that the positive correlation relationship in the present application refers to the relationship between the independent variable and the dependent variable. The positive correlation relationship means that the independent variable increases (decreases) as the dependent variable increases (decreases), and it can be an additive relationship, a multiplicative relationship, etc.

[0073] Preferably, in an embodiment of the present application, the association degree between the initial building data item sets A and B is the product of the first correlation coefficient and the association value between the initial building data item sets A and B.

[0074] Determine the support degree between the initial building data item sets A and B according to the occurrence frequency and the association degree of the binary building data items. The support degree between the initial building data item sets A and B is positively correlated with the occurrence frequencies and the association degree of all binary building data items determined by the initial building data item sets A and B.

[0075] Preferably, as an embodiment of the present application, denote the sum of the occurrence frequencies of all binary building data items determined by the initial building data item sets A and B as the second frequency sum, and the support degree between the initial building data item sets A and B is the normalized value of the product of the second frequency sum and the association degree between the initial building data item sets A and B.

[0076] Set a support degree threshold, and the value range of the support degree threshold is greater than or equal to 0.6 and less than or equal to 0.8. In this embodiment, the value of the support degree threshold is 0.7.

[0077] When the support of two initial building data item sets is greater than or equal to the support threshold, the item set formed by combining the two initial building data item sets is denoted as the frequent building data item set.

[0078] So far, the frequent building data item set is determined.

[0079] Step S003: In special cases, define the candidate building data item set, obtain the first special correlation degree, assign values to the correlation degree according to the first special correlation degree, and complete the processing of special cases.

[0080] After obtaining the frequent building data item set, it is necessary to perform connection pruning on the frequent building data item set to obtain three initial building data item sets. In this embodiment, the frequent building data item set formed by combining the initial building data item sets A and B is taken as an example for analysis. The three initial building data item sets obtained by performing connection pruning on the frequent building data item set are denoted as A, B, and C. For convenience of description, the initial building data item sets A, B, and C are denoted as candidate building data item sets.

[0081] When the item set formed by combining the initial building data item sets A and B is a frequent building data item set, the item set formed by combining the initial building data item set C and the initial building data item set B is a frequent building data item set, and the item set formed by combining the initial building data item set A and the initial building data item set C is not a frequent building data item set, it is denoted as a special case, and the following analysis is continued; when the conditions for continuing the following analysis are not met, the following optimization process is not performed.

[0082] For example: The concrete quality inspection data item set B will directly affect the concrete pouring quality inspection data item set C, and the concrete material supplier data item set A will directly affect the concrete quality inspection data item set B and indirectly affect the concrete pouring quality inspection data item set C. At this time, the correlation degree calculated between the item set A and the item set B is relatively large, the correlation degree calculated between the item set B and the item set C is relatively large, and the correlation degree calculated between the item set A and the item set C is relatively small. For the specific value of the support, for example, when the support of the item set formed by combining the initial building data item sets A and B is 0.8, the support of the item set formed by combining the initial building data item set C and the initial building data item set B is 0.95, and the support of the item set formed by combining the initial building data item set A and the initial building data item set C is 0.5, it is the above special case.

[0083] The existing Apriori algorithm will analyze the three initial item sets of building data obtained by connection pruning, that is, taking the co-occurrence frequency of the initial item sets A, B, and C of building data as the updated support degree. However, a special situation may occur, that is, the correlation between the initial item set A of building data and the initial item set B of building data is relatively large, the correlation between the initial item set C of building data and the initial item set B of building data is relatively large, but the correlation between the initial item set A of building data and the initial item set C of building data is relatively small (this special situation is the condition for the following analysis). At this time, there is an indirect correlation between the initial item set A of building data and the initial item set C of building data. The initial item set A of building data can affect the initial item set C of building data by influencing the initial item set B of building data. If the Apriori algorithm is used to take the co-occurrence frequency of the initial item sets A, B, and C of building data as the method of updated support degree, the actual effect of constructing the updated support degree will be poor. To avoid this problem, in this embodiment, the correlation between the three initial item sets of building data obtained by connection pruning is optimized, and then the updated support degree is constructed according to the optimized result to improve the actual effect of constructing the updated support degree.

[0084] Combine any item in the initial item set A of building data, any item in the initial item set B of building data, and any item in the initial item set C of building data to obtain a ternary item. Obtain all ternary items of the three initial item sets of building data that continue the following analysis. Calculate the occurrence frequency of each ternary item in all ternary items, and sort the occurrence frequencies of all ternary items determined by the initial item sets A, B, and C of building data in descending order of the occurrence frequency of the ternary items to obtain a ternary item sequence.

[0085] Take each ternary item in the ternary item sequence as the ternary item to be analyzed. Denote the absolute value of the difference between the ternary item to be analyzed and the item corresponding to the initial item set C of the next ternary item in the ternary item to be analyzed as the first absolute value; denote the absolute value of the difference between the ternary item to be analyzed and the item corresponding to the initial item set A of the next ternary item in the ternary item to be analyzed as the second absolute value; denote the absolute value of the difference between the ternary item to be analyzed and the item corresponding to the initial item set B of the next ternary item in the ternary item to be analyzed as the third absolute value; denote the ratio of the second absolute value to the third absolute value as the first ratio; denote the absolute value of the difference between the first absolute value and the first ratio as the item absolute value of the ternary item to be analyzed.

[0086] The greater the item absolute value of the ternary item to be analyzed, the greater the influence of the change of the item in the initial item set A of building data on the change of the item in the initial item set C of building data.

[0087] Obtain the updated support degree of the candidate building data item set according to the absolute value of the item of the ternary item determined by the correlation degree of different initial building data item sets in the candidate building data item set and the initial building data item set. Denote the correlation degree between the initial building data item set A and the initial building data item set B as the first correlation degree, and denote the correlation degree between the initial building data item set B and the initial building data item set C as the second correlation degree. The updated support degree of the candidate building data item set is positively correlated with the first correlation degree, the second correlation degree, and the absolute value of the item of the ternary item determined by the initial building data item set.

[0088] Preferably, in an embodiment of the present application, the normalized value of the product of the sum of the absolute values of all items of the ternary item determined by the initial building data item set and the first correlation degree and the second correlation degree is denoted as the updated support degree of the candidate building data item set.

[0089] Obtain the first special correlation degree according to the correlation degree between the initial building data item sets A and B and the updated support degree of the candidate building data item set. The first special correlation degree is the correction value of the correlation degree obtained by processing the above special situation.

[0090] The first special correlation degree is positively correlated with both the correlation degree between the initial building data item sets A and B and the updated support degree of the candidate building data item set.

[0091] Preferably, in an embodiment of the present application, the sum of the updated support degree of the candidate building data item set and the number 1 is denoted as the first sum value, and the product of the first sum value and the correlation degree between the initial building data item sets A and B is denoted as the first special correlation degree.

[0092] Take the first special correlation degree as the processing result of the above special situation, assign the value of the correlation degree to the first special correlation degree, and complete the processing of the above special situation.

[0093] Thus, the processing of the special situation is completed.

[0094] Step S004: Determine the frequent building data item set, repeat the process of obtaining the frequent building data item set until no new frequent building data item set is generated, and use the data mining algorithm to determine the mining result of the association rules of the construction-related data based on all the obtained frequent building data item sets, so as to realize the early warning of potential quality hazards in construction projects based on data mining technology.

[0095] Denote the sum of the occurrence frequencies of all items included in the ternary item sequence as the second sum value, and denote the sum of the correlation degrees between all candidate building data item sets as the third sum value.

[0096] Obtain the final support degree of the candidate building data item set according to the third sum value and the second sum value. The final support degree of the candidate building data item set is positively correlated with both the second sum value and the third sum value.

[0097] Preferably, in an embodiment of the present application, the final support of the candidate building data item set is the product of the second sum value and the third sum value.

[0098] Compare the final support of the candidate building data item set with the support threshold. When the final support of the candidate building data item set is greater than or equal to the support threshold, the item set composed of all candidate building data item sets is denoted as the building data frequent item set.

[0099] Continuously repeat the above process for processing the building data frequent item set, where the number of candidate building data item sets is updated, and the number of candidate building data item sets is the number of the original candidate building data item sets plus 1, until no new building data frequent item sets are generated.

[0100] Use the Apriori algorithm to generate association rules based on all obtained building data frequent item sets and obtain the confidence of candidate association rules. Each candidate association rule has a corresponding confidence, and the confidence of the candidate association rule represents the probability that the consequent appears when the antecedent appears.

[0101] Perform normalization processing on all confidences, and use the average value of the normalized values of the confidences as the interference threshold. When the normalized value of the confidence is greater than or equal to the interference threshold, the candidate association rule is considered as an actual association rule; otherwise, the candidate association rule is considered as an interference association rule. Among them, the threshold setter can adjust it.

[0102] It should be understood that the calculation method of the interference threshold is only one implementation manner of the present application. As other implementation manners, the implementer can determine the setting method of the interference threshold according to the actual application situation, and the present application does not make special restrictions.

[0103] All actual association rules are the mining results of the association rules of construction-related data. Construction management personnel perform early warning of potential quality hazards in construction projects based on the mining results of construction-related data, so as to realize early warning of potential quality hazards in construction projects based on data mining technology.

[0104] In the identification and assessment of construction project quality risks, early warning of potential quality hazards in construction projects based on data mining technology can efficiently process a large amount of data, mine potential risk factors, and can realize rapid identification and accurate assessment of project quality risks, providing strong support for risk assessment.

[0105] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for warning of potential quality hazards in construction projects based on data mining technology, characterized in that, The method includes the following steps: Collect construction-related data and convert it into numerical data, and establish an initial item set of building data; Arbitrarily select an item from each of the two initial item sets of building data for combination to obtain binary building data items, calculate the initial support of the initial item set of building data and the occurrence frequency of the binary building data items, obtain the correlation value between the two initial item sets of building data according to the differences between the items in the initial item set of building data, combine the occurrence frequencies of the items in the initial item set of building data to obtain the correlation degree between the two initial item sets of building data, and determine the support of the two initial item sets of building data according to the occurrence frequency and correlation degree of the binary building data items, and construct a frequent item set of building data; When, among the three initial item sets of building data, there are item sets constructed from two initial item sets of building data that are not frequent item sets of building data, and the item sets constructed from every other two initial item sets of building data are frequent item sets of building data, use the initial item set of building data as a candidate item set of building data, and obtain a frequent item set of building data by adopting the following process for obtaining a frequent item set of building data: For any item selected from each of the three initial item sets of building data for combination to obtain a ternary item of the initial item set of building data, obtain the updated support of the candidate item set of building data according to the correlation degree between two different initial item sets of building data in the candidate item set of building data and the absolute value of the item of the ternary item of the initial item set of building data, obtain the first special correlation degree according to the correlation degree between two different initial item sets of building data and the updated support of the candidate item set of building data, and assign a value to the correlation degree of the initial item set of building data according to the first special correlation degree; obtain the final support of the candidate item set of building data according to the ternary item of the candidate item set of building data and the correlation degree between the candidate item sets of building data, and determine the frequent item set of building data according to the final support; Define the candidate item set of building data again, and obtain the frequent item set of building data again according to the process for obtaining the frequent item set of building data, wherein the number of candidate item sets of building data is updated, and the number of candidate item sets of building data is the number of the original candidate item set of building data plus 1, until no new frequent item sets of building data are generated; Use a data mining algorithm to determine the mining result of the association rules of the construction-related data based on all the obtained frequent item sets of building data, and the mining result is used for early warning of potential quality hazards in construction projects.

2. The early warning method for potential quality hazards in construction projects based on data mining technology according to claim 1, characterized in that, The specific method for obtaining the occurrence frequency of the binary building data items is as follows: Obtain all binary building data items corresponding to all different initial item sets of building data, and calculate the occurrence frequency of each binary building data item among all binary building data items.

3. A method for warning of potential quality hazards in construction projects based on data mining technology according to claim 1, characterized in that, The specific method included in obtaining the correlation value between two initial item sets of building data according to the differences between the items in the initial item set of building data is as follows: Sort the occurrence frequencies of all binary building data items determined by the two initial item sets of building data in descending order according to the occurrence frequency of the binary building data items to obtain a binary building data item sequence; Record the absolute value of the difference between two adjacent occurrence frequencies in the binary building data item sequence as the first occurrence frequency absolute value of the two adjacent occurrence frequencies; Denote the latter construction data initial item set among the two construction data initial item sets as the post-construction data initial item set, and denote the absolute value of the difference between the corresponding data of the adjacent two occurrence frequencies in the post-construction data initial item set as the second occurrence frequency absolute value; Denote the product of the first occurrence frequency absolute value and the second occurrence frequency absolute value as the first occurrence frequency product, and there is a negative correlation between the association value of the two construction data initial item sets and the first occurrence frequency product.

4. A method for warning of potential quality hazards in construction projects based on data mining technology according to claim 3, characterized in that, The specific method for obtaining the association degree of the two construction data initial item sets by combining the occurrence frequencies of the items in the construction data initial item set includes: Set any one of the two construction data initial item sets as the reference construction data initial item set; Arrange the items included in the reference construction data initial item set in ascending order to obtain the first sequence; arrange the occurrence frequencies of each item in the first sequence in all construction-related data converted into numerical data in the order of the items in the first sequence to obtain the second sequence, and perform curve fitting on the first sequence and the second sequence to obtain the first curve of the reference construction data initial item set; Denote the correlation coefficient between the first curves of the two construction data initial item sets as the first correlation coefficient; The association degrees of the two construction data initial item sets are positively correlated with the first correlation coefficient and the association value of the two construction data initial item sets respectively.

5. A method for warning of potential quality hazards in construction projects based on data mining technology according to claim 1, characterized in that, The specific method for determining the support degrees of the two construction data initial item sets according to the occurrence frequencies and association degrees of the binary construction data items and constructing the construction data frequent item set includes: The support degrees of the two construction data initial item sets are positively correlated with the occurrence frequencies and association degrees of all binary construction data items determined by the two construction data initial item sets; Set a support threshold. When the support degree of the two construction data initial item sets is greater than or equal to the support threshold, denote the item set formed by combining the two construction data initial item sets as the construction data frequent item set.

6. A method for warning of potential quality hazards in construction projects based on data mining technology according to claim 1, characterized in that, The specific method for calculating according to the association degree between two different construction data initial item sets in the candidate construction data item set and the item absolute value of the ternary item of the construction data initial item set includes: Denote the construction data initial item sets A, B, and C as the candidate construction data item set; Combine any one item in the construction data initial item set A, any one item in the construction data initial item set B, and any one item in the construction data initial item set C to obtain the ternary item, obtain all ternary items of the three construction data initial item sets in all special cases, calculate the occurrence frequency of each ternary item in all ternary items, and sort the occurrence frequencies of all ternary items determined by the construction data initial item sets A, B, and C in descending order of the occurrence frequency of the ternary item to obtain the ternary item sequence; Take any one of the ternary items in the ternary item sequence as the ternary item to be analyzed. Calculate the absolute values of the differences between the corresponding items of the ternary item to be analyzed and the next ternary item in the initial item sets C, A, and B of the building data, and record them as the first absolute value, the second absolute value, and the third absolute value respectively. Denote the ratio of the second absolute value to the third absolute value as the first ratio. Denote the absolute value of the difference between the first absolute value and the first ratio as the item absolute value of the ternary item to be analyzed.

7. A method for warning of potential quality hazards in construction projects based on data mining technology according to claim 6, characterized in that, The specific method for obtaining the updated support degree of the candidate building data item set according to the correlation degree between two different initial item sets of the building data and the item absolute value of the ternary item of the initial item set of the building data includes: Denote the correlation degree between the initial item set A of the building data and the initial item set B of the building data as the first correlation degree, and denote the correlation degree between the initial item set B of the building data and the initial item set C of the building data as the second correlation degree. The updated support degree of the candidate building data item set is positively correlated with the first correlation degree, the second correlation degree, and the item absolute value of the ternary item determined by the initial item set of the building data respectively.

8. A method for warning of potential quality hazards in construction projects based on data mining technology according to claim 1, characterized in that, The specific method for obtaining the final support degree of the candidate building data item set according to the ternary item of the candidate building data item set and the correlation degree between the candidate building data item sets includes: Denote the sum of the occurrence frequencies of all ternary items included in the ternary item sequence as the second sum value, and denote the sum of the correlation degrees between all candidate building data item sets as the third sum value. The final support degree of the candidate building data item set is positively correlated with the second sum value and the third sum value respectively.

9. A method for warning of potential quality hazards in construction projects based on data mining technology according to claim 1, characterized in that, The specific method for determining the frequent item set of the building data according to the final support degree includes: When the final support degree of the candidate building data item set is greater than or equal to the support degree threshold, denote the item set composed of all candidate building data item sets as the frequent item set of the building data.

10. A method for warning of potential quality hazards in construction projects based on data mining technology according to claim 1, characterized in that, The specific method for using the data mining algorithm to determine the mining result of the association rule of the construction-related data based on all the obtained frequent item sets of the building data includes: Use the data mining algorithm to generate association rules based on all the obtained frequent item sets of the building data and obtain the confidence degree of the candidate association rules; Set the interference threshold. When the normalized value of the confidence degree is greater than or equal to the interference threshold, determine the candidate association rule as the actual association rule, otherwise, determine the candidate association rule as the interference association rule; Determine all the actual association rules as the mining result of the association rules of the construction-related data.

Citation Information

Patent Citations

  • Electric power engineering project cooperative relationship feature identification method and system

    CN114971562A

  • Railway vehicle RAMS data correlation analysis method based on data mining

    CN117194995A