A model construction method and system based on data hierarchical classification

By constructing a basic knowledge graph and defining relationships based on semantic analysis and changes in project status, the problem of scattered and complex data in construction projects was solved, achieving efficient and accurate data classification and processing.

CN121580147BActive Publication Date: 2026-05-19GUIZHOU BAISHENG CONSTR ENG CONSULTING CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUIZHOU BAISHENG CONSTR ENG CONSULTING CO LTD
Filing Date
2026-01-28
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In the digital transformation of the construction industry, construction project data is scattered, diverse in type, and lacks a unified classification and coding standard, resulting in high data processing and analysis costs, complex integration solutions, and low data accuracy and efficiency.

Method used

By collecting construction project data, extracting status keywords using semantic analysis technology, constructing a basic knowledge graph, defining relationships based on changes in project status, tracking the change paths of adjacent nodes, and conducting iterative feedback, a structured data hierarchical model is formed.

Benefits of technology

It improves the reliability of data association and the efficiency of hierarchical processing, ensures the consistency and accuracy of data hierarchical results, supports dynamic updates and scenario-based classification of project data, and improves the efficiency and accuracy of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121580147B_ABST
    Figure CN121580147B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, in particular to a model construction method and system based on data hierarchical classification. The method comprises the following steps: preliminarily marking building project data according to the data types of the building project, and setting project entities; extracting state keywords from each project entity, and performing semantic clustering according to the scenes to which the state keywords belong, to construct a basic knowledge graph containing state labels; based on the generated basic knowledge graph, defining the correlation of each project entity under the project state change as a trigger condition; determining the adjacent nodes when the project state changes, tracking the change path of the adjacent nodes, and updating the change path to the edge attribute of the basic knowledge graph; and determining the expected classification after each update according to the updated data of the change path in the basic knowledge graph; and the efficiency and accuracy of the data hierarchical classification processing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically a model construction method and system based on data hierarchical classification. Background Technology

[0002] In the process of digital transformation in the construction industry, the deep application of technologies such as BIM (Building Information Modeling) and the Internet of Things has driven an explosive growth in project data. How to effectively classify, store, and process data is crucial for enterprise development. Construction project data is scattered from different sources, diverse in type, and belongs to different business systems. The lack of a unified classification and coding standard results in high costs for subsequent data processing and analysis, and makes the development of integrated solutions complex and expensive, failing to meet the needs of industry development.

[0003] For example, Chinese Patent Publication No. CN118227595A discloses a data classification and storage method and apparatus based on edge empowerment. The method includes: training a preset edge node object model based on event features, attribute features, and service features in the preset object model data to obtain a preprocessed edge node object model; receiving first state information sent by a first device in the Internet of Things, and determining whether the first state information meets the triggering conditions and execution conditions in the preset first scene linkage rules according to the preprocessed edge node object model. If it meets the conditions, the corresponding first execution action information is output; determining the corresponding cold database and hot database based on the first state information, the first execution action information, and the preset cold and hot data separation rules; and obtaining the updated cold database and updated hot database according to the preset cold and hot data separation rules at a preset time frequency.

[0004] For example, Chinese Patent Publication No. CN117332133A discloses a data classification method based on expert scoring, which relates to the field of data processing technology. The method includes the following steps: Step S1: Determine the initial data classification standard for the target enterprise and establish a first data classification table; Step S2: Establish a classification reference standard group for outputting reference data classification results; Step S3: Calculate the expected total score and the expert total score of the reference standard group according to the classification weight values, and compare the expected total score with the expert total score according to a preset expected score threshold; Step S4: Update the initial data classification standard when the expert total score meets the expected score threshold range.

[0005] Existing technologies involve receiving data, analyzing corresponding cold and hot data, and updating the data using the cold and hot data format; and quantifying the current data classification based on industry similarity scores after expert evaluation. Existing technologies tend to record the status of data updates and complete industry-related classification processing, ignoring the group correlation between individual data and other data when data changes, as well as the progress of each data as the project changes, which reduces the accuracy and processing efficiency of data classified according to construction projects. Summary of the Invention

[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a model construction method based on data hierarchical classification, including: S1, collecting building project data, initially labeling the building project data according to the data type of the building project, and setting up project entities that include project status factors and project status interpretation.

[0007] S2 uses semantic analysis technology to extract state keywords from each project entity, and performs semantic clustering based on the scenario to which the state keywords belong, to construct a basic knowledge graph containing state labels.

[0008] S3, based on the generated basic knowledge graph, uses changes in project status as a trigger condition to compare and analyze the identified project entities and define the relationships between project entities under changes in project status.

[0009] S4. Based on the relationships between project entities, determine the adjacent nodes in the basic knowledge graph that are connected by relationships when the project state changes; track the change paths of adjacent nodes according to the time sequence of project state changes, and update the change paths to the edge attributes of the basic knowledge graph.

[0010] S5 determines the expected level after each update based on the updated data in the basic knowledge graph according to the change path, and determines the final output basic knowledge graph through iterative feedback.

[0011] A model building system based on data hierarchical classification includes: an entity configuration module, a graph construction module, an association analysis module, an attribute update module, and an iterative feedback module; wherein, the output of the entity configuration module is connected to the graph construction module, the output of the graph construction module is connected to the association analysis module, the output of the association analysis module is connected to the attribute update module, and the output of the attribute update module is connected to the iterative feedback module.

[0012] The entity configuration module is used to collect building project data, initially label the building project data according to the data type of the building project, and set up project entities that include project status factors and project status interpretations.

[0013] The graph construction module is used to extract state keywords from each project entity using semantic analysis technology, and to perform semantic clustering based on the scenario to which the state keywords belong, thereby constructing a basic knowledge graph containing state labels.

[0014] The association analysis module is used to compare and analyze the identified project entities based on the generated basic knowledge graph, with changes in project status as the trigger condition, and to define the association relationships between the project entities under the change in project status.

[0015] The attribute update module is used to determine the adjacent nodes in the basic knowledge graph that are connected by the relationships between the project entities based on the relationships between them when the project status changes; it tracks the change paths of the adjacent nodes according to the time sequence of the project status changes, and updates the change paths to the edge attributes of the basic knowledge graph.

[0016] The iterative feedback module is used to update the data in the basic knowledge graph according to the change path, determine the expected level after each update, and determine the final output basic knowledge graph through iterative feedback.

[0017] The beneficial effects of this invention are as follows: First, this invention classifies and grades the identified building project data through project entity integration → knowledge graph construction → association identification → graph change analysis → data grading, and addresses the problems of ambiguous association relationships and low grading accuracy in data grading and classification, thereby improving the reliability of data association and the efficiency of data grading processing.

[0018] Second, this invention initially labels data according to the type of construction project, mapping data distribution characteristics to project attributes, configuring project status factors, and semantically labeling them in the time and quality / safety dimensions to form a unified project entity containing project status factors and project status interpretations. Simultaneously, semantic analysis technology is used to extract keywords and generate keyword pairs. Frequent itemsets containing status-related keywords are selected and clustered, with status labels added. Finally, a basic knowledge graph is constructed using project entities as nodes, status keywords as attributes, and association rules as edges. This improves the efficiency of data scenario-based classification, giving the processed data a clear structural basis and enhancing data processing efficiency and hierarchical classification effectiveness.

[0019] Third, this invention uses changes in project status keywords as the main triggering condition and stages and time nodes as auxiliary triggering conditions. It employs hard constraints and soft constraints to divide directly and indirectly related entities, compares the commonalities and differences of project entities, defines the classification of relationships, and uses the corresponding relationship classification as the output relationship to complete the processing of collaborative relationships between project entities. This makes the hierarchical results present a holistic and consistent picture, ensuring the consistent effect of each project under multiple processing methods.

[0020] IV. This invention uses project entities undergoing state changes as starting nodes, extracts adjacent nodes and tracks the change paths, and uses the update time difference between adjacent nodes as a sorting index. The change paths, update time differences, and difference data are then updated to the edge attributes of the basic knowledge graph to complete data updates during the dynamic progress of the project. This ensures that the classification criteria are updated synchronously with the project status, improving the consistency between the knowledge graph and project progress. Finally, based on the relationships and state keywords in the graph, data mapping is performed to set the expected classification. For classification results that do not meet the expectations, levels are set according to the scenarios where the state labels are located and corrections are made. Iterative verification continues until all data meets the expected results, at which point iteration stops and the final graph is output. This ensures that the descriptions of each scenario are clearly defined during project updates, facilitating staff verification of project data and improving the accuracy of data classification. Attached Figure Description

[0021] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0022] Figure 1 This is a flowchart illustrating a model building method based on data hierarchical classification.

[0023] Figure 2 This is a flowchart illustrating step S1 of a model construction method based on data hierarchical classification.

[0024] Figure 3 This is a flowchart illustrating step S2 of a model building method based on data hierarchical classification.

[0025] Figure 4 This is a flowchart illustrating step S3 of a model building method based on data hierarchical classification.

[0026] Figure 5 This is a flowchart illustrating step S4 of a model building method based on data hierarchical classification.

[0027] Figure 6 This is a flowchart illustrating step S5 of a model construction method based on data hierarchical classification.

[0028] Figure 7 This is a system framework diagram for a model building system based on data hierarchical classification. Detailed Implementation

[0029] The embodiments of the present invention are described in detail below. The embodiments described below are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. Where specific techniques or conditions are not specified in the embodiments, they shall be performed in accordance with the techniques or conditions described in the literature in the art or in accordance with the product manual.

[0030] See Figure 1A model construction method based on data hierarchical classification includes: S1, collecting building project data, initially labeling the building project data according to the data type of the building project, and setting up project entities that include project status factors and project status interpretation.

[0031] S2 uses semantic analysis technology to extract state keywords from each project entity, and performs semantic clustering based on the scenario to which the state keywords belong, to construct a basic knowledge graph containing state labels.

[0032] S3, based on the generated basic knowledge graph, uses changes in project status as a trigger condition to compare and analyze the identified project entities and define the relationships between project entities under changes in project status.

[0033] S4. Based on the relationships between project entities, determine the adjacent nodes in the basic knowledge graph that are connected by relationships when the project state changes; track the change paths of adjacent nodes according to the time sequence of project state changes, and update the change paths to the edge attributes of the basic knowledge graph.

[0034] S5 determines the expected level after each update based on the updated data in the basic knowledge graph according to the change path, and determines the final output basic knowledge graph through iterative feedback.

[0035] like Figure 2 As shown, the implementation of step S1 includes: S11, viewing the data distribution characteristics of the current construction project data according to the data type to which the current construction project data belongs; the data type includes unstructured data (such as construction log PDF) and structured data (data stored in relational databases, etc.) at the time of storage.

[0036] Data distribution characteristics represent data stored in different ways. For example, structured data is concentrated in relational databases such as schedule tables and cost tables, while unstructured data is distributed in on-site safety monitoring logs and quality acceptance reports. These stored data are recorded as data distribution characteristics using semantic tags. These data distribution characteristics will serve as the semantic basis for interpreting project status and explaining the data format received in the current scenario.

[0037] S12, map the acquired data distribution characteristics to project attributes, and configure the project status factors associated with the project attributes.

[0038] Project attributes include indicators such as project ID, project type, and region. Project status factors include indicators such as schedule management factors associated with schedule data and security factors associated with security logs. Each project status factor represents an indicator of a project entity in a corresponding scenario, mapped to data such as quantified schedule delay days.

[0039] S13, semantically annotate the current project attributes in the time and quality and safety dimensions to obtain a preliminary annotated interpretation of the project status.

[0040] The current construction project data will be initially labeled according to the time dimension and the quality and safety dimension. The time dimension includes the project stage (project establishment, design, construction, acceptance, operation and maintenance, etc.), time nodes (construction schedule, actual completion time, etc.) and data update records (data generation time, update frequency, data validity period, etc.).

[0041] The quality and safety dimensions include quality status (acceptance results, compliance status, etc.), safety status (production status, risk level, etc.), and approval and acceptance status. These data will be combined in the form of digital annotation to determine the safety status of the current construction project data in the hierarchical and classified processing process.

[0042] Project status factors record indicators described in terms of time and quality and safety dimensions, enabling project entities to record data under the corresponding factors in an entity-attribute-association format. Based on the semantic annotations of the time and quality and safety dimensions, project status interpretations are generated, such as recording the corresponding project's qualification, potential hazards, delays, etc. These data are bound to the content recorded by the data distribution characteristics, forming a data combination of data-annotation-interpretation, providing a data foundation for subsequent status keyword extraction.

[0043] Based on the above dimensions, the status factors and status keywords for interpreting project status generally include: time schedule keywords and quality and safety keywords.

[0044] Time-related keywords can be described in positive or negative terms, such as positive terms like normal, on schedule, ahead of schedule, ahead of schedule, meeting standards, completed, completed, and passed acceptance, and negative terms like delayed, lagging, postponed, suspended, interrupted, and stagnant. As for quality and safety-related keywords, they can be divided into positive terms like qualified, meeting standards, no defects, zero rectification, and no accidents, and negative terms like flaws, repairs, reconstruction, hidden dangers, abnormalities, and rework.

[0045] These keywords represent the specific progress of the current project, as well as the efficiency and quality of project completion in various regions and locations. A knowledge graph is then constructed based on these states to distinguish the classification and level of building project updates, such as normal and good, warning and alert, abnormal risk and serious accident.

[0046] In one embodiment of the present invention, step S2 constructs a graph with project entities as nodes, status keywords as attributes, and association rules as relational edges, providing its data foundation for subsequent analysis.

[0047] like Figure 3As shown, the implementation of step S2 includes: S21, performing semantic parsing on the fields corresponding to the project status factors and project status interpretation, and extracting at least one set of keywords.

[0048] When extracting keywords, first identify the project entity, such as project ID + project name as the identifier of the project entity, and extract the fields corresponding to the project status factor and project status interpretation.

[0049] Specific indicators such as stage (construction), quality status (acceptance results), and material properties (concrete strength C30) are extracted from the project status factors. Descriptions such as concrete strength meeting standards during the construction stage and frame structure passing acceptance are extracted from the project status interpretation. Keywords such as project ID, construction, acceptance results, and concrete strength meeting standards are extracted from the fields corresponding to the project status factors and project status interpretation.

[0050] When extracting keywords, the extracted terms include not only the project's construction status but also words related to its construction scenarios and attributes. These terms are all used as keywords in this extraction process. The keyword extraction method employs NLP semantic analysis algorithms, such as Stanford NER (Named Entity Recognition algorithm), to extract relevant data.

[0051] S22, perform association rule identification on the extracted keywords, regard the keywords associated when the project changes as status keywords, and generate an association list.

[0052] Typically, the Apriori algorithm or FP-Growth algorithm is used to mine association rules for keywords, find frequently occurring keyword pairs, and mark the keywords associated with project changes. For example, if a design change causes a project delay, the keyword pair corresponding to the design change + project delay will be used as the output status keywords.

[0053] The keyword pairs corresponding to the status keywords are semantically clustered to generate clusters and assigned status labels, such as schedule delay cluster, cost overrun cluster, etc. Each cluster shares similar semantic features, such as project delay and schedule delay belonging to the same cluster. Semantic differences between clusters are distinguished by Euclidean distance to explain the semantic differences between the data corresponding to different status keywords. Finally, the related data under each cluster are combined into a list for output to obtain the association list.

[0054] The implementation of step S22 includes: S221, generating keyword pairs corresponding to each group of keywords in the form of keyword pairs based on the preset relationship between keywords.

[0055] The pre-defined relationships between keywords are inherent connections summarized from historical data in the construction field. These terms represent the forms in which keywords can form word pairs, such as related word combinations like design change → construction delay, concrete strength C30 → quality compliance, etc. In step S22, based on the pre-defined relationships between keywords, all possible keyword pairs combined in attribute-state, scenario-state, and other ways will be generated to prevent unrelated keyword pairs from being combined, thus avoiding the waste of resources in subsequent frequent itemset mining.

[0056] S222 performs frequent itemset mining on keyword pairs, and considers keyword pairs that satisfy minimum support and minimum confidence as the output frequent itemsets.

[0057] When mining frequent itemsets, the minimum support can be set to 20%, meaning the frequency of keyword pairs is ≥ 20% of the total number of occurrences in the corresponding scenario; the minimum confidence can be set to 70%, meaning the ratio of the frequency of occurrence of keyword pairs A and B to the total number of occurrences of keyword A is ≥ 70%.

[0058] The 20% and 70% values ​​set here are for illustrative purposes only, representing the values ​​for filtering frequent itemsets under normal circumstances. Alternatively, they can be set based on the average of the minimum support and minimum confidence scores when filtering frequent itemsets in historical data, to adjust the number of frequent itemsets to be filtered.

[0059] S223, filter out keyword pairs that contain state keywords in the frequent item set, take the keyword pair as state keywords, perform cluster analysis on each state keyword, and set state labels based on the clusters after cluster analysis.

[0060] The aforementioned status keywords represent terms used to evaluate the specific status of a project in the project status interpretation, such as terms like "meeting standards" and "delay." These terms will describe the status of each project entity in the knowledge graph, depending on the project entity they refer to.

[0061] When performing cluster analysis, the keyword pairs are converted into numerical vectors, and the K-means algorithm can be used for clustering to minimize the intra-cluster squared error and find the number of groups K. After that, the clusters are divided into multiple clusters. The semantic similarity between different keyword pairs is described by the intra-cluster Euclidean distance, and a state label is set for each cluster according to the scenario described in each cluster.

[0062] If a cluster contains "construction period delay, schedule delay, rainy season construction → postponement", it is labeled "schedule delay cluster"; if a cluster contains "rectification, non-compliance, rebar spacing not up to standard → rectification", it is labeled "quality rectification cluster".

[0063] S23. Using the project entities in the association list as nodes, the status keywords as node attributes, and the association rules corresponding to the status keywords as edges, a basic knowledge graph corresponding to each project entity is formed.

[0064] The resulting knowledge graph can quickly locate the current status of a construction project and trace its causes, generating structured and interpretable data to achieve adaptive processing of construction data classification and grading in various scenarios. For example, it can update real-time safety logs during the construction phase and adjust quality reports during the acceptance phase, ensuring that the classification model is always synchronized with the actual status of the project and avoiding decision-making biases caused by lagging classification. At the same time, the knowledge graph will progressively display the data classification of construction data at each time node and stage in the current scenario, allowing each data point to be traced back based on its node, improving the accuracy and response speed of the classification model in tracing back data.

[0065] In one embodiment of the present invention, step S3 will detect whether the state of the project entity has changed. Only when the triggering condition is met will the subsequent entity comparison and association identification proceed. The triggering condition will be based on the keyword pair corresponding to the state keyword as the main triggering condition, and the stage and time node represented by the project entity as the auxiliary triggering condition.

[0066] When the keywords corresponding to the status keywords change, it indicates that the current project has entered a status such as progress, delay, or completion, which is a change in the quality and safety dimension and needs to be updated in a timely manner. When the stages or time nodes corresponding to the auxiliary trigger conditions change, it indicates that the project has changed in the time dimension and the project status needs to be tracked in a timely manner to help identify the update status of the project entities.

[0067] like Figure 4 As shown, the implementation method of step S3 includes: S31, using the change of the status keyword corresponding to the project entity as the main triggering condition, and the stage and time node corresponding to the project entity as auxiliary triggering conditions, and configuring the triggering conditions corresponding to the current project entity.

[0068] When a project entity experiences a change in status, the primary triggering condition takes precedence, and its fulfillment triggers data updates and relationship identification; secondary triggering conditions only have the effect of triggering when there is a phase change and the time node deviation is ≥3 days.

[0069] S32, query all project entities in the basic knowledge graph that correspond to the triggering condition, obtain the project entities that are directly related to the current project entity and the project entities that are indirectly related to it, and generate a candidate list of related entities.

[0070] Directly related project entities refer to entities that share attributes with the current project entity. Direct association requires that their core attributes match completely, and other attributes are matched in a similar way. For example, if two construction projects completely match in terms of construction stage, quality acceptance results, safety risk level, region, and building type in the time and quality and safety dimensions, and have similar content such as expected completion time, effective period, and data update frequency, these two projects and their related data are considered directly related project entities. At this time, based on the specific scenario of each construction project, the main core attributes and other attributes with permissible differences will be selected from its data. These attributes can be searched based on the labeled parts of the pre-set hard and soft constraints in the database, thereby obtaining the core attributes required for the current situation.

[0071] Directly associated project entities are associated through explicit matching of attribute fields, without relying on changes in project status. This is a direct association based on relatively static attributes.

[0072] Indirectly related project entities refer to project entities that belong to the same association rule as the current project entity. For example, if two projects are delayed due to design changes, they belong to the same association rule combination and are considered indirectly related. Indirect associations are combined through association rules, indicating logically connected related entities. Indirect associations can find entities similar to the current project that belong to the same association rule. When judging the similarity of project attributes, the similarity threshold set for indirect and direct associations is the same, both using 80% as the threshold. The association between projects is explained by converting the corresponding project entities into vectors and using cosine similarity to calculate the similarity.

[0073] In other words, direct association refers to a complete match of hard constraints and a similarity of soft constraints ≥80%, while indirect association refers to the same association rule and a similarity of soft constraints ≥80%, with the state changing due to transmission / parallel changes.

[0074] These associations tend to trigger a change in the project's state, leading to a reconstruction of the data in the original knowledge graph and an evolution of the data portion after the state change.

[0075] To determine the directly and indirectly related project entities, step S32 further includes: determining at least one constraint in a project entity, dividing the project entities according to hard and soft constraints, and determining the project entities directly and indirectly related to the current project entity based on the hard and soft constraints corresponding to each project entity.

[0076] Hard constraints represent key attributes that affect the classification and grading, and are the parts that must be completely consistent when directly related, such as project type, region, and core quality and safety indicators (such as concrete strength standards). Soft constraints represent the parts that allow for differences, and are the parts that determine similarity between directly and indirectly related project entities, such as data update frequency and estimated completion time.

[0077] In directly related project entities, hard constraints represent attribute descriptions configured in the database, while soft constraints represent other data components. In indirectly related project entities, hard constraints represent project entities belonging to the same association rule, while soft constraints represent the data components corresponding to each project entity.

[0078] S33, compare the project entities whose status has changed with the project entities included in the candidate list of related entities, and extract the common features and differences after comparison.

[0079] Since the project entities in the indirect and direct association settings are obtained by using similarity, when comparing common features and differences, the parts with similarity ≥80% will be set as common features, and the parts with similarity <80% will be set as differences. The extracted data dimensions will cover attribute dimensions (such as region, building type), state dimensions (such as original state, new state), and rule dimensions (such as association rule tags).

[0080] S34. Based on the common and difference features after comparison, the association relationship classification of each project entity is set according to the scenario and association rules of the corresponding project entity.

[0081] The association classification will generate associations such as those based on common features (same scene, same rule, positive synchronous association), associations based on differences (same scene, same rule, reverse difference association), associations using both common and difference features (cross scene, same rule, state collaborative association), and the remaining cross scene, cross rule, no obvious association, thus obtaining the associations under state changes.

[0082] For example, the same scenario-same rule-positive synchronous association can be represented as: scenario similarity ≥ 80% + consistent association rules + common features accounting for ≥ 70% of the total features of common and different features. That is, the current project entity describes a scenario with a partial similarity of 80%, the project entities described belong to the same association rule, and there are many common features obtained from these project entities.

[0083] The same scenario-same rule-reverse difference association and its positive synchronous association differ only in the ratio of common features. It mainly uses the proportion of difference features ≥70% to identify the association between project entities in scenarios with many difference features.

[0084] Cross-scene-same-rule-state collaborative association is represented by scene similarity of 50-79% + consistent association rules + common feature ratio ≥60% to obtain project entities under state collaboration. When the common feature ratio is small or the scene similarity is low, it means that the two project entities are too different and there is no obvious relationship. If classification is performed, it will cause data recognition error.

[0085] The association classification constructed at this time will serve as the output association, recording the indirect and direct associations between nodes in the basic knowledge graph. This data will also serve as the basis for subsequent expected classification to complete the data classification of the current basic knowledge graph.

[0086] In one embodiment of the present invention, in step S4, nodes in the basic knowledge graph are connected according to the association relationship of each project entity, and the data configured in the basic knowledge graph is updated through the relationship edges after the nodes are connected.

[0087] like Figure 5 As shown, the implementation of step S4 includes: S41, taking the project entity that triggers the state change as the starting node, extracting the adjacent nodes connected to it through relational edges from the basic knowledge graph based on the association relationship of each project entity; the adjacent nodes are project entities that have a direct or indirect association with the project entity whose current state has changed.

[0088] S42, collect the state change times of the starting node and adjacent nodes to form a change path from the starting node to the adjacent nodes.

[0089] When generating a change path, the timestamp of the state change of the starting node is first collected. Then, the data of the adjacent nodes when the state changes is queried. If the state of the adjacent node changes after the state of the starting node changes, it is recorded as a propagated change. A propagated change indicates that there is a direct relationship between the starting node and the adjacent node. If the adjacent node changes independently, it is recorded as a parallel change. A parallel change indicates that there is an indirect relationship between the starting node and the adjacent node. The changes of the two are independent of each other but belong to the same association rule.

[0090] The identified transmission and parallel change paths are then sorted in ascending order of timestamps to output a set of path sequence data. In this data, the state change time represents the timestamp of the change, and the change paths are sorted in chronological order.

[0091] For example, there is a starting node P001, and adjacent nodes P002, P003, and P004; among which P001 is directly related to P003 and P004, and indirectly related to P002. P002 is the earliest, followed by P003, and then P004. In this case, the output method of sorting P001 and its adjacent nodes can be P001→P002, P001→P003, P001→P004. Each change path is labeled in this form to obtain the data of each edge attribute mapping.

[0092] S43 updates the relationship edges between project entities based on the update time difference between adjacent nodes in the change path, thus completing the edge attribute update of each relationship edge in the basic knowledge graph.

[0093] When updating edge attributes, the start time, end time, and duration of change for each change path will be recorded. These attributes will be mapped to the corresponding relationship edges according to the previously identified relationships, the propagation of changes in path labels, and parallel changes, in order to explain the data change process of each relationship edge as the project progresses.

[0094] Preferably, when updating edge attributes, an incremental update method can be used, with the difference before and after each update as the data to be updated, in order to determine the main changes during the project's progress.

[0095] Therefore, the implementation of step S43 also includes: retrieving the update time difference of adjacent nodes in the change path, using the update time difference as the sorting index, calculating the difference data at each update, and updating the difference data to the corresponding relation edge. The difference data represents the incremental content updated each time the change path changes.

[0096] It should be noted that the update time difference between adjacent nodes represents the order in which two nodes under a connection update. The smaller the update time difference, the faster the change is transmitted and the more likely it is to be processed first.

[0097] In one embodiment of the present invention, step S5 involves extracting the relationships and state keywords of each node in the basic knowledge graph, mapping these data to the desired hierarchy, and verifying the updated data of each node to obtain the basic knowledge graph output after iteration.

[0098] like Figure 6 As shown, the implementation of step S5 includes: S51, based on the association relationship and state keywords of each node in the basic knowledge graph during data update, data mapping is performed on each node, and expected hierarchical levels are set.

[0099] During data mapping, assuming the status keywords are positive statuses in the time dimension, such as "on-time completion" status keywords, and the association relationship is "same scenario - same rule - positive synchronous association," then the status keywords can be mapped to public level or other levels of description using the association relationship. These quantified expected levels will be recorded in the database. Based on the combination of status keywords and association relationships, the expected level is set for each node when updating data. Each node will use its own status label as the core and the association relationship of all connected nodes as an auxiliary factor to look up the corresponding expected level from the database to complete the configuration of each node.

[0100] At this point, preliminary expectation classification needs to be completed. This requires quickly and accurately matching the expectation classification to each node. At this time, we can rely on the status keywords that can directly reflect the description on the node, and use the status keywords as the basis for expectation classification. The status labels are the identifiers in the corresponding scenarios after aggregation, and are general descriptions. When there are conflicts during expectation classification iteration, the scenario where the status label is located will be selected to set its level, so as to explain the expectation classification configured for the current node and to define the data classification situation in different scenarios.

[0101] S52 iterates through each node in the basic knowledge graph. If the expected hierarchical structure does not meet the expected results, the hierarchy is set according to the scenario where the state label is located, and the data where the state label is located is labeled and added to the basic knowledge graph.

[0102] In step S52, an iterative method will be used for verification. If there are no conflicts or abnormal parts in the verified data, it is considered to meet the expected results; otherwise, it is considered to not meet the expected results. The verification content includes, but is not limited to, situations such as inconsistent node levels with the same relationship, and mismatch between the relationship and status label and the expected level.

[0103] In one scenario, the expected result can be that the consistency of the hierarchical classification of related nodes is ≥95% and the matching degree between the status label and the classification is ≥90%. The hierarchical consistency represents the proportion of nodes with the same classification result to the total number of nodes in the group. The standard for setting hierarchical consistency is used to determine that there are certain similarities among the nodes in the group when they are expected to be classified. For example, if the preset expected classification of the nodes in the group is the normal and good level, then after checking the classification one by one, the hierarchical consistency means that the vast majority of the whole is in this level, which can help identify data updates and abnormal situations.

[0104] As for the hierarchical matching degree, it represents the degree of conformity between the status label and the preset expected hierarchy, that is, the proportion of the number of nodes whose status labels match the expected hierarchy to the total number of nodes in the graph. This value is used to check the specific situation of each part of the current project, to ensure that the status labels containing the scenario description correspond to the hierarchical logic, so that the final graph hierarchy can accurately correspond to the scenario implemented in the current project.

[0105] When the node hierarchy of the same relationship is inconsistent, it means that the status label of each node describes an independent scenario. At this time, the nodes under the same relationship are mostly indirectly connected. It is necessary to set the hierarchy through the scenario in which the status label is located. The hierarchy can be set to multiple scenarios such as security risk scenario, quality control scenario, progress management scenario, and general information scenario. The hierarchy under the scenario is used as the basis for selecting the expected hierarchy. The expected hierarchy mapped by the hierarchy is selected for output. For nodes that do not meet the expectations, scenario hierarchy label and reason label for not meeting the expectations are added.

[0106] For example, if the status label of node C is "unqualified" and it falls under the category of a security risk scenario, then the scenario level and related data corresponding to the status label will be labeled to complete the iterative processing of the basic knowledge graph.

[0107] When the association and status labels do not match the expected classification, it means that the expected classification for the corresponding data does not exist in the database. This may be due to inaccurate extraction of status labels or associations, resulting in a missed judgment of the current scenario. In this case, it is necessary to set the expected classification according to the scenario and explain the specific situation of each node with the description of the scenario.

[0108] S53: When all the data corresponding to the expected level meet the expected results, stop iterating on the basic knowledge graph and use the iterated data as the final output basic knowledge graph.

[0109] The final output knowledge graph will respond to changes in project status by synchronizing the expected hierarchy, relationships, and edges of each node through change path tracking and incremental updates, thus keeping it consistent with the actual project status and improving the display effect and processing efficiency of hierarchical data distribution.

[0110] like Figure 7 As shown, the present invention also provides a model building system based on data hierarchical classification, including: an entity configuration module, a graph construction module, an association analysis module, an attribute update module, and an iterative feedback module; wherein, the output end of the entity configuration module is connected to the graph construction module, the output end of the graph construction module is connected to the association analysis module, the output end of the association analysis module is connected to the attribute update module, and the output end of the attribute update module is connected to the iterative feedback module.

[0111] The entity configuration module is used to collect building project data, initially label the building project data according to the data type of the building project, and set up project entities that include project status factors and project status interpretations.

[0112] The graph construction module is used to extract state keywords from each project entity using semantic analysis technology, and to perform semantic clustering based on the scenario to which the state keywords belong, thereby constructing a basic knowledge graph containing state labels.

[0113] The association analysis module is used to compare and analyze the identified project entities based on the generated basic knowledge graph, with changes in project status as the trigger condition, and to define the association relationships between the project entities under the change in project status.

[0114] The attribute update module is used to determine the adjacent nodes in the basic knowledge graph that are connected by the relationships between the project entities based on the relationships between them when the project status changes; it tracks the change paths of the adjacent nodes according to the time sequence of the project status changes, and updates the change paths to the edge attributes of the basic knowledge graph.

[0115] The iterative feedback module is used to update the data in the basic knowledge graph according to the change path, determine the expected level after each update, and determine the final output basic knowledge graph through iterative feedback.

[0116] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention, which are still covered within the protection scope of the present invention.

Claims

1. A model construction method based on data hierarchical classification, characterized in that, include: S1, collect construction project data, initially label the construction project data according to the data type of the construction project, and set up project entities that include project status factors and project status interpretation; S2 utilizes semantic analysis technology to extract state keywords from each project entity, and performs semantic clustering based on the scenario to which the state keywords belong, constructing a basic knowledge graph containing state labels; Step S2 can be implemented in the following ways: S21, perform semantic parsing on the fields corresponding to the project status factors and project status interpretation, and extract at least one set of keywords; S22, perform association rule identification on the extracted keywords, regard the keywords associated when the project changes as status keywords, and generate an association list; S23. Using the project entities in the association list as nodes, the status keywords as node attributes, and the association rules corresponding to the status keywords as edges, a basic knowledge graph corresponding to each project entity is formed. The implementation methods of step S22 include: S221, Based on the preset relationship between keywords, generate keyword pairs corresponding to each group of keywords in the form of keyword pairs; S222, perform frequent itemset mining on keyword pairs, and consider keyword pairs that satisfy minimum support and minimum confidence as the output frequent itemsets; S223, filter out keyword pairs that contain state keywords in the frequent item set, take the keyword pair as state keywords, perform cluster analysis on each state keyword, and set state labels based on the clusters after cluster analysis; S3, based on the generated basic knowledge graph, takes the change of project status as the trigger condition, compares and analyzes the identified project entities, and defines the relationship between the project entities under the change of project status; S4. Based on the relationships between project entities, determine the adjacent nodes in the basic knowledge graph that are connected by relationships when the project status changes; track the change paths of adjacent nodes according to the time sequence of project status changes, and update the change paths to the edge attributes of the basic knowledge graph. S5: Based on the updated data in the basic knowledge graph according to the change path, determine the expected level after each update, and determine the final output basic knowledge graph through iterative feedback. Step S5 can be implemented in the following ways: S51, based on the relationship and status keywords of each node in the basic knowledge graph during data updates, perform data mapping on each node and set expected levels; S52 iterates through each node in the basic knowledge graph. If the expected hierarchical structure does not meet the expected results, the hierarchy is set according to the scenario where the state label is located, and the state label is added to the basic knowledge graph after data annotation. S53: When all the data corresponding to the expected level meet the expected results, stop iterating on the basic knowledge graph and use the iterated data as the final output basic knowledge graph.

2. The model construction method based on data hierarchical classification according to claim 1, characterized in that, The implementation methods for step S1 include: S11, based on the data type to which the current building project data belongs, view the data distribution characteristics of the current building project data; S12, map the acquired data distribution characteristics to project attributes, and configure the project status factors associated with the project attributes; S13, semantically annotate the current project attributes in the time and quality and safety dimensions to obtain a preliminary annotated interpretation of the project status.

3. The model construction method based on data hierarchical classification according to claim 1, characterized in that, Step S3 can be implemented in the following ways: S31, use the change of the status keyword corresponding to the project entity as the main triggering condition, and the stage and time node corresponding to the project entity as auxiliary triggering conditions, and configure the triggering conditions corresponding to the current project entity. S32, query all project entities in the basic knowledge graph that correspond to the triggering condition, obtain the project entities that are directly related to the current project entity and the project entities that are indirectly related to the current project entity, and generate a candidate list of related entities; S33, compare the project entities whose status has changed with the project entities included in the candidate list of related entities, and extract the common features and differences after comparison in turn. S34. Based on the common and difference features after comparison, the association relationship classification of each project entity is set according to the scenario and association rules of the corresponding project entity.

4. The model construction method based on data hierarchical classification according to claim 3, characterized in that, The implementation of step S32 also includes: Identify at least one constraint in a project entity, divide the project entity into hard and soft constraints, and determine the project entities directly and indirectly related to the current project entity based on the hard and soft constraints corresponding to each project entity.

5. The model construction method based on data hierarchical classification according to claim 1, characterized in that, Step S4 can be implemented in the following ways: S41, taking the project entity that triggers the state change as the starting node, and extracting the adjacent nodes connected to it through relational edges from the basic knowledge graph based on the association relationship of each project entity; S42, collect the state change times of the starting node and adjacent nodes to form a change path from the starting node to the adjacent nodes; S43 updates the relationship edges between project entities based on the update time difference between adjacent nodes in the change path, thus completing the edge attribute update of each relationship edge in the basic knowledge graph.

6. The model construction method based on data hierarchical classification according to claim 5, characterized in that, The implementation of step S43 also includes: Retrieve the update time difference of adjacent nodes in the change path, use the update time difference as the sorting index, calculate the difference data at each update, and update the difference data to the corresponding relation edge.

7. A model building system based on data hierarchical classification, characterized in that, include: The entity configuration module is used to collect building project data, initially label the building project data according to the data type of the building project, and set up project entities that include project status factors and project status interpretation. The knowledge graph construction module is used to extract state keywords from each project entity using semantic analysis technology, and to perform semantic clustering based on the context to which the state keywords belong, thereby constructing a basic knowledge graph containing state labels. Specific implementation methods include: Perform semantic parsing on the fields corresponding to project status factors and project status interpretation, and extract at least one set of keywords; The extracted keywords are used to identify association rules, and keywords associated with changes in the project are regarded as status keywords, and an association list is generated. Using the project entities in the association list as nodes, the status keywords as node attributes, and the association rules corresponding to the status keywords as edges, a basic knowledge graph corresponding to each project entity is formed. The methods for generating the association list include: Based on the preset relationships between keywords, keyword pairs are generated for each group of keywords. Frequent itemset mining is performed on keyword pairs, and keyword pairs that satisfy the minimum support and minimum confidence are considered as the output frequent itemsets; Select keyword pairs that contain state-related keywords from the frequent item set, use these keyword pairs as state keywords, perform cluster analysis on each state keyword, and set state labels based on the clusters obtained from the cluster analysis. The association analysis module is used to compare and analyze the identified project entities based on the generated basic knowledge graph, with the change of project status as the trigger condition, and to define the association relationship of each project entity under the change of project status. The attribute update module is used to determine the adjacent nodes in the basic knowledge graph that are connected by the relationships between the project entities based on the relationships between them when the project status changes; it tracks the change path of the adjacent nodes according to the time sequence of the project status changes and updates the change path to the edge attributes of the basic knowledge graph. The iterative feedback module is used to update data in the basic knowledge graph according to the change path, determine the expected level after each update, and determine the final output basic knowledge graph through iterative feedback. Specific implementation methods include: Based on the relationships and status keywords of each node in the basic knowledge graph during data updates, data mapping is performed on each node, and expected hierarchies are set. Iterate through each node in the basic knowledge graph. If the desired hierarchy does not meet the expected results, set the hierarchy based on the scenario where the state label is located, and add the state label to the basic knowledge graph after data annotation. When all data corresponding to the expected level meet the expected results, stop iterating on the basic knowledge graph and use the iterated data as the final output basic knowledge graph.