A building regulation-based vertical field large model training method and system

By constructing a dynamic legal knowledge base and a structured knowledge graph, combined with a constraint relationship matrix and an incremental knowledge pool, the problem of insufficient understanding and adaptability of large models in building regulations is solved. This enables deep semantic analysis and dynamic data management of building regulations, improving the compliance and adaptability of the model.

CN120633797BActive Publication Date: 2026-03-24浙江蓝宸数联科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing large models are inadequate in understanding and adapting to building regulations, making it difficult to guarantee design compliance. Furthermore, the lack of a dynamic matching mechanism in data management leads to low model training efficiency.

Method used

By constructing a dynamic legal knowledge base and structured knowledge graph, and combining a constraint relationship matrix and an incremental knowledge pool, a deep understanding of building regulations and dynamic data management can be achieved through semantic parsing and multi-dimensional parameter association.

Benefits of technology

It achieves deep semantic analysis and multi-dimensional correlation modeling of building regulations, can quantitatively assess the risk of conflict between regional regulations, has continuous evolution capabilities, and improves the compliance and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633797B_ABST
    Figure CN120633797B_ABST
Patent Text Reader

Abstract

The application discloses a kind of vertical field large model training method and system based on building regulations, it is related to large model training technical field, constructs dynamic regulation knowledge base including clause semantic atlas, constraint relationship matrix, and generates structured knowledge atlas based on clause semantic atlas and constraint relationship matrix;While constructing multiple type databases to store building data, setting up matching unit and incremental knowledge pool to connect multiple type databases;Training data is provided through structured knowledge atlas and incremental knowledge pool, large model is trained based on training data, and building vertical field large model is generated.The application significantly improves the intelligent capacity of model by constructing structured knowledge atlas and incremental knowledge pool, and using continuous feedback mechanism, effectively cope with the complexity of building regulations, realize the high compliance and practicality of building vertical field large model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of large model training, in particular to a large model training method and system based on building regulations in the vertical field. BACKGROUND

[0002] In the field of architecture, strictly following various building regulations and standards is the basis for ensuring building safety, functionality and sustainability; with the rapid development of artificial intelligence technology, especially large language models, people have begun to explore how to use these powerful tools to assist or even automate building design, approval and compliance checks;

[0003] Although existing general large models perform well in text understanding and generation, they lack a deep understanding of the complex logic behind building regulations, the constraints between entities and the differences in specific regions. Simply inputting the regulation text into the model cannot guarantee that the output design or recommendations strictly comply with the regulations, and there is a serious compliance risk. At the same time, existing systems generally use offline batch data management mechanisms, and when new regulations are released or local standards are revised, manual re-labeling of training data and full-scale reconstruction of the model are required, making it difficult to adapt to the update frequency of building regulations. In addition, the existing technical system has a data island problem, and professional databases such as regional specification libraries, acceptance case libraries and construction method libraries in the construction industry are often stored separately, using heterogeneous data structures and lacking unified search logic. Although the mainstream solution attempts to physically aggregate data through a data warehouse, it lacks a dynamic matching mechanism, making it difficult to effectively integrate multi-source data during model training.

[0004] Therefore, it is of great significance to construct a large model training method based on building regulations in the vertical field. SUMMARY

[0005] The purpose of the present application is to provide a large model training method and system based on building regulations in the vertical field to solve the problems in the background art.

[0006] In order to achieve the above-mentioned purpose, the present application provides the following technical solution: a large model training method based on building regulations in the vertical field, comprising:

[0007] A dynamic regulation knowledge base is constructed, including a clause semantic graph and a constraint relationship matrix;

[0008] A structured knowledge graph is generated based on the clause semantic graph and the constraint relationship matrix;

[0009] A plurality of type databases are constructed to store building data, and a matching unit and an incremental knowledge pool are connected to the plurality of type databases;

[0010] Training data is provided through the structured knowledge graph and the incremental knowledge pool, a large model is trained based on the training data, and a large model in the building vertical field is generated.

[0011] In a preferred embodiment, the step of constructing the dynamic regulatory knowledge base includes clause semantic graph and constraint relationship matrix is:

[0012] Obtaining the original regulatory text for semantic analysis to obtain the clause semantic graph includes building regulation entities, core semantic structure and semantic relationship;

[0013] Based on the core semantic structure and semantic relationship in the clause semantic graph, a constraint relationship matrix is constructed, which is composed of multi-dimensional parameters including constraint strength, spatial influence, time decay, regional adaptation and conflict risk.

[0014] In a preferred embodiment, the step of generating a structured knowledge graph based on the clause semantic graph and the constraint relationship matrix is:

[0015] Associating building regulation entities with multi-dimensional parameters in the constraint relationship matrix to generate enhanced knowledge nodes;

[0016] Convert multi-dimensional parameters to directed edges through matrix mapping algorithm;

[0017] Generate a structured knowledge graph by connecting enhanced knowledge nodes through directed edges.

[0018] In a preferred embodiment, the step of constructing multiple type databases to store building data, setting a matching unit and an incremental knowledge pool to connect multiple type databases is:

[0019] Construct multiple type databases to store building data, including regional specification database, building case database, material parameter database, construction method database and acceptance problem database, and define the retrieval tags of each type database;

[0020] Set a dynamic window on the type database, which is used to receive data and generate a dynamic transmission channel;

[0021] Construct an incremental knowledge pool, which contains constraint relationship matrix, for reorganizing building data to generate incremental training data;

[0022] Construct a matching unit between multiple type databases and incremental knowledge pool, which is used to store retrieval tags and perform data matching;

[0023] Multiple type databases, matching unit and incremental knowledge pool are connected through dynamic transmission channel.

[0024] In a preferred embodiment, the step of providing training data through structured knowledge graph and incremental knowledge pool, training large model based on training data, and generating building vertical field large model is:

[0025] The structured knowledge graph and the incremental knowledge pool are connected to the large model through a data transmission channel, and the large model includes a feature extractor and a dynamic learning engine;

[0026] The structured knowledge graph generates structured training data, which is transmitted to the large model through a data transmission channel, and the large model is trained based on the dynamic learning engine;

[0027] The feature extractor extracts training features in the large model training process, including attention hotspots, regional omissions, and prediction biases, and transmits the training features to the incremental knowledge pool based on the data transmission channel;

[0028] After receiving the training features, the incremental knowledge pool activates multiple type databases to generate incremental training data;

[0029] The incremental training data is transmitted to the large model, and the dynamic learning engine is trained to generate a large model in the building vertical field.

[0030] In a preferred embodiment, after the incremental knowledge pool receives the training features, the step of activating multiple type databases to generate incremental training data is:

[0031] After receiving the training features, the incremental knowledge pool transmits them to the matching unit through a dynamic transmission channel;

[0032] The matching unit matches the search labels and the training features based on a semantic matching algorithm to generate a call data table, and calls the building data based on the call data table and transmits it to the incremental knowledge pool;

[0033] The incremental data pool receives the building data, reorganizes the building data based on a constraint relationship matrix to generate incremental training data, and transmits it to the large model through a data transmission channel.

[0034] In a preferred embodiment, the matching unit matches the search labels and the training features based on a semantic matching algorithm to generate a call data table, and transmits the building data to the incremental knowledge pool based on the call data table.

[0035] A preset matching threshold is set, the matching unit receives the training features, and a semantic matching algorithm is used to calculate the matching degree of the training features and the search labels;

[0036] The search labels with a matching degree exceeding the matching threshold are sorted according to the size of the matching degree to generate a call data table, and the call data table is transmitted to a dynamic window;

[0037] The dynamic window calls the building data in the type database based on the call data table and transmits it to the incremental knowledge pool.

[0038] The present application also provides a large model training system in the vertical field based on building regulations, comprising:

[0039] Regulation segmentation module: constructing a dynamic regulation knowledge base including clause semantic atlas and constraint relationship matrix;

[0040] Knowledge graph generation module: connected with the regulation segmentation module, generating a structured knowledge graph based on the clause semantic atlas and the constraint relationship matrix;

[0041] Incremental knowledge pool module: connected with the knowledge graph generation module, constructing multiple type databases to store building data, setting a matching unit and connecting multiple type databases with the incremental knowledge pool;

[0042] Large model training module: connected with the incremental knowledge pool module, providing training data through the structured knowledge graph and the incremental knowledge pool, training a large model based on the training data, and generating a building vertical field large model.

[0043] In the above technical solution, the technical effects and advantages provided by the present application are as follows:

[0044] 1. By constructing a dynamic regulation knowledge base and a structured knowledge graph, the present application realizes deep semantic analysis and multi-dimensional association modeling of building regulations. Traditional methods mostly use static text matching or simple rule bases, which are difficult to handle complex constraint relationships and temporal and spatial dynamic characteristics between regulation clauses. By introducing a constraint relationship matrix, the clause semantics are decomposed into five dynamic parameters such as constraint strength and spatial influence, and are converted into weighted directed edges using a matrix mapping algorithm to construct a three-dimensional knowledge graph with temporal and spatial attributes. The three-dimensional knowledge graph enables the model to quantitatively evaluate the conflict risk between different regional specifications, actively identify outdated clauses that decay over time, and dynamically adjust the priority weight of clause application. In specific case reasoning, the system can realize intelligent compliance review across clauses through multi-dimensional parameter linkage of enhanced knowledge nodes.

[0045] 2. The present application breaks through the bottleneck of traditional training data staticization by constructing a double-layer data supply system of "incremental knowledge pool + dynamic matching unit". By setting up five types of professional databases such as regional specification library and acceptance problem library, and designing a dynamic window mechanism to access the latest cases in real time, the system forms a data ecosystem covering the whole life cycle of buildings. More importantly, the matching unit establishes an intelligent mapping between training features and database tags through semantic matching algorithm. When detecting that the attention hotspot deviates or the regional characteristics are missing, the data reorganization mechanism of the incremental knowledge pool can be automatically triggered. This closed-loop feedback mechanism enables the model to have continuous evolution ability. BRIEF DESCRIPTION OF DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description only represent some embodiments described in the present application, and other drawings can also be obtained by those skilled in the art based on these drawings.

[0047] Figure 1 The method flowchart of the present application.

[0048] Figure 2 The system block diagram of the present application. DETAILED DESCRIPTION

[0049] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0050] Embodiment 1, please refer to Figure 1 The vertical field large model training method based on building regulations described in this embodiment includes:

[0051] S1, constructing a dynamic regulation knowledge base including clause semantic map and constraint relationship matrix;

[0052] S2, generating a structured knowledge graph based on the clause semantic map and the constraint relationship matrix;

[0053] S3, constructing a plurality of type databases to store building data, and setting a matching unit and an incremental knowledge pool connected to the plurality of type databases;

[0054] S4, providing training data through the structured knowledge graph and the incremental knowledge pool, training a large model based on the training data, and generating a building vertical field large model;

[0055] As described in steps S1-S4 above, in the field of architecture, strictly following various building regulations and standards is the basis for ensuring building safety, functionality and sustainability; with the rapid development of artificial intelligence technology, especially large language models, people have begun to explore how to use these powerful tools to assist and even automate building design, approval and compliance checks;

[0056] Although the existing general large models perform well in text understanding and generation, they lack a deep understanding of the complex logic behind building regulations, the constraints between entities, and the differences in specific regions. Simply inputting the regulation text into the model cannot guarantee that the output design or recommendations strictly comply with the regulations, and there is a serious compliance risk. Meanwhile, the existing systems generally use an offline batch update data management mechanism. When new regulations are released or local standards are revised, manual re-labeling of training data and full reconstruction of the model are required, making it difficult to adapt to the update frequency of building regulations. In addition, the existing technical system has a data island problem. Professional databases specific to the construction industry, such as regional specification libraries, acceptance case libraries, and construction method libraries, are often stored separately, use heterogeneous data structures, and lack unified search logic. Although the mainstream solution attempts to physically aggregate data through a data warehouse, the lack of a dynamic matching mechanism makes it difficult to effectively integrate multi-source data during model training.

[0057] The present application realizes deep semantic analysis and multi-dimensional association modeling of building regulations by constructing a dynamic regulation knowledge base and a structured knowledge graph. Traditional methods mostly use static text matching or simple rule libraries, which are difficult to handle complex constraint relationships and temporal and spatial dynamic characteristics between regulation clauses. By introducing a constraint relationship matrix, the clause semantics are decomposed into five dynamic parameters such as constraint strength and spatial influence, and are converted into weighted directed edges using a matrix mapping algorithm to construct a three-dimensional knowledge graph with temporal and spatial attributes. The three-dimensional knowledge graph enables the model to quantitatively evaluate the conflict risk between different regional specifications, actively identify outdated clauses that decay over time, and dynamically adjust the priority weight of clause application. During specific case reasoning, the system can realize intelligent compliance review across clauses through multi-dimensional parameter linkage of enhanced knowledge nodes.

[0058] By constructing a dual-layer data supply system of "incremental knowledge pool + dynamic matching unit", the bottleneck of traditional static training data is broken. By setting up five types of professional databases such as regional specification library and acceptance problem library, and designing a dynamic window mechanism to access the latest cases in real time, the system forms a data ecosystem covering the entire life cycle of construction. More importantly, the matching unit establishes an intelligent mapping between training features and database labels through a semantic matching algorithm. When attention hotspots deviate or regional characteristics are missing, the data reorganization mechanism of the incremental knowledge pool can be automatically triggered. This closed-loop feedback mechanism enables the model to have continuous evolution capability.

[0059] In one embodiment, the step S1 of constructing a dynamic regulation knowledge base includes clause semantic graph and constraint relationship matrix, comprising:

[0060] S11, obtaining the original regulation text for semantic analysis to obtain a clause semantic graph including building regulation entities, core semantic structures, and semantic relationships;

[0061] S12, constructing a constraint relationship matrix based on the core semantic structure and semantic relationship in the clause semantic graph, the constraint relationship matrix being composed of multiple dimensions including constraint strength, spatial influence, time decay, regional adaptation, and conflict risk;

[0062] As described in steps S11-S12 above, first, the core semantic structure such as the "subject-act-condition" triple in the clause semantic graph is parameterized modeled by the semantic analysis engine, and for the predicate structure of each clause (such as "high-rise buildings must be provided with a ring-shaped fire lane"), the act verb strength coefficient is extracted, for example, "must" is mapped to the constraint strength parameter 1.0, and "recommended clause" is 0.6, and the spatial influence parameter is associated through the geographic encoder, and the building type is matched with the regional characteristics based on GIS data, for example, the corrosion clause in the coastal area is 0.9; the time decay parameter adopts a hyperbolic decay function wherein is the clause validity coefficient, which is generally taken as 0.3; the conflict risk parameter calculates the path conflict probability between clause nodes through a graph neural network, for example, the correlation edge weight of the fire separation distance and the plot ratio clause; the regional adaptation parameter is dynamically generated in combination with the semantic similarity of the regional specification node in the knowledge graph; finally, a five-dimensional tensor matrix is constructed, and a multi-dimensional feature fusion is performed using an attention mechanism to realize the quantitative representation and dynamic update of the clause constraint relationship.

[0063] In one embodiment, the step S2 of generating a structured knowledge graph based on the clause semantic graph and the constraint relationship matrix comprises:

[0064] S21, associating the building regulation entity with the multi-dimensional parameters in the constraint relationship matrix to generate an enhanced knowledge node;

[0065] S22, converting the multi-dimensional parameters into directed edges through a matrix mapping algorithm;

[0066] S23, generating a structured knowledge graph by connecting the enhanced knowledge nodes through directed edges;

[0067] As described in steps S21-S23 above, in generating the structured knowledge graph, first, the building regulation entity and the constraint relationship matrix are associated with multi-dimensional parameters through the enhanced knowledge node construction engine. Each entity corresponds to a unique URI in the clause semantic graph, which is dynamically bound with the constraint relationship matrix. The attribute fusion algorithm is used to encode the five-dimensional parameters into the extended attribute vector of the node. The constraint strength is normalized by the Sigmoid function to the weight value in the [0, 1] interval. The space influence is bound with the GIS coordinate to generate the GeoHash space code. The time decay is converted into the time stamp. At the same time, the conflict risk is calculated by the graph attention network to obtain the edge weight abnormal value of the entity and other associated entities. The regional adaptation is dynamically linked to the local standard vector in the regional specification library. In the edge generation stage, the matrix parameters are converted into directed edges by the multi-dimensional tensor decomposition algorithm. For any two entities with semantic relationship, the five-dimensional parameters in the constraint relationship matrix are extracted to form a feature tensor, which is processed by the tensor decomposition technology. Then, the parameterized mapping mechanism is used to convert the processed core parameters into the directionality, weight value and dynamic attribute of the edge. Finally, the structured knowledge graph is generated by connecting the enhanced knowledge nodes through the directed edges.

[0068] In one embodiment, the step S3 of setting the matching unit and the incremental knowledge pool connected to the plurality of type databases for storing building data comprises:

[0069] S31, constructing a plurality of type databases for storing building data, wherein the type databases include a regional specification library, a building case library, a material parameter library, a construction method library, and an acceptance problem library, and the retrieval tags of the type databases are defined;

[0070] S32, setting a dynamic window on the type databases for receiving data and generating a dynamic transmission channel;

[0071] S33, constructing an incremental knowledge pool, wherein the incremental knowledge pool contains a constraint relationship matrix for reorganizing building data to generate incremental training data;

[0072] S34, constructing a matching unit between the plurality of type databases and the incremental knowledge pool, wherein the matching unit is used for storing retrieval tags and performing data matching;

[0073] The plurality of type databases, the matching unit, and the incremental knowledge pool are connected through the dynamic transmission channel;

[0074] As described in steps S31-S34 above, a multi-source heterogeneous building data storage system is established, and a mechanism for dynamic data acquisition and incremental knowledge generation is designed, providing rich and on-demand reorganized training data for large model training. The core is to build multiple specialized type databases: regional specification database stores supplementary or differentiated regulations in specific regions; building case database contains detailed information of a large number of completed projects, design drawings, approval records and compliance analysis; material parameter database collects physical, chemical and performance parameters of various building materials and their application scope; construction method database records common construction methods, construction process and technical requirements of different building components; and acceptance problem database collects common compliance problems, non-conformities and their cause analysis in engineering acceptance. Each database is pre-defined with detailed retrieval tags. In order to realize dynamic flow and on-demand access of data, dynamic windows are set on each type database. These dynamic windows are not only data interfaces, but also can flexibly open data flow according to external requests (such as calls from matching units) or internal strategies (such as regular updates), and are responsible for generating dynamic transmission channels connecting the database and external components (matching units, incremental knowledge pool). Further, the type database can be dynamically expanded to store updated building regulations, so that dynamic updated building regulations can be introduced in the large model training process without the need to repeatedly train the model after the building regulations are updated. For the incremental knowledge pool, it is an intelligent data processing center that reorganizes the original building data from the type database using the built-in constraint relationship matrix. For example, when receiving parameter data of a certain material and specification requirements of a certain region, the incremental knowledge pool will filter and combine according to the constraint relationship between material parameters and regional specifications, derive new building data, and generate incremental training data highly related to specific regulatory constraints. This reorganization ensures that the training data provided to the large model is not chaotic, but is a knowledge fragment with higher information density organized around regulatory constraints. In order to coordinate the process of data flowing from the type database to the incremental knowledge pool, a matching unit is designed. The matching unit stores the retrieval tags of all type databases and receives training features generated from the large model training process, and uses matching algorithms to match the retrieval tags and training features. The entire system connects multiple type databases, matching units and incremental knowledge pools through dynamic transmission channels, forming a closed loop that can dynamically provide, process and reorganize building data according to the training needs of the large model.

[0075] In one embodiment, the step S4 of providing training data through the structured knowledge graph and the incremental knowledge pool, training the large model based on the training data, and generating the building vertical field large model comprises:

[0076] S41, the structured knowledge graph and the incremental knowledge pool are connected to the large model through a data transmission channel, the large model includes a feature extractor and a dynamic learning engine;

[0077] S42, the structured knowledge graph generates structured training data and transmits it to the large model through a data transmission channel, and the large model is trained based on the dynamic learning engine;

[0078] S43, the feature extractor extracts training features in the training process of the large model, including attention hotspots, regional omissions, and prediction biases, and transmits the training features to the incremental knowledge pool based on the data transmission channel;

[0079] S44, after receiving the training features, the incremental knowledge pool activates multiple type databases to generate incremental training data;

[0080] S45, the incremental training data is transmitted to the large model, and the building vertical field large model is trained and generated based on the dynamic learning engine;

[0081] As described in steps S41-S45 above, first, the large model is connected to the structured knowledge graph and the incremental knowledge pool through a data transmission channel, and is equipped with a dynamic learning engine for processing and learning data and a feature extractor for monitoring the learning state. The model initially mainly receives training data from the structured knowledge graph, which is highly extracted from building regulations, providing the core logic and constraint relationship of the regulations, so that the model establishes a basic framework for complying with the regulations. During the training process, the feature extractor of the large model continuously monitors the internal performance of the model, captures key "training features", such as parts of the model that allocate high attention when processing specific regulatory provisions or data (attention hotspots), knowledge gaps for specific regional specifications (regional omissions), and errors in prediction tasks (prediction bias); These training features are fed back to the incremental knowledge pool. Based on these feedbacks, the incremental knowledge pool activates various types of building databases connected to it, retrieves actual building data highly related to the current learning difficulties of the model, and uses the built-in constraint relationship matrix to intelligently reorganize these retrieved raw data to generate incremental training data that can solve the model's knowledge deficiency. These reorganized data are then sent back to the large model for more detailed and targeted incremental training by the dynamic learning engine; This iterative training cycle based on model feedback and actual data reorganization enables the large model to continuously optimize its understanding and application of building regulations, and ultimately forms a highly specialized and closely integrated building vertical field large model.

[0082] In one embodiment, the step S44 of activating multiple type databases to generate incremental training data after the incremental knowledge pool receives the training features, comprises:

[0083] S441、The incremental knowledge pool receives the training features and transmits them to the matching unit through a dynamic transmission channel;

[0084] S442、The matching unit generates a call data table based on the semantic matching algorithm matching the search labels and the training features, and calls the building data based on the call data table and transmits it to the incremental knowledge pool;

[0085] S443、The incremental data pool receives the building data, reorganizes the building data based on the constraint relationship matrix to generate incremental training data, and transmits it to the large model through the data transmission channel;

[0086] As described in steps S441-S443 above, after receiving the training features, the incremental knowledge pool transmits the data to the matching unit through a dynamic transmission channel. The matching unit uses a semantic matching algorithm to compare the training features with the pre-set search labels, generates a call data table containing relevant data indexes, and lists the labels with high relevance to the current training features to extract relevant building data from different types of databases. The matching unit sends a request according to the call data table, and the dynamic window extracts the required building data, such as regional specifications, past cases, and material parameters, after receiving these requests. These original building data are transmitted back to the incremental knowledge pool. Then, the incremental knowledge pool reorganizes the received data according to the built-in constraint relationship matrix. This reorganization process ensures that the logic and constraint relationships of the building data are reasonably applied to generate clear incremental training data.

[0087] In one embodiment, the step S442 of the matching unit generating a call data table based on the semantic matching algorithm matching the search labels and the training features, and calling the building data based on the call data table and transmitting it to the incremental knowledge pool, comprises:

[0088] S4421、The matching unit receives the training features and calculates the matching degree of the training features and the search labels using the semantic matching algorithm;

[0089] S4422、Sort the search labels with a matching degree exceeding the matching threshold according to the size of the matching degree to generate a call data table, and transmit the call data table to the dynamic window;

[0090] S4423、The dynamic window calls the building data in the type database based on the call data table and transmits it to the incremental knowledge pool;

[0091] As described in steps S4421-S4423 above, the matching unit first receives the training features from the incremental knowledge pool, sets a predetermined matching threshold, calculates the matching degree between the training features and the stored search labels using a semantic matching algorithm, and identifies high-matching labels related to the features. When the matching degree of a certain search label exceeds the set threshold, it will be included in the generated call data table, and in this process, the matching labels will be sorted according to their matching degrees with the training features. The generated call data table is then sent to the dynamic window through the dynamic transmission channel. After receiving the call data table, the dynamic window quickly retrieves the corresponding building data from each type of database according to the labels in the table. These data are then transmitted back to the incremental knowledge pool to provide basic materials for subsequent data restructuring and model training, thereby improving the learning efficiency of the model while ensuring data relevance. Furthermore, the dynamic window regulates the bandwidth of the dynamic transmission channel based on the call data table, with higher bandwidth dynamic transmission channels generated for type databases with higher matching degrees, and lower bandwidth dynamic transmission channels generated for type databases with lower matching degrees. Based on this method, efficient transmission of building data can be achieved.

[0092] Embodiment 2, please refer to Figure 2 As shown in the figure, the building regulation-based vertical field large model training system described in this embodiment includes:

[0093] The regulation segmentation module: constructs a dynamic regulation knowledge base including a clause semantic graph and a constraint relationship matrix.

[0094] The knowledge graph generation module: connected with the regulation segmentation module, generates a structured knowledge graph based on the clause semantic graph and the constraint relationship matrix.

[0095] The incremental knowledge pool module: connected with the knowledge graph generation module, constructs multiple type databases to store building data, and sets a matching unit connected to multiple type databases.

[0096] The large model training module: connected with the incremental knowledge pool module, provides training data through the structured knowledge graph and the incremental knowledge pool, trains a large model based on the training data, and generates a building vertical field large model.

[0097] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for training large-scale models in a vertical domain based on building regulations, characterized in that: Constructing a dynamic legal knowledge base includes clause semantic graphs and constraint relationship matrices; A structured knowledge graph is generated based on the clause semantic graph and the constraint relationship matrix; Build multiple types of databases to store building data, and set up matching units and incremental knowledge pools to connect the multiple types of databases; Training data is provided through structured knowledge graphs and incremental knowledge pools, and large models are trained based on the training data to generate large models for the vertical field of architecture. The steps for generating a structured knowledge graph based on the clause semantic graph and constraint relation matrix are as follows: By associating building code entities with multi-dimensional parameters in the constraint relationship matrix, enhanced knowledge nodes are generated; Multidimensional parameters are transformed into directed edges using a matrix mapping algorithm; Structured knowledge graphs are generated by connecting enhanced knowledge nodes with directed edges; The steps for constructing multiple types of databases to store building data and setting up an incremental knowledge pool to connect to multiple types of databases are as follows: Multiple types of databases are constructed to store building data. These databases include regional code databases, building case databases, material parameter databases, construction method databases, and acceptance issue databases. Search tags are defined for each type of database. Set up a dynamic window on the type database to receive data and generate dynamic transmission channels; Construct an incremental knowledge pool, which contains a constraint relationship matrix, to reorganize building data and generate incremental training data; A matching unit is built among multiple types of databases and an incremental knowledge pool. The matching unit is used to store search tags and perform data matching. Multiple types of databases, matching units, and incremental knowledge pools are connected through a dynamic transport channel.

2. The method for training a large vertical domain model based on building regulations according to claim 1, characterized in that: The steps for constructing a dynamic legal knowledge base, including clause semantic graphs and constraint relationship matrices, are as follows: Obtaining the original regulatory text and performing semantic parsing yields a semantic graph of the clauses, including building regulation entities, core semantic structures, and semantic relationships. Based on the core semantic structure and semantic relationships in the clause semantic graph, a constraint relationship matrix is ​​constructed. The constraint relationship matrix consists of multi-dimensional parameters, including constraint strength, spatial influence, time decay, regional adaptation, and conflict risk.

3. The method for training a large vertical domain model based on building regulations according to claim 1, characterized in that: The steps for providing training data through structured knowledge graphs and incremental knowledge pools, training a large model based on the training data, and generating a large model for the architectural vertical domain are as follows: The structured knowledge graph and the incremental knowledge pool are connected to the large model through a data transmission channel. The large model includes a feature extractor and a dynamic learning engine. The structured knowledge graph generates structured training data, which is then transmitted to the large model via a data transmission channel. The large model is then trained based on a dynamic learning engine. The feature extractor extracts training features from the large model training process, including attention hotspots, regional missing features, and prediction bias, and transmits the training features to the incremental knowledge pool via a data transmission channel. After receiving the training features, the incremental knowledge pool activates multiple types of databases to generate incremental training data. Incremental training data is transferred to a large model, which is then trained and generated using a dynamic learning engine to create a large model for the architectural vertical domain.

4. The method for training a large vertical domain model based on building regulations according to claim 3, characterized in that: After receiving the training features, the incremental knowledge pool activates multiple types of databases to generate incremental training data in the following steps: After receiving the training features, the incremental knowledge pool transmits them to the matching unit through a dynamic transmission channel; The matching unit uses a semantic matching algorithm to match and retrieve tags and training features to generate a call data table, and then uses the call data table to transfer building data to the incremental knowledge pool. The incremental data pool receives building data, reorganizes the building data based on the constraint relationship matrix to generate incremental training data, and transmits it to the large model through the data transmission channel.

5. The method for training a large vertical domain model based on building regulations according to claim 4, characterized in that: The matching unit generates a call data table based on semantic matching algorithms by matching retrieval tags and training features, and then calls the building data to the incremental knowledge pool based on the call data table. A preset matching threshold is set, and the matching unit receives training features and uses a semantic matching algorithm to calculate the matching degree between the training features and the search tags. Search tags with a matching degree exceeding the matching threshold are sorted according to their matching degree to generate a call data table, and the call data table is transmitted to the dynamic window; The dynamic window retrieves building data from the type database based on the data table and transmits it to the incremental knowledge pool.

6. A large-scale vertical domain model training system based on building regulations, used to implement the large-scale vertical domain model training method based on building regulations as described in any one of claims 1-5, characterized in that: Regulatory segmentation module: Constructs a dynamic regulatory knowledge base, including clause semantic graphs and constraint relationship matrices; Knowledge graph generation module: Connected to the regulation segmentation module, it generates a structured knowledge graph based on the clause semantic graph and constraint relationship matrix; Incremental Knowledge Pool Module: Connects to the knowledge graph generation module, constructs multiple types of databases to store building data, and sets up matching units and incremental knowledge pools to connect multiple types of databases; Large Model Training Module: Connected to the Incremental Knowledge Pool Module, it provides training data through structured knowledge graphs and the incremental knowledge pool, trains large models based on the training data, and generates large models for the architectural vertical domain.

Citation Information

Patent Citations

  • Safety propaganda and education training knowledge graph and data management method and system based on AI

    CN118585658A

  • Construction management method based on digitization

    CN118964440A