A historical building ontology modeling system for generative artificial intelligence

By constructing a historical building ontology modeling system oriented towards generative artificial intelligence, the problem of insufficient processing capability for complex construction knowledge in existing technologies has been solved, realizing intelligent reasoning and dynamic generation of historical buildings, and improving the historical authenticity and generation efficiency of ancient building design.

CN120541930BActive Publication Date: 2025-12-02TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510648390.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-12-02
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

Existing technologies cannot effectively handle the complex construction knowledge in the field of architectural engineering, lack dynamic reasoning and generation capabilities, and are difficult to apply to the design and digital restoration of historical buildings.

Method used

A historical building ontology modeling system oriented towards generative artificial intelligence is adopted, including a data acquisition and preprocessing layer, an attribute graph structure establishment layer, an attribute graph structure reasoning layer, and a generative artificial intelligence interaction layer. It constructs a knowledge graph of historical building construction, combines rule-based reasoning and data-driven reasoning to complete the knowledge, and generates ancient building design images and knowledge enhancement feedback through the GenAI model.

Benefits of technology

It achieves accurate expression, intelligent reasoning, and dynamic generation of historical building construction knowledge, solves the problems of digital resource redundancy and low efficiency of traditional parametric design, and improves the historical authenticity and generation efficiency of ancient building design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

This invention relates to the field of architectural engineering, and particularly to a historical building ontology modeling system oriented towards generative artificial intelligence. The system includes a data acquisition and preprocessing layer for collecting, parsing, and standardizing the storage of multi-source data on historical building construction; an attribute graph structure building layer for establishing a hierarchical semantic model and constructing a knowledge graph of historical building construction; an attribute graph structure reasoning layer for knowledge completion by combining rule-based reasoning and data-driven reasoning; and a generative artificial intelligence interaction layer responsible for generating semantic tags based on the knowledge graph and driving the stable diffusion model in the GenAI model to generate images of antique building designs and knowledge enhancement feedback. This invention constructs a structured knowledge representation model oriented towards the characteristics of historical building construction, solving the problem of digital resource redundancy and difficulty in coordinating with artificial intelligence technology in the process of digital protection of historical building heritage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of architectural engineering, and in particular relates to a historical building ontology modeling system oriented towards generative artificial intelligence. Background Technology

[0002] Historical building design refers to the process of designing and creating buildings with historical value, cultural significance, and artistic features. It not only focuses on the practical function of the building, but also emphasizes showcasing the style and craftsmanship of a specific historical period through elements such as form, materials, and decoration.

[0003] Existing technologies cannot handle the complex construction knowledge in the field of architectural engineering, lack dynamic reasoning and generation capabilities, and are difficult to apply to the design and digital restoration of historical buildings. Summary of the Invention

[0004] The purpose of this invention is to provide a historical building ontology modeling system for generative artificial intelligence, which aims to solve the problems that existing technologies cannot handle the complex construction knowledge in the field of architectural engineering, lack dynamic reasoning and generation capabilities, and are difficult to apply to the design and digital restoration of historical buildings.

[0005] This invention is implemented as follows: a historical building ontology modeling system oriented towards generative artificial intelligence, the system comprising:

[0006] Data Acquisition and Preprocessing Layer: Used for the acquisition, parsing, and standardized storage of multi-source data on the construction of historical buildings;

[0007] Attribute graph structure layer: used to build a hierarchical semantic model and construct a knowledge graph of historical building construction;

[0008] Attribute graph structure reasoning layer: used to combine rule-based reasoning and data-driven reasoning for knowledge completion;

[0009] Generative AI Interaction Layer: Responsible for generating semantic tags based on knowledge graphs and driving the stable diffusion model in the GenAI model to generate images of ancient architectural designs and knowledge-enhancing feedback.

[0010] Preferably, in the process of collecting multi-source data on the construction of historical buildings, two types of historical building data are collected: architectural drawing information and measured data. The architectural drawing information includes floor plans and section views. The floor plans include the coordinate positions of column points, and the measured data includes pixel images of the exterior of the historical buildings taken from multiple angles.

[0011] Preferred steps for parsing and standardizing the storage of multi-source data on historical building construction include:

[0012] Drawing information processing: After establishing a Cartesian coordinate system with the center point of the main entrance of the building as the origin (0,0,0), the plane (X,Y) coordinates of each plane column point are marked and stored in the attribute list of the column entity in float data format. The height H of each column is marked and stored in the attribute list of the column entity in float data format.

[0013] Actual data processing: Historical building images were collected using a multi-angle standardization method to unify image tone, lighting and resolution. All images were uniformly cropped to the resolution required for training, and an image training dataset for LoRA fine-tuning was constructed. A caption / tag mapping table was established for LoRA training reference.

[0014] Preferably, the steps for establishing a hierarchical semantic model and constructing a knowledge graph of historical building construction include:

[0015] Establish a knowledge classification system for the construction of historical buildings and establish an attribute tag generation system;

[0016] Define the entity layer structure and attribute names of the historical building construction knowledge ontology;

[0017] Define the relational layer structure of the knowledge ontology of historical buildings.

[0018] Preferably, in the steps of establishing a knowledge classification system for the construction of historical buildings and establishing an attribute tag generation system, the knowledge entities of the construction of historical buildings are abstracted into four types: functional, layout and composition, materials and construction, and culture and symbolism. The corresponding attribute type tags are automatically generated using a Python program.

[0019] Preferably, the process of defining the entity layer structure and attribute names of the historical building construction knowledge ontology includes: extracting construction knowledge by type classification, dividing it into 8 entity categories: roof, interface, space, column grid, proportion, wall, detail, and component; defining attribute names and setting labels for these 8 entity categories; the entity name label information for the 8 entity categories of roof, interface, space, column grid, proportion, wall, detail, and component are respectively: _Building, _Interface, _Space, _Post, _Ratio, _WallPlan, _Detail, and _Element; different definition methods are used to determine the attributes of entities corresponding to different labels; and the attributes of the next higher level are then expanded into attribute entity modeling.

[0020] Preferably, the process of defining the relational layer structure of the knowledge ontology of historical buildings includes defining the core semantic relationships between conceptual entities. These core semantic relationships include:

[0021] Relative relationship refers to the relative positional relationship between spaces or components in a historical building;

[0022] Composition relationship indicates that one entity is a component of another entity;

[0023] It has a relational relationship and is used to express that a component or space has certain detailed features or subordinate elements;

[0024] Equivalent to a classification relationship, indicating that an entity is equivalent to or subordinate to a certain category.

[0025] Preferably, the attribute graph structure reasoning layer uses Neo4j's LPG graph database as the underlying platform. Utilizing attribute graph modeling and Cypher-based rule queries, it achieves analogy, reasoning, and structural relationship identification of the semantics of historical building construction, providing three reasoning methods:

[0026] Class-level hierarchical structural inheritance reasoning: By modeling the inheritance relationship between building space types and their constituent components, after inputting a specified space type, the typical exterior components and parameters are called to support the automatic facade construction logic based on type recognition.

[0027] Rule-based reasoning based on attributes and tags: By combining and judging the spatial location and functional semantic tags of building components, when several spatial tag and location tag conditions are met, the generation rules of specific component combinations are triggered.

[0028] Spatial logic reasoning based on component topology: By analyzing the coordinate data and construction logic of components, the spatial topological relationship between components is analyzed. When different types of column components have the same position or spatial adjacency, the system automatically establishes semantic connection relationship.

[0029] Preferably, the generative AI interaction layer uses the reasoning results based on the component-space-facade relationship in Neo4j to automatically identify the semantic component partitions that should be included in the generated target view. It reads the geometric position information and classification labels of entities through Python scripts, performs two-dimensional projection coordinate transformation, and draws ControlNet semantic segmentation mask images. Based on the reasoning results of the graph database, it extracts the style tags of the target building instance and automatically splices them together with the space type and component composition to form a natural language description prompt with high semantic accuracy.

[0030] Preferably, the steps for obtaining the ControlNet semantic segmentation mask image specifically include:

[0031] Raw data preparation and export: Export the spatial location attributes of components obtained from the atlas database in CSV format to form a coordinate matrix;

[0032] Coordinate normalization and pixel mapping: Transformation T using the translation matrix shift Align the smallest coordinate point with the origin:

[0033]

[0034] In the above formula, X and Z correspond to coordinate values ​​in the original space, where X represents the horizontal position and Z represents the height value. min and Z min This represents the minimum value in the X direction and the minimum value in the Z direction in the original coordinate system.

[0035] Using scaling matrix transformation T scale Scaling the coordinates to the [0,1] interval facilitates mapping after normalization:

[0036]

[0037] In the above formula, X max and X min Represents the extreme values ​​in the X direction of the original coordinate system; Z represents the extreme values ​​in the X direction. max and Z min This represents the extreme value in the Z direction of the original coordinate system;

[0038] Using combinatorial normalization T norm A matrix is ​​used to perform a translation followed by scaling on the coordinate positions, thus achieving a one-step transformation from the world coordinate system to the normalized coordinate system:

[0039] T norm =T scal e·T shift

[0040] Map the normalized coordinates to the image pixel space:

[0041]

[0042] In the above formula, W and H are the width and height of the image, in pixels; m is the pixel margin reserved at the image edge, used to avoid the coordinates falling at the boundary and being cropped, often used for white space at the image edge; u is the X coordinate mapped to the horizontal pixel coordinate of the image; v is the Z coordinate mapped to the vertical pixel coordinate of the image; multiplying by (W-2m) or (H-2m) means enlarging the normalized result to the actual resolution range of the image; adding m means to retain the margin at the image edge.

[0043] The beneficial effects of the historical building ontology modeling system for generative artificial intelligence provided by this invention are:

[0044] A structured knowledge representation model oriented towards the construction characteristics of historical buildings was constructed, which solved the problem of redundancy of digital resources and difficulty in coordinating with artificial intelligence technology in the process of digital protection of historical building heritage;

[0045] By combining the knowledge graph of historical building construction with the knowledge-enhanced reasoning model, we can improve GenAI's knowledge completion and expression capabilities. By enhancing GenAI's knowledge constraints, we can solve the problem that design elements that do not conform to historical experience and objective laws are prone to appear in the automated generation design process of ancient buildings.

[0046] It solves the problems of low efficiency, long time consumption, and structural errors that are prone to occur in the process of manual rule writing in traditional parametric design. Attached Figure Description

[0047] Figure 1 Flowcharts of each structural layer of the historical building ontology modeling system for generative artificial intelligence provided in this embodiment of the invention;

[0048] Figure 2 A flowchart illustrating the conceptual entity, entity attribute classification, and modeling method of historical building ontology provided in this embodiment of the invention;

[0049] Figure 3 A flowchart illustrating the method for classifying and modeling the relational attributes of historical building entities provided in this embodiment of the invention;

[0050] Figure 4-1 A flowchart of the first reasoning mechanism provided in an embodiment of the present invention;

[0051] Figure 4-2 A flowchart illustrating the second reasoning mechanism provided in this embodiment of the invention;

[0052] Figure 4-3 A flowchart illustrating the third reasoning mechanism provided in this embodiment of the invention;

[0053] Figure 5 A flowchart for generating ControlNet semantic segmentation mask images driven by semantic graph inference, provided in an embodiment of the present invention;

[0054] Figure 6 A multi-channel generation flowchart provided for embodiments of the present invention;

[0055] Figure 7 This is a schematic diagram illustrating the generation of entity labels and entity attribute labels in the process of modeling the ontology of historical buildings in Huizhou based on the present invention, provided as an embodiment of the present invention.

[0056] Figure 8 This is a schematic diagram illustrating the comparison of multi-channel generated image results provided in an embodiment of the present invention;

[0057] Figure 9 This is a schematic diagram of the FID check and expert scoring results of the multi-channel generated results provided in this embodiment of the invention;

[0058] Figure 10This is a schematic diagram of the scoring results of four expert evaluation indicators provided in an embodiment of the present invention;

[0059] Figure 11 This is a semantic encoding diagram provided for an embodiment of the present invention. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0061] This invention provides a historical building ontology modeling system for generative artificial intelligence, the system comprising:

[0062] Data Acquisition and Preprocessing Layer: Used for the acquisition, parsing, and standardized storage of multi-source data on the construction of historical buildings;

[0063] Attribute graph structure layer: used to build a hierarchical semantic model and construct a knowledge graph of historical building construction;

[0064] Attribute graph structure reasoning layer: used to combine rule-based reasoning and data-driven reasoning for knowledge completion;

[0065] Generative AI Interaction Layer: Responsible for generating semantic tags based on knowledge graphs and driving the stable diffusion model in the GenAI model to generate images of ancient architectural designs and knowledge-enhancing feedback.

[0066] This invention provides a historical building ontology modeling system for generative artificial intelligence. Through structured knowledge representation, knowledge-enhanced reasoning, semantic retrieval optimization, and GenAI model integration, it achieves accurate representation, intelligent reasoning, dynamic generation, and semantic retrieval of historical building construction knowledge. This promotes the development of digital protection of historical buildings and intelligent ancient building design, and realizes the triple goals of historical authenticity, style consistency, and intelligent generation efficiency in the process of digital representation of historical buildings.

[0067] The overall architecture of the historical building ontology modeling system for generative artificial intelligence is divided into three main parts: a data acquisition and preprocessing layer, a knowledge graph construction and reasoning layer, and a generative artificial intelligence interaction layer. The system adopts a technical approach of "layered modeling + multimodal fusion + dynamic reasoning + generative enhancement" to achieve in-depth analysis, reasoning, and dynamic generation of historical building construction knowledge. Subsystems are established according to the four layers. The system logical architecture is as follows: Figure 1 As shown.

[0068] Data Acquisition and Preprocessing Layer: Used for the acquisition, parsing, and standardized storage of multi-source data on the construction of historical buildings.

[0069] In this system, step 1: multi-source data acquisition:

[0070] In this invention, two types of data related to historical buildings need to be collected: 1) architectural drawing information, which needs to include floor plans and cross-sectional views, and the floor plans need to include the coordinate positions of column points; 2) measured data, which requires pixel images of the exterior of the historical buildings taken from multiple angles.

[0071] Step 2: Data Standardization and Structure:

[0072] 1) Drawing Information Processing: After establishing a Cartesian coordinate system with the center point of the main entrance of the building as the origin (0,0,0), the planar (X,Y) coordinates of each planar column point are calibrated and stored in the "x:coordinate" and "y:coordinate" attribute lists of the column entity in float data format. Then, the height H of each column is calibrated and stored in the "h:height" attribute list of the column entity in float data format. 2) Measured Data Processing: Historical building images are collected using a multi-angle standardization method to unify image tone, lighting, and resolution to enhance sample consistency. All images are further cropped to the required training resolution (512x512) to construct a high-quality image training dataset for LoRA fine-tuning. Subsequently, a caption / tag mapping table is automatically created for LoRA training reference.

[0073] Attribute graph structure layer: used to build a hierarchical semantic model and construct a knowledge graph of historical building construction.

[0074] In this system, Step 1: Establish a classification system for historical building construction knowledge and an attribute tag generation system. Abstract historical building construction knowledge entities into four types: 1) Functional type (space usage type); 2) Layout and composition type (plan elements, facade elements); 3) Material and construction type (structure, components, construction details); 4) Cultural and symbolic type (ritual, hierarchy, theme, narrative, regional style). A Python program is used to automatically generate the corresponding attribute type tags: "Function_", "Layout_", "Tectonic_", and "Symbolism_".

[0075] Step 2: Define the entity layer structure and attribute names of the historical building construction knowledge ontology:

[0076] The spatial layout of Chinese historical buildings is constrained by a basic structural framework: the horizontal coordinates and height of the columns. This reflects the timber structure organization pattern, namely the "large timberwork" system in the Chinese architectural system, exhibiting strong regional stylistic characteristics. The "bay" module facing the entrance determines the compositional features of the front facade elements, while the "depth" module perpendicular to the entrance restricts the compositional features of the side facade elements. The method proposed in this invention, in entity extraction, first classifies and extracts construction knowledge into eight categories based on a rule-based extraction model: "roof," "interface," "space," "column grid," "proportion," "wall," "detail," and "component." The entity name labels for these entity nodes are set as: "_Building," "_Interface," "_Space," "_Post," "_Ratio," "_WallPlan," "_Detail," and "_Element."

[0077] Define attribute names and set labels for the above 8 types of entities respectively.For entities tagged "_Building", they are defined as "Hall" and "Entrance" attributes according to the building's usage type, stored using the tags "_TingTang" and "_MenDi"; for entities tagged "_Interface", they are defined as "External Interface" and "Internal Interface" according to the location of the enclosed space (the difference between the outer perimeter interface and the courtyard interface), stored using the tags "_ExternalFacade" and "_InternalFacade"; for entities tagged "_Space", they are defined as "Hall", "Hall", "Side Room", "Courtyard", and "Corridor" attributes according to the space's functional location, stored using the tags "_Tang" and "_Ting". For entities tagged "_CeShi", "_Yuan", and "_Lang", the attributes are defined as "Corridor Column", "Step Column", "Spine", and "Door Post" according to the component function category, and stored with the tag information "_LangZhu", "_BuZhu", "_JiZhu", and "_MenZhu". For entities tagged "_Ratio", the attributes are defined as "Width" and "Depth" according to the spatial ratio, and stored with the tag information "_KaiJian" and "_JinShen". For entities tagged "_WallPlan", the attributes are defined as "Front Wall" and "Side Wall" according to the location of the wall, and stored with the tag information "_MainWall". For entities tagged "_Detail", they are stored as "_WallCap", "_Relief", and "_Base" according to their construction detail categories; for entities tagged "_Element", they are defined as "Door Set" and "Window Set" according to their door and window component types; for entities tagged "_Ratio", they are defined as "Horizontal Scale" and "Vertical Scale" according to their spatial scale planar dimensions; for entities tagged "_Detail", they are stored as "_WallCap", "_Relief", and "_Base"; for entities tagged "_Detail", they are defined as "_WallCap", "_Relief", and "_Base" according to their construction detail categories; for entities tagged "_Element", they are defined as "Door Set" and "_Window Set" according to their door and window component types; for entities tagged "_Ratio", they are defined as "Horizontal Scale" and "Vertical Scale" according to their spatial scale planar dimensions; for entities tagged "_Detail", they are stored as "_Detail" and "_SideWall"; for entities tagged "_Detail", they are defined as "_Detail" and "_SideWall" according to their construction detail categories. Entities labeled "_WallPlan" are defined as "Front Wall" and "Side Wall" according to their facade orientation, and stored using the labels "_MainWall" and "_SideWall". Entities labeled "_Detail" are defined as "Top Decoration", "Middle Decoration", and "Base Decoration" according to their vertical hierarchy, and stored using the labels "_WallCap", "_Relief", and "_Base". Entities labeled "_Element" are defined as "Door Set" and "Window Set" according to the component type of the facade opening, and stored using the labels "_DoorSet" and "_WindowSet".

[0078] Expand the attributes at the previous level into entity models. Add "_position" attribute entities to describe the positional relationships of spatial elements in front, behind, left, right, and center.

[0079] Add an entity with the "_component" attribute that describes the component ID;

[0080] Added an entity with the attribute "_coordinate" to describe the planar coordinates of a component; added an entity with the attribute "_height" to describe the height of a component.

[0081] Add an entity with the attribute "_Length" to describe the width of the "bay" in traditional Chinese architecture;

[0082] Add an entity with the attribute "_Span" to describe the "depth" span in traditional Chinese architecture;

[0083] Add an entity with the attribute "_SanKaiJian" to describe the "open-plan" style of the front facade wall structure;

[0084] Add an entity with the attribute "_Ting-Tang" to describe the "hall-style" form of the side facade walls;

[0085] Add attribute entities for describing the material, style, symbolism, and symbolic characteristics of decorative details;

[0086] Add an entity with the attribute "_orientation" to describe the positioning of decorative details and the door and window groups that make up the facade.

[0087] (like Figure 2 (As shown)

[0088] Step 3: Define the relational layer structure of the historical building construction knowledge ontology:

[0089] The purpose of this step is to define the relationships between conceptual entities, forming a knowledge network. The entity relationship extraction pattern invented in this method is based on construction logic, focusing on the organization of architectural elements, the order of component connections, and spatial layout patterns. The constructed relationship layer is based on four core semantic relationships:

[0090] 1) Relative relationship, which indicates the relative positional relationship between spaces or components in a historical building, such as "the porch is located in front of the main hall" or "the side room is located on the east side of the main hall". It is used to express the orientation, direction and alignment logic between components. The relationship is defined as "relative_to".

[0091] 2) Composition relationship, indicating that one entity is a component of another entity, such as "columns are part of the porch" or "window group belongs to the facade", used to construct the whole-part structure in the building system, and the relationship name is defined as "part_of";

[0092] 3) It has relevance and is used to express that a component or space has certain detailed features or auxiliary elements, such as "the wall has wall lines" or "the door group includes door head decoration". This relationship is applicable to the connection of material, construction, decoration and other elements in the construction level. The relationship name is defined as "has_".

[0093] 4) Equivalent to a classification relation, indicating an equivalence or subordinate relationship between an entity and a category, such as "an entity is a 'column'" or "a space is a 'side hall'", used to describe the type affiliation or semantic classification of an entity, and the relation name is defined as "is_". (e.g.) Figure 3 (As shown)

[0094] The aforementioned relational information is stored in the format of <entity, relation, entity> using triplet data format. After establishing relational tags in the Neo4j database, a well-structured and hierarchical semantic network of historical buildings can be constructed to support graph construction and GenAI model access tasks.

[0095] Attribute graph structure reasoning layer: used to combine rule-based reasoning and data-driven reasoning for knowledge completion.

[0096] In this system, such as Figure 4-1 , Figure 4-2 as well as Figure 4-3 As shown, the inference rule is defined as follows:

[0097] To achieve a structured understanding and computational representation of the spatial composition and component relationships of historical buildings, this invention constructs a knowledge reasoning system based on a graph database. The system uses Neo4j's LPG graph database as its underlying platform, leveraging attribute graph modeling and Cypher-based rule queries to achieve analogy, reasoning, and structural relationship identification of the semantics of historical building construction. Through the three rule and path computation mechanisms proposed in this invention, semantic recognition with "class-reasoning" characteristics and reconstruction of semantic associations between structural components can be achieved. The system reasoning mechanism proposed in this invention mainly includes the following three methods (such as... Figure 4-1 , Figure 4-2 as well as Figure 4-3 As shown):

[0098] 1) Class-Hierarchy Inference. By modeling the inheritance relationships between architectural space types and their constituent components, after inputting a specified space type (such as hall or pavilion), it automatically calls upon the typical exterior component composition and parameters (such as wall height, column spacing, etc.) to support automatic facade construction logic based on type recognition. The following is an example of its operation: Input (space node attributes)<spacetype:Space{name:‘Tang’}> ); Reasoning logic construction (performing "category-component" mapping reasoning based on inheritance semantics); Output (reasoning out component attributes)<partof:WallPlan{name:‘Tang’}> ).

[0099] 2) Attribute & Label Rule Inference. This method uses the spatial location and functional semantic labels of building components for combination judgment. When certain spatial and location label conditions are met, a generation rule for a specific component combination is triggered. For example, when the space is "Ting" or "Tang" and the location is "front" or "rear," its compositional feature is automatically inferred to be a "Ting-Tang" type wall combination. The following is an example of its operation: Input (attribute judgment, such as...)<Position{position:'front'> and<Space{name:'Ting'}> ); Reasoning logic construction (IF ATHENB type rule logic); Output (generating labels)<WallPlan{type:‘Ting-Tang’}> ).

[0100] 3) Component-Level Topological Inference. Utilizing component coordinate data and construction logic, the system analyzes the spatial topological relationships between components. When different types of column components have the same location or spatial adjacency, the system automatically establishes semantic connections, such as "adjacent," "subordinate," or "combined." The following is an example of its operation: Input (e.g., ...)<Post:'Tang'> {x=value} and<Post:'Lang'> {x=value}); Reasoning logic construction (spatial logic judgment based on component position and construction rules); Output (automatically generate graph relation (Tang)-[:NEXT_TO]->(Lang) and automatically add it to the graph structure).

[0101] This invention further interfaces with the GenAI (Stable Diffusion) model, constructing a semantic-driven generation mechanism that combines structure control (ControlNet) and style learning (LoRA). The core technical path includes automatically reasoning from the graph database and generating image structure constraints (Masks) and descriptive prompts (Prompts) for dual constraint input in image generation within the stable diffusion model. This interface mechanism significantly improves generation efficiency and the accuracy of semantic expression in architectural images, providing a novel intelligent interface mechanism for high-quality AIGC generation driven by traditional architectural image data. Specific implementation details are as follows:

[0102] Semantic graph reasoning-driven structural constraint generation: Based on the reasoning results of the "component-space-facade" relationship in Neo4j, the system automatically identifies and generates semantic component partitions that should be included in the target view (such as front elevation and side view). The system uses a Python script to read the geometric location information and classification labels of entities (such as "door," "window," "wall," etc.), performs two-dimensional projective coordinate transformation, and draws a 512×512 pixel ControlNet semantic segmentation mask image, such as... Figure 5 As shown.

[0103] This method skips the traditional modeling step and directly generates regions through pixel fragment synthesis, thus automating the generation of control constraints for building structures. The core process in this step includes: data export, 2D coordinate normalization transformation, pixel-level region drawing, and semantic coloring mapping, as detailed below (this process is as follows...). Figure 5 (As shown): 1) Raw data preparation and export. The spatial location attributes of the components (such as X, Z coordinates and height) obtained from the atlas database are exported in CSV format to form a coordinate matrix; 2) Coordinate normalization and pixel mapping. To map the actual spatial coordinates to the pixel region in the 512×512 image, the minimum coordinates are first aligned to the origin through a translation matrix transformation (T_shift):

[0104]

[0105] In the above formula, X and Z correspond to coordinate values ​​in the original space, with X representing the horizontal position (east-west and north-south) and Z representing the height value. Where X... min and Z min This represents the minimum value in the X direction and the minimum value in the Z direction in the original coordinate system.

[0106] Next, the coordinates are scaled to the image size using a scaling matrix transformation (T_scale):

[0107]

[0108] In the above formula, X max and X min Represents the extreme values ​​in the X direction of the original coordinate system; Z represents the extreme values ​​in the X direction. max and Z min This represents the extreme value in the Z direction of the original coordinate system.

[0109] Then use the combined normalization matrix (T_norm):

[0110] This ultimately achieves the coordinate-to-pixel position transformation (u,v), mapping the normalized coordinates to the image pixel space, where W = 512, H = 512, and m = edge-preserving pixels.

[0111] T norm =T scale ·T shift

[0112]

[0113] In the above formula, W and H are the width and height of the image, in pixels (512×512 in this case); m is the pixel margin reserved at the image edge, used to avoid the coordinates falling at the boundary and being cropped, often used for leaving white space at the image edge; u is the X coordinate mapped to the horizontal pixel coordinate of the image; v is the Z coordinate mapped to the vertical pixel coordinate of the image, and note that the Y-axis direction is reversed (the y-axis increases downward in the image coordinate system, so 1-· is used); multiplying by (W-2m) or (H-2m) means enlarging the normalized result to the actual resolution range of the image; adding m means to retain the margin at the image edge.

[0114] Image rendering and semantic encoding: The transformed component coordinates are mapped to corresponding pixel blocks in the image, and color coding is performed according to semantic categories to achieve automatic rendering of semantic mask diagrams. For example... Figure 11 As shown in the diagram, each type of component in the algorithm corresponds to a type of geometric block (which can be a rectangle, line segment, circle, etc.), which is then filled into the corresponding region of the image to form a semantic segmentation map. The final output is a ControlNet constraint image with a clear structure and distinct semantics, which is used as input for AI models.

[0115] The system automatically generates a prompt based on attribute combinations. Building upon the inference results from the graph database, the system extracts style tags (such as "three bays," "white walls and black tiles," and "horse-head walls") from the target building instance, combining them with the space type (such as "Huizhou traditional dwelling main entrance"). The system automatically combines elements such as “three-segment symmetrical facade” and “structure” (e.g., “three-segment symmetrical facade”) to form a natural language description with high semantic accuracy. Its generation logic is: [Quantity + Configuration] + [Orientation / Spatial Type] + [Style Feature Vocabulary] + [Building Type] + [Semantic Suffix]. An example is generated as follows: “SanKaiJian Huizhou traditional dwelling main entrance” "with MaTouQiang and black tiles". This Prompts generator is used to guide the generation of styles and component layouts in accordance with ControlNet input.

[0116] Generative AI Interaction Layer: Responsible for generating semantic tags based on knowledge graphs and driving the stable diffusion model in the GenAI model to generate images of ancient architectural designs and knowledge-enhancing feedback.

[0117] In this system, step 1: optimization of joint input generation using LoRA and ControlNet:

[0118] During generation, the structural control mask and semantic prompt generated by the above process are simultaneously input into the StableDiffusion model. The LoRA model handles style fine-tuning, while the ControlNet input provides geometric distribution control. Together, they achieve semantic-driven architectural image generation. This strategy not only improves the structural accuracy and style consistency of the generated results but also significantly reduces the cost of manual modeling and cleaning.

[0119] Step 2: As Figure 6 As shown, multi-channel generation and evaluation:

[0120] Channel 1: Based on historical drawings and measured image data from the data input terminal, an initial database is formed and...

[0121] The LoRA model is trained step by step, and the LoRA model path and naming are simultaneously modeled as entity nodes in Neo4j.

[0122] Channel 2: Uses SD+LoRA generation mode;

[0123] Channel 3: Using SD+LoRA+ControlNet generation mode

[0124] Channel 4: Complete the semantics based on the LPG inference results constructed in this invention;

[0125] Channel 5: Consistency check of generated results to form a scoring mechanism. The image with the highest overall score is used as the generated result, and the knowledge graph is updated in reverse, with two entity nodes, Prompt and ControlNet, added, and the file path generated as an attribute label, and stored in the Neo4j database.

[0126] like Figure 7 and Figure 8 As shown, this invention uses a historical residential building in the Huizhou region as an example to verify the versatility and application effectiveness of the method proposed in this invention. The historical building construction knowledge ontology modeling method proposed in this invention significantly improves the performance of Huizhou historical residential building images generated based on the StableDiffusion model in terms of structural accuracy, stylistic consistency, and semantic expression. Compared with existing general image generation methods, this invention has the following technical advantages and outstanding effects:

[0127] (1) The quality and structural accuracy of the generated images are significantly improved:

[0128] By introducing the ControlNet control mechanism and the LoRA style adjustment module, combined with the graph semantic reasoning interface proposed in this invention, KA-SDM achieved a minimum value of 16.9 in the image quality evaluation index FID (Fréchet Inception Distance), far superior to the original Stable Diffusion (39.2) and the ordinary SD+LoRA / ControlNet scheme (24.7), indicating that the generated result has a higher perceptual similarity to real Huizhou-style images (e.g., Figure 9 (As shown).

[0129] (2) The expert semantic scoring performance is excellent, which enhances the cognitive credibility of the generated content.

[0130] In the expert scoring stage, the proposed method achieved an average expert semantic score of 4.5 / 5, which is better than SD (2.8 / 5) and SD+LoRA / ControlNet (3.6 / 5). This demonstrates that the system of the present invention can more accurately express the semantic features of building structures and traditional typological logic, especially in the two generation tasks of "front facade" and "overall contextual perspective".

[0131] (3) The generated results are more consistent and expressive in terms of spatial logic and cultural expression.

[0132] Among the four expert evaluation indicators (such as...) Figure 10 As shown):

[0133] The typology fidelity score reached 4.6 / 5;

[0134] Spatial Logic scored as high as 4.8 / 5;

[0135] The Material Plausibility score is 4.6 / 5.

[0136] The symbolism score is 4.5 / 5.

[0137] These results indicate that the images generated by this method not only possess visual appeal in form but also accurately convey the traditional spatial hierarchy, proportions, and structural system of Huizhou architecture, demonstrating a high degree of design expressiveness and cultural adaptability.

[0138] (4) Saves training resources and enhances model adaptability

[0139] This invention achieves effective fine-tuning of the LoRA model using a small-scale, high-quality image set (approximately 30 images), demonstrating low training sample requirements and strong knowledge transfer capabilities. It reduces the generative model's dependence on large-scale training data and significantly saves computational resources and data preparation.

[0140] (5) Supports a two-way fusion generation mechanism for structural control and style adjustment

[0141] By using graph-driven semantic completion and pixel-level structural control (ControlNet mask), this invention can finely guide the structure of architectural images while maintaining stylistic consistency. It supports personalized image generation tasks from dimensions such as specified angles, component combinations, and spatial types, and has wide applicability to various scenarios.

[0142] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A historical building ontology modeling system for generative artificial intelligence, characterized in that, The system includes: Data Acquisition and Preprocessing Layer: Used for the acquisition, parsing, and standardized storage of multi-source data on the construction of historical buildings; Attribute graph structure layer: used to build a hierarchical semantic model and construct a knowledge graph of historical building construction; Attribute graph structure reasoning layer: used to combine rule-based reasoning and data-driven reasoning for knowledge completion; Generative AI Interaction Layer: Responsible for generating semantic tags based on knowledge graphs and driving the stable diffusion model in the GenAI model to generate images of ancient architectural designs and knowledge-enhancing feedback; The attribute graph structure reasoning layer uses Neo4j's LPG graph database as its underlying platform. Utilizing attribute graph modeling and Cypher-based rule queries, it achieves analogy, reasoning, and structural relationship identification of the semantics of historical building construction, providing three reasoning methods: Class-level hierarchical structural inheritance reasoning: By modeling the inheritance relationship between building space types and their constituent components, after inputting a specified space type, the typical exterior components and parameters are called to support the automatic facade construction logic based on type recognition. Rule-based reasoning based on attributes and tags: By combining and judging the spatial location and functional semantic tags of building components, when several spatial tag and location tag conditions are met, the generation rules of specific component combinations are triggered. Spatial logic reasoning based on component topology: By analyzing the coordinate data and construction logic of components, the spatial topological relationship between components is analyzed. When different types of column components have the same position or spatial adjacency, the system automatically establishes semantic connection relationship.

2. The historical building ontology modeling system for generative artificial intelligence according to claim 1, characterized in that, In the process of collecting multi-source data on the construction of historical buildings, two types of historical building data are collected: architectural drawing information and measured data. The architectural drawing information includes floor plans and section views. The floor plans include the coordinate positions of column points. The measured data includes pixel images of the exterior of historical buildings taken from multiple angles.

3. The historical building ontology modeling system for generative artificial intelligence according to claim 1, characterized in that, The steps involved in parsing and standardizing the storage of multi-source data on historical building construction include: Drawing information processing: After establishing a Cartesian coordinate system with the center point of the main entrance of the building as the origin (0,0,0), the plane (X,Y) coordinates of each plane column point are marked and stored in the attribute list of the column entity in float data format. The height H of each column is marked and stored in the attribute list of the column entity in float data format. Actual data processing: Historical building images were collected using a multi-angle standardization method to unify image tone, lighting and resolution. All images were uniformly cropped to the resolution required for training, and an image training dataset for LoRA fine-tuning was constructed. A caption / tag mapping table was established for LoRA training reference.

4. The historical building ontology modeling system for generative artificial intelligence according to claim 1, characterized in that, The steps for establishing a hierarchical semantic model and constructing a knowledge graph of historical building construction include: Establish a knowledge classification system for the construction of historical buildings and establish an attribute tag generation system; Define the entity layer structure and attribute names of the historical building construction knowledge ontology; Define the relational layer structure of the knowledge ontology of historical buildings.

5. The historical building ontology modeling system for generative artificial intelligence according to claim 4, characterized in that, In the steps of establishing a knowledge classification system for the construction of historical buildings and establishing an attribute tag generation system, the knowledge entities of the construction of historical buildings are abstracted into four types: functional, layout and composition, materials and construction, and culture and symbolism. The corresponding attribute type tags are automatically generated using a Python program.

6. The historical building ontology modeling system for generative artificial intelligence according to claim 4, characterized in that, The process of defining the entity layer structure and attribute names of the historical building construction knowledge ontology includes: extracting construction knowledge by type classification and dividing it into eight entity categories: roof, interface, space, column grid, proportion, wall, detail, and component. Attribute names and labels are defined for these eight entity categories. The entity name label information for the eight entity categories of roof, interface, space, column grid, proportion, wall, detail, and component are _Building, _Interface, _Space, _Post, _Ratio, _WallPlan, _Detail, and _Element, respectively. Different definition methods are used to determine the attributes of entities corresponding to different labels, and the attributes of the next higher level are then expanded into attribute entity models.

7. The historical building ontology modeling system for generative artificial intelligence according to claim 4, characterized in that, The process of defining the relational layer structure of the knowledge ontology for historical buildings includes defining the core semantic relationships between conceptual entities. These core semantic relationships include: Relative relationship refers to the relative positional relationship between spaces or components in a historical building; Composition relationship indicates that one entity is a component of another entity; It has a relational relationship and is used to express that a component or space has certain detailed features or subordinate elements; Equivalent to a classification relationship, indicating that an entity is equivalent to or subordinate to a certain category.

8. The historical building ontology modeling system for generative artificial intelligence according to claim 1, characterized in that, The generative AI interaction layer uses the reasoning results based on the component-space-facade relationship in Neo4j to automatically identify the semantic component partitions that should be included in the generated target view. It reads the geometric position information and classification labels of entities through Python scripts, performs two-dimensional projection coordinate transformation, and draws ControlNet semantic segmentation mask images. Based on the reasoning results of the graph database, it extracts the style tags of the target building instance and automatically splices them together with the space type and component composition to form a natural language description prompt with high semantic accuracy.

9. The historical building ontology modeling system for generative artificial intelligence according to claim 8, characterized in that, The steps to obtain the ControlNet semantic segmentation mask image specifically include: Raw data preparation and export: Export the spatial location attributes of components obtained from the atlas database in CSV format to form a coordinate matrix; Coordinate normalization and pixel mapping: Transformation T using the translation matrix shift Align the smallest coordinate point with the origin: In the above formula, X and Z correspond to coordinate values ​​in the original space, where X represents the horizontal position and Z represents the height value. min and Z min This represents the minimum value in the X direction and the minimum value in the Z direction in the original coordinate system; Using scaling matrix transformation T scale Scaling the coordinates to the [0,1] interval facilitates mapping after normalization: In the above formula, X max and X min Represents the extreme values ​​in the X direction of the original coordinate system; Z represents the extreme values ​​in the X direction. max and Z min This represents the extreme value in the Z direction of the original coordinate system; Using combinatorial normalization T norm A matrix is ​​used to perform a translation followed by scaling on the coordinate positions, thus achieving a one-step transformation from the world coordinate system to the normalized coordinate system: T norm =T scale ·T shift Map the normalized coordinates to the image pixel space: In the above formula, W and H are the width and height of the image, in pixels; m is the pixel margin reserved at the image edge, used to prevent the coordinates from being cropped when they fall on the boundary, often used for white space at the image edge; u is the X coordinate mapped to the horizontal pixel coordinate of the image; v is the Z coordinate mapped to the vertical pixel coordinate of the image; multiplying by (W-2m) or (H-2m) means enlarging the normalized result to the actual resolution range of the image; adding m means to retain the margin at the image edge; The transformed component coordinates are mapped to corresponding pixel blocks in the image, and color encoding is performed according to semantic categories to obtain the ControlNet semantic segmentation mask image.

Citation Information

Patent Citations

  • Product appearance image generation method and device and similar image retrieval method and device

    CN118097374A

  • SysML model knowledge graph semantic relationship reasoning method and device

    CN118627626A