Traditional Arcade Facade Component Intelligent Recognition and Language Conversion Methods Based on Hierarchical Coding System

By constructing a six-level hierarchical coding system and a Mask R-CNN model, the systematization and standardization problems of traditional arcade facade component identification have been solved, achieving efficient and automated identification and detailed description, and supporting large-scale building surveys and digital protection.

CN120544203BActive Publication Date: 2025-10-28XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511028374.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-10-28
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

Traditional arcade building facade component identification lacks a systematic coding system, relies on human experience, lacks standardized language expression, and is difficult to automate and accurately analyze by computers.

Method used

A six-level hierarchical coding system is constructed, and a component recognition model combining Mask R-CNN, ResNet-101, and FPN is adopted. Spatial relationships are modeled through graph convolutional networks, and knowledge graphs are constructed for automated processing by combining multi-dimensional feature extraction and standardized description generation.

Benefits of technology

It achieves standardized representation and efficient automatic recognition of arcade facade components, improving the recognition efficiency of a single arcade building by 100 times, generating detailed description text, and supporting large-scale building surveys and digital archiving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544203B_ABST
    Figure CN120544203B_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent recognition and language conversion method for traditional arcade facade components based on a hierarchical coding system. By constructing a six-level hierarchical coding system, training a specialized component recognition model, and achieving automatic coding and standardized description generation, it lays a data foundation for large-scale arcade building evolution analysis. The method includes: S1. Constructing a six-level standard coding system; S2. Collecting, preprocessing, and labeling images to construct a labeled dataset; S3. Constructing an arcade component recognition model, training and recognizing the model using existing and labeled datasets; S4. Assigning coding information from the standard coding system to the recognition results of S3, extracting multi-dimensional features, performing component association analysis, and constructing a component feature database; S5. Automatically generating three levels of standardized component description text based on S3 and S4; S6. Constructing a structured knowledge graph of arcade facade components based on S3 to S5; and S7. Integrating the processing capabilities of S2 to S6 into a complete system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of digital protection of architectural cultural heritage and artificial intelligence technology. In particular, it refers to a method for intelligent identification and language conversion of traditional arcade facade components based on a hierarchical coding system. By constructing a standardized component coding system and training an artificial intelligence model, the automatic identification, coding and standardized description generation of arcade facade components can be achieved. Background Technology

[0002] As an important cultural heritage, the facade features of traditional arcade buildings are a key basis for distinguishing arcade building types from different periods and regions. However, existing technologies have the following shortcomings in the identification and standardized representation of arcade building facade components:

[0003] (1) Lack of a systematic component coding system

[0004] Different researchers use different naming and classification methods for the same component. For example, "longevity beam" is called "lintel beam" or "door lintel beam" in different regions, which makes it difficult to systematically organize architectural knowledge and support large-scale computer processing and evolution analysis.

[0005] (2) Facade component identification relies on human experience

[0006] Traditional methods require professionals to identify each component on the facade one by one, which is inefficient and highly subjective. In particular, when facing a large-scale survey of hundreds of arcade buildings, manual identification methods cannot guarantee consistency and accuracy.

[0007] (3) Lack of standardized language expression for component characteristics

[0008] Existing records are mostly descriptive texts, lacking structured coding systems and cannot be directly used for computer analysis. This severely restricts subsequent work such as architectural evolution analysis, conservation, and restoration.

[0009] (4) Insufficient application of artificial intelligence technology

[0010] Although deep learning has made breakthroughs in the field of image recognition, in the area of ​​building component recognition, due to the lack of standardized training data and specialized recognition models, existing technologies struggle to accurately identify the complex components of arcade facades.

[0011] Therefore, there is an urgent need for a method based on a standardized coding system to automatically identify and code components of arcade facades by training an artificial intelligence model, thus providing a foundation for subsequent architectural evolution analysis. Summary of the Invention

[0012] The main objective of this invention is to provide an intelligent identification and language conversion method for traditional arcade facade components based on a hierarchical coding system, which solves the problems existing in the prior art. By constructing a six-level hierarchical coding system, training a specialized component identification model, and realizing automatic coding and standardized description generation, it lays a data foundation for large-scale arcade building evolution analysis.

[0013] To achieve the above objectives, the solution of the present invention is:

[0014] A method for intelligent recognition and language conversion of traditional arcade facade components based on a hierarchical coding system includes:

[0015] Step S1. Construct a standard coding system with six levels: building system level coding, regional subsystem coding, category coding, subcategory coding, specific building coding, and component coding. Component coding adopts a position-type-attribute sub-coding composite structure, with position and type constituting the basic coding of the component.

[0016] Step S2. Systematically collect images according to the preset image acquisition standards to obtain sufficient images of the arcade facade. Then, preprocess the collected images of the arcade facade and manually annotate them to build an annotated dataset.

[0017] Step S3. Construct a model for recognizing arcade components. Train and recognize arcade components using existing datasets and labeled datasets. The network architecture of the arcade component recognition model adopts Mask R-CNN as the basic framework. Its backbone network uses ResNet-101 combined with the Feature Pyramid Network (FPN). Multi-scale features are extracted through bottom-up and top-down feature fusion to adapt to the size differences of arcade components, from small decorations to large gables. In addition, a spatial relationship module is added to the basic framework of Mask R-CNN. The spatial relationship module models the spatial constraint relationships between different components through graph convolutional networks.

[0018] Step S4. For the identification results of step S3, assign coding information of the standard coding system and extract multi-dimensional features, perform component association analysis, and construct a component feature database;

[0019] Step S5. Based on the recognition results of step S3 and the encoding information of step S4, automatically generate standardized component description texts at three levels to achieve accurate conversion from visual features to natural language. These include a basic layer description of no more than 20 characters, a standard layer description of no more than 50 characters, and a professional layer description of no more than 100 characters.

[0020] Step S6. Integrate the recognition results of step S3, the encoding information of step S4, and the description text of step S5 to construct a structured knowledge graph of the arcade facade components;

[0021] Step S7. Integrate the processing capabilities in steps S2 to S6 into a complete system, which specifically includes: realizing the image preprocessing capability of step S2, realizing the intelligent component recognition capability of step S3, realizing the automatic encoding and feature extraction capability of step S4, realizing the multi-level description generation capability of step S5, and realizing the knowledge graph construction capability of step S6, thereby realizing end-to-end automated processing through a unified system architecture.

[0022] The standard coding system in step S1 is specifically as follows:

[0023] (i) Building system-level coding, using CTQL as the top-level identifier for traditional arcade buildings;

[0024] (ii) Regional sub-codes are set according to the geographical areas where the arcade buildings are distributed;

[0025] (iii) Group coding, used to distinguish regional characteristics;

[0026] (iv) Subgroup coding, accurate to specific historical streets;

[0027] (v) Specific building codes, assigning a number of digits to each arcade building;

[0028] (vi) Component coding, the specific sub-codes are:

[0029] (1) Location: Based on the construction logic of the arcade facade, R represents the roof part, B represents the body part, and F represents the base part;

[0030] (2) Type, used to reflect the detailed classification of location;

[0031] (3) Attributes, including at least one of the component’s material, age and functional characteristics, with M representing material, T representing age and U representing function.

[0032] The preset image acquisition standards in step S2 are as follows:

[0033] (1) When taking the picture, the axis of the camera lens should be perpendicular to the main plane of the building facade;

[0034] (2) The shooting distance should be controlled at 1.5 to 2 times the building height;

[0035] (3) Use a standard lens with an equivalent focal length of 35-50 mm;

[0036] (4) Select uniform lighting conditions for shooting;

[0037] (5) The image resolution should be no less than 4000×3000 pixels.

[0038] The preprocessing process in step S2 includes:

[0039] First, perspective correction is performed by detecting vertical and horizontal lines in the facade and calculating the perspective transformation matrix to correct the tilted image to a standard orthographic projection.

[0040] Secondly, color standardization is performed by using a standard color chart placed during shooting and calculating a color correction matrix.

[0041] Then, adaptive histogram equalization is used to enhance image contrast;

[0042] Finally, scale normalization is performed to adjust all images to a uniform pixel density.

[0043] In step S2, when constructing the labeled dataset, at least 5 architectural history experts with more than 10 years of experience are invited to participate in the labeling. The preprocessed arcade facade image is labeled on the labeling software. The component outline is accurately selected on the image, the component type is selected, the material properties and preservation status are labeled, and finally the labeled dataset is obtained. In order to ensure the labeling quality, a cross-validation mechanism is adopted during the labeling. Each component is independently labeled by at least 2 experts. When there is a disagreement, a consensus is reached through discussion.

[0044] The training strategy for the arcade component recognition model in step S3 adopts the following three-stage progressive method:

[0045] In the first stage, the COCO dataset was used for pre-training, allowing the arcade component recognition model to learn basic visual feature extraction capabilities.

[0046] In the second stage, the architectural image dataset is used for domain adaptation, enabling the arcade component recognition model to understand the unique visual patterns of architectural components.

[0047] In the third stage, the labeled dataset from step S2 is used for fine-tuning, with a focus on optimizing the ability of the arcade component recognition model to recognize unique arcade components.

[0048] Preferably, the training process of the arcade component recognition model in step S3 adopts the following optimization measures:

[0049] (1) Use the focus loss function to solve the problem of component category imbalance, and give higher loss weight to rare component categories with fewer numbers;

[0050] (2) Introduce an online difficult example mining mechanism, and pre-set easily confused component pairs in the arcade components according to practical experience. During the training process, automatically identify these component pairs and increase the training frequency of these component pairs according to the preset number or proportion value.

[0051] (3) By using data augmentation technology, the training samples are expanded by 3 times to improve the generalization ability of the arcade component recognition model.

[0052] Preferably, a post-processing mechanism is established in step S3 to ensure the reasonableness of the recognition results, specifically:

[0053] (1) Filter the detection boxes output by the arcade component recognition model by confidence level and retain only the results with a confidence level higher than 0.8;

[0054] (2) Use non-maximum suppression algorithm to eliminate overlapping detection boxes and avoid the same part being repeatedly identified;

[0055] (3) Construct a logical rule base based on architectural common sense, and verify the rationality of the identification results, such as checking whether the relative positional relationship of the components conforms to the architectural construction logic.

[0056] The process of allocating standard codes in step S4 is as follows:

[0057] First, the basic code is determined based on the identified component type and its location on the arcade facade;

[0058] By analyzing the visual features of the components, we can infer their possible construction date;

[0059] By combining the geographical information of the image's shooting location, the relevant regional codes are automatically filled in;

[0060] The material properties of a component are determined by a pre-trained material classification network.

[0061] Preferably, the multidimensional features extracted in step S4 cover the following dimensions:

[0062] (1) Geometric features;

[0063] (2) Visual characteristics;

[0064] (3) Stylistic characteristics;

[0065] (4) Material characteristics.

[0066] Preferably, the component association analysis step in step S4 is as follows:

[0067] (1) Spatial relationship modeling: Based on the bounding box coordinates of the component recognition results, calculate the spatial distance matrix and relative positional relationship between components, and establish a component spatial relationship diagram;

[0068] (2) Statistical co-occurrence probability: Based on a large number of samples, the co-occurrence frequency of different component types is statistically analyzed to construct a component co-occurrence probability matrix;

[0069] (3) Co-occurrence pattern mining: Statistical analysis of the co-occurrence frequency of different component types in the arcade sample, and the use of association rule mining algorithm to discover the combination pattern of strongly associated components;

[0070] (4) Verification of consistency of age: Based on the inferred age attributes of components, examine the consistency of age of different components of the same building and identify possible later renovations;

[0071] (5) Construct a component relationship network: Use components as nodes and association strength as edge weights to construct a weighted undirected graph for subsequent architectural style analysis and integrity verification.

[0072] Preferably, in step S4, the constructed component feature database is stored using MongoDB; each identified component has an independent record created in the database, containing a unique identifier, a complete six-level code, precise coordinates in the original image, various extracted feature vectors, identification confidence score, and processing timestamp information; by establishing a multi-dimensional index, fast retrieval based on encoding, feature similarity, and spatial location is supported.

[0073] The following quality control mechanism is introduced in step S5:

[0074] (1) Architectural syntax checker: Based on traditional dependency parsing, it adds a check for architectural expression norms;

[0075] (2) Building Terminology Consistency Check Module: Construct a mapping table of traditional arcade building professional terms and map synonyms to standard terms;

[0076] (3) Architectural Logic Verifier: Establishes architectural domain-specific rules to verify the correctness of the historical logic described;

[0077] (4) Architectural text readability evaluation algorithm: Based on the characteristics of architectural professional texts, the algorithm sets the index that the density of professional terms does not exceed 30% and the average sentence length does not exceed 25 characters.

[0078] In step S6, the ontology of the knowledge graph defines the following six types of core entities:

[0079] (1) The building entity, recording the basic information of the entire arcade building;

[0080] (2) Component entities, storing detailed attributes of various building components;

[0081] (3) Material entities, describing the physicochemical properties of different building materials;

[0082] (4) Period entities, dividing historical periods and associating them with the characteristics of the era;

[0083] (5) Geographic entities, recording geospatial information;

[0084] (6) Craftsmanship entities, recording traditional and modern construction techniques;

[0085] The relationships between entities include:

[0086] (1) Compositional relationships, describing the overall and partial relationships between the building and its components;

[0087] (2) Positional relationship, recording the spatial distribution of components on the elevation;

[0088] (3) Material relationships, connecting components and the building materials used;

[0089] (4) Due to time constraints, indicate the construction and modification history of the components;

[0090] (5) Technological relationships, explaining the manufacturing techniques of the components;

[0091] (6) Evolutionary relationships, tracing the historical changes in component styles.

[0092] Preferably, in step S6, the knowledge graph integrates multi-source data through a knowledge extraction and fusion process, specifically:

[0093] (1) Automatically extract component entities and basic attributes from the identification results of step S3 to create the initial nodes of the knowledge graph;

[0094] (2) Establish connection edges between nodes by analyzing the spatial adjacency and functional association between components;

[0095] (3) Supplement the historical and cultural information of components by integrating external data sources;

[0096] (4) Use entity alignment technology to identify and merge different representations that point to the same object to ensure the consistency and integrity of knowledge.

[0097] Preferably, in step S6, the knowledge graph has a reasoning rule engine that supports complex knowledge discovery, specifically:

[0098] (1) Reasoning mechanism based on descriptive logic;

[0099] (2) Using temporal reasoning, analyze the evolution path of component styles;

[0100] (3) Spatial reasoning to discover the distribution pattern of components.

[0101] Preferably, in step S6, the knowledge graph has a knowledge service interface that provides diverse access methods, specifically:

[0102] (1) SPARQL query interface;

[0103] (2) Natural language query interface;

[0104] (3) Visual interface;

[0105] (4) RESTful API.

[0106] The system in step S7 adopts an end-to-end processing flow and an asynchronous processing mechanism, supporting parallel processing of batch images and ensuring efficient system operation. Specifically:

[0107] When a user uploads an image of the arcade facade, the system automatically starts the processing pipeline. First, it performs image quality checks and preprocessing. Then, it calls the trained arcade component recognition model to identify the components inside the arcade. Next, it assigns coding information to each component and extracts multi-dimensional features. Subsequently, it generates three levels of component description text and finally stores it in the knowledge graph and updates the index.

[0108] Preferably, the system in step S7 employs the following technical means to optimize system performance:

[0109] (1) Convert the 32-bit floating-point model parameters into 8-bit integer representations using model quantization techniques;

[0110] (2) Use TensorRT to optimize the inference engine and make full use of the parallel computing capabilities of the GPU;

[0111] (3) Implement a distributed architecture design to support multi-node collaborative processing;

[0112] (4) Establish an intelligent caching mechanism to cache frequently accessed feature data and query results.

[0113] Preferably, the system in step S7 continuously improves system performance through a continuous optimization mechanism, specifically:

[0114] (1) Establish a user feedback collection system;

[0115] (2) Organize expert reviews regularly;

[0116] (3) Expand the training dataset;

[0117] (4) Update and optimize the recognition model;

[0118] (5) Improve the content of the knowledge graph.

[0119] After adopting the above technical solution, the present invention has the following technical effects:

[0120] (1) The standardization level has been greatly improved: the standardized expression of arcade components has been realized through the six-level coding system, which has completely solved the problems of confusing terminology and inconsistent classification in traditional research; the standardized coding has laid the foundation for comparative research across regions and periods.

[0121] (2) Revolutionary improvement in identification efficiency: The complete identification, coding and description generation of a single arcade building takes only 30 seconds, which is more than 100 times more efficient than traditional manual methods. This makes large-scale building surveys and digital archiving possible.

[0122] (3) More comprehensive and accurate knowledge expression: The automatically generated three-level descriptive text contains both concise basic information and professional and in-depth detailed analysis, meeting the needs of users at different levels. Standardized language expression facilitates the dissemination and sharing of knowledge.

[0123] (4) Provide a data foundation for subsequent research: Standardized coded data can be directly input into the evolution analysis system of arcade buildings, supporting large-scale research on the evolution of architectural styles; knowledge graphs provide a rich library of components and combination rules for the intelligent design and reconstruction of other related systems.

[0124] (5) Promote the digital transformation of cultural heritage protection: The technical solution provided by this invention can be applied to the digital protection of other types of traditional buildings, accelerating the digital process of the entire cultural heritage protection field. Attached Figure Description

[0125] Figure 1 This is a system architecture diagram of a specific embodiment of the present invention. Detailed Implementation

[0126] To further explain the technical solution of the present invention, the present invention will be described in detail below through specific embodiments.

[0127] refer to Figure 1 As shown, this invention discloses a method for intelligent recognition and language conversion of traditional arcade facade components based on a hierarchical coding system, including:

[0128] Step S1. Construct a standard coding system with six levels: building system level coding, regional subsystem coding, category coding, subcategory coding, specific building coding, and component coding. The component coding adopts a sub-coding composite structure of location-type-attribute. Location and type constitute the basic coding of the component, which is used to reflect the actual location of the component in the arcade (similar to coordinates).

[0129] This step establishes a standard coding system from macro to micro to achieve standardized expression of arcade facade components; the design of the coding system follows the principles of hierarchical progression, unique identification, and easy expansion.

[0130] Specifically, in the above standard coding system:

[0131] (i) Building system-level coding, using CTQL as the top-level identifier for traditional arcade buildings;

[0132] (ii) Regional sub-system coding: Based on the geographical area where the arcade buildings are distributed, MN represents the Minnan sub-system, LN represents the Lingnan sub-system, XJP represents the Singapore sub-system, MLXY represents the Malaysian sub-system, etc.

[0133] (iii) Group coding, used to distinguish regional characteristics. For example, under the Minnan subgroup, XM represents the Xiamen group, QZ represents the Quanzhou group, and ZZ represents the Zhangzhou group. Under the Lingnan subgroup, CS represents the Chaoshan group, GZ represents the Guangzhou group, HN represents the Hainan group, and GX represents the Guangxi group, etc.

[0134] (iv) Subgroup coding, accurate to specific historical streets, for example, LCQ-ZS represents Quanzhou Licheng District-Zhongshan Road, AX-LY represents Anxi-Luoyang Street, etc.;

[0135] (v) Specific building code: assign a number of digits to each arcade building (usually three digits, but not limited to this).

[0136] (vi) Component coding, the specific sub-codes are:

[0137] (1) Location: Based on the construction logic of the arcade facade, R represents the roof part, B represents the body part, and F represents the base part;

[0138] (2) Type, used to reflect the detailed classification of the location, for example: in the roof part, R01 represents the parapet wall R01, R02 represents the gable, R03 represents the eaves, etc.; in the body part, B01 represents the arcade column, B02 represents the arcade lintel, B03 represents the second floor wall, B04 represents the window, B05 represents the door, etc.; in the base part, F01 represents the column base, F02 represents the ground, F03 represents the steps, etc.

[0139] (3) Attributes, including at least one of the material, age and functional characteristics of the component. M represents material, T represents age and U represents function. For example, in the material attribute code, M01 represents wood, M02 represents brick and stone, M03 represents concrete, M04 represents metal and M05 represents ceramic, etc. In the age attribute code, T01 represents early Qing Dynasty, T02 represents mid-Qing Dynasty, T03 represents late Qing Dynasty, T04 represents early modern period, T05 represents late modern period and T06 represents contemporary period. In the functional attribute code, U01 represents load-bearing, U02 represents decoration, U03 represents maintenance, U04 represents use and U05 represents connection, etc.

[0140] By combining the above-mentioned codes at various levels, each component of the arcade can obtain a globally unique standard code. For example, the standard code CTQL-MN-QZ-LCQ-ZS-028-R02-M03-T04-U02 can fully express the complex information of "traditional arcade - Minnan sub-system - Quanzhou group - Lichengzi group - Zhongshan Road - Building No. 028 - gable component - concrete material - early modern period - decorative function".

[0141] To ensure the understandability and readability of the code, a bidirectional mapping dictionary between the code and natural language can be established according to certain rules. This dictionary can be used to include professional terms related to arcade buildings in existing technologies. Each code entry can contain information such as a standard Chinese name, a set of local colloquialisms, an English equivalent, a functional description, and the usage context. This allows for seamless conversion between machine code and human language.

[0142] Step S2. Systematically collect images according to the preset image acquisition standards to obtain sufficient images of the arcade facade. Then, preprocess the collected images and manually annotate them to build an annotated dataset.

[0143] This step aims to establish a standardized image acquisition and preprocessing process to obtain high-quality images of the arcade facade, which will then be annotated by experienced and knowledgeable personnel to obtain an annotated dataset that can be used for training. This will ensure that the model trained in subsequent steps can accurately identify facade components.

[0144] Specifically, in step S2 above, the preset image acquisition standards are as follows:

[0145] (1) When taking pictures, the axis of the camera lens should be perpendicular to the main plane of the building facade to ensure that a standard orthographic projection image is obtained;

[0146] (2) The shooting distance should be controlled at 1.5 to 2 times the building height to ensure that the facade is fully captured in the shot while maintaining sufficient detail clarity;

[0147] (3) Use a standard lens with an equivalent focal length of 35-50 mm to avoid perspective distortion caused by a wide-angle lens;

[0148] (4) Choose to shoot under uniform lighting conditions such as early morning / evening on cloudy or sunny days to avoid strong shadows interfering with component recognition;

[0149] (5) The image resolution should be no less than 4000×3000 pixels to ensure the recognizability of the details of the parts.

[0150] In step S2 above, during systematic data collection: a certain number of arcade buildings are selected from different regions with arcade architecture, such as the Minnan region, Lingnan region, and Hainan region, for photography. For example, in the Minnan region, 150 arcade buildings are selected from 6 representative historical blocks, including Quanzhou Zhongshan Road, West Street, East Street, Xiamen Zhongshan Road, Gulangyu Island, and Zhangzhou Ancient City. In the Lingnan region, 150 buildings are selected from 10 arcade blocks, including Guangzhou Beijing Road, Shangxiajiu Street, Jiangmen Thirty-Three Market Street, Haikou Zhongshan Road, and Wenchang Wennan Old Street. At least 3 panoramic photos of the front facade and 10-15 close-up photos of each arcade building are collected. The close-up photos focus on characteristic components such as gables, column heads, arches, and decorative patterns. Each arcade building is photographed at multiple time points to obtain diverse lighting conditions.

[0151] In step S2 above, the preprocessing process includes:

[0152] First, perspective correction is performed by detecting vertical and horizontal lines in the facade and calculating the perspective transformation matrix to correct the tilted image to a standard orthographic projection.

[0153] Secondly, color standardization is performed by using a standard color chart placed during shooting and calculating a color correction matrix to ensure that images shot under different lighting conditions have consistent color performance.

[0154] Then, adaptive histogram equalization is used to enhance image contrast, making the details of parts in the shadow areas clearer;

[0155] Finally, scale normalization is performed to adjust all images to a uniform pixel density, which facilitates subsequent feature extraction and comparison.

[0156] In step S2 above, when constructing the labeled dataset, at least 5 architectural history experts with more than 10 years of experience are invited to participate in the labeling. The preprocessed arcade facade image is labeled on the labeling software. The component outline is accurately selected on the image, the component type is selected, the material properties and preservation status are labeled, and finally the labeled dataset is obtained. In order to ensure the quality of the labeling, a cross-validation mechanism is adopted during the labeling. Each component is labeled independently by at least 2 experts. When there is a disagreement, a consensus is reached through discussion.

[0157] Step S3. Construct a model for recognizing arcade components. Train and recognize arcade components using existing datasets and labeled datasets; refer to... Figure 1 As shown, the network architecture design of the arcade component recognition model adopts Mask R-CNN as the basic framework, which can simultaneously complete object detection, classification, and instance segmentation tasks. The backbone network of the arcade component recognition model uses ResNet-101 combined with a Feature Pyramid Network (FPN). Through bottom-up and top-down feature fusion, it can effectively extract multi-scale features and adapt to the size differences of arcade components, from small decorations to large gables. Furthermore, the basic framework of Mask R-CNN innovatively adds a spatial relationship module. The spatial relationship module models the spatial constraint relationships between different components through a Graph Convolutional Network (GCN), such as "the gable must be located at the top of the facade" and "the column base must be located at the bottom of the column." Thus, the introduction of spatial relationships can significantly improve the ability of the arcade component recognition model to judge the rationality of component positions, thereby reducing false recognitions.

[0158] Specifically, in step S3 above, the training strategy for the arcade component recognition model adopts a three-stage progressive method:

[0159] In the first stage, the COCO dataset was used for pre-training, allowing the arcade component recognition model to learn basic visual feature extraction capabilities.

[0160] In the second stage, the architectural image dataset is used for domain adaptation, enabling the arcade component recognition model to understand the unique visual patterns of architectural components.

[0161] In the third stage, the labeled dataset from step S2 is used for fine-tuning, with a focus on optimizing the ability of the arcade component recognition model to recognize unique arcade components.

[0162] In step S3 above, the training process of the arcade component recognition model employs several optimization measures to address the technical challenges in training, specifically:

[0163] (1) Use the focus loss function to solve the problem of component category imbalance, and give higher loss weight to rare component categories with fewer numbers;

[0164] (2) Introduce an online difficult example mining mechanism, and pre-set easily confused component pairs in the arcade components according to practical experience (such as "Baroque pediment" and "Neoclassical pediment"). During the training process, automatically identify these component pairs and increase the training frequency of these component pairs according to the preset number or proportion value.

[0165] (3) By using data augmentation techniques (including random cropping, color jitter, geometric transformation, etc.), the training samples are expanded by 3 times to improve the generalization ability of the arcade component recognition model.

[0166] In step S3 above, a comprehensive post-processing mechanism was established to ensure the reasonableness of the recognition results, specifically:

[0167] (1) Filter the detection boxes output by the arcade component recognition model by confidence level and retain only the results with a confidence level higher than 0.8;

[0168] (2) Use non-maximum suppression algorithm to eliminate overlapping detection boxes and avoid the same part being repeatedly identified;

[0169] (3) Construct a logical rule base based on architectural common sense, and verify the rationality of the identification results, such as checking whether the relative positional relationship of the components conforms to the architectural construction logic.

[0170] Step S4. For the identification results of step S3, assign coding information of the standard coding system and extract multi-dimensional features, perform component association analysis, and construct a component feature database;

[0171] The process of assigning standard codes is as follows:

[0172] First, basic codes are determined based on the identified component types and their locations on the arcade facade. Then, by analyzing the visual features of the components, such as decorative complexity, geometric shape, and texture, their possible construction dates are inferred. Finally, geographical information from the image capture location is combined to automatically fill in region-related codes (i.e., regional subgroup codes, class group codes, subclass group codes, and specific building codes). A pre-trained material classification network is used to determine the material properties of the components. This material classification network employs a general deep learning architecture and has been specifically trained and optimized for the unique material systems of traditional arcade buildings (including Minnan red bricks, cement plaster, washed stone, and wood carvings), constructing a dataset of over 1000 arcade building materials to accurately identify the unique material types of arcade buildings.

[0173] Specifically, the multidimensional features extracted in step S4 above cover the following dimensions:

[0174] (1) Geometric features, including the absolute dimensions and relative proportions of the components, are determined by referring to the known dimensions of the building or the dimensions of standard components;

[0175] (2) Visual features are extracted using a variety of operators, including local binary mode for texture analysis, scale-invariant feature transformation for key point detection, and HSV color histogram for color feature extraction.

[0176] (3) Style features are extracted through a deep style network to generate a 128-dimensional feature vector, which can effectively distinguish different decorative styles such as traditional Minnan, Baroque, Neoclassical, and Art Deco.

[0177] (4) Material characteristics are determined by a pre-trained material classification network, which is based on general texture recognition technology but has been adapted to the domain:

[0178] a. A dedicated dataset for traditional arcade building materials was constructed, containing 87 common arcade building materials, covering traditional materials (Minnan red bricks, bluestone, fir wood, etc.) and modern materials (cement, reinforced concrete, washed stone, etc.).

[0179] b. Material aging identification is introduced, which distinguishes the state of new and old materials by characteristics such as texture roughness and color fading, to assist in the determination of age;

[0180] c. Establish a knowledge base linking materials and eras, such as rules like "the stone washing technique was introduced in the 1920s" and "mosaic decoration was popular in the 1950s," to improve the accuracy of era judgments;

[0181] d. A material coding mapping table was established to directly output standardized material codes (M01-M05, etc.), achieving seamless integration with the six-level coding system.

[0182] In step S4 above, the steps for component association analysis are as follows:

[0183] (1) Spatial relationship modeling: Based on the bounding box coordinates of the component recognition results, calculate the spatial distance matrix and relative positional relationships (up, down, left, right, containment, adjacent, etc.) between components, and establish a component spatial relationship diagram;

[0184] (2) Statistical co-occurrence probability: Based on a large number of samples, the co-occurrence frequency of different component types is statistically analyzed, and a component co-occurrence probability matrix is ​​constructed, such as the probability of "Baroque pediment" and "Corinthian capital" appearing at the same time;

[0185] (3) Co-occurrence pattern mining: Statistical analysis of the co-occurrence frequency of different component types in the arcade sample, and the use of association rule mining algorithm (Apriori algorithm, minimum support 0.3, minimum confidence 0.7) to discover the combination pattern of strongly associated components;

[0186] (4) Verification of consistency of age: Based on the inferred age attributes of components, examine the consistency of age of different components of the same building and identify possible later renovations;

[0187] (5) Construct a component relationship network: Use components as nodes and association strength as edge weights to construct a weighted undirected graph for subsequent architectural style analysis and integrity verification.

[0188] The component association analysis described above is a key innovation of this step. This invention not only identifies individual components but also analyzes the combination patterns of different components on the same facade. Statistical analysis revealed strong correlations between certain component combinations; for example, the co-occurrence probability of "red brick square columns" and "arched arcades" reached 85%, and the co-occurrence probability of "Corinthian capitals" and "Baroque pediments" reached 73%. These association rules not only validated the rationality of the identification results but also provided important evidence for subsequent architectural style evolution analysis.

[0189] In step S4 above, the constructed component feature database is stored using MongoDB, fully utilizing its excellent support for unstructured data. Each identified component has an independent record created in the database, containing a unique identifier, complete six-level encoding, precise coordinates in the original image, extracted feature vectors, identification confidence score, processing timestamp, and other complete information. By establishing a multi-dimensional index, rapid retrieval based on encoding, feature similarity, spatial location, and other methods is supported.

[0190] Step S5. Based on the recognition results of Step S3 and the encoding information of Step S4, three levels of standardized component description text are automatically generated to achieve accurate conversion from visual features to natural language. These include a basic layer description of no more than 20 characters, a standard layer description of no more than 50 characters, and a professional layer description of no more than 100 characters. The basic layer description organizes the language in the order of "component name - material properties - location information" and is mainly used for rapid identification and retrieval. The standard layer adds more elements such as era characteristics, regional style, size specifications, and decorative features on the basis of the basic layer. When generating the description, key information is automatically extracted from the feature data and the language is organized according to the preset word order rules. The professional layer adds deeper information such as process technology, cultural connotation, and historical value on the basis of the standard layer to generate a description text with academic depth. For example, for a balcony column component, the basic layer generates a concise and clear description such as "wooden square column - east side of the balcony," while the standard layer generates a detailed description such as "mid-Qing Dynasty Quanzhou-style wooden square column, 20×20 cm in cross-section, with simple molding decoration on the column capital, well preserved." For a Baroque pediment component, the professional layer generates an in-depth description such as "This pediment was built in 1923, using cement plastering techniques, decorated with scroll patterns and wreath reliefs in the center, and symmetrical scroll decorations on both sides. The overall style presents a typical Nanyang Baroque style, reflecting the historical characteristics of commercial buildings in the Minnan region influenced by Western culture in the early 20th century, and has important architectural history research value."

[0191] In step S5 above, the process of generating descriptive text integrates two methods: rule templates and deep learning. For common standard components, predefined language templates are used to ensure the standardization and accuracy of the description. For complex or special components, a large language model that has been fine-tuned in the field of architecture is called. Domestic mainstream models such as Alibaba Tongyi Qianwen, Baidu Wenxin Yiyan, Zhipu ChatGLM, or DeepSeek are used as the basis. LoRA fine-tuning is performed through 100,000 collected architectural professional documents to enhance the model's understanding of architectural professional knowledge and generate more flexible and expressive descriptive text.

[0192] Step S5 above introduces a quality control mechanism to ensure the professionalism and readability of the generated description. While employing some existing natural language processing techniques, this invention innovatively combines and improves these techniques, specifically adapting them to the characteristics of architectural descriptive texts. It constructs an architectural terminology database, a domain-specific logical rule base, and readability evaluation standards, forming a complete quality control system for architectural descriptive texts. Specifically:

[0193] (1) Architectural syntax checker: Based on traditional dependency parsing, it adds architectural expression standard checks, such as spatial logic verification of "the column head is located above the column body", to ensure that the description conforms to common sense of architectural construction.

[0194] (2) Building Terminology Consistency Check Module: A mapping table of traditional arcade professional terms containing more than 3,000 entries was constructed, and synonyms such as "longevity beam / lintel beam / door lintel beam" were uniformly mapped to standard terms to ensure the standardization of descriptions;

[0195] (3) Architectural Logic Verifier: More than 50 architectural field-specific rules have been established, such as "Baroque decoration appeared no earlier than 1890" and "RC structure appeared no earlier than 1910", to verify the correctness of the historical logic described.

[0196] (4) Architectural text readability evaluation algorithm: Based on the characteristics of architectural professional texts, indicators such as the density of professional terms not exceeding 30% and the average sentence length not exceeding 25 characters are set to ensure that the professionalism is maintained while being easy to understand.

[0197] Step S6. Integrate the recognition results from Step S3, the encoding information from Step S4, and the descriptive text from Step S5 to construct a structured knowledge graph of the arcade facade components, providing support for complex queries and knowledge reasoning; the ontology of the knowledge graph defines six core entities:

[0198] (1) The building entity, recording the basic information of the entire arcade building;

[0199] (2) Component entities, storing detailed attributes of various building components;

[0200] (3) Material entities, describing the physicochemical properties of different building materials;

[0201] (4) Period entities, dividing historical periods and associating them with the characteristics of the era;

[0202] (5) Geographic entities, recording geospatial information;

[0203] (6) Craftsmanship entities, recording traditional and modern construction techniques;

[0204] The relationships between entities include:

[0205] (1) Compositional relationships, describing the overall and partial relationships between the building and its components;

[0206] (2) Positional relationship, recording the spatial distribution of components on the elevation;

[0207] (3) Material relationships, connecting components and the building materials used;

[0208] (4) Due to time constraints, indicate the construction and modification history of the components;

[0209] (5) Technological relationships, explaining the manufacturing techniques of the components;

[0210] (6) Evolutionary relationships, tracing the historical changes in component styles. Each relationship has a clear attribute definition, such as the composition relationship which includes attributes such as quantity, position, and importance.

[0211] In step S6 above, the knowledge graph integrates multi-source data through knowledge extraction and fusion processes, specifically:

[0212] (1) Automatically extract component entities and basic attributes from the identification results of step S3 to create the initial nodes of the knowledge graph;

[0213] (2) Establish connection edges between nodes by analyzing the spatial adjacency and functional association between components;

[0214] (3) Supplement the historical and cultural information of the components by integrating external data sources such as historical documents and expert knowledge;

[0215] (4) Use entity alignment technology to identify and merge different representations that point to the same object to ensure the consistency and integrity of knowledge.

[0216] In step S6 above, the knowledge graph has a reasoning rule engine that supports complex knowledge discovery, specifically:

[0217] (1) Based on the reasoning mechanism of descriptive logic, implicit information can be derived from explicit knowledge; for example, the rule "if the component is decorated with Baroque style and is located in the Minnan region, the construction period may be 1890-1930" can help infer the missing date information.

[0218] (2) Temporal reasoning to analyze the evolution path of component styles, such as the evolutionary sequence of "Doric column → Corinthian column → Composite column";

[0219] (3) Spatial reasoning to discover the distribution pattern of components, such as "the arcade buildings in commercial streets mostly adopt an open design, while the arcade buildings in residential areas mostly adopt a closed design".

[0220] In step S6 above, the knowledge graph has a knowledge service interface that can provide diverse access methods, specifically:

[0221] (1) SPARQL query interface, which supports complex graph query operations, and professional users can write precise query statements;

[0222] (2) Natural language query interface, which uses semantic parsing technology to convert users' natural language questions into structured queries;

[0223] (3) Visual interface, using force-guided diagrams and other technologies to intuitively display the relationship network between components;

[0224] (4) RESTful API, which supports the integration of third-party applications and provides standardized data output in JSON format.

[0225] Step S7. System Integration and Practical Application

[0226] The processing capabilities in steps S2 to S6 above are integrated into a complete system, which specifically includes: image preprocessing capability (step S2), intelligent component recognition capability (step S3), automatic encoding and feature extraction capability (step S4), multi-level description generation capability (step S5), and knowledge graph construction capability (step S6). End-to-end automated processing is achieved through a unified system architecture.

[0227] Specifically, the system in step S7 above adopts an end-to-end processing flow and an asynchronous processing mechanism to support parallel processing of batch images and ensure efficient system operation. Specifically:

[0228] When a user uploads an image of the arcade facade, the system automatically starts the processing pipeline. First, it performs image quality checks and preprocessing. Then, it calls the trained arcade component recognition model to identify the components inside the arcade. Next, it assigns coding information to each component and extracts multi-dimensional features. Subsequently, it generates three levels of component description text and finally stores it in the knowledge graph and updates the index.

[0229] The system in step S7 above employs the following technical means to optimize system performance:

[0230] (1) The 32-bit floating-point model parameters are converted into 8-bit integer representations by model quantization technology, which improves the inference speed by 3 times while maintaining the recognition accuracy;

[0231] (2) TensorRT is used to optimize the inference engine and make full use of the parallel computing power of the GPU. The complete processing time of a single image is controlled within 30 seconds.

[0232] (3) Implement a distributed architecture design to support multi-node collaborative processing and meet the needs of large-scale data processing;

[0233] (4) Establish an intelligent caching mechanism to cache frequently accessed feature data and query results, which significantly improves the system response speed.

[0234] The system in step S7 above continuously improves its performance through a continuous optimization mechanism:

[0235] (1) Establish a user feedback collection system to record identification errors and improvement suggestions found during use;

[0236] (2) Regularly organize expert reviews to conduct professional evaluations of the system output results and formulate optimization plans;

[0237] (3) Expand the training dataset, especially by adding newly discovered rare component samples and edge cases;

[0238] (4) Update and optimize the recognition model, use new data for incremental learning, and improve the ability to handle complex situations;

[0239] (5) Improve the content of the knowledge graph and continuously supplement new component relationships and reasoning rules.

[0240] In practical applications, a comprehensive evaluation of the system's performance was achieved through large-scale verification testing. Specifically, 10 representative historical arcade blocks were selected, and a complete identification, coding, and description generation test was conducted on 1000 arcade buildings. Test results showed that the component recognition accuracy reached 92.5%, with the accuracy rate for common component types exceeding 95%. The coding accuracy reached 96%, with the vast majority of components accurately assigned to six-level codes. The reasonableness of the generated descriptions was assessed by experts to be over 90%, and the generated text conformed to professional expression standards.

[0241] This invention achieves significant technological innovations in the following aspects:

[0242] (I) Innovation of the six-level hierarchical coding system

[0243] For the first time, a complete coding system has been established, from the building system down to specific components, with each component receiving a globally unique identifier. This coding system not only includes location and type information but also integrates multi-dimensional attributes such as materials, age, and function, achieving a structured expression of architectural knowledge. Compared to traditional simple classification methods, the coding system of this invention has stronger expressive power and scalability.

[0244] (II) Innovation of specialized component identification model

[0245] To address the unique characteristics of arcade facade components, a spatial relationship module was innovatively added to the standard object detection network. This module models the spatial constraints between components using a graph convolutional network, significantly improving the accuracy and plausibility of the identification. This deep learning approach, which integrates domain knowledge, provides new insights for object recognition in other professional fields.

[0246] (III) Innovation in intelligent description generation mechanism

[0247] This system enables the automatic conversion of image recognition results into standardized text descriptions, resolving the long-standing "semantic gap" problem in architectural digitization. By combining rule-based templates and deep learning, it ensures both the standardization of descriptions and their flexibility and expressiveness. The three-tiered description system meets the needs of different users.

[0248] (IV) Knowledge Graph-Driven Intelligent Analysis Innovation

[0249] A complete knowledge graph of arcade building components was constructed, which not only stores static component information but also supports dynamic knowledge discovery through a reasoning rule engine. This knowledge organization method breaks through the limitations of traditional databases and provides a powerful analytical tool for architectural history research and conservation.

[0250] (V) Domain-specific innovation in general technologies

[0251] This invention innovatively adapts general computer vision technology to a specific domain:

[0252] -The material classification is not a simple application of existing models, but rather a construction of a knowledge system of traditional arcade building materials, including a complete mapping of material types, process characteristics, and era characteristics;

[0253] - A reasoning chain of "visual features → material identification → historical dating → code generation" has been established, realizing the deep integration of technology and domain knowledge;

[0254] - Through training with 1000+ dedicated samples, the general technology is able to identify traditional processes such as washing stones, chopped stones, and polished stones;

[0255] - The material identification results directly serve the generation of Level 6 codes and the inference of building age, reflecting the innovation of system integration.

[0256] Through the above-described solution, the present invention has the following significant advantages over the prior art:

[0257] (1) The degree of standardization has been greatly improved.

[0258] The six-level coding system has enabled the standardized expression of arcade components, completely solving the problems of confusing terminology and inconsistent classification in traditional research; the standardized coding has laid the foundation for comparative research across regions and periods.

[0259] (2) Revolutionary improvement in recognition efficiency

[0260] The complete identification, coding, and description generation of a single arcade building takes only 30 seconds, more than 100 times more efficient than traditional manual methods. This makes large-scale building surveys and digital archiving possible.

[0261] (3) Knowledge expression is more comprehensive and accurate

[0262] The automatically generated three-tiered descriptive text provides both concise basic information and in-depth professional analysis, meeting the needs of users at different levels. Standardized language facilitates the dissemination and sharing of knowledge.

[0263] (4) Provide a data foundation for subsequent research

[0264] Standardized coded data can be directly input into the evolution analysis system of arcade buildings, supporting large-scale research on the evolution of architectural styles; knowledge graphs provide a rich library of components and combination rules for the intelligent design and reconstruction of other related systems.

[0265] (5) Promote the digital transformation of cultural heritage protection

[0266] The technical solution provided by this invention can be extended to the digital preservation of other types of traditional buildings, accelerating the digitalization process of the entire cultural heritage preservation field.

[0267] The following illustrates specific embodiments of the present invention.

[0268] Taking the Li Gan Sesame Oil Shop arcade at No. 287, Section 1, Dihua Street, Taipei City as an example, this invention will be explained in detail in terms of its implementation process and effects. The building housing the Li Gan Sesame Oil Shop is located at No. 287, Section 1, Dihua Street, Datong District, Taipei City. Built in the 1870s, it is a Western-style arcade building with approximately 150 years of history. The first owner purchased the building around the 1920s and moved the shop from Hsinchu to Dihua Street, specializing in sesame oil and bitter tea oil. Their sesame oil received an award from the Hsinchu Prefecture Industrial Advanced Association in 1926. The building is a two-story arcade structure with a facade approximately 4.5 meters wide, decorated with Minnan bricks and cement, exhibiting typical arcade features from the early late Qing Dynasty, and is an important historical building on Dihua Street.

[0269] I. Image Acquisition and Preprocessing

[0270] Using a DSLR camera with a 35mm F1.8 lens, the photos were taken from a distance of 7 meters from the building facade. A cloudy day with even lighting was chosen to avoid strong shadows. Three high-resolution 4000×3000 pixel images of the facade were captured, along with 12 close-up shots, focusing on key features such as the cement plaque-style shop sign, vase-shaped decorative railings, round columns, arched pavilion legs, and wooden window frames.

[0271] In the image preprocessing stage, perspective correction was first performed by detecting the vertical lines of the facade, eliminating a tilt angle of approximately 2.3 degrees. Color correction was then performed using a standard color chart to ensure accurate reproduction of the reddish-brown of the Minnan bricks and the grayish-white of the cement decorations. Adaptive histogram equalization enhanced the visibility of details in the shadowed areas inside the pavilion's legs. Finally, all images were uniformly adjusted to a standard resolution of 3000×2000 pixels.

[0272] II. Component Identification and Automatic Coding

[0273] The pre-processed image was input into the trained arcade component recognition model, and the system completed the entire recognition task within 24 seconds. The main recognized components and their automatically generated codes are as follows:

[0274] The cement plaque shop sign was identified by the code "CTQL-TW-NTW-TBS-DH-287-B07-M03-T03-U02", indicating that it is a component of the shop sign located at No. 287 Dihua Street, Taipei, in the arcade building. It is made of concrete, was built in the late Qing Dynasty, and served a decorative function. The identification confidence level is 0.95.

[0275] The two circular columns were coded "CTQL-TW-NTW-TBS-DH-287-F04a-M02-T03-U01" and "CTQL-TW-NTW-TBS-DH-287-F04b-M02-T03-U01" respectively, accurately identifying their masonry structure and load-bearing function. The identification confidence level was 0.96.

[0276] The arched pavilion's base was coded "CTQL-TW-NTW-TBS-DH-287-F05-M02-T03-U01", and the system correctly identified its arched structure as being constructed of Minnan bricks. The recognition confidence level was 0.94.

[0277] The vase-shaped decorative railing was assigned the code "CTQL-TW-NTW-TBS-DH-287-B08-M03-T03-U02", which accurately reflects its decorative attributes.

[0278] The three wooden windows on the second floor were assigned codes: the central window "CTQL-TW-NTW-TBS-DH-287-B04a-M01-T03-U04" (opens to the left and right), and the side windows "CTQL-TW-NTW-TBS-DH-287-B04b-M01-T03-U04" and "CTQL-TW-NTW-TBS-DH-287-B04c-M01-T03-U04" (opens up and down).

[0279] III. Feature Extraction Results

[0280] The geometric features extracted by the system show that the building facade is 4.5 meters wide, with a total height of 8.5 meters across two floors, and the pavilion's base is 2.8 meters deep. A circular column, 0.35 meters in diameter, divides the facade into three equal parts. A cement plaque, 1.8 meters wide and 0.6 meters high, is located in the center of the facade. The vase-shaped balustrade is 0.9 meters high, with each unit measuring 0.3 meters wide.

[0281] Visual feature analysis shows that the overall facade color is predominantly reddish-brown, characteristic of Minnan bricks (75%), with cement decorations in grayish-white (15%) and wood parts in dark brown (10%). Texture features reveal regular brick joint patterns and a smooth cement surface. Style feature vectors, extracted using a deep network, scored 0.87 in the imitation Western-style dimension and 0.72 in the traditional Minnan style dimension, reflecting a blend of Chinese and Western characteristics.

[0282] The material identification accurately determined the use of traditional materials such as Minnan red bricks, cement plaster, and wooden doors and windows, which are consistent with the architectural techniques of the late Qing Dynasty.

[0283] The material recognition module, based on a training set specifically for arcade buildings, accurately identifies various materials and their characteristics.

[0284] - Southern Fujian Red Brick (Confidence 0.96): Identified as a traditional kiln-fired red brick by its unique orange-red hue (HSV hue 15-25°) and regular brick joint pattern;

[0285] - Cement plaster decoration (confidence level 0.94): The fine texture of early cement materials can be identified, and there are slight cracks on the surface, which is consistent with the cement craftsmanship characteristics of the late Qing Dynasty and early Republic of China.

[0286] - Traditional Chinese fir doors and windows (confidence level 0.92): Based on the vertical grain direction and surface patina characteristics, they are identified as Fujian fir wood, and are in good condition.

[0287] These material identification results automatically generate corresponding codes (M02, M03, M01) and are linked to the architectural technology system of the late Qing Dynasty, providing a reliable basis for dating buildings.

[0288] IV. Standardized Description Generation

[0289] Basic floor description: A two-story brick and wood structure resembling a Western-style oil shop arcade;

[0290] Standard floor description: Late Qing Dynasty Taipei Dihua Street arcade, Minnan brick facade, arched pavilion legs, cement plaque signboard, vase-shaped railing decoration;

[0291] Professional Description: Built in the 1870s, this arcade-style building, reminiscent of Western-style buildings, is now the site of the Li Gan Sesame Oil Shop, with a history of approximately 150 years. The facade is constructed of Minnan red bricks. Two round columns on the first floor divide the facade into three equal parts, with flat stone beams supporting the columns. The pavilion's base employs traditional Minnan brick arch techniques, showcasing exquisite late Qing Dynasty construction craftsmanship. The central cement plaque and vase-shaped balustrade decorations reflect the architectural characteristics of a blend of Chinese and Western styles characteristic of traditional arcade buildings. The second-floor wooden windows feature a central left-right opening and side-up-down openings, reflecting traditional living habits. This building bears witness to the historical transformation of Dihua Street from a late Qing Dynasty commercial port to a modern commercial street.

[0292] V. Knowledge Graph Construction

[0293] The system automatically created the "Li Gan Sesame Oil Shop Arcade" architectural node and established its compositional relationships with 26 component nodes. Each component node is associated with corresponding material nodes, period nodes, and process nodes. For example, the arched pavilion foot node is connected to the "Southern Fujian Brick" node through the "Materials Used" relationship, to the "Late Qing Dynasty" node through the "Construction Period" relationship, and to the "Brick Arch Construction" node through the "Process Adopted" relationship.

[0294] Through knowledge-based reasoning, the system discovered that the building shares similarities with other late Qing Dynasty commercial buildings on Dihua Street in its imitation Western-style architecture, inferring that this style was popular in Taipei after the city opened to foreign trade in the late Qing Dynasty. Simultaneously, the system identified the unique functional characteristics of oil shop buildings, such as the need for large storage spaces and ventilation, which explains the rationality of its internal spatial layout.

[0295] VI. Verification and Evaluation

[0296] Five architectural history experts were invited to evaluate the system's output. The component recognition accuracy score was 93%, with experts believing the system accurately identified all major components, particularly the arched pavilion legs and vase-shaped railings. The coding accuracy score was 96%, with the coding system accurately reflecting the building's historical characteristics. The descriptive text reasonableness score was 91%, with experts believing the generated description accurately reflected the building's historical value and architectural features.

[0297] The entire processing takes 24 seconds, including 4 seconds for image preprocessing, 14 seconds for part recognition, 4 seconds for encoding and feature extraction, and 2 seconds for description generation and knowledge graph update. Compared to the 2-3 hours required by traditional manual methods, this represents an efficiency improvement of over 100 times.

[0298] VII. Statistics on System Application Effectiveness

[0299] In a large-scale test involving 1,000 arcade buildings, the system demonstrated stable and high-performance characteristics.

[0300] - Component identification accuracy: 92.5% (common component types >95%, rare component types >85%)

[0301] - Encoding accuracy: 96% (98% accuracy for location encoding, 94% accuracy for attribute encoding)

[0302] - Description generation accuracy rate: 90% (Basic layer 95%, Standard layer 92%, Professional layer 85%)

[0303] - Average processing time: 28.3 seconds per building (fastest 22 seconds, slowest 35 seconds)

[0304] - Knowledge graph size: 15,000+ nodes, 30,000+ relationship edges

[0305] - System availability: 99.5% (stable operation 24 / 7)

[0306] The practical application of the method in the arcade of Li Gan Sesame Oil Shop fully verifies the effectiveness and practicality of the invention. The system not only accurately identifies various traditional architectural components, such as arched pavilion legs and vase-shaped railings, but also achieves intelligent conversion from images to knowledge through standardized coding and multi-level description. In particular, the system's accurate grasp of the cultural characteristics of traditional arcade architecture is demonstrated in its identification of the combination of late Qing Dynasty imitation Western-style buildings and traditional Minnan architectural elements. Furthermore, the construction of a knowledge graph reveals the common characteristics of traditional commercial buildings on Dihua Street, providing digital tools to support comparative research on traditional arcade architecture.

[0307] The above embodiments and figures are not intended to limit the product form and style of the present invention. Any appropriate changes or modifications made by those skilled in the art should be considered as not departing from the patent scope of the present invention.

Claims

1. A method for intelligent recognition and language conversion of traditional arcade facade components based on a hierarchical coding system, characterized in that... include: Step S1. Construct a standard coding system with six levels: building system level coding, regional subsystem coding, category coding, subcategory coding, specific building coding, and component coding. Component coding adopts a position-type-attribute sub-coding composite structure, with position and type constituting the basic coding of the component. Step S2. Systematically collect images according to the preset image acquisition standards to obtain sufficient images of the arcade facade. Then, preprocess the collected images of the arcade facade and manually annotate them to build an annotated dataset. Step S3. Construct a model for recognizing arcade components. Train and recognize arcade components using existing datasets and labeled datasets. The network architecture of the arcade component recognition model adopts Mask R-CNN as the basic framework. Its backbone network uses ResNet-101 combined with the Feature Pyramid Network (FPN). Multi-scale features are extracted through bottom-up and top-down feature fusion to adapt to the size differences of arcade components, from small decorations to large gables. In addition, a spatial relationship module is added to the basic framework of Mask R-CNN. The spatial relationship module models the spatial constraint relationships between different components through graph convolutional networks. Step S4. For the identification results of step S3, assign coding information of the standard coding system and extract multi-dimensional features, perform component association analysis, and construct a component feature database; Step S5. Based on the recognition results of step S3 and the encoding information of step S4, automatically generate standardized component description texts at three levels to achieve accurate conversion from visual features to natural language. These include a basic layer description of no more than 20 characters, a standard layer description of no more than 50 characters, and a professional layer description of no more than 100 characters. Step S6. Integrate the recognition results of step S3, the encoding information of step S4, and the description text of step S5 to construct a structured knowledge graph of the arcade facade components; Step S7. Integrate the processing capabilities in steps S2 to S6 into a complete system, which specifically includes: realizing the image preprocessing capability of step S2, realizing the intelligent component recognition capability of step S3, realizing the automatic encoding and feature extraction capability of step S4, realizing the multi-level description generation capability of step S5, and realizing the knowledge graph construction capability of step S6, thereby realizing end-to-end automated processing through a unified system architecture.

2. The method for intelligent recognition and language conversion of traditional arcade facade components based on a hierarchical coding system as described in claim 1, characterized in that... The standard coding system in step S1 is specifically as follows: (i) Building system-level coding, using CTQL as the top-level identifier for traditional arcade buildings; (ii) Regional sub-codes are set according to the geographical areas where the arcade buildings are distributed; (iii) Group coding, used to distinguish regional characteristics; (iv) Subgroup coding, accurate to specific historical streets; (v) Specific building codes, assigning a number of digits to each arcade building; (vi) Component coding, the specific sub-codes are: (1) Location: Based on the construction logic of the arcade facade, R represents the roof part, B represents the body part, and F represents the base part; (2) Type, used to reflect the detailed classification of location; (3) Attributes, including at least one of the component’s material, age and functional characteristics, with M representing material, T representing age and U representing function.

3. The method for intelligent recognition and language conversion of traditional arcade facade components based on a hierarchical coding system as described in claim 1, characterized in that... In step S2: The preset image acquisition standards are as follows: (1) When taking the picture, the axis of the camera lens should be perpendicular to the main plane of the building facade; (2) The shooting distance should be controlled at 1.5 to 2 times the building height; (3) Use a standard lens with an equivalent focal length of 35-50 mm; (4) Select uniform lighting conditions for shooting; (5) The image resolution must be no less than 4000×3000 pixels; The preprocessing process includes: First, perspective correction is performed by detecting vertical and horizontal lines in the facade and calculating the perspective transformation matrix to correct the tilted image to a standard orthographic projection. Secondly, color standardization is performed by using a standard color chart placed during shooting and calculating a color correction matrix. Then, adaptive histogram equalization is used to enhance image contrast; Finally, scale normalization is performed to adjust all images to a uniform pixel density. When constructing the labeled dataset, at least five architectural history experts with more than 10 years of experience were invited to participate in the annotation. The pre-processed images of the arcade facade were annotated using annotation software. The outlines of the components were precisely selected on the images, the component types were selected, and the material properties and preservation status were annotated to obtain the final labeled dataset. In order to ensure the quality of annotation, a cross-validation mechanism was adopted during the annotation process. Each component was annotated independently by at least two experts, and when there were disagreements, a consensus was reached through discussion.

4. The method for intelligent recognition and language conversion of traditional arcade facade components based on a hierarchical coding system as described in claim 1, characterized in that... The training strategy for the arcade component recognition model in step S3 adopts the following three-stage progressive method: In the first stage, the COCO dataset was used for pre-training, allowing the arcade component recognition model to learn basic visual feature extraction capabilities. In the second stage, the architectural image dataset is used for domain adaptation, enabling the arcade component recognition model to understand the unique visual patterns of architectural components. In the third stage, the labeled dataset from step S2 is used for fine-tuning, with a focus on optimizing the ability of the arcade component recognition model to recognize unique arcade components.

5. The method for intelligent recognition and language conversion of traditional arcade facade components based on a hierarchical coding system as described in claim 4, characterized in that... The training process of the arcade component recognition model in step S3 adopted the following optimization measures: (1) Use the focus loss function to solve the problem of component category imbalance, and give higher loss weight to rare component categories with fewer numbers; (2) Introduce an online difficult example mining mechanism, and pre-set easily confused component pairs in the arcade components according to practical experience. During the training process, automatically identify these component pairs and increase the training frequency of these component pairs according to the preset number or proportion value. (3) By using data augmentation technology, the training samples are expanded by 3 times to improve the generalization ability of the arcade component recognition model.

6. The method for intelligent recognition and language conversion of traditional arcade facade components based on a hierarchical coding system as described in claim 4, characterized in that... In step S3, a post-processing mechanism is established to ensure the reasonableness of the recognition results, specifically: (1) Filter the detection boxes output by the arcade component recognition model by confidence level and retain only the results with a confidence level higher than 0.8; (2) Use non-maximum suppression algorithm to eliminate overlapping detection boxes and avoid the same part being repeatedly identified; (3) Construct a logical rule base based on architectural common sense, and verify the rationality of the identification results, such as checking whether the relative positional relationship of the components conforms to the architectural construction logic.

7. The method for intelligent recognition and language conversion of traditional arcade facade components based on a hierarchical coding system as described in claim 1, characterized in that... The process of allocating standard codes in step S4 is as follows: First, the basic code is determined based on the identified component type and its location on the arcade facade; By analyzing the visual features of the components, we can infer their possible construction date; By combining the geographical information of the image's shooting location, the relevant regional codes are automatically filled in; The material properties of a component are determined by a pre-trained material classification network.

8. The method for intelligent recognition and language conversion of traditional arcade facade components based on a hierarchical coding system as described in claim 7, characterized in that... The multidimensional features extracted in step S4 cover the following dimensions: (1) Geometric features; (2) Visual characteristics; (3) Stylistic characteristics; (4) Material characteristics; The steps of component association analysis in step S4 are as follows: (1) Spatial relationship modeling: Based on the bounding box coordinates of the component recognition results, calculate the spatial distance matrix and relative positional relationship between components, and establish a component spatial relationship diagram; (2) Statistical co-occurrence probability: Based on a large number of samples, the co-occurrence frequency of different component types is statistically analyzed to construct a component co-occurrence probability matrix; (3) Co-occurrence pattern mining: Statistical analysis of the co-occurrence frequency of different component types in the arcade sample, and the use of association rule mining algorithm to discover the combination pattern of strongly associated components; (4) Verification of consistency of age: Based on the inferred age attributes of components, examine the consistency of age of different components of the same building and identify possible later renovations; (5) Construct a component relationship network: Use components as nodes and association strength as edge weights to construct a weighted undirected graph for subsequent architectural style analysis and integrity verification.

9. The method for intelligent recognition and language conversion of traditional arcade facade components based on a hierarchical coding system as described in claim 7, characterized in that: In step S4, the constructed component feature database is stored using MongoDB; each identified component has an independent record created in the database, containing a unique identifier, a complete six-level code, precise coordinates in the original image, various extracted feature vectors, identification confidence score, and processing timestamp information; by establishing a multi-dimensional index, it supports fast retrieval based on encoding, feature similarity, and spatial location.

10. The method for intelligent recognition and language conversion of traditional arcade facade components based on a hierarchical coding system as described in claim 1, characterized in that... The following quality control mechanism is introduced in step S5: (1) Architectural syntax checker: Based on traditional dependency parsing, it adds a check for architectural expression norms; (2) Building Terminology Consistency Check Module: Construct a mapping table of traditional arcade building professional terms and map synonyms to standard terms; (3) Architectural Logic Verifier: Establishes architectural domain-specific rules to verify the correctness of the historical logic described; (4) Architectural text readability evaluation algorithm: Based on the characteristics of architectural professional texts, the algorithm sets the index that the density of professional terms does not exceed 30% and the average sentence length does not exceed 25 characters.

11. The method for intelligent recognition and language conversion of traditional arcade facade components based on a hierarchical coding system as described in claim 1, characterized in that... In step S6, the ontology of the knowledge graph defines the following six types of core entities: (1) The building entity, recording the basic information of the entire arcade building; (2) Component entities, storing detailed attributes of various building components; (3) Material entities, describing the physicochemical properties of different building materials; (4) Period entities, dividing historical periods and associating them with the characteristics of the era; (5) Geographic entities, recording geospatial information; (6) Craftsmanship entities, recording traditional and modern construction techniques; The relationships between entities include: (1) Compositional relationships, describing the overall and partial relationships between the building and its components; (2) Positional relationship, recording the spatial distribution of components on the elevation; (3) Material relationships, connecting components and the building materials used; (4) Due to time constraints, indicate the construction and modification history of the components; (5) Technological relationships, explaining the manufacturing techniques of the components; (6) Evolutionary relationships, tracing the historical changes in component styles.

12. The method for intelligent recognition and language conversion of traditional arcade facade components based on a hierarchical coding system as described in claim 11, characterized in that... In step S6, the knowledge graph integrates multi-source data through a knowledge extraction and fusion process, specifically: (1) Automatically extract component entities and basic attributes from the identification results of step S3 to create the initial nodes of the knowledge graph; (2) Establish connection edges between nodes by analyzing the spatial adjacency and functional association between components; (3) Supplement the historical and cultural information of components by integrating external data sources; (4) Use entity alignment technology to identify and merge different representations that point to the same object to ensure the consistency and integrity of knowledge; In step S6, the knowledge graph has a reasoning rule engine that supports complex knowledge discovery, specifically: (1) Reasoning mechanism based on descriptive logic; (2) Using temporal reasoning, analyze the evolution path of component styles; (3) Spatial reasoning to discover the distribution pattern of components.

13. The method for intelligent recognition and language conversion of traditional arcade facade components based on a hierarchical coding system as described in claim 1, characterized in that... The system in step S7 adopts an end-to-end processing flow and an asynchronous processing mechanism, supporting parallel processing of batch images and ensuring efficient system operation. Specifically: When a user uploads an image of the arcade facade, the system automatically starts the processing pipeline. First, it performs image quality checks and preprocessing. Then, it calls the trained arcade component recognition model to identify the components inside the arcade. Next, it assigns coding information to each component and extracts multi-dimensional features. Subsequently, it generates three levels of component description text and finally stores it in the knowledge graph and updates the index.

14. The method for intelligent recognition and language conversion of traditional arcade facade components based on a hierarchical coding system as described in claim 13, characterized in that... The system in step S7 employs the following technical means to optimize system performance: (1) Convert the 32-bit floating-point model parameters into 8-bit integer representations using model quantization techniques; (2) Use TensorRT to optimize the inference engine and make full use of the parallel computing capabilities of the GPU; (3) Implement a distributed architecture design to support multi-node collaborative processing; (4) Establish an intelligent caching mechanism to cache frequently accessed feature data and query results.

15. The method for intelligent recognition and language conversion of traditional arcade facade components based on a hierarchical coding system as described in claim 13, characterized in that... The system in step S7 continuously improves its performance through a continuous optimization mechanism, specifically: (1) Establish a user feedback collection system; (2) Organize expert reviews regularly; (3) Expand the training dataset; (4) Update and optimize the recognition model; (5) Improve the content of the knowledge graph.

Citation Information

Patent Citations

  • Building three-dimensional model semantization method and system

    CN113569331A

  • Southern Fujian historical block typical culture gene identification method

    CN119169624A