Historical building archive data processing method based on association query and clustering analysis
By constructing knowledge graphs and advanced text analysis, the multimodal, massive and complex correlation problems of historical architectural archival data have been solved, efficient correlation query and cluster analysis have been achieved, and the scientificity and efficiency of cultural heritage protection and utilization have been improved.
Patent Information
- Application Number
- CN202510816456.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies find it difficult to effectively process multimodal, massive, and complexly correlated historical architectural archival data, and are unable to achieve complex correlation queries across decades, regions, and types, as well as effective clustering of high-dimensional, mixed features, limiting their application value in cultural heritage protection, cultural heritage research, and tourism development.
Through data preprocessing, association query, feature extraction and cluster analysis methods, a knowledge graph is constructed, and multi-hop complex association query is performed using a graph database. Combined with advanced text analysis and clustering algorithms, multi-dimensional features are extracted for intelligent clustering, generating building clusters and mining deep knowledge.
It has achieved in-depth mining and structuring of historical architectural archival data, provided accurate data support, offered a scientific basis for cultural heritage protection, restoration strategy formulation, and cultural tourism development, and improved the scientific nature and efficiency of research and utilization.
Smart Images

Figure CN120670486A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a method for processing historical building archive data based on association query and cluster analysis. Background Art
[0002] In the context of today's accelerating urbanization process, historical buildings, as key carriers of urban memory and cultural heritage, have extremely rich historical and cultural information in their archival data. These archival data not only cover basic information such as the structural characteristics and construction years of the buildings themselves, but also contain a large amount of important information related to the historical background, repair records, changes in usage functions, etc.
[0003] However, with the growing awareness of historical building preservation and the deepening of related work, these archival data face a series of daunting challenges. On the one hand, historical buildings are numerous and widely distributed, and their archival data comes from a variety of sources, including field surveys, document consultation in archives and libraries, searches in specialized databases, and web crawlers. This has led to an explosive growth in the amount of data collected, and the data formats are extremely complex and diverse, including text archives and multimodal data such as images, videos, and drawings. On the other hand, historical buildings have experienced the vicissitudes of time, covering a wide range of historical periods. The archival recording methods and standards vary from era to era, resulting in complex data correlations. The relationships between different buildings and the connections between buildings and historical events and figures are intertwined, forming a vast and difficult-to-understand data network.
[0004] Traditional data processing methods primarily focus on the simple organization and classification of single-type data. Faced with multimodal, massive, and complexly correlated data, they can often only perform superficial processing. In terms of correlation queries, they struggle to deeply explore the potential connections between different data sets, and are unable to implement complex, multi-hop correlation queries across eras, regions, and building types. This makes it impossible to provide comprehensive and accurate data support for studying the inheritance and development of historical architecture. In terms of cluster analysis, traditional methods also struggle to effectively cluster high-dimensional, mixed-feature data, making it impossible to accurately identify historical building clusters with similar characteristics, and thus unable to provide a scientific basis for the formulation of targeted protection and utilization strategies. This greatly limits the application value of historical architectural archival data in fields such as cultural heritage protection, cultural heritage research, and tourism development, and makes it difficult to meet modern society's urgent need for in-depth research and rational utilization of historical architecture. Summary of the Invention
[0005] To achieve the above objectives, the present invention proposes a method for processing historical architectural archive data based on association query and cluster analysis, comprising the following steps:
[0006] Step 1: Data preprocessing, collecting historical architectural archival data, including data collection, data cleaning, data conversion and data integration;
[0007] Step 2: Implementation of association query, building an association model for historical building archive data and defining the association relationship between data, including association model construction and association query algorithm application;
[0008] Step 3: Feature extraction implementation: for historical building archival data, extract features for cluster analysis, including basic information feature extraction, historical and cultural feature extraction, and repair and protection feature extraction;
[0009] Step 4: Cluster analysis implementation, cluster analysis is performed on the extracted features to divide the historical building archive data into different clusters, including clustering algorithm selection, clustering parameter setting, and clustering implementation and evaluation;
[0010] Step 5: Result analysis and application implementation. Analyze the results of cluster analysis and explore the characteristics and patterns of historical architectural archival data in different clusters.
[0011] In one example, in step one, data is collected through diversified channels to collect multimodal archival data covering the entire life cycle of historical buildings.
[0012] In one example, in step 1, data cleaning utilizes a data cleaning tool or script to detect and remove duplicate data, perform integrity checks, and perform accuracy checks and repairs.
[0013] In one example, in step one, data conversion deeply applies NLP technology to text data and stores it in a structured manner, and applies computer vision technology to image and drawing data to quantify and encode visual information.
[0014] In one example, in step 1, data integration integrates the cleansed and transformed heterogeneous data into a hybrid database system.
[0015] In one example, in step 2, the association model is constructed based on a knowledge graph in a graph database, and the association query algorithm is applied using the graph database's native query language to execute multi-hop complex association queries.
[0016] In one example, in step three, basic information feature extraction directly extracts structured attributes, historical and cultural feature extraction applies advanced text analysis to mine implicit value, and repair and protection feature extraction analyzes professional documents to quantify protection status.
[0017] In one example, in step four, the clustering algorithm is selected based on the data characteristics and objectives, the clustering parameters are set to scientifically determine the key parameters, and the clustering implementation and evaluation are iteratively optimized to optimize the clustering results.
[0018] In one example, in step five, the result analysis deeply interprets the clustering connotation, and the application implementation drives decision-making and knowledge services.
[0019] In one example, the result analysis and application implementation further includes identifying significant features by comparing and analyzing the differences between clusters based on the generated building clusters and their core feature distribution, and exploring potential patterns and trends. Based on the deep knowledge mined, data intelligence support is provided for the hierarchical protection decision-making, repair strategy formulation, cultural tourism value development and academic research of cultural heritage.
[0020] The method for processing historical building archive data based on association query and cluster analysis proposed by the present invention can bring the following beneficial effects:
[0021] 1. Through deep data governance and graph database modeling, the present invention transforms the originally scattered, isolated and unstructured massive historical architectural archives into a highly correlated, machine-readable and deeply mineable structured knowledge network, breaking through the bottleneck of low information retrieval efficiency and difficulty in discovering correlations under the traditional archive management model. Based on the complex correlation query capability of this knowledge graph and the intelligent clustering analysis driven by multi-dimensional features, it can automatically mine deep laws and implicit knowledge such as the evolution of architectural style, the influence of regional culture, and the correlation between the effectiveness of protection measures, so that the dormant archival data can truly be transformed into high-value knowledge assets that can support research, protection and utilization, realizing a qualitative leap from "data storage" to "knowledge discovery".
[0022] 2. This invention proposes cluster analysis results and knowledge graph insights, providing cultural heritage managers with objective, quantitative, and visual decision-making basis, which can be directly applied to accurately identify protection priorities, optimize the formulation of repair strategies, and design differentiated cultural tourism development paths. At the same time, the laws of architectural style dissemination and technical craftsmanship inheritance revealed by it also provide new clues and new perspectives for academic research. Ultimately, this solution significantly improves the scientific nature, foresight, and resource utilization efficiency of cultural heritage protection work, and provides strong support for the research, protection, and utilization of historical buildings. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0024] Figure 1 This is a flow chart of a method for processing historical architectural archive data based on association query and cluster analysis. DETAILED DESCRIPTION
[0025] In order to more clearly illustrate the overall concept of the present invention, a detailed description is given below in an exemplary manner in conjunction with the accompanying drawings.
[0026] In the description of the present invention, it should be understood that the terms "center", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "axial", "radial", "circumferential", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the present invention.
[0027] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature identified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0028] In the present invention, unless otherwise expressly specified or limited, terms such as "mounted," "connected," "connect," and "fixed" should be understood broadly. For example, they may refer to fixed connection, detachable connection, or integration; mechanical connection, electrical connection, or communication; direct connection or indirect connection through an intermediate medium; and internal communication between two components or interaction between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0029] In the present invention, unless otherwise clearly specified and limited, a first feature "above" or "below" a second feature may be that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. In the description of this specification, the descriptions with reference to the terms "one scheme", "some schemes", "examples", "specific examples", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the scheme or example are included in at least one scheme or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same scheme or example. Moreover, the specific features, structures, materials or characteristics described may be combined in an appropriate manner in any one or more schemes or examples.
[0030] like Figure 1 As shown, the present invention proposes a method for processing historical building archive data based on association query and cluster analysis, comprising the following steps:
[0031] Step 1: Data preprocessing:
[0032] Data collection: Through diversified channels (field research and investigation, archives / library literature review, professional database search, web crawler to capture relevant research materials), multimodal archival data covering the entire life cycle of historical buildings are collected, including architectural design drawings (blueprints, manuscripts), historical archival texts (local chronicles, engineering records), photos (historical and current), videos (documentary, aerial photography), research reports, repair records, etc.
[0033] Data cleaning: Systematically improve data quality: Use data cleaning tools or scripts (such as Python Pandas, OpenRefine) to detect and remove duplicate data (precisely match unique identifiers, fuzzy match key fields to calculate string similarity, and compare file hash values), perform integrity checks (verify that key fields are not empty / value ranges / associated references, and verify file readability), and perform accuracy checks and repairs (text data uses spelling checks + professional terminology library error correction, NLP syntax analysis, cross-source consistency checks, and external knowledge base verification; image / drawing data uses OpenCV and other libraries to calculate clarity indicators, detect integrity / damaged areas, and attempt image enhancement / super-resolution repair, and re-acquire in serious cases).
[0034] Duplicate data detection and removal, exact matching, finding and deleting duplicate data through unique identifiers (such as building numbers), fuzzy matching using string similarity algorithms (such as Levenshtein distance) to detect and merge similar names (such as "Palace Museum" and "Palace Museum (Forbidden City)"), file hash value comparison, calculating file hash values (such as MD5), identifying and deleting identical images or documents.
[0035] Integrity check: key fields are not empty, ensure that important fields (such as building name, construction year) are not empty, delete data with missing key fields, value range verification, check whether the field value is within a reasonable range, for example, the construction year should be consistent with the historical period range, associated reference verification, ensure that associated fields (such as building number) are correctly referenced in the database, and delete incorrectly referenced data.
[0036] Accuracy check and repair, text data repair, using spelling check, professional terminology library error correction (such as the spelling error of "Palace Museum") and NLP syntax analysis, combined with cross-source data verification, to correct text errors, image data repair, through clarity detection (such as Laplace operator) and image enhancement technology (such as super-resolution repair) to improve image quality, severely damaged images attempt to re-acquire,
[0037] Data conversion: Structured unstructured information: For text data (archives, reports), NLP technology is deeply applied (named entity recognition to extract building names / eras / places / people / events, relationship extraction to establish entity connections, key information extraction such as protection level / value description) and structured storage as database tables or graph node attributes; for image / drawing data, computer vision technology is used (target detection to identify building components / roof forms, image classification to determine architectural style, feature extraction to analyze material texture, OCR to identify drawing annotation text / size) to quantify and encode visual information into structured feature vectors or labels.
[0038] Relationship extraction establishes entity connections using relation extraction in NLP technology, implemented using the Python language and NLP libraries. Named entity recognition is used to first accurately locate key entities such as historical buildings, historical events, and related figures from text data. For example, information such as building names, years, locations, and people can be identified from archival text. Relationship extraction is then used to analyze the text after entity recognition based on preset rules or models to determine the relationships between entities. For example, in the sentence "This building was designed by the famous architect XX during the reign of Emperor Guangxu of the Qing Dynasty," relationship extraction clarifies that "building" and "famous architect XX" have a "designed" relationship, and that "building" and "during the reign of Emperor Guangxu of the Qing Dynasty" have a "built" relationship.
[0039] Data integration: Build a unified data view: Integrate cleaned and transformed heterogeneous data into a hybrid database system (such as relational database MySQL / PostgreSQL to store structured attributes, and graph database Neo4j to store entity relationship networks) based on unique building identifiers to establish strong associations across data types.
[0040] Step 2: Implementation of associated query:
[0041] Association model construction: Build a knowledge graph based on a graph database: model core entities (historical buildings, historical events, related figures, other buildings, geographical locations, architectural styles) as nodes, and rich relationships between entities (such as "built in a certain era", "involved in a certain event", "designed by someone", "adopting a certain style", "located in a certain place", "affected / affected by a certain building", "belonging to a certain protection unit") as edges, and give nodes and edges detailed attributes.
[0042] Application of association query algorithms: Deeply mine associations using graph query languages: Using graph database native query languages (such as Cypher) to execute multi-hop complex association queries. For example, a query for "all organizations involved in a building's renovations, materials used, and their suppliers" can be efficiently performed by traversing paths such as "building - renovation event - participating organizations" and "renovation event - used materials - material suppliers," revealing deep connections.
[0043] Step 3: Feature extraction implementation:
[0044] Extraction of basic information features: Direct extraction of structured attributes: directly obtain from the integrated database or extract clear fields through simple parsing, such as building name, construction year (parsed date format), precise geographic location (latitude and longitude), building area / height (extract values from drawing metadata or report tables), and building material list (obtained from structured descriptions or material tags).
[0045] Extraction of historical and cultural features: Apply advanced text analysis to mine hidden value: For text archives and research reports, use NLP technology for topic modeling (such as LDA) to automatically discover hot topics of discussion, sentiment analysis to evaluate the emotional tendency of descriptions (praise or criticism of the value of the building), event extraction to identify key historical event nodes, and combine web crawlers and document organization to structuredly store relevant historical stories and celebrity anecdotes.
[0046] Extraction of repair and protection features: Analyze professional documents to quantify protection status: Extract structured information from repair reports, protection plans and other documents through rule matching and information extraction models: the number of repairs and time series, specific protection measures (such as "roof tile removal overhaul", "structural reinforcement"), official protection level (national protection / provincial protection / municipal protection), current preservation status score (such as intact / slightly damaged / endangered) and evaluation basis.
[0047] Step 4: Cluster analysis implementation:
[0048] Clustering algorithm selection: Select the best algorithm based on data characteristics and goals: For high-dimensional feature vectors, K-means or its optimized variants (such as K-means++) are preferred to achieve efficient large-scale clustering; if hierarchical relationships (such as architectural style genealogy) need to be displayed, hierarchical clustering (Agglomerative Clustering) is used; for mixed feature types (numerical + categorical), Gower distance + K-prototypes or specific hybrid clustering algorithms can be considered.
[0049] Clustering parameter setting: Scientifically determine key parameters: Use the elbow method (ElbowMethod) to analyze the inflection point of the sum of squares within the cluster or the silhouette coefficient (SilhouetteCoefficient) maximization principle to determine the optimal number of clusters K; carefully select the initial center (such as K-means++); clearly define the distance metric (Euclidean distance for numerical features, Hamming distance / Jaccard similarity for categorical features, or use Gower distance after standardization).
[0050] Clustering Implementation and Evaluation: Iteratively optimize clustering results: Input the standardized / normalized feature matrix into the selected algorithm for clustering to generate building clusters. Rigorously evaluate performance: Calculate metrics such as the silhouette coefficient (closer to 1 is optimal), the Calinski-Harabasz index (higher values are better), and the Davies-Bouldin index (lower values are better). If performance is poor, adjust strategies: reselect the algorithm, optimize feature engineering (such as PCA / t-SNE for dimensionality reduction), adjust the K parameter or distance metric, and remove outliers.
[0051] Step 5: Result analysis and application implementation:
[0052] Results analysis: In-depth interpretation of clustering connotations: statistical and visual analysis of the core feature distribution of each cluster (such as architectural style proportion pie charts, construction age distribution histograms / box plots, geographical distribution heat maps, and protection level stacked bar charts); comparative analysis of differences between clusters to identify significant features; and exploration of potential patterns and trends (such as the evolution of a certain style over time / region, and the correlation between the effectiveness of specific protection measures and the status of the building).
[0053] Application implementation: Driving decision-making and knowledge services: Applying analysis results to conservation priority ranking (identifying endangered and high-value clusters), repair strategy recommendations (referring to successful cases in the same cluster), tourist route planning (recommendations by style / regional clusters), academic research (discovering new clues to the spread of architectural schools) and building an intelligent knowledge question-answering system (answering complex inquiries based on graphs and clustering results).
[0054] This solution constructs a multimodal, knowledge-driven intelligent processing framework for historical architectural archives. Through systematic data governance (integrating multi-source heterogeneous data, combining precise matching, fuzzy algorithms, and domain knowledge bases for deep cleaning, and utilizing NLP entity / relationship extraction, CV target detection, and image classification technology to accurately convert unstructured text / drawings into structured data), a knowledge graph centered on a graph database is established (entity association modeling supports multi-hop complex queries). Multidimensional features (basic attributes, cultural themes, and conservation status) are then extracted to drive intelligent clustering analysis (based on K-means++ / hierarchical clustering algorithms, with parameters optimized using the elbow method and silhouette coefficient). Ultimately, deep knowledge such as architectural style lineages, regional distribution patterns, and conservation effectiveness correlations is mined from massive archives, providing data intelligence support for hierarchical protection decisions, restoration strategy formulation, cultural tourism value development, and academic research on cultural heritage, achieving a full-chain transformation from original archives to actionable knowledge.
[0055] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0056] The foregoing is merely an embodiment of the present invention and is not intended to limit the present invention. It will be apparent to those skilled in the art that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A method for processing historical architectural archive data based on association query and cluster analysis, characterized by: The following steps are involved: Step 1: Data preprocessing, collecting historical architectural archival data, including data collection, data cleaning, data conversion and data integration; Step 2: Implementation of association query, building an association model for historical building archive data and defining the association relationship between data, including association model construction and association query algorithm application; Step 3: Feature extraction implementation: for historical building archival data, extract features for cluster analysis, including basic information feature extraction, historical and cultural feature extraction, and repair and protection feature extraction; Step 4: Cluster analysis implementation, cluster analysis is performed on the extracted features to divide the historical building archive data into different clusters, including clustering algorithm selection, clustering parameter setting, and clustering implementation and evaluation; Step 5: Result analysis and application implementation. Analyze the results of cluster analysis and explore the characteristics and patterns of historical architectural archival data in different clusters.
2. The method for processing historical architectural archive data based on association query and cluster analysis according to claim 1, characterized in that: In step one, data collection involves collecting multimodal archival data covering the entire life cycle of historical buildings through diversified channels.
3. The method for processing historical architectural archive data based on association query and cluster analysis according to claim 1, characterized in that: In step 1, data cleaning utilizes data cleaning tools or scripts to detect and remove duplicate data, perform integrity checks, and perform accuracy checks and repairs.
4. The method for processing historical architectural archive data based on association query and cluster analysis according to claim 1, characterized in that: In the step 1, data conversion deeply applies NLP technology to text data and stores it in a structured manner, and applies computer vision technology to quantify and encode visual information in image and drawing data.
5. The method for processing historical architectural archive data based on association query and cluster analysis according to claim 1 is characterized by: In the step 1, data integration integrates the cleaned and transformed heterogeneous data into the hybrid database system.
6. The method for processing historical architectural archive data based on association query and cluster analysis according to claim 1, characterized in that: In the step 2, the association model is constructed based on the knowledge graph of the graph database, and the association query algorithm is applied to execute multi-hop complex association queries using the native query language of the graph database.
7. The method for processing historical architectural archive data based on association query and cluster analysis according to claim 1, characterized in that: In the step three, basic information feature extraction directly extracts structured attributes, historical and cultural feature extraction uses advanced text analysis to mine implicit value, and repair and protection feature extraction analyzes professional documents to quantify the protection status.
8. The method for processing historical architectural archive data based on association query and cluster analysis according to claim 1, characterized in that: In the step 4, the clustering algorithm is selected based on the data characteristics and the target optimization model, the clustering parameters are set to scientifically determine the key parameters, and the clustering implementation and evaluation are iteratively optimized to optimize the clustering results.
9. The method for processing historical architectural archive data based on association query and cluster analysis according to claim 1, characterized in that: In step five, the result analysis deeply interprets the clustering connotation, and the application implementation drives decision-making and knowledge services.
10. The method for processing historical architectural archive data based on association query and cluster analysis according to claim 9, characterized in that: The analysis and application of the results further include identifying significant features based on the generated building clusters and their core feature distribution, and exploring potential patterns and trends through comparative analysis of differences between clusters. Based on the deep knowledge mined, data intelligence support is provided for hierarchical protection decisions, restoration strategy formulation, cultural tourism value development, and academic research of cultural heritage.