A deep learning-based urban space visual cultural heritage element perception evaluation algorithm

By using deep learning-based multi-source data processing and a dual-stream model, the subjectivity and fragmentation issues in the assessment of urban visual cultural heritage are resolved. This enables the quantitative assessment of both tangible and intangible elements, providing scientific assessment results and forward-looking analysis to support urban planning and management.

CN122114740APending Publication Date: 2026-05-29SOUTHWEST FORESTRY UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHWEST FORESTRY UNIVERSITY
Filing Date
2026-03-12
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies for assessing urban visual cultural heritage suffer from problems such as subjective bias, difficulty in quantification, traditional data sources, fragmented single-modal analysis, and lack of an end-to-end intelligent assessment framework, resulting in a lack of horizontal comparability and scientific rigor in the assessment results.

Method used

We employ deep learning-based multi-source heterogeneous data acquisition and preprocessing to construct a dual-stream deep learning model for material entity recognition and intangible atmosphere perception. We calculate the Urban Spatial Visual Cultural Heritage Comprehensive Value Index (UVCH-VCI) and perform visualization and decision support on the WebGIS platform.

Benefits of technology

It enables the objective quantification of urban visual cultural heritage assessment, accurately measures material characteristics and reveals intangible intrinsic value, provides forward-looking analytical capabilities, provides a scientific basis for urban planning, and supports dynamic management and timely intervention.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The application discloses a kind of urban space visual culture heritage element perception evaluation algorithm based on deep learning, comprising the following steps: S1: multi-source heterogeneous data acquisition and pretreatment: the multi-source data of target city space is collected, and data at least includes sub-meter resolution optical remote sensing image, close-range or panoramic street view image, space-time history map and text description data containing cultural semantics;Geometric correction, space-time registration, data fusion and enhancement preprocessing operation is carried out to the data, and the standardized three-dimensional space-time data set is constructed.By first systematically introducing cross-modal deep learning into the field of urban visual cultural heritage evaluation, the abstract concept of "intangible perception", which has long relied on subjective feelings, is transformed into objective indicators that can be calculated and compared through mathematical models.This fundamentally solves the core problem of subjectivity of traditional methods.The evaluation process is like "chemical analysis", which decomposes complex "cultural perception" into measurable "elements".
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image or video recognition technology, specifically to a deep learning-based algorithm for perceiving and evaluating urban spatial visual cultural heritage elements. Background Technology

[0002] Urban Visual Cultural Heritage (UVCH) is a complex of tangible and intangible landscapes that have accumulated in a city throughout its history, are perceptible to human vision, and carry specific cultural significance and collective memory. It is not only a core identifier of a city's unique identity but also an important resource for sustainable urban development and cultural tourism innovation. A scientific and systematic assessment of UVCH is a prerequisite for its effective protection, rational utilization, and sustainable inheritance.

[0003] Currently, the evaluation methods for UVCH in industry and academia can be mainly divided into the following categories, but all of them have significant limitations:

[0004] Qualitative assessment methods based on expert experience:

[0005] This is the most traditional method, in which a panel of experts, including urban planners, historians, and architects, assesses the historical, artistic, and scientific value of a heritage site based on relevant regulations (such as the "Guidelines for the Protection of Cultural Relics and Historic Sites in China") and academic theories, through on-site inspections, document review, and conference discussions.

[0006] Detailed description of the defects:

[0007] Subjectivity bias: Evaluation results are highly dependent on the individual cognition, aesthetic preferences, knowledge background, and even personal experience of experts. For example, an expert with a focus on architecture may give high praise to a modern imitation of an ancient building with innovative structure, while a historian may reject it due to its lack of historical authenticity, leading to "different opinions for different people" and a lack of cross-comparison of conclusions.

[0008] Cognitive Load & Inconsistency: Urban space is a complex mega-system. Experts are prone to fatigue and cognitive bias when processing massive amounts of visual information in a short period of time, which can lead to inconsistent judgments on the importance of the same element or difficulty in converging the evaluation results of different expert groups.

[0009] Difficult to quantify and communicate: Evaluation results are mostly presented in textual descriptions and vague levels such as "excellent", "good", and "average", which makes it impossible to conduct precise quantitative comparisons and spatial statistical analysis, and is not conducive to establishing a long-term supervision mechanism.

[0010] Quantitative evaluation method based on indicator system:

[0011] To overcome subjectivity, scholars have attempted to construct a multi-factor weighted scoring system. For example, they have established indicators and assigned weights to aspects such as "historical age," "scarcity," "completeness," and "coherence," and calculated the scores using a scoring table.

[0012] Detailed description of the defects:

[0013] The dilemma in constructing an indicator system: How to scientifically and comprehensively select indicators is itself a difficult problem. Existing systems often focus on "hard indicators" of physical entities (such as building age and integrity rate), while insufficient consideration or inability to quantify "soft" cultural perception values ​​such as "atmosphere," "artistic conception," and "genius loci" (the spirit of place).

[0014] The challenge of determining weights: Weights are usually determined using the expert Delphi method or a simple analytic hierarchy process (AHP), which is essentially still subjective. Furthermore, once determined, weights tend to be fixed and difficult to adapt to the evaluation needs of different cities and cultural backgrounds.

[0015] The data sources remain traditional: the scoring of the indicators is still based on experts' interpretation of limited photos and literature, failing to make full use of modern technology to obtain more comprehensive and objective data;

[0016] Existing attempts to combine technology with this field and their shortcomings:

[0017] In recent years, GIS spatial analysis and machine learning have begun to be introduced into heritage assessment. For example, GIS buffer analysis is used to assess the impact of heritage sites on the surrounding environment, or simple image classification models are used to identify architectural styles;

[0018] Detailed description of the defects:

[0019] "Emphasis on form over substance, emphasis on material over perception": Current approaches often remain at a superficial level of identifying and spatially analyzing material forms, failing to delve into the deconstruction and quantification of the core value of cultural heritage—the "perceptual experience." They cannot answer crucial questions such as, "Does this place evoke a sense of tranquility and history?"

[0020] Fragmented "Single-Modal" Analysis: The analysis process is often fragmented. Image analysis is separated from text analysis, and spatial analysis is disconnected from semantic analysis. For example, architectural style is analyzed using only images, and historical value is analyzed using only text, ignoring the inherent connections and mutual corroboration between images and text, and between form and meaning;

[0021] The lack of an end-to-end intelligent evaluation framework: A complete closed-loop system has not been formed, from raw data input to automatic multi-dimensional feature deconstruction, and then to comprehensive value intelligent computing and decision-making. Most research focuses on fragmented technological applications rather than systematic methodological innovation.

[0022] To address this, we propose a deep learning-based algorithm for perceiving and evaluating urban spatial visual cultural heritage elements. Summary of the Invention

[0023] To achieve the above objectives, the present invention provides the following technical solution: a deep learning-based algorithm for perceiving and evaluating urban spatial visual cultural heritage elements, comprising the following steps:

[0024] S1: Multi-source heterogeneous data acquisition and preprocessing: Acquire multi-source data of the target urban space, including at least sub-meter resolution optical remote sensing images, close-up or panoramic street view images, spatiotemporal historical maps, and textual description data containing cultural semantics; perform geometric fine correction, spatiotemporal registration, data fusion and enhancement preprocessing operations on the data to construct a standardized three-dimensional spatiotemporal dataset;

[0025] S2: Multi-scale visual element deconstruction based on deep learning: Construct a dual-stream deep learning model to extract features from preprocessed data. One stream is the material entity recognition stream used for instance-level semantic segmentation to identify core material elements, and the other stream is the non-material atmosphere perception stream used to generate non-material atmosphere comprehensive perception embedding vectors.

[0026] S3: Calculation of the Comprehensive Value Index of Urban Spatial Visual Cultural Heritage (UVCH-VCI): Based on the output of step S2, calculate the Material Element Integrity and Typicality Index (PITI), Visual Landscape Aesthetics Index (VLAI), and Non-material Perceptual Resonance Index (IPRI) respectively, and obtain the final UVCH-VCI through weighted fusion calculation;

[0027] S4: Visualization of assessment results and decision support: The UVCH-VCI and its sub-indices are visualized in a multi-dimensional interactive manner on the WebGIS platform. Based on the assessment results, targeted protection and revitalization decision recommendations are automatically generated through rule-based reasoning and generative AI models.

[0028] Preferably, the material entity recognition stream in step S2 adopts an improved U-Net++ model, which integrates a CBAM attention module in the encoder part. The model is configured to perform pixel-level instance-level semantic segmentation on remote sensing images and street view images to identify and label specific categories of objects that constitute visual cultural heritage. These specific categories of objects include traditional building components, distinctive paving, ancient trees, stone inscriptions and sculptures, and water system bridges.

[0029] 3. The deep learning-based urban spatial visual cultural heritage element perception and evaluation algorithm according to claim 1, characterized in that: the non-material atmosphere perception flow in step S2 specifically includes:

[0030] S1: Use a pre-trained Vision Transformer (ViT) or Swin Transformer model to extract global visual depth features from panoramic street view images;

[0031] S2: The BERT model, which is pre-trained in the field of natural language processing (NLP) and adaptively fine-tuned in the field of cultural heritage, is used to encode the text description data and generate deep text semantic vectors.

[0032] S3: Using a cross-modal contrastive learning framework, image-text pairs spatially-co-located at the same location are used as positive samples for training to learn a shared embedding space. Based on this, a fixed-dimensional non-material atmosphere comprehensive perception embedding vector representing the overall non-material atmosphere of the location is generated through feature fusion.

[0033] Preferably, the formula for calculating the Material Element Integrity and Typicality Index (PITI) in step S3 is: PITI = α * CI + β * TD; where CI (Component Integrity) is the identification integrity score of the key material element, which is determined by calculating the ratio of the actual identified area to the theoretical total area; TD (Typicity Degree) is the matching score between the identified element combination and the preset cultural heritage element prototype library, which is determined by calculating the Fraser distance or cosine similarity between the feature vector of the element combination and the best matching prototype and then normalizing it; α and β are weighting coefficients and α + β = 1.

[0034] Preferably, the Visual Landscape Aesthetic Index (VLAI) mentioned in step S3 is calculated based on the segmentation results of the material entity recognition flow in step S2, and its landscape pattern index includes at least the clustering index (AI), the spread index (CONTAG), the Shannon diversity index (SHDI), and the average fractal dimension (FRAC_MN). The landscape pattern index is then determined by weighted fusion with the predicted score of a domain aesthetic quality regression model pre-trained and fine-tuned on an aesthetic dataset.

[0035] Preferably, the Intangible Perception Resonance Index (IPRI) mentioned in step S3 is determined by calculating the maximum cosine similarity between the intangible atmosphere integrated perception embedding vector generated in step S2 and one or more preset "ideal cultural heritage site" perception prototype vectors.

[0036] Preferably, the final comprehensive index calculation formula in step S3 is: UVCH-VCI = γ*PITI_norm+δ*VLAI_norm +ε* IPRI_norm; where PITI_norm, VLAI_norm, and IPRI_norm are the normalized values ​​of each sub-index; γ, δ, and ε are dynamic weight coefficients, and satisfy γ+δ+ε= 1; the weight coefficients can be dynamically adjusted according to different evaluation target modes.

[0037] Preferably, the decision suggestion generation process in step S4 includes: diagnosing the causes of low-value areas through a rule-based reasoning engine, locating specific material element deficiencies or non-material perception biases; and calling a generative AI model to generate specific repair or optimization suggestions in natural language form based on the diagnostic results.

[0038] Preferably, the visualization in step S4 includes generating an interactive heatmap of UVCH-VCI and its sub-indices on a WebGIS platform, and supporting users to query feature identification details, perceptual embedding vectors and their similarity analysis with the ideal prototype by clicking.

[0039] Preferably, when the program is executed by the processor, it implements all the steps of the deep learning-based urban spatial visual cultural heritage element perception and evaluation algorithm as described in any one of claims 1 to 9.

[0040] Compared with existing technologies, this invention provides a deep learning-based algorithm for perceiving and evaluating urban spatial visual cultural heritage elements, which has the following beneficial effects:

[0041] 1. This deep learning-based algorithm for assessing the perception of urban spatial visual cultural heritage elements is the first to systematically introduce cross-modal deep learning into the field of urban visual cultural heritage assessment. It transforms the abstract concept of "intangible perception," which has long relied on subjective feelings, into an objective indicator (IPRI) that can be calculated and compared through mathematical models. This fundamentally solves the core pain point of traditional methods: subjectivity. The assessment process is like "chemical analysis," breaking down complex "cultural perception" into measurable "elements," thus achieving a leap in the assessment paradigm from "empiricism" to "data-driven scientism."

[0042] 2. This deep learning-based algorithm for assessing the perception of cultural heritage elements in urban spatial vision differs from traditional methods, which often resemble a city's "physical examination" by only measuring its temperature and blood pressure (material integrity). Instead, it performs a comprehensive "CT scan + MRI + gene sequencing," not only accurately measuring material physical characteristics (PITI, VLAI) but also revealing its inherent "healthy character" and "spiritual outlook" (IPRI). For example, a newly built imitation ancient town might have a high PITI (complete material elements) but an extremely low IPRI (lacking a genuine perception of historical accumulation). This algorithm can keenly capture this significant difference in value—a "formal resemblance but not spiritual resemblance"—thus avoiding "protective destruction" or "constructive imitation."

[0043] 3. This deep learning-based algorithm for perceiving and assessing urban spatial visual cultural heritage elements, taking a 100-square-kilometer urban built-up area as an example, requires months or even years for an expert team to conduct a preliminary assessment. However, deploying the algorithm of this invention on a server with sufficient computing power allows data collection and processing to be completed within days, and model inference and assessment calculations within hours. This makes it possible to conduct annual or even quarterly routine "visual health monitoring" of the entire city, providing technical support for dynamic management and timely intervention.

[0044] 4. This invention's deep learning-based algorithm for assessing the perception of urban spatial visual cultural heritage elements provides not just a post-hoc "health check report," but a pre-emptive "surgical plan" and "risk warning." During the planning phase of urban renewal projects, decision-makers can render the 3D model of the design scheme into street view images and input them into the system for "digital sand table" simulation, predicting in advance the scope and extent of the new building's impact on the surrounding UVCH-VCI. For example, the system can warn: "The proposed 200-meter high-rise residential building on plot A will reduce the average IPRI within a 500-meter radius by 15%, significantly reducing the perception of 'openness' and 'historical continuity.'" This forward-looking analytical capability is completely lacking in traditional methods, providing unprecedented scientific basis for urban planning approval and safeguarding the city's cultural roots from the source. Detailed Implementation

[0045] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0046] Example

[0047] An Example of a Deep Learning-Based Algorithm for Perceiving and Evaluating Urban Spatial Visual Cultural Heritage Elements

[0048] A deep learning-based algorithm for perceiving and evaluating urban spatial visual cultural heritage elements includes the following steps:

[0049] S1: Multi-source heterogeneous data acquisition and preprocessing: Acquire multi-source data of the target urban space, including at least sub-meter resolution optical remote sensing images, close-up or panoramic street view images, spatiotemporal historical maps, and textual description data containing cultural semantics; perform geometric fine correction, spatiotemporal registration, data fusion and enhancement preprocessing operations on the data to construct a standardized three-dimensional spatiotemporal dataset.

[0050] S2: Multi-scale visual element deconstruction based on deep learning: Construct a dual-stream deep learning model to extract features from preprocessed data. One stream is the material entity recognition stream used for instance-level semantic segmentation to identify core material elements, and the other stream is the non-material atmosphere perception stream used to generate non-material atmosphere comprehensive perception embedding vectors.

[0051] S3: Calculation of the Comprehensive Value Index of Urban Spatial Visual Cultural Heritage (UVCH-VCI): Based on the output of step S2, calculate the Material Element Integrity and Typicality Index (PITI), Visual Landscape Aesthetics Index (VLAI), and Non-material Perceptual Resonance Index (IPRI) respectively, and obtain the final UVCH-VCI through weighted fusion calculation;

[0052] S4: Visualization of assessment results and decision support: The UVCH-VCI and its sub-indices are visualized in a multi-dimensional interactive manner on the WebGIS platform. Based on the assessment results, targeted protection and revitalization decision recommendations are automatically generated through rule-based reasoning and generative AI models.

[0053] Specifically, the material entity recognition flow in step S2 adopts an improved U-Net++ model, which integrates a CBAM attention module in the encoder part. The model is configured to perform pixel-level instance-level semantic segmentation on remote sensing images and street view images to identify and label specific categories of objects that constitute visual cultural heritage. These specific categories of objects include traditional building components, distinctive paving, ancient trees, stone inscriptions and sculptures, and water system bridges.

[0054] 3. A deep learning-based algorithm for perceiving and evaluating urban spatial visual cultural heritage elements according to claim 1, characterized in that: the non-material atmosphere perception flow in step S2 specifically includes:

[0055] S1: Use a pre-trained Vision Transformer (ViT) or Swin Transformer model to extract global visual depth features from panoramic street view images;

[0056] S2: The BERT model, which is pre-trained in the field of natural language processing (NLP) and adaptively fine-tuned in the field of cultural heritage, is used to encode the text description data and generate deep text semantic vectors.

[0057] S3: Using a cross-modal contrastive learning framework, image-text pairs spatially-co-located at the same location are used as positive samples for training to learn a shared embedding space. Based on this, a fixed-dimensional non-material atmosphere comprehensive perception embedding vector representing the overall non-material atmosphere of the location is generated through feature fusion.

[0058] Specifically, the formula for calculating the Material Element Integrity and Typicality Index (PITI) in step S3 is: PITI = α * CI + β * TD; where CI (Component Integrity) is the identification integrity score of the key material element, determined by calculating the ratio of the actual identified area to the theoretical total area; TD (Typicity Degree) is the matching score between the identified element combination and the preset cultural heritage element prototype library, determined by calculating the Fraser distance or cosine similarity between the feature vector of the element combination and the best matching prototype and then normalizing it; α and β are weighting coefficients and α + β = 1.

[0059] Specifically, the Visual Landscape Aesthetic Index (VLAI) in step S3 is calculated based on the segmentation results of the material entity recognition flow in step S2. The landscape pattern index includes at least the clustering index (AI), the spread index (CONTAG), the Shannon diversity index (SHDI), and the average fractal dimension (FRAC_MN). The landscape pattern index is then determined by weighted fusion with the predicted score of a domain aesthetic quality regression model pre-trained and fine-tuned on an aesthetic dataset.

[0060] Specifically, the Immaterial Perception Resonance Index (IPRI) in step S3 is determined by calculating the maximum cosine similarity between the immaterial atmosphere integrated perception embedding vector generated in step S2 and one or more preset "ideal cultural heritage site" perception prototype vectors.

[0061] Specifically, the final comprehensive index calculation formula in step S3 is: UVCH-VCI = γ*PITI_norm+δ*VLAI_norm +ε* IPRI_norm; where PITI_norm, VLAI_norm, and IPRI_norm are the normalized values ​​of each sub-index; γ, δ, and ε are dynamic weight coefficients, and satisfy γ+δ+ε= 1; the weight coefficients can be dynamically adjusted according to different evaluation target modes.

[0062] Specifically, the decision suggestion generation process in step S4 includes: diagnosing the causes of low-value areas through a rule-based reasoning engine, locating specific material element deficiencies or non-material perception biases; and calling a generative AI model to generate specific repair or optimization suggestions in natural language form based on the diagnostic results.

[0063] Specifically, the visualization in step S4 includes generating an interactive heatmap of UVCH-VCI and its sub-indices on the WebGIS platform, and allowing users to query feature identification details, perceptual embedding vectors, and similarity analysis with the ideal prototype by clicking.

[0064] Specifically, when the program is executed by the processor, it implements all the steps of the deep learning-based urban spatial visual cultural heritage element perception and evaluation algorithm as described in any one of claims 1 to 9.

[0065] Through the aforementioned technical solution, this invention, for the first time, systematically introduces cross-modal deep learning into the field of urban visual cultural heritage assessment, transforming the abstract concept of "intangible perception," which has long relied on subjective feelings, into an objective indicator (IPRI) that can be calculated and compared through mathematical models. This fundamentally solves the core pain point of traditional methods—subjectivity. The assessment process is like "chemical analysis," breaking down complex "cultural perception" into measurable "elements," achieving a leap in the assessment paradigm from "empiricism" to "data-driven scientism." Traditional methods are like giving a city a "physical examination," only measuring its temperature and blood pressure (material integrity). This invention, however, is like performing a comprehensive "CT + MRI + gene sequencing," not only accurately measuring material physical characteristics (PITI, VLAI) but also revealing its inherent "healthy temperament" and "spiritual outlook" (IPRI). For example, a newly built imitation ancient town may have a high PITI (complete material elements) but an extremely low IPRI (lacking a genuine perception of historical accumulation). The algorithm of this invention can keenly capture this huge value difference of "formally similar but not spiritually similar," thereby avoiding "protective destruction" or "constructive imitation." Taking a 100-square-kilometer urban built-up area as an example, assembling an expert team for preliminary assessment might take months or even years. However, deploying the algorithm of this invention, on a server with sufficient computing power, allows data collection and processing to be completed within days, and model inference and evaluation calculations within hours. This makes it possible to conduct annual or even quarterly routine "visual health monitoring" of the entire city, providing technical support for dynamic management and timely intervention. This invention no longer provides a post-event "health check report," but rather a pre-event "surgical plan" and "risk warning." During the planning stage of urban renewal projects, decision-makers can render the 3D model of the design scheme into street view images and input them into this system for "digital sand table" simulation, predicting in advance the scope and extent of the impact of new buildings on the surrounding UVCH-VCI. For example, the system can warn: "The proposed 200-meter high-rise residential building in Block A will reduce the average IPRI within a 500-meter radius by 15%, significantly reducing the perception of 'openness' and 'historical continuity.'" This forward-looking analytical capability is completely lacking in traditional methods, providing unprecedented scientific basis for urban planning approval and protecting the city's cultural roots from the source.

[0066] S1: Data Acquisition and Preprocessing

[0067] Data sources: 0.2m resolution drone oblique photography model of Pingjiang Street, Baidu panoramic street view covering the entire street, detailed survey map from the 1940s provided by Suzhou Municipal Archives, rubbing of the "Pingjiang Map" stele, and 100,000 tourist reviews crawled from Douban and Mafengwo.

[0068] Preprocessing: Oblique photogrammetry was modeled using ContextCapture Center to generate a textured 3D mesh. The 1940s map was vectorized and spatiotemporally overlaid with the modern 3D model on the Cesium platform, clearly showing historical changes such as river filling and building additions. High-precision OCR recognition of the ancient text was performed using the Transkribus tool within the PyTorch framework, and Chinese word segmentation was performed using the Jieba word segmentation tool.

[0069] S2: Multi-scale visual element deconstruction

[0070] Material Entity Recognition Flow: A U-Net++ dataset containing 5,000 labeled samples of Pingjiang Street and the surrounding ancient city of Suzhou was constructed, labeled with categories such as "small blue tile sloping roof", "white-painted wall", "granite plinth", "stone bridge", "river wharf", "camphor tree", and "osmanthus". After model training, it successfully identified that 92% of the traditional texture of "two roads and one river" (Pingjiang Road, Northeast Street, and Pingjiang River) in the core area of ​​the street is well preserved, but it was found that three key historical revetments have been replaced by modern cement slope protection.

[0071] Non-material atmospheric perception flow:

[0072] Visual feature extraction: The depth features of the panoramic image were extracted using the Swing Transformer Base model.

[0073] Text semantic analysis: A BERT model finely tuned on a "cultural heritage corpus" was used to process tourist reviews. The model successfully captured frequently co-occurring sentiment words and thematic words, such as "serene," "poetic," "small bridge, flowing water," "local life," and "commercialized."

[0074] Cross-modal fusion: After the cross-modal contrastive learning model was trained, perceptual embedding vectors were generated for each segment of the block. Analysis revealed that the vectors of the area near the river were highly similar to the prototype of "tranquil poetic dwelling"; while in the middle section of the main street (Pingjiang Road), due to the abundance of shops, the feature weights of "commercial vitality" and "crowding" in its vectors were significantly increased.

[0075] S3: UVCH-VCI Calculation

[0076] PITI calculation: The CI score of the key material element in the core area is 0.92. The TD calculation of the element combination and the prototype of "Suzhou Ancient City" is 0.95. Let α=0.5, β=0.5, then PITI = 0.5 * 0.92 + 0.5 * 0.95 = 0.935.

[0077] VLAI calculation: The landscape pattern index shows that the core area has a CONTAG index as high as 85% (high concentration of elements), an SHDI of 1.2 (simple and orderly landscape type), and a FRAC_MN of 1.48 (morphology close to natural river network). The aesthetic scoring model gives it a high average score of 4.8 / 5.0. The overall VLAI is 0.94.

[0078] IPRI calculations show that the cosine similarity between the perceptual embedding vectors of most locations in the core area and the "tranquil poetic dwelling" prototype exceeds 0.9. However, on the main street of Pingjiang Road, we calculated its similarity to the negative prototype of "over-commercialization" (trained by aggregating negative comments), and found that 20% of the sampling points on this section had a similarity higher than 0.6, which lowered the IPRI of this area.

[0079] Final Index and Zoning Assessment: A protection priority weighting of γ=0.5, δ=0.25, and ε=0.25 was adopted. The average UVCH-VCI for the entire region was approximately 0.90, belonging to the "extremely high value" category. The system further identified two low-value anomaly areas: one was the previously discovered cement revetment area (with locally extremely low PITI), and the other was the middle section of Pingjiang Road main street (with locally low IPRI).

[0080] S4: Visualization and Decision Support

[0081] Visualization: On the Cesium platform, a clear five-color heat map of red, orange, yellow, green, and blue is presented to decision-makers. The core protected areas in green and blue contrast sharply with the abnormal areas in orange and red.

[0082] Intelligent suggestion generation:

[0083] For the cement revetment area: The system rule engine triggered the "material element repair" strategy and suggested: "Immediately start the archaeological restoration of the section of revetment, and restore its original appearance of stepped stone revetment and stepped river wharf by referring to the Pingjiang Map and historical photos."

[0084] Regarding the middle section of the main street: The system-generated AI module received a diagnostic report ("The IPRI is low, with a high similarity to the 'over-commercialization' prototype, mainly due to excessive and cluttered advertising signs and excessive noise from pedestrians"), and generated the following suggestions: "It is recommended to implement a 'noise reduction and quality improvement' plan for the middle section of Pingjiang Road. First, visual burden reduction: Formulate 'Shop Sign Guidelines' to limit the size, brightness, and color saturation of signs, and promote the use of more traditional materials such as wood and fabric. Second, business format guidance: Through policy incentives, encourage the introduction of slow-paced businesses related to 'Suzhou-style life' (such as Pingtan storytelling venues, Suzhou embroidery studios, and boutique tea and bookstores) to dilute the proportion of pure retail businesses. Third, intelligent traffic diversion: Develop a mini-program to display the real-time pedestrian density of each attraction, and alleviate local congestion through virtual queuing and route recommendations, reshaping the tourist experience of 'changing scenery with every step and finding tranquility amidst the hustle and bustle'."

[0085] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A deep learning-based algorithm for perceiving and evaluating urban spatial visual cultural heritage elements, characterized in that: Includes the following steps: S1: Multi-source heterogeneous data acquisition and preprocessing: Acquire multi-source data of the target urban space, including at least sub-meter resolution optical remote sensing images, close-up or panoramic street view images, spatiotemporal historical maps, and textual description data containing cultural semantics; perform geometric fine correction, spatiotemporal registration, data fusion and enhancement preprocessing operations on the data to construct a standardized three-dimensional spatiotemporal dataset; S2: Multi-scale visual element deconstruction based on deep learning: Construct a dual-stream deep learning model to extract features from preprocessed data. One stream is the material entity recognition stream used for instance-level semantic segmentation to identify core material elements, and the other stream is the non-material atmosphere perception stream used to generate non-material atmosphere comprehensive perception embedding vectors. S3: Calculation of the Comprehensive Value Index of Urban Spatial Visual Cultural Heritage (UVCH-VCI): Based on the output of step S2, calculate the Material Element Integrity and Typicality Index (PITI), Visual Landscape Aesthetics Index (VLAI), and Non-material Perceptual Resonance Index (IPRI) respectively, and obtain the final UVCH-VCI through weighted fusion calculation; S4: Visualization of assessment results and decision support: The UVCH-VCI and its sub-indices are visualized in a multi-dimensional interactive manner on the WebGIS platform. Based on the assessment results, targeted protection and revitalization decision recommendations are automatically generated through rule-based reasoning and generative AI models.

2. The deep learning-based urban spatial visual cultural heritage element perception and evaluation algorithm according to claim 1, characterized in that: The material entity recognition flow in step S2 adopts an improved U-Net++ model, which integrates a CBAM attention module in the encoder part. The model is configured to perform pixel-level instance-level semantic segmentation on remote sensing images and street view images to identify and label specific categories of objects that constitute visual cultural heritage. These specific categories of objects include traditional building components, distinctive paving, ancient trees, stone inscriptions and sculptures, and water system bridges.

3. The deep learning-based urban spatial visual cultural heritage element perception and evaluation algorithm according to claim 1, characterized in that: The non-material atmosphere sensing flow mentioned in step S2 specifically includes: S1: Use a pre-trained Vision Transformer (ViT) or Swin Transformer model to extract global visual depth features from panoramic street view images; S2: The BERT model, which is pre-trained in the field of natural language processing (NLP) and adaptively fine-tuned in the field of cultural heritage, is used to encode the text description data and generate deep text semantic vectors. S3: Using a cross-modal contrastive learning framework, image-text pairs spatially-co-located at the same location are used as positive samples for training to learn a shared embedding space. Based on this, a fixed-dimensional non-material atmosphere comprehensive perception embedding vector representing the overall non-material atmosphere of the location is generated through feature fusion.

4. The deep learning-based urban spatial visual cultural heritage element perception and evaluation algorithm according to claim 1, characterized in that: The formula for calculating the Material Element Integrity and Typicality Index (PITI) in step S3 is: PITI = α * CI + β * TD; where CI (Component Integrity) is the identification integrity score of the key material element, which is determined by calculating the ratio of the actual identified area to the theoretical total area; TD (Typicity Degree) is the matching score between the identified element combination and the preset cultural heritage element prototype library, which is determined by calculating the Fraser distance or cosine similarity between the feature vector of the element combination and the best matching prototype and then normalizing it; α and β are weighting coefficients and α + β = 1.

5. The deep learning-based algorithm for perceiving and evaluating urban spatial visual cultural heritage elements according to claim 1, characterized in that: The Visual Landscape Aesthetic Index (VLAI) mentioned in step S3 is calculated based on the segmentation results of the material entity recognition flow in step S2. The landscape pattern index includes at least the clustering index (AI), the spread index (CONTAG), the Shannon diversity index (SHDI), and the average fractal dimension (FRAC_MN). The landscape pattern index is then determined by weighted fusion with the predicted score of a domain aesthetic quality regression model pre-trained and fine-tuned on an aesthetic dataset.

6. The deep learning-based urban spatial visual cultural heritage element perception and evaluation algorithm according to claim 1, characterized in that: The Immaterial Perception Resonance Index (IPRI) mentioned in step S3 is determined by calculating the maximum cosine similarity between the immaterial atmosphere integrated perception embedding vector generated in step S2 and one or more preset "ideal cultural heritage site" perception prototype vectors.

7. The deep learning-based urban spatial visual cultural heritage element perception and evaluation algorithm according to claim 1, characterized in that: The final comprehensive index calculation formula mentioned in step S3 is: UVCH-VCI = γ*PITI_norm+δ*VLAI_norm +ε* IPRI_norm; where PITI_norm, VLAI_norm, and IPRI_norm are the normalized values ​​of each sub-index; γ, δ, and ε are dynamic weight coefficients, and satisfy γ+δ+ε= 1; the weight coefficients can be dynamically adjusted according to different evaluation target modes.

8. The deep learning-based algorithm for perceiving and evaluating urban spatial visual cultural heritage elements according to claim 1, characterized in that: The decision suggestion generation process in step S4 includes: diagnosing the causes of low-value areas through a rule-based reasoning engine, locating specific material element deficiencies or non-material perception biases; and calling a generative AI model to generate specific repair or optimization suggestions in natural language form based on the diagnostic results.

9. The deep learning-based algorithm for perceiving and evaluating urban spatial visual cultural heritage elements according to claim 1, characterized in that: The visualization in step S4 includes generating an interactive heatmap of UVCH-VCI and its sub-indices on the WebGIS platform, and supporting users to query feature identification details, perceptual embedding vectors and their similarity analysis with the ideal prototype by clicking.

10. The deep learning-based urban spatial visual cultural heritage element perception and evaluation algorithm according to claim 1, characterized in that: When the program is executed by the processor, it implements all the steps of the deep learning-based urban spatial visual cultural heritage element perception and evaluation algorithm as described in any one of claims 1 to 9.