Historical and cultural block building style coordination image evaluation method and system

CN122289266BActive Publication Date: 2026-08-07XIAN CONSTR SCI & TECH UNIV ARCHITECTURAL DESIGN INST
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAN CONSTR SCI & TECH UNIV ARCHITECTURAL DESIGN INST
Filing Date
2026-05-25
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]针对现有技术中历史文化街区建筑风貌协调性审查依赖主观判断、不同专家评判标准不统一以及现有评价方法特征维度不完整的技术问题,本发明提供历史文化街区建筑风貌协调性图像评估方法及系统

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122289266B_ABST
    Figure CN122289266B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image processing and urban planning, and provides a historical and cultural block building style coordination image evaluation method and system, which comprises the following steps: collecting street view panoramic images and generating multi-view expanded elevation image sequences; obtaining each building elevation area mask by an instance segmentation network; extracting style features from five dimensions of color, material, proportion, contour and detail; constructing a style benchmark image based on the statistical distribution of the five-dimensional features of historical protected buildings; and calculating the coordination deviation degree of each dimension by Mahalanobis distance and giving a weighted comprehensive score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of image processing and urban planning technology, and in particular to an image evaluation method and system for the harmony of architectural style in historical and cultural blocks. Background Technology

[0002] Historic and cultural districts, as spatial carriers of urban heritage, bear irreplaceable historical information and cultural value. In the process of rapid urbanization, these districts face the dual pressures of protection and development, with the issue of stylistic harmony between newly constructed or renovated buildings and existing historical building complexes becoming increasingly prominent. Ensuring that new and renovated buildings maintain harmony with surrounding historical buildings in terms of roof form, facade color, material texture, and building proportions is one of the core tasks of the protective planning review of historic and cultural districts. However, currently, this review process mainly relies on subjective judgments by planning approval experts through on-site inspections and comparison of renderings. Different experts have significantly different focuses in their assessment of harmony. Some experts focus on whether the roof form and cornice continue the traditional skyline characteristics of the district, while others emphasize whether the facade materials and colors are harmoniously unified with the surrounding historical buildings. This dispersion of evaluation standards leads to diametrically opposed conclusions for the same renovation plan in different reviews, forcing developers to repeatedly revise their plans and often extending the project cycle by several months.

[0003] Chinese invention application CN119477804A discloses a method and system for evaluating the integrity of historical urban landscape based on deep learning. This method acquires urban streetscape images from multiple locations, uses a pre-trained convolutional neural network model to extract sky ratio and green view rate as spatial features at the street scale and as spatial features of natural elements, calculates the number of colors and color harmony of shop signs using a color baseline database as spatial features of building color, calculates the proportion of traditional materials and material consistency using a sample set of building materials as spatial features of material mechanism, and obtains expert scores based on the ELO scoring algorithm to train an artificial neural network evaluation model to achieve the evaluation of historical urban landscape integrity. This method has some reference value in the multi-dimensional feature extraction of urban street scene images, but it has the following shortcomings: First, this method is aimed at the macro-evaluation of the integrity of the historical urban landscape, focusing on the overall preservation of the landscape at the street level, rather than the coordination review between a single newly built or renovated building and the surrounding historical building complex. It cannot provide a quantitative basis for case-by-case assessment of landscape coordination in planning approval. Second, the feature extraction dimensions of this method do not include key elements that have a significant impact on landscape coordination, such as building outline, facade proportion, and decorative details, and cannot comprehensively depict the multi-dimensional features of architectural landscape. Third, this method relies on expert scoring to train the evaluation model, and is still affected by differences in expert subjective judgment. It lacks a numerical indicator system that quantifies the landscape elements of each dimension into an objectively comparable system. Different experts have significantly different focuses in their assessment of coordination, and the same renovation plan may obtain completely opposite conclusions in different reviews.

[0004] Furthermore, existing planning control methods primarily rely on qualitative textual descriptions, such as regulations stipulating that building height should not exceed the eaves height of adjacent historical buildings and that facade colors should use warm gray tones. These lack quantifiable numerical indicators that can be objectively compared, making it difficult to guarantee the objectivity and consistency of planning approvals. Developers are often forced to repeatedly revise plans due to inconsistent expert reviews, extending project cycles by several months and increasing the consumption of social resources. Meanwhile, while existing deep learning-based urban streetscape evaluation methods have made some progress in automatic image feature extraction, these methods mostly focus on overall street quality scoring, using macro-scale spatial features as evaluation input. They fail to establish an instance-level analysis framework with individual buildings as evaluation units, nor do they form a coordinated comparison mechanism based on the statistical distribution of historical buildings. Regarding building outline analysis, existing methods typically extract overall indicators such as sky ratio or building coverage through semantic segmentation, but they lack frequency domain feature description methods capable of finely depicting differences in rooftop skyline morphology. Their ability to differentiate between pitched and flat roofs and to quantify the continuity of eaves lines is insufficient. In terms of evaluation dimensions, existing methods lack effective quantitative means for elements that significantly influence the perception of architectural style harmony, such as the proportional relationships of building facades and the richness of decorative details. Therefore, there is an urgent need for an objective, multi-dimensional, and quantifiable evaluation method for the architectural style harmony of historical and cultural blocks based on image processing technology, in order to solve the technical problems of strong subjectivity in harmony review, inconsistent evaluation standards, and incomplete feature dimensions in existing technologies. Summary of the Invention

[0005] To address the technical problems in existing technologies, such as reliance on subjective judgment in the review of architectural style harmony in historical and cultural blocks, inconsistent evaluation standards among different experts, and incomplete characteristic dimensions in existing evaluation methods, this invention provides an image-based evaluation method and system for architectural style harmony in historical and cultural blocks.

[0006] This invention discloses an image evaluation method for the architectural style harmony of historical and cultural blocks, comprising the following steps: Step S1, Street view panoramic image acquisition and multi-view facade image sequence generation: A panoramic camera is used to acquire street view panoramic images segment by segment along each street of the historical and cultural district to obtain a panoramic image dataset containing information on the facades of buildings on both sides of each street; perspective projection transformation is performed on each frame of the panoramic image dataset to generate a multi-view unfolded facade image sequence along each street direction.

[0007] Step S2, Building facade instance segmentation and facade region mask extraction: The instance segmentation network is used to segment the building facade instances in each frame of the multi-view unfolded facade image sequence. The facade of each building in each frame is segmented independently to obtain the facade region mask of each building.

[0008] Step S3, Automatic Quantization Extraction of Five-Dimensional Features of Building Appearance: Based on the facade area mask, the appearance features of each building facade are quantitatively extracted in five dimensions: color, material, proportion, outline and detail.

[0009] Step S4, Construction of the Streetscape Benchmark Profile: Using all the identified historical buildings in the street as the benchmark building group, the characteristic value distribution of the benchmark building group in five dimensions is statistically analyzed to construct the streetscape benchmark profile.

[0010] Step S5, Calculation and comprehensive evaluation of style coordination deviation: The renderings of the new or renovated plan are processed by steps S2 to S3 to obtain their five-dimensional feature values. The Mahalanobis distance between the feature values ​​of each dimension and the style benchmark image is calculated as the coordination deviation of each dimension. The weighted comprehensive score of the five-dimensional coordination deviation is used as the overall style coordination index.

[0011] The aforementioned method, driven by a combination of multi-view streetscape image acquisition, automatic extraction of multi-dimensional architectural features, and quantitative assessment of harmony, achieves an objective review of architectural style harmony. The five-dimensional feature extraction covers key elements affecting style harmony, such as color, material, proportion, outline, and details. Using the statistical distribution of historical buildings in the block as a baseline portrait of the style, and employing Mahalanobis distance to measure the degree of harmony deviation in each dimension, it eliminates interference from differences in the dimensions of different dimensions and correlations between features, providing objective, unified, and quantifiable technical support for planning approval. Compared with existing technologies, this invention has at least the following technical advantages: First, by establishing an instance-level analysis framework with individual buildings as the evaluation unit, the object of style harmony review is refined from the macro-street level to the individual building level, so that the evaluation conclusions can directly correspond to the specific construction project approval requirements; Second, by using a five-dimensional feature quantification system, qualitative concepts in architectural style are transformed into calculable numerical indicators, eliminating the problem of inconsistent conclusions caused by differences in the evaluation focus of different review experts; Third, by providing information on the deviation direction and magnitude of each dimension through radar chart visualization analysis, designers can make targeted adjustments to the scheme, significantly shortening the cycle of scheme iteration and optimization.

[0012] This invention also provides an image evaluation system for the architectural style harmony of historical and cultural blocks, including an image acquisition and projection module, a facade segmentation module, a feature extraction module, a baseline profile construction module, and a harmony evaluation module. Each module corresponds to the function of steps S1 to S5 in the above method. This system achieves decoupling and flexible combination of processing links through a modular architecture design, supports one-time construction of baseline profiles and batch evaluation of multiple schemes, and is suitable for the daily business scenarios of planning approval agencies and architectural design agencies. Attached Figure Description

[0013] Figure 1 This is a flowchart of the image evaluation method for the architectural style harmony of historical and cultural blocks provided in the embodiments of the present invention.

[0014] Figure 2 This is an architecture diagram of the image evaluation system for the architectural style coordination of historical and cultural blocks provided in this embodiment of the invention. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort should fall within the scope of protection of the present invention.

[0016] See Figure 1 This invention provides an image-based assessment method for the architectural style harmony of historical and cultural blocks. This method is driven by a combination of multi-view streetscape image acquisition, automatic extraction of multi-dimensional architectural style features, and quantitative assessment of harmony, enabling an objective review of the stylistic harmony between newly built or renovated buildings and existing historical building complexes. In one embodiment of this invention, a historical and cultural block (containing a core protected area of ​​approximately 120 identified historical buildings) is used as an application scenario for illustration. The method specifically includes the following steps: Step S1: Street view panoramic image acquisition and multi-view facade image sequence generation. The goal of this step is to obtain complete facade visual information of each street in the historical and cultural district, providing a high-quality image data foundation for subsequent independent segmentation of building facades and multi-dimensional feature extraction.

[0017] Specifically, a vehicle-mounted or handheld panoramic camera is used to collect panoramic street view images segment by segment along the streets of the historical and cultural district. Preferably, the panoramic camera is mounted on the roof of a low-speed acquisition vehicle, which travels at a constant speed along the streets of the district, and the panoramic camera automatically triggers shooting at fixed intervals. In one embodiment of the invention, the acquisition interval is set to 8m, the panoramic image resolution is 8192×4096 pixels, the image is stored in an equidistant cylindrical projection format, covering a 360° field of view horizontally and a 180° field of view vertically. In another embodiment of the invention, the acquisition interval can be selected within the range of 5m to 15m based on the street width and building facade density: when the acquisition interval is less than 5m, the visual overlap between adjacent panoramic images is too high, resulting in a significant increase in data redundancy and a significant increase in the overall acquisition and subsequent processing time; when the acquisition interval is greater than 15m, there are blind spots in the visual perspective of building facades between adjacent acquisition points, which may result in the omission of some building facade information or insufficient pixel coverage of distant buildings. Preferably, for alleys or lanes less than 6m wide, a smaller spacing of 5m to 8m (e.g., 5m, 6m, or 8m) is used for handheld walking data collection, while for main streets wider than 6m, a larger spacing of 10m to 15m (e.g., 10m, 12m, or 15m) is used for vehicle-mounted data collection. In this embodiment, the vehicle-mounted data collection spacing is preferably 8m, and the handheld data collection spacing is preferably 5m. In one embodiment of the present invention, the resolution of the panoramic image is not less than 4096×2048 pixels: when the panoramic image resolution is less than 4096×2048 pixels, the number of pixels occupied by a single building facade in the panoramic image is insufficient to support the subsequent feature extraction requirements for the recognition of facade texture details and decorative components, and the accuracy of gray-level co-occurrence matrix calculation in the material dimension and the detection accuracy of decorative components in the detail dimension will decrease significantly; the higher the resolution, the more sufficient the pixel coverage area of ​​a single building, but correspondingly, the image storage and processing overhead increases. Preferably, the panoramic image resolution is selected as 4096×2048 pixels, 6144×3072 pixels, or 8192×4096 pixels, with 8192×4096 pixels being preferred in this embodiment. The panoramic camera is preferably installed at a height of 2.5m to 3.0m above the ground. This height range avoids obstruction of the building facade by pedestrians and vehicles that are too close, while maintaining a viewing angle that is basically consistent with the human eye's observation of the building facade, ensuring that the visual characteristics of the acquired images match the on-site observations during planning review. For narrow alleys or lanes where vehicles cannot pass, handheld panoramic cameras can be used for supplementary data collection on foot, with a handheld collection interval preferably of 5m. Data collection should ideally be conducted during sunny, evenly lit periods to reduce the impact of shadows and glare on the accuracy of subsequent color feature extraction. Preferably, data collection is conducted between 9:00 AM and 3:00 PM when the cloud cover does not exceed 30%.In one embodiment of the present invention, panoramic images of approximately 3.6 km of streets in the historical and cultural district were acquired, resulting in approximately 450 panoramic images.

[0018] After acquiring the panoramic image dataset, perspective projection transformation is performed on each frame of the panoramic image to generate a sequence of multi-view unfolded facade images for each street direction. Panoramic images in equidistant cylindrical projection format suffer from severe perspective distortion, and directly using them for building facade analysis can lead to significant feature extraction errors. Therefore, this step performs perspective projection transformation on each frame of the panoramic image according to the street direction, extracting multiple perspective projection images from different viewpoints sequentially at preset angle intervals in the horizontal direction. In one embodiment of this invention, the preset angle interval is 60°, meaning that six perspective projection images are extracted from each frame of the panoramic image along the horizontal direction. Each perspective projection image has a horizontal field of view of 90° and a vertical field of view of 90°, with an output resolution of 1024×1024 pixels. In one embodiment of the present invention, the preset angle interval can be selected within the range of 30° to 90°. When the preset angle interval is less than 30°, the number of perspective projection images generated for each frame of panoramic image is too large (more than 12), the image overlap between adjacent viewpoints is too high, the data redundancy is significantly increased, and the processing time is increased. When the preset angle interval is greater than 90°, in order to cover the visual area between adjacent viewpoints, the horizontal field of view of a single perspective projection image needs to be increased to more than 90°. At this time, the edge distortion of the perspective projection will increase sharply, affecting the accuracy of subsequent building facade instance segmentation and landscape feature extraction. Preferably, the preset angle interval is 30°, 45°, 60°, or 90°, corresponding to the extraction of 12, 8, 6, or 4 perspective projection images per frame of panoramic image, respectively. In this embodiment, the preset angle interval is preferably 60°. The mathematical essence of perspective projection transformation is to map the panoramic image in the spherical coordinate system to the perspective image in the planar coordinate system. During the transformation process, three parameters of the virtual camera need to be specified: focal length, pitch angle, and yaw angle. Preferably, the pitch angle of the virtual camera is set to 0° to prevent tilt distortion of the vertical lines of the building facade. In one embodiment of the invention, the focal length of the virtual camera is automatically calculated based on the output resolution and field of view. The calculation relationship is that the focal length equals half the width of the output image divided by the tangent of half the field of view. When the field of view is 90°, the focal length is exactly equal to half the width of the output image, i.e., 512 pixels. The yaw angle is set according to the street direction to ensure that the generated perspective projection image is directly facing the building facades on both sides of the street. To obtain the street direction information, in one embodiment of the invention, the street direction angle at the location of the panoramic camera is calculated using road network geographic information data, and then the yaw angle is set to the perpendicular direction of the street direction angle to face the building facades on both sides. After the above processing, approximately 2700 frames of multi-view unfolded facade images are generated in this embodiment, forming a facade image sequence covering the entire block.

[0019] It is worth noting that the data acquisition quality in step S1 directly affects the processing accuracy of all subsequent steps. The resolution of the panoramic image determines the pixel coverage area of ​​a single building facade, while the field of view parameter of the perspective projection affects the degree of geometric distortion of the facade image. In one embodiment of the present invention, when the street width is less than 6m, the horizontal field of view is adjusted to 120° to ensure that the facades of buildings on both sides can be fully included in the field of view, and the output resolution is increased accordingly to 1280×1280 pixels to compensate for the decrease in pixel density caused by the increase in the field of view. Preferably, after the perspective projection transformation, image quality screening is also performed to mark and remove image frames with severe motion blur (with a Laplacian variance of less than 100 as the judgment threshold), overexposure, or underexposure (with an average pixel brightness exceeding the range of 30 to 225 as the judgment threshold), ensuring that the facade images entering the subsequent processing flow all have the clarity and color fidelity required for feature extraction.

[0020] Furthermore, to address the common problem of street trees obscuring building facades in historical and cultural districts, one embodiment of this invention introduces vegetation occlusion marking processing in step S1. Specifically, a semantic segmentation network is used to detect vegetation areas in each frame of the unfolded facade image. When the occlusion area of ​​a building facade exceeds 50%, the system automatically selects the frame with the smallest occlusion ratio from other viewpoints of that building as a replacement. If the occlusion ratio of all available frames exceeds 50%, the building is marked as partially occluded, and only the unoccluded areas are used for calculation in subsequent feature extraction. This processing strategy effectively reduces the interference of street trees on the accuracy of building facade feature extraction and improves the applicability and robustness of street view data in real-world scenarios.

[0021] Step S2, Building Facade Instance Segmentation and Facade Region Mask Extraction. The goal of this step is to perform instance-level segmentation of the building facades in each frame of the multi-view unfolded facade image sequence generated in Step S1, and generate an independent facade region mask for each building, thereby providing pixel-level spatial positioning basis for subsequent five-dimensional feature quantization extraction of each building.

[0022] Specifically, an instance segmentation network is used to process the facade images of each frame. In one embodiment of the invention, an instance segmentation network based on the Mask R-CNN architecture is employed, with ResNet-101 as the backbone feature extraction network. After pre-training on the COCO dataset, fine-tuning training is performed using a dedicated historical district dataset annotated with building facade instance masks. This dedicated dataset contains approximately 2000 annotated images, with each building facade in each image annotated with an independent instance mask. During fine-tuning training, the learning rate is set to 0.001, the training epochs are 30, and the batch size is 4. To improve the robustness of the instance segmentation network in complex historical street scenes, one embodiment of this invention employs multiple data augmentation strategies during the training phase, including random horizontal flipping, random brightness and contrast adjustment (within ±20% of the original values), random Gaussian blur (with a standard deviation randomly selected between 0.5 and 2.0), and random cropping and scaling (with a scaling ratio between 0.8 and 1.2). These data augmentation strategies effectively increase the diversity of training samples and mitigate the risk of overfitting the model on limited labeled data. After training, the instance segmentation network achieves an average precision of 0.78 on the test set, with an AP value of 0.86 at an IoU threshold of 0.5.

[0023] In practical reasoning, the facade images of buildings in historical and cultural districts possess unique characteristics that distinguish them from typical urban streetscapes, requiring attention during segmentation. First, the facades of historical buildings are often irregular, with some featuring curved gables, projecting balconies, or protruding porches, resulting in complex facade mask shapes that necessitate sophisticated edge tracking capabilities from the instance segmentation network. Second, historical districts often have high building density, with adjacent buildings sometimes sharing walls. In such cases, the instance segmentation network must accurately distinguish between closely adjacent building instances rather than merging them into a single continuous facade region. Third, common street-front shop signs, awnings, and air conditioner units in historical districts can partially obscure building facades, requiring the instance segmentation network to correctly identify the complete outline of the facade even under partial occlusion. In one embodiment of this invention, by incorporating samples of these complex scenarios into a dedicated training dataset and explicitly requiring annotators to exclude attachments from the facade mask in the annotation specifications, the adaptability of the trained instance segmentation network to the unique scenarios of historical districts is ensured.

[0024] During the inference phase, the unfolded facade image of each frame generated in step S1 is input into the instance segmentation network. The network output includes the bounding box coordinates, category label, and pixel-level instance mask for each building. Preferably, detection results with a confidence score below 0.7 are filtered to reduce the interference of false detections on subsequent feature extraction. In one embodiment of the present invention, an inter-frame building instance association mechanism is also introduced to perform instance matching and mask fusion on the same building detected in adjacent frame images. Specifically, instances with an overlap of bounding boxes (IoU) greater than 0.5 in adjacent frames are matched, and the multi-frame masks of successfully matched instances are jointly processed. The frame with the largest pixel coverage area is selected as the main view facade area mask for that building. This inter-frame association mechanism effectively avoids the problem of the same building being repeatedly calculated at adjacent acquisition points, ensuring the accuracy of subsequent feature statistics.

[0025] After the above processing, this embodiment performs instance segmentation on approximately 2700 frames of unfolded facade images of the block, detecting facade instances of approximately 860 independent buildings, of which approximately 120 are identified historical protected buildings. Each building obtains a corresponding facade area mask, which is stored in the form of a binary image, with pixel values ​​inside the mask being 1 and pixel values ​​outside the mask being 0.

[0026] It should be further noted that the quality of the training data for the instance segmentation network has a crucial impact on segmentation accuracy. In one embodiment of this invention, the annotation of the dedicated training dataset employs a multi-level quality control strategy: the first level involves annotators annotating the building facade boundaries pixel by pixel; the second level involves quality inspectors cross-checking the annotation results and correcting annotations with boundary deviations exceeding 5 pixels. Special attention needs to be paid to the boundary demarcation between adjacent buildings during the annotation process, especially in row-style historical buildings where adjacent buildings often share gable walls. In such cases, visual cues such as house numbers, building height differences, and changes in facade materials are required to determine the segmentation boundaries. Furthermore, for buildings with large areas of signage or advertising obstructing the facade, the signage and advertising areas are excluded from the facade area mask during annotation to ensure that subsequent color and material feature extraction reflects the facade features of the building itself rather than the features of the attached structures.

[0027] Step S3: Automatic Quantization and Extraction of Five-Dimensional Features of Architectural Style. This step, based on the masks of each building facade area obtained in Step S2, performs automatic quantization and extraction of architectural style features for each building facade from five dimensions: color, material, proportion, outline, and detail. These five dimensions cover key visual elements affecting the harmony of architectural style. The specific feature extraction methods for each dimension are as follows: (1) Color Dimension: The color dimension aims to describe the color tone and color richness of the building facade. In one embodiment of the present invention, the pixels in the facade area mask are first converted from the RGB color space to the CIE-Lab color space. The L channel of the CIE-Lab color space represents lightness, the a channel represents red-green hue, and the b channel represents yellow-blue hue. Its design conforms to the uniformity of human visual perception. The technical reason for using the CIE-Lab color space instead of directly using the RGB color space is that there is a strong correlation between the three channels in the RGB color space, and the Euclidean distance in the RGB space is not a linear correspondence with the color difference perceived by the human eye. This will lead to a deviation between the measurement of color difference in the subsequent coordination deviation calculation and the human visual judgment. The CIE-Lab color space is a color space with uniform perception. In this space, color differences at equal distances correspond to equal color differences perceived by the human eye, so that the color feature values ​​based on the CIE-Lab space can more accurately reflect the human visual perception of the color coordination of the building. Preferably, before the color space conversion, histogram equalization processing is performed on the pixels in the facade area mask to eliminate the influence of uneven illumination. Furthermore, in one embodiment of the present invention, shadow pixels within the mask of the facade area are detected and excluded. When the L channel value of a certain pixel is less than 20, it is determined to be a deep shadow pixel and excluded from subsequent color statistics, so as to avoid the dark pixels in the shadow area from causing a shift in the mean value of the main color tone.

[0028] In the CIE-Lab color space, calculate the mean value of the a channel of all pixels within the mask of the facade area. and b channel mean This serves as a description of the building's primary color scheme. The physical meaning of the primary color scheme is: A higher value indicates a redder hue. The smaller the value, the more greenish the hue. A higher value indicates a more yellow hue. A smaller value indicates a blue tint. Traditional buildings in historical and cultural districts are typically predominantly warm gray, which is represented in the CIE-Lab color space as... The value is between 2 and 8. The value ranges from 8 to 18.

[0029] Color dispersion is used to describe the richness and complexity of colors on a facade. Its calculation formula is: ,in: This refers to color dispersion, a dimensionless numerical value with a range of values. The larger the value, the richer or more mixed the colors of the facade. , , These are the standard deviations of all pixels within the mask of the facade area in the L, a, and b channels, respectively, with units consistent with the numerical units of each channel; , , These are the weighting coefficients for channel L, channel a, and channel b, respectively. Preferably, , , The basis for this weighting is that the human eye is more sensitive to changes in chroma (a and b channels) than to changes in lightness (L channel), and in the assessment of appearance harmony, the deviation of chroma has a more significant impact on visual perception.

[0030] (2) Material dimension: The material dimension aims to distinguish different material types of building facades, such as painted walls, exposed brick walls, and stone curtain walls. In one embodiment of the present invention, the image within the facade area mask is converted into a grayscale image, and texture features are extracted using a grayscale co-occurrence matrix.

[0031] In the calculation of the gray-level co-occurrence matrix (GLCM), the gray levels are quantized into 64 levels to balance computational efficiency and feature resolution. The computation window size is set to 64×64 pixels, the stride (pixel offset) is 1 pixel, and the computation directions include four directions: 0°, 45°, 90°, and 135°. Two core features are extracted from the GLCM: Contrast as an indicator of roughness: ,in: This is a roughness index value, dimensionless, with a range of values ​​of [value missing]. , In this embodiment, the number of gray levels is... ; For angle The first gray-level co-occurrence matrix in the direction The normalized value of each element; and These are the row and column indices of the gray-level co-occurrence matrix, respectively. A larger roughness value indicates a rougher texture. The roughness of exposed brick walls is usually between 150 and 300, while the roughness of painted walls is usually between 20 and 80.

[0032] The second moment of an angle is used as an index of regularity: ,in: The regularity index is dimensionless and its value ranges from [value missing]. The higher the value, the more regular and uniform the texture. Stone curtain walls have a higher regularity value (usually between 0.05 and 0.12) due to the regular distribution of joints, while old painted walls have a lower regularity value (usually between 0.008 and 0.025) due to surface peeling and mottled.

[0033] The roughness and regularity mentioned above together constitute the feature vector of the material dimension. It is worth noting that the technical effect of averaging the gray-level co-occurrence matrices in four directions is to eliminate the influence of texture directionality on feature values, making the features robust to the shooting angle of the building facade. In actual historical and cultural district scenarios, the material type of building facades is closely related to their construction period and building function. For example, traditional residences built in the late Qing Dynasty and early Republic of China often use exposed brick facades, which have high roughness and regularity; while public buildings built in the 1950s and 1960s often use terrazzo or faux stone facades, which have medium roughness and low regularity. This invention, through a two-dimensional joint representation of roughness and regularity, can effectively distinguish the above-mentioned different material types and provide a reliable material similarity measurement basis for subsequent coordination assessment. In addition, in one embodiment of this invention, the gray-level quantization parameters of the gray-level co-occurrence matrix were optimized. Through comparative experiments on different gray-level quantization parameters (including 16 levels, 32 levels, 64 levels, and 128 levels), it was found that 64-level gray quantization achieved the best balance between material discrimination ability and computational efficiency. Too few gray levels will cause the texture differences between different materials to be masked by quantization errors, while too many gray levels will cause the gray-level co-occurrence matrix to be too sparse, making the statistical features unstable.

[0034] (3) Proportional dimension: The proportional dimension aims to describe the overall shape proportions and window density of the building facade. In one embodiment of the present invention, the proportional dimension includes two indicators: the width-to-height ratio of the building facade and the window-to-wall area ratio.

[0035] The method for calculating the width-to-height ratio of a building facade is as follows: Calculate the minimum bounding rectangle based on the facade area mask obtained in step S2, and take the ratio of the width to the height of the minimum bounding rectangle as the width-to-height ratio. Preferably, when calculating the minimum bounding rectangle, a morphological closing operation is performed on the mask edges to fill the small holes, and the structuring element of the closing operation is a 5×5 rectangle. The aspect ratio of traditional buildings in historical and cultural districts is typically between 1.2 and 3.0, while the aspect ratio of modern high-rise buildings is typically less than 0.8.

[0036] The window-to-wall area ratio is calculated as follows: a window detection model is used to detect windows within the facade area. In one embodiment of this invention, the window detection model adopts the YOLOv8 architecture and is trained using approximately 5000 facade images labeled with window locations, achieving a detection accuracy (AP@0.5) of 0.82. The areas of all detected window regions are summed, and the ratio of the total window area to the total area of ​​the facade area mask is calculated as the window-to-wall area ratio. The range of values ​​is The window-to-wall area ratio of traditional historical buildings is usually between 0.15 and 0.35, while the window-to-wall area ratio of modern buildings that use glass curtain walls extensively can reach 0.6 or higher.

[0037] The feature vector of the proportional dimension is In one embodiment of the present invention, it should be noted that the calculation results of the aspect ratio and window-to-wall area ratio are affected by the projection angle and cropping range of the building facade in the image. To eliminate the interference of shooting distance and angle differences on the proportion calculation, the present invention performs perspective correction processing on the facade area mask before calculating the aspect ratio, and uses the coordinates of the four corner points of the facade area mask to perform affine transformation to correct the facade to a front view state. In addition, when the facade area mask range of the same building differs in different frame images (for example, only a local area of ​​the facade is captured when shooting at close range), the present invention uses the proportion calculation result of the frame with the largest mask area as the representative value of the building, to ensure that the proportion features reflect the overall form of the building facade rather than a local cropped segment. In the practice of protection planning of historical and cultural blocks, the aspect ratio of buildings is related to the continuity and sense of enclosure of the street interface, and the window-to-wall area ratio is related to the contrast rhythm of the facade. Both are important visual parameters that affect pedestrians' subjective perception of the quality of street space.

[0038] (4) Outline dimension: The outline dimension aims to describe the morphological characteristics of the building roof skyline and the undulating rhythm of the eaves line. These two elements are key factors affecting the harmony of the skyline of historical and cultural districts.

[0039] The method for extracting the roof skyline morphology features is as follows: The upper edge contour of the facade area mask obtained in step S2 is tracked to obtain the roof skyline curve. Specifically, the facade area mask is scanned row by row from top to bottom, and the coordinates of the highest point of the mask area in each column of pixels are recorded. The coordinates of the highest points in each column are then connected to form the skyline curve. Subsequently, a one-dimensional discrete Fourier transform is performed on the skyline curve to obtain the Fourier descriptor sequence. In one embodiment of the present invention, the preceding part is truncated. The Fourier descriptor of order 1 is then normalized (i.e., divided by 1 / 2). ), to obtain the normalized Fourier descriptor vector In this vector, lower-order descriptors reflect the overall shape of the skyline (e.g., pitched roofs are triangular, flat roofs are horizontal), while higher-order descriptors reflect the detailed undulations of the skyline (e.g., the influence of ridge ornaments, chimneys, and other protrusions). Amplitude normalization of the Fourier descriptors makes the features invariant to scale variations of the building facade in the image.

[0040] The method for calculating the undulation rhythm value of the eaves line is as follows: extract the horizontal line segment of the upper edge of the mask in the facade area within a specified height range as the eaves line, and calculate the standard deviation of the vertical coordinate sequence of the eaves line. This value serves as a description of the undulating rhythm of the eaves line. In one embodiment of the invention, the specified height range is the area 10% above the height of the facade mask. A smaller value indicates a straighter eaves line, while a larger value indicates a more pronounced undulation in the eaves line. Well-preserved traditional blocks in historical and cultural districts typically have continuous, straight eaves lines. The value is usually less than 8 pixels (corresponding to a height deviation of about 0.3m in the actual building).

[0041] The feature vector of the contour dimension is jointly constructed by the normalized Fourier descriptor vector and the undulation rhythm value of the eaves line. The technical advantage of using Fourier descriptors to describe the roof skyline morphology in this invention lies in the fact that Fourier descriptors possess translation invariance, rotation invariance, and scale invariance (after normalization), making the skyline morphological features unaffected by the building's position, orientation, and scale in the image. Furthermore, the truncation order of the Fourier descriptor provides a flexible mechanism for controlling the granularity of morphological description; lower truncation orders focus on the macroscopic morphological features of the skyline (such as overall slope and symmetry), while higher truncation orders can capture the details of local undulations in the skyline. The basis for selecting the first 16 orders of Fourier descriptors in this invention is that, after analyzing the skyline morphology of approximately 120 historical buildings in the block, it was found that the first 16 orders of descriptors could explain more than 95% of the total variance of the skyline morphology; further increasing the descriptor order would bring negligible morphological information gain while significantly increasing computational overhead. In one embodiment of the present invention, N can be selected within the integer range of 8 to 32: when N is less than 8, the truncated Fourier descriptor cannot fully describe the main frequency domain features of the roof skyline morphology, and its ability to distinguish between different roof forms such as pitched roofs and flat roofs is significantly insufficient, making it difficult to support the assessment of the harmony of the outline dimension; when N is greater than 32, the proportion of high-frequency noise components introduced by the higher-order Fourier descriptor increases, which in turn interferes with the stability of the skyline morphology features and increases the computational cost of subsequent Mahalanobis distance calculation. Preferably, for blocks dominated by a single roof form (e.g., blocks that are all flat roofs or all pitched roofs), N is taken as 8 to 12; for blocks with diverse roof forms (e.g., blocks that simultaneously include multiple roof forms such as pitched roofs, flat roofs, and hip roofs), N is taken as 16 to 32. In this embodiment, N is preferably taken as 16. In real-world scenarios, the Fourier descriptor of the skyline of pitched roof buildings exhibits a clear low-order component dominance (the amplitudes of the 1st to 3rd orders account for more than 80% of the total energy), while the skyline descriptor of flat roof buildings exhibits a relatively uniform distribution of components of each order. This frequency domain difference provides a reliable quantitative basis for accurately measuring the coordination of roof shape in subsequent coordination deviation calculations.

[0042] (5) Detail dimension: The detail dimension aims to describe the richness and stylistic features of historical decorative components on the building facade. In one embodiment of the present invention, the feature extraction of the detail dimension includes two aspects: the detection density of decorative components and the style classification.

[0043] Decorative component detection employs a target detection network to detect five typical historical decorative components within the facade area: window frames, lintels, railings, brick carvings, and eaves brackets. In one embodiment of this invention, the target detection network uses a Faster R-CNN architecture and is trained using approximately 3000 historical building facade images labeled with the categories and locations of decorative components. The detection density of decorative components is... Defined as the ratio of the total number of detected decorative components to the masked area of ​​the facade region, expressed in units per 10,000 pixels. Preferably, only detection results with a confidence level higher than 0.6 are counted. The detection density of decorative components in historical buildings is typically between 3 and 15 per 10,000 pixels, while that in minimalist modern buildings is usually less than 1 per 10,000 pixels.

[0044] In terms of style classification, image patches are extracted for each detected decorative component and input into a pre-trained style classification network for classification. In one embodiment of the present invention, the style classification network adopts the ResNet-50 architecture, and the classification categories include four types: traditional Chinese style, Western classical style, Sino-Western fusion style, and modern minimalist style.

[0045] For each building, all decorative components are categorized by style, and the percentage of each category is calculated to form a style distribution vector. The sum of all elements is 1.

[0046] The feature vector for the detailed dimension is jointly constructed by the decorative component detection density and the style distribution vector. Preferably, when no decorative components are detected on a building facade, the decorative component detection density is set to 0, and the style distribution vector is uniformly distributed. In the actual scenario of historical and cultural districts, the difficulty of detecting decorative components is closely related to the building's age and preservation condition. For older buildings that have undergone multiple renovations, their decorative components may have partial defects or be covered by later plaster layers, leading to a decrease in detection recall. To address this, in one embodiment of the present invention, the training data of the decorative component detection network is specifically enhanced by introducing data augmentation strategies that simulate local defects and color degradation of components. These strategies include randomly occluding 10% to 40% of the component area and superimposing Gaussian noise and color shift on the component image, making the trained detection network more robust to decorative components in poor preservation conditions.

[0047] Through feature extraction across the five dimensions described above, each building acquires a five-dimensional feature description, covering key visual elements affecting architectural harmony such as color, material, proportion, outline, and details. These five feature extraction methods are independent yet complementary in their evaluation objectives, collectively forming a comprehensive, multi-dimensional representation of architectural style. It is worth emphasizing that the selection of these five dimensions and the design of specific feature parameters for each dimension are not arbitrary but rather a systematic quantitative mapping based on the theoretical framework of architectural harmony in architecture and urban planning. In classic architectural theory, the core elements influencing visual harmony between buildings are typically summarized as color, material, volume proportion, outline, and decorative details. This invention maps the qualitative descriptive elements within this theoretical framework one by one into calculable numerical features, forming a complete transformation chain from qualitative concepts to quantitative indicators. Specifically, the mean of the dominant hue and color dispersion in the color dimension correspond to the concepts of tone consistency and color richness in architectural color harmony theory; roughness and regularity in the material dimension correspond to the two core attributes of the visual texture of building materials; the width-to-height ratio and window-to-wall area ratio in the proportion dimension correspond to the classic parameters describing the density and opening / closing ratio of facades in architectural morphology; the Fourier descriptor and the undulating rhythm of the cornice line in the outline dimension correspond to the morphological frequency domain analysis method in urban skyline coordination research; and the detection density and style classification in the detail dimension correspond to the quantitative evaluation requirements of architectural decoration richness and style consistency. This systematic feature design based on disciplinary theory ensures the completeness and scientific nature of the method of this invention in terms of evaluation dimensions.

[0048] It should be further explained that the five-dimensional feature extraction results in step S3 form the data foundation for subsequent steps S4 and S5. The feature extraction accuracy in step S3 directly affects the reliability of the final coordination degree assessment. In one embodiment of the present invention, range checks and outlier marking processing were performed on the feature extraction results of all five dimensions: when the color dispersion of a building exceeds a preset upper limit threshold (set to 50 in this embodiment), the building is marked as potentially having facade tiling or severe color difference problems; when the window-to-wall area ratio of a building exceeds 0.9, the building is marked as potentially having a glass curtain wall structure, and the representativeness of its feature values ​​needs to be treated with caution. The above outlier marking information is transmitted to the baseline profile construction process in step S4 for outlier removal decisions.

[0049] Step S4: Construction of the Streetscape Benchmark Profile. The goal of this step is to establish a benchmark profile of the streetscape based on the statistical distribution of the five-dimensional characteristics of all identified historical buildings within the street, providing a quantitative reference for the coordination assessment of subsequent new construction or renovation plans.

[0050] Specifically, approximately 120 historical buildings identified in step S2 are used as a benchmark building group. Statistical analysis is performed on the five-dimensional feature values ​​extracted from each building in the benchmark building group in step S3. Before performing the statistical analysis, in one embodiment of the invention, the feature value data of the benchmark building group is first checked for completeness and consistency. The completeness check confirms that all five dimensions of feature values ​​have been successfully extracted from each building in the benchmark building group. Buildings whose feature values ​​for a certain dimension could not be successfully extracted due to severe facade obstruction or poor image quality are excluded from the statistical analysis for that dimension. The consistency check detects whether there are data points in the benchmark building group whose feature values ​​deviate significantly from the normal range due to acquisition errors or processing anomalies. Specifically, the median and median absolute deviation are calculated for each feature component of each dimension. Data points deviating from the median by more than three times the median absolute deviation are marked as suspicious outliers and removed in subsequent statistics. The above data quality assurance measures ensure the statistical robustness of the landmark benchmark portrait and avoid the risk of individual outliers causing undue deviations in the benchmark portrait.

[0051] In one embodiment of the present invention, the mean vector is calculated for the feature vector of each dimension. Covariance Matrix ( (Each corresponds to one of the five dimensions), and the mean vector and covariance matrix are used to jointly represent the baseline distribution of the streetscape in that dimension.

[0052] Taking color as an example, each of the approximately 120 historical buildings in the benchmark building complex has its own color feature vector. ( The mean vector for the color dimension is: ,in: In this embodiment, the number of buildings in the baseline building complex is used as an example. ; For the first The color feature vector of the benchmark building. The covariance matrix of the color dimension is: , where: superscript This represents the matrix transpose operation. The covariance matrix has dimensions of . (Corresponding to the three components of the color feature vector), its diagonal elements reflect the variance of each component, and its off-diagonal elements reflect the correlation between the components.

[0053] The mean vector described above describes the typical characteristic values ​​of the historical building complex in the color dimension, while the covariance matrix describes the variation magnitude and correlation structure of each characteristic component. Preferably, before calculating the covariance matrix, outliers in the characteristic values ​​that deviate from the mean by more than three standard deviations are removed to eliminate the interference of individual special buildings on the baseline profile.

[0054] The mean vectors and covariance matrices of the other four dimensions (material, proportion, outline, and detail) are calculated separately using the same method. The mean vectors and covariance matrices of the five dimensions together constitute a complete baseline profile of the neighborhood's character. .

[0055] The physical meaning of the benchmark profile of a neighborhood's character lies in its statistical distribution, which depicts the common characteristics and range of variation of the historical building complex across various character dimensions. If the characteristic value of a newly built or renovated building in a certain dimension falls within the high probability density area of ​​the benchmark distribution for that dimension, it is considered to be in harmony with the neighborhood's character in that dimension; if the characteristic value deviates significantly from the center of the benchmark distribution, there is a risk of insufficient harmony in that dimension.

[0056] In one embodiment of this invention, the statistical quality of the benchmark architectural portrait was verified. Specifically, a multivariate normality test was used to verify the normality of the eigenvalue distribution of the benchmark architectural complex across various dimensions. In the color dimension, the test found that the color feature distribution of the benchmark architectural complex approximately follows a multivariate normal distribution, thus satisfying the applicability conditions of Mahalanobis distance. In the contour dimension, since the higher-order components of the Fourier descriptor may exhibit non-normal distribution characteristics, in one embodiment of this invention, a logarithmic transformation was performed on the eigenvalues ​​of the contour dimension to improve their normality, and the mean vector and covariance matrix were recalculated after the transformation. This processing strategy ensures the statistical rationality of the subsequent Mahalanobis distance calculation.

[0057] Furthermore, in practical applications, different historical and cultural blocks may possess drastically different architectural benchmarks. For example, blocks dominated by traditional Chinese architecture and concession-style blocks dominated by modern Western architecture exhibit significant differences in color schemes, material types, and decorative styles. The architectural benchmark portrait construction method of this invention naturally adapts to these differences between blocks because the benchmark portrait is entirely determined by the statistical distribution of the block's own historical buildings, without relying on any preset absolute standards. This adaptive characteristic allows the method of this invention to be directly applied to historical and cultural blocks of different architectural types without additional parameter tuning.

[0058] It is worth noting that the construction of the baseline profile is a one-time process; once established, it can be reused to evaluate all new construction or redevelopment projects in the neighborhood. When the list of historic buildings in the neighborhood is updated, only the mean vector and covariance matrix of the corresponding dimensions need to be updated incrementally.

[0059] Step S5: Calculation and Comprehensive Evaluation of Landscape Harmony Deviation. The goal of this step is to quantitatively compare the renderings of the new or renovated plan with the landscape baseline image constructed in Step S4, calculate the harmony deviation in each dimension, and provide a comprehensive evaluation conclusion and visual diagnostic information.

[0060] First, the architectural renderings of the new or renovated design are input into the instance segmentation network in step S2 to obtain a mask for the facade area. Then, the five-dimensional feature values ​​are obtained through the five-dimensional feature extraction process in step S3. Preferably, when the design provides renderings from multiple perspectives, the average of the feature values ​​from each perspective is taken as the comprehensive feature value of the design.

[0061] The coordination deviation of each dimension was calculated using Mahalanobis distance. Taking the first dimension as an example... For example, in a dimension, the feature value of the solution to be evaluated in that dimension is The baseline profile for this dimension is Then the first The degree of dimensional coordination deviation is: ,in: For the first The degree of coordination deviation of the dimension, dimensionless, with a range of values ​​of... The larger the value, the farther it deviates from the streetscape benchmark in that dimension; For the scheme to be evaluated in the Feature vectors in a specific dimension; For the first A baseline mean vector of dimension; For the first The inverse of the covariance matrix of a dimension. The technical advantage of Mahalanobis distance over Euclidean distance lies in the fact that Mahalanobis distance automatically eliminates the correlation and dimensional differences between feature components through the inverse of the covariance matrix, making the degree of deviation between different dimensions comparable. For example, in the color dimension, the value range and variance of channel a and channel b may differ significantly. Directly using Euclidean distance would cause the component with the larger variance to dominate the distance calculation, while Mahalanobis distance avoids this problem through standardization.

[0062] Preferably, the reference covariance matrix is ​​regularized before calculating the Mahalanobis distance to ensure its invertibility. In one embodiment of the invention, the covariance matrix is ​​regularized using shrinkage estimation, specifically as follows: ,in The shrinkage coefficient is 0.1. It is the diagonal matrix of the covariance matrix.

[0063] The weighted composite score for the five-dimensional coordination deviation is: ,in: It is a dimensionless index for overall style coordination. The larger the value, the greater the degree of overall deviation (i.e., the worse the coordination). For the first Dimension weight coefficients, In one embodiment of the present invention, the weights of each dimension are determined by combining the Analytic Hierarchy Process (AHP) with expert surveys in the field of historical and cultural district protection planning. Preferably, the weight of the color dimension is... Material dimension weight Proportional dimension weight Contour dimension weight Detailed dimension weights The basis for this weighting is that, in the practice of historical and cultural district protection planning, color harmony is the most intuitive element in terms of the overall appearance, while the influence of detailed decorations is relatively secondary.

[0064] Based on the overall landscape harmony index The evaluation conclusions are divided into three levels: when When the scheme is deemed to be at a coordination level, it indicates that the scheme is highly consistent with the streetscape benchmark in all dimensions; when When the deviation is significant, it is judged as a basic coordination level, indicating that the plan has some deviations in individual dimensions but is generally acceptable. Fine-tuning of the dimensions with larger deviations is recommended. If the design is deemed inconsistent, it indicates that the design deviates significantly from the streetscape benchmark and requires substantial design modifications.

[0065] In addition, this step also displays the degree of coordination deviation of the new or renovated plan in five dimensions using a radar chart. The five axes of the radar chart correspond to the color dimension, material dimension, proportion dimension, outline dimension, and detail dimension, respectively, and the value on each axis represents the degree of coordination deviation in that dimension. Simultaneously, the radar chart uses different colors to mark the deviation direction information for each dimension; for example, color dimension is marked as warmer or cooler, and material dimension is marked as rougher or smoother. The method for determining the deviation direction is: calculating the difference between the feature value of the scheme to be evaluated and the benchmark mean vector. The direction of deviation is determined by the sign of the difference between each component.

[0066] Preferably, the radar chart simultaneously plots the typical fluctuation range of the baseline profile (using the 75th percentile of the coordination deviation of each dimension in the baseline building group as the reference circle), allowing designers to intuitively determine whether the deviation of the proposed scheme in each dimension exceeds the normal variation range of the historical buildings within the block. The radar chart provides designers with intuitive multidimensional diagnostic information, enabling them to quickly identify the dimensions and directions requiring key adjustments, significantly improving the efficiency of iterative modification of the scheme.

[0067] In one embodiment of the present invention, when a designer resubmits an evaluation after adjusting the design based on the evaluation report, the system supports overlaying the evaluation results of multiple versions on the same radar chart, distinguishing different versions by color depth or line type differences. This allows the designer to track the trajectory of iterative optimization of the design and the changes in deviation in various dimensions brought about by each adjustment. This function provides visualized and quantitative feedback support for the progressive optimization of architectural design harmony.

[0068] Furthermore, this step provides specific adjustment suggestions for each dimension. When the color dimension shows a large deviation in coordination, and the deviation is towards cooler tones, the system suggests that designers increase the proportion of warm colors in the building facade color scheme. When the outline dimension shows a large deviation in coordination, mainly due to differences in low-order Fourier descriptors, the system suggests that designers adjust the basic form of the roof (e.g., changing a flat roof to a pitched roof to match the traditional roof form of the block). When the detail dimension shows a large deviation in coordination, and the density of decorative components is far below the baseline value, the system suggests that designers appropriately increase decorative elements on the facade to enrich the expression of historical details. The generation of the above adjustment suggestions is based on the component analysis results of the coordination deviation in each dimension, transforming abstract numerical deviations into design language that designers can directly understand and implement, thus achieving a complete closed loop from quantitative assessment to design guidance.

[0069] See Figure 2 This invention also provides an image evaluation system for the architectural style harmony of historical and cultural blocks. This system is used to implement all the functions described in the above method embodiments. The system includes five functional modules: an image acquisition and projection module, a facade segmentation module, a feature extraction module, a baseline portrait construction module, and a harmony evaluation module. Each module corresponds one-to-one with steps S1 to S5 in the method embodiments.

[0070] The image acquisition and projection module is used to execute the street view panoramic image acquisition and multi-view facade image sequence generation functions described in step S1. This module receives the raw panoramic image data acquired by the panoramic camera as input, performs perspective projection transformation processing, and outputs a multi-view unfolded facade image sequence. Preferably, the image acquisition and projection module includes a panoramic image preprocessing subunit and a perspective projection transformation subunit. The panoramic image preprocessing subunit is used to perform color consistency correction and distortion correction preprocessing on the raw panoramic image to eliminate the influence of varying lighting conditions at different acquisition times on image quality. In one embodiment of the invention, color consistency correction uses a histogram matching method, using the facade color histogram of a specified standard building within the block as a reference template to align the color distribution of all panoramic images. The perspective projection transformation subunit is used to unfold the preprocessed equidistant cylindrical projection format panoramic image according to the street orientation direction using multi-view perspective projection, generating a planar facade image that conforms to human visual habits. This subunit internally maintains a street orientation angle database, which stores the orientation angle information of each street in the block to ensure that the direction of the perspective projection is accurately aligned with the street orientation.

[0071] The facade segmentation module is used to perform the building facade instance segmentation and facade region mask extraction functions described in step S2. This module receives a multi-view unfolded facade image sequence output by the image acquisition and projection module as input, processes each frame image using a pre-trained instance segmentation network, and outputs facade region masks for each building. In one embodiment of the invention, the facade segmentation module further includes an inter-frame association subunit, used to perform instance matching and deduplication processing on the same building in adjacent frame images, ensuring that each building retains only one optimal viewpoint facade region mask in the final output result. Preferably, the inter-frame association subunit uses a Hungarian matching algorithm based on the bounding box center coordinates to achieve cross-frame instance association, and the matching cost matrix is ​​calculated by weighting the intersection-union ratio of the bounding boxes and the cosine similarity of the appearance features. When a building is detected in multiple consecutive frame images, the frame with the largest and unobstructed facade region mask pixel coverage area is selected as the representative frame of that building to obtain the most complete facade information for subsequent feature extraction.

[0072] The feature extraction module is used to perform the automatic quantization extraction function of five-dimensional architectural features described in step S3. This module receives the facade area masks output by the facade segmentation module as input, extracts architectural feature values ​​from five dimensions: color, material, proportion, outline, and detail, and outputs the extraction results to the benchmark portrait construction module or the harmony evaluation module. Preferably, the feature extraction module has five parallel feature extraction sub-units, each corresponding to the feature calculation of the above five dimensions. Each sub-unit operates independently to improve processing efficiency. In one embodiment of the invention, the five feature extraction sub-units can be executed simultaneously using the parallel computing power of the graphics processor, so that the five-dimensional feature extraction time for a single building is controlled within 0.8 seconds. The output of the feature extraction module is a standardized set of five-dimensional feature vectors, with each building corresponding to a set of five-dimensional feature vectors. The dimensions of each feature vector are 3 for color, 2 for material, 2 for proportion, 16 for outline, and 5 for detail.

[0073] The benchmark profile construction module is used to execute the streetscape benchmark profile construction function described in step S4. This module receives the five-dimensional feature vector set of the benchmark building group output by the feature extraction module as input, calculates the mean vector and covariance matrix for each dimension, and outputs a complete streetscape benchmark profile data structure. Preferably, the benchmark profile construction module internally maintains a historical protected building list database, which stores the identification information of all identified historical protected buildings within the street, and is used to automatically select a feature subset of the benchmark building group from the feature vector set of all buildings for statistical analysis. Furthermore, the benchmark profile construction module also supports incremental update functionality; when the list of historical protected buildings within the street changes, the streetscape benchmark profile can be partially updated without recalculating all statistics. The incremental update is implemented by using an online statistical update algorithm to recursively correct the mean vector and covariance matrix for the feature values ​​of newly added or removed buildings. Specifically, when a feature vector of a new building is added, the updated mean vector and covariance matrix are directly calculated using the existing mean vector and covariance matrix as well as the new feature vector, through a recursive formula, thus avoiding the need to recalculate the feature values ​​of all benchmark buildings.

[0074] The coordination evaluation module is used to perform the landscape coordination deviation calculation and comprehensive evaluation functions described in step S5. This module receives the five-dimensional feature values ​​of the scheme to be evaluated and the landscape baseline portrait output by the baseline portrait construction module as input. It calculates the Mahalanobis distance coordination deviation for each dimension and generates an overall landscape coordination index and evaluation level conclusion through weighted comprehensive scoring. Preferably, the coordination evaluation module also includes a radar chart visualization sub-unit and an evaluation report generation sub-unit. The radar chart visualization sub-unit visually displays the coordination deviation and deviation direction of each dimension in the form of a five-axis radar chart, providing designers with targeted adjustment guidance. The evaluation report generation sub-unit automatically generates a structured evaluation report containing the evaluation level, deviation values ​​for each dimension, deviation direction analysis, and adjustment suggestions. This report can be submitted to the approval agency as an objective basis for planning approval. Furthermore, the coordination evaluation module also provides a batch evaluation function. When designers iterate and modify the same scheme multiple times, they can input the renderings of each version sequentially for batch evaluation. The system automatically generates a trend chart of the coordination index changes for each version, allowing designers to intuitively grasp the optimization effect and convergence trend of the scheme iteration modifications.

[0075] The five modules described above form a deep data flow coupling relationship: the output of the image acquisition and projection module serves as the input of the facade segmentation module, the output of the facade segmentation module serves as the input of the feature extraction module, the output of the feature extraction module simultaneously flows to the baseline portrait construction module and the coordination degree evaluation module, and the output of the baseline portrait construction module serves as the reference benchmark for the coordination degree evaluation module. Preferably, the evaluation results of the coordination degree evaluation module can be fed back to the image acquisition and projection module. When the evaluation results show that the confidence level of certain dimensions is low, the system automatically suggests that the acquisition party perform supplementary acquisition to improve the reliability of the evaluation. This feedback mechanism ensures that there is not only a positive data flow transmission relationship between the modules, but also forms a closed-loop collaborative optimization system architecture.

[0076] In terms of data storage and management, in one embodiment of the present invention, the system maintains a unified spatial database to store all intermediate and final data, including panoramic image data, unfolded facade image sequences, facade area masks, five-dimensional feature vectors, and landscape benchmark profiles. The spatial database uses street block numbers and building numbers as primary indexes, supporting rapid retrieval and filtering based on multiple dimensions such as geographical location, building type, and protection level. Preferably, the spatial database also stores the geographical coordinates of each building, supporting the display of the five-dimensional feature values ​​and coordination deviation distribution of each building in map form on a geographic information system platform, providing spatialized landscape control data support for urban planning and management departments.

[0077] Furthermore, in one embodiment of the present invention, the system also provides a parameter configuration interface, allowing planning and management departments to customize system parameters according to the protection requirements and style characteristics of different blocks. Configurable parameters include, but are not limited to: weighting coefficients for five-dimensional coordination deviation, thresholds for coordination level classification, confidence filtering thresholds for various detection networks in feature extraction, outlier removal criteria in baseline profile construction, and display styles and color coding schemes for radar charts. Through the flexible settings of the parameter configuration interface, the system can adapt to the evaluation requirements of historical and cultural blocks with different management needs and different levels of protection. For example, for national-level historical and cultural blocks, a more stringent coordination level threshold can be set (e.g., reducing the threshold for incoordination level from 4.0 to 3.0), while for general historical and cultural areas, the threshold standard can be appropriately relaxed.

[0078] Regarding system operating efficiency, one embodiment of this invention optimizes the processing time of each module. In the image acquisition and projection module, perspective projection transformation is accelerated using a graphics processor, with an average time of 0.12s to generate unfolded facade images from six perspectives from a single panoramic image. In the facade segmentation module, instance segmentation network inference uses a half-precision floating-point inference mode, with a single-frame image segmentation processing time of 0.35s. In the feature extraction module, five-dimensional feature calculations employ a parallel processing architecture, with a total time of 0.8s for five-dimensional feature extraction of a single building. In the benchmark profile construction module, the calculation of the mean vector and covariance matrix takes no more than 2s (for a feature set of 120 buildings). In the coordination evaluation module, the Mahalanobis distance calculation and comprehensive scoring for a single scheme take no more than 0.05s. The processing efficiency of these modules ensures that the system can achieve near real-time evaluation response in practical application scenarios.

[0079] To verify the effectiveness and superiority of the method of this invention, experimental tests were conducted in the aforementioned historical and cultural district. The test environment consisted of a workstation equipped with an NVIDIA RTX 4090 graphics processor (24GB VRAM) and an Intel i9-13900K processor (64GB RAM), running Ubuntu 22.04, and using PyTorch 2.0 as the deep learning framework. Both the instance segmentation network and the object detection network ran in half-precision floating-point mode to accelerate the inference process, and the calculation of the gray-level co-occurrence matrix and Fourier transform was implemented using the NumPy scientific computing library.

[0080] The test dataset contains approximately 450 panoramic images of the aforementioned neighborhood and renderings of eight proposed new construction or redevelopment plans for planning approval. These eight plans cover different types of projects, including two redevelopment plans for street-front commercial buildings, three new residential building plans, and three redevelopment plans for public service facilities. Each plan provides three to five architectural renderings from different perspectives. The renderings were generated by the architectural design firm using professional rendering software, with an image resolution of 2048×1536 pixels. The rendered scenes include the surrounding environment of the building site to ensure the visual characteristics of the renderings are comparable to the actual photographed images. To eliminate the impact of color deviation from the rendering software on the evaluation results, this embodiment performs color calibration preprocessing on the renderings, using the colors of actual photographed images of existing buildings in the surrounding area as a reference to correct the global color shift in the renderings.

[0081] In the consistency evaluation test, five experts with experience in the protection and planning of historical and cultural blocks were invited to subjectively evaluate the eight proposed schemes. Each expert independently provided a harmony score for each scheme across five dimensions: color, material, proportion, outline, and detail (using a Likert scale of 1 to 5, with 5 indicating complete harmony). They also provided an overall harmony rating (harmonious, basically harmonious, or disharmonious). The expert evaluation results showed that the five experts had some disagreement on the overall harmony rating of the same scheme. For four of the eight schemes, the consensus rate was only 60% (i.e., only three out of the five experts gave the same rating). This result directly confirms the problem of inconsistent subjective evaluation standards in existing technologies.

[0082] The method of this invention was used to automatically evaluate the same eight schemes, and the evaluation results were completely repeatable and deterministic. Statistical analysis showed that the automatic evaluation conclusions of the method of this invention completely matched the majority opinion of five experts (or the conclusions of three or more experts who agree) for six schemes (a consistency rate of 75%). The automatic evaluation conclusions of the remaining two schemes differed from the majority opinion of the experts by one level (e.g., the automatic evaluation was "basically harmonious" while the majority opinion of the experts was "harmonious"), which is within an acceptable range of deviation. More importantly, in the four schemes where the experts' judgments differed significantly, the radar chart dimensional analysis of the method of this invention clearly revealed the degree and direction of deviation in each dimension, providing an objective quantitative reference for the formation of expert consensus.

[0083] In terms of evaluation efficiency, the method of this invention takes approximately 4.2 seconds to complete the evaluation of a single scheme (including 3 to 5 perspective renderings) (excluding panoramic image acquisition time and the initial construction time of the benchmark profile). Compared with the several days required for expert on-site inspection and review, the evaluation efficiency is improved by more than three orders of magnitude. The initial construction of the benchmark profile takes approximately 180 seconds (processing the five-dimensional feature statistics of approximately 120 benchmark buildings), and can then be reused for all subsequent evaluation tasks of the block.

[0084] In terms of feature extraction accuracy, the average deviation between the CIE-Lab color space mean and the manually labeled reference value in the color dimension is 1.8 Lab units; the accuracy of gray-level co-occurrence matrix texture features in material dimension for material type classification is 89.2% (classification tests were conducted on painted walls, exposed brick walls, and stone curtain walls); the aspect ratio calculation error in the proportion dimension is within 5%; the Fourier descriptor in the contour dimension has an accuracy of 94.6% in distinguishing between pitched roofs and flat roofs; and the recall rate for decorative component detection in the detail dimension is 78.3%. These feature extraction accuracies demonstrate that the method of this invention has high reliability in multi-dimensional feature quantification and can meet the practical application needs of evaluating the architectural style harmony of historical and cultural blocks.

[0085] Furthermore, to verify the superiority of Mahalanobis distance in measuring coordination deviation, this embodiment also conducted a comparative experiment between Mahalanobis distance and Euclidean distance. The experimental results show that when using Euclidean distance as a measure of coordination deviation, the agreement rate between the evaluation conclusions in the color and contour dimensions and the majority opinion of experts is only 62.5%, lower than the 75% when using Mahalanobis distance. The reason for this is that Euclidean distance cannot eliminate the dimensional differences and correlations between the feature components, leading to an overemphasis on the contribution of brightness component changes to deviation in the color dimension, and in the contour dimension, the magnitude of low-order Fourier descriptors is much larger than that of high-order descriptors, thus dominating the weight in Euclidean distance. Both of these deviations deviate from the human eye's perception of landscape coordination. Mahalanobis distance achieves adaptive component weighting through the inverse of the covariance matrix, making the contribution of each component to deviation inversely proportional to its variation amplitude in the benchmark building group, which is more consistent with the statistical logic of abnormal deviation judgment.

[0086] In the comparative test of distance metrics, cosine distance and Manhattan distance were also evaluated. The results show that cosine distance focuses on the direction of the eigenvectors while ignoring amplitude information, and cannot effectively distinguish cases where color tones are similar but color dispersions are significantly different in the color dimension. While Manhattan distance is simple to calculate and insensitive to outliers, its characteristic of treating each component independently while ignoring the correlation between components makes it perform poorly in the evaluation of material and contour dimensions. Based on the above comparative analysis, Manhattan distance exhibits the best overall performance in measuring architectural style harmony deviation.

[0087] The experimental results above demonstrate that the method of this invention, by establishing a five-dimensional feature quantification system and a Mahalanobis distance coordination deviation calculation framework, effectively solves the technical problems of strong subjectivity, inconsistent evaluation standards, and incomplete feature dimensions in existing technologies regarding coordination review. This method provides objective, quantifiable, and repeatable evaluation results, and through multi-dimensional visualization analysis using radar charts, it offers designers clear guidance for adjustments. This transforms architectural style coordination review from qualitative judgments relying on personal experience to data-driven quantitative assessments, significantly promoting the scientific rigor and efficiency of historical and cultural district protection planning approvals.

[0088] Furthermore, this embodiment also tested the evaluation guidance effect during the iterative optimization process of the scheme. Two schemes from the eight schemes were selected as incompatible in the evaluation, and the designer modified the schemes based on the radar chart analysis and adjustment suggestions provided by the system. After two to three rounds of iterative modifications, the overall style coordination index of the two schemes dropped below 2.0, reaching the coordination level. Before using the method of this invention, designers needed to undergo an average of 4.5 rounds of expert review and scheme modification for the same eight schemes to pass planning approval. However, after using the method of this invention, designers can independently optimize and adjust the schemes through the system's real-time feedback during the scheme design stage, reducing the number of scheme modification rounds to an average of 2.3 rounds, and shortening the scheme optimization cycle from an average of 45 days to an average of 12 days, significantly reducing the time and communication costs of project development.

[0089] Furthermore, the applicability of the method of this invention in historical and cultural blocks of different sizes and architectural styles has been preliminarily verified. In addition to the aforementioned historical and cultural blocks primarily featuring traditional Chinese architecture, baseline profile construction and scheme evaluation tests were also conducted in a concession-style block (containing approximately 80 historical buildings) primarily featuring modern Western architecture and a commercial block (containing approximately 60 historical buildings) primarily featuring a blend of Chinese and Western architectural styles. The test results show that the method of this invention can adaptively establish baseline profiles for blocks of different architectural styles, and can provide reasonable coordination assessment conclusions in all three different types of blocks. This test result verifies the versatility and portability of the method of this invention, indicating that it is applicable to architectural style coordination assessment scenarios in various historical and cultural blocks.

[0090] The embodiments of the present invention are not limited to the specific embodiments described above. Those skilled in the art can make various equivalent changes or substitutions based on the technical solutions of the present invention, and all such changes or substitutions should be included within the protection scope of the present invention.

Claims

1. A method for evaluating the harmony of architectural style in historical and cultural blocks using images, characterized in that... Includes the following steps: Step S1, Street view panoramic image acquisition and multi-view facade image sequence generation: A panoramic camera is used to acquire street view panoramic images segment by segment along each street of the historical and cultural district to obtain a panoramic image dataset containing information on the facades of buildings on both sides of each street; perspective projection transformation is performed on each frame of panoramic image in the panoramic image dataset to generate a multi-view unfolded facade image sequence along each street direction. Step S2, Building facade instance segmentation and facade region mask extraction: The instance segmentation network is used to segment the building facade instances in each frame of the multi-view unfolded facade image sequence, and the facade of each building in each frame is segmented independently to obtain the facade region mask of each building. Step S3, Automatic Quantization and Extraction of Five-Dimensional Features of Architectural Appearance: Based on the facade area mask, five dimensions of architectural appearance features are quantitatively extracted for each building facade. These five dimensions include color, material, proportion, outline, and detail. Among them, the color dimension describes the building's color tone using the mean value and color dispersion of the dominant color channel of pixels within the facade area mask in the CIE-Lab color space; the material dimension uses the gray-level co-occurrence matrix to extract the roughness and regularity of the facade texture to distinguish different material types; the proportion dimension describes the facade's proportional relationship using the width-to-height ratio and window-to-wall area ratio; the outline dimension extracts the roof shape features and the undulating rhythm of the eaves line using the Fourier descriptor of the roof skyline; and the detail dimension describes the richness of historical details using the detection density and style classification of decorative components. Step S4, Construction of the Streetscape Benchmark Profile: Using all the identified historical buildings in the street as the benchmark building group, the characteristic value distribution of the benchmark building group in five dimensions is statistically analyzed to construct the streetscape benchmark profile. Step S5, Calculation and comprehensive evaluation of style coordination deviation: The renderings of the new or renovated plan are processed by steps S2 to S3 to obtain their five-dimensional feature values. The Mahalanobis distance between the feature values ​​of each dimension and the style benchmark image is calculated as the coordination deviation of each dimension. The weighted comprehensive score of the five-dimensional coordination deviation is used as the overall style coordination index.

2. The image evaluation method for the architectural style harmony of historical and cultural blocks according to claim 1, characterized in that, In step S1, the acquisition interval of the panoramic camera is 5m to 15m, and the resolution of the acquired panoramic image is not less than 4096×2048 pixels; in the perspective projection transformation, multiple perspective projection images from different angles are extracted for each frame of panoramic image according to the street direction at preset angle intervals, with the preset angle intervals being 30° to 90°.

3. The image evaluation method for the architectural style harmony of historical and cultural blocks according to claim 1, characterized in that, In step S3, the quantization extraction of color dimensions includes: converting the pixels in the facade area mask from the RGB color space to the CIE-Lab color space, calculating the mean of the a channel and the mean of the b channel as the main color tone description, calculating the standard deviation of each of the three Lab channels and taking the weighted sum of the standard deviations of the three channels as the color dispersion.

4. The image evaluation method for the architectural style harmony of historical and cultural blocks according to claim 1, characterized in that, In step S3, the quantization extraction of the material dimension includes: converting the image within the mask of the facade area into a grayscale image, calculating the contrast as a roughness index and the second moment of the angle as a regularity index using the grayscale co-occurrence matrix; the calculation window size of the grayscale co-occurrence matrix is ​​64×64 pixels, the step size is 1 pixel, and the direction is the average value of four directions: 0°, 45°, 90° and 135°.

5. The image evaluation method for the architectural style harmony of historical and cultural blocks according to claim 1, characterized in that, In step S3, the quantization extraction of the contour dimension includes: performing contour tracking on the upper edge of the building facade area mask to obtain the roof skyline curve, performing Fourier transform on the skyline curve to obtain the Fourier descriptor sequence, and truncating the first N Fourier descriptors as the roof shape feature vector, where N is an integer from 8 to 32; at the same time, extracting the vertical coordinate sequence of the eaves line and calculating its standard deviation as the eaves line undulation rhythm value.

6. The image evaluation method for the architectural style harmony of historical and cultural blocks according to claim 1, characterized in that, In step S3, the quantitative extraction of the proportional dimension includes: calculating the aspect ratio of the building facade based on the minimum bounding rectangle of the facade area mask; using the window detection model to detect the windows in the facade area and calculating the ratio of the total window area to the total facade area as the window-to-wall area ratio.

7. The image evaluation method for the architectural style harmony of historical and cultural blocks according to claim 1, characterized in that, In step S4, the construction of the landmark baseline profile includes: calculating the mean vector and covariance matrix of the feature values ​​of each historical building in the landmark building complex in five dimensions, and using the mean vector and covariance matrix to jointly represent the landmark baseline distribution of the block in that dimension.

8. The image evaluation method for the architectural style harmony of historical and cultural blocks according to claim 1, characterized in that, In step S5, the coordination deviation of each dimension is calculated using Mahalanobis distance; in the weighted comprehensive score of the five-dimensional coordination deviation, the weight of each dimension is determined by combining the analytic hierarchy process with expert surveys on the protection planning of historical and cultural blocks; the overall style coordination index is divided into three levels: coordinated, basically coordinated, and uncoordinated by comparing with preset thresholds.

9. The image evaluation method for the architectural style harmony of historical and cultural blocks according to claim 1, characterized in that, Step S5 also includes: displaying the coordination deviation of the new or renovated plan in five dimensions using a radar chart. Each axis of the radar chart corresponds to a dimension, and the value on the axis represents the coordination deviation of that dimension. At the same time, the direction of deviation for each dimension is marked to guide designers to make targeted adjustments to dimensions whose coordination deviation is greater than or equal to the preset dimension threshold.

10. A system for evaluating the harmony of architectural style in historical and cultural blocks, used to implement the image evaluation method for evaluating the harmony of architectural style in historical and cultural blocks as described in any one of claims 1-9, characterized in that, include: The image acquisition and projection module is used to acquire panoramic street scene images segment by segment along each street of the historical and cultural district using a panoramic camera, and to perform perspective projection transformation on the panoramic images to generate a multi-view unfolded facade image sequence along each street direction. The facade segmentation module is used to perform building facade instance segmentation on each frame of the multi-view unfolded facade image sequence using an instance segmentation network. It independently segments the facade of each building in each frame and obtains the facade area mask of each building. The feature extraction module is used to perform quantitative extraction of the appearance features of each building facade in five dimensions: color, material, proportion, outline, and detail, based on the facade area mask. Among them, the color dimension describes the building color tone by the mean of the main color channel and the color dispersion of the pixels in the facade area mask in the CIE-Lab color space; the material dimension extracts the roughness and regularity of the facade texture by the gray-level co-occurrence matrix; the proportion dimension describes the facade proportion relationship by the width-to-height ratio and the window-to-wall area ratio; the outline dimension extracts the roof shape features and the undulation rhythm of the eaves line by the Fourier descriptor of the roof skyline; and the detail dimension describes the richness of historical details by the detection density and style classification of decorative components. The baseline profile construction module is used to construct a baseline profile of the neighborhood's appearance based on the statistical distribution of the five-dimensional feature values ​​of all identified historical buildings within the neighborhood. The coordination evaluation module uses the Mahalanobis distance between the five-dimensional feature values ​​of the scheme to be evaluated and the landscape benchmark image as the coordination deviation of each dimension. The weighted comprehensive score of the five-dimensional coordination deviation is used as the overall landscape coordination index, and the direction and magnitude of the deviation of each dimension are displayed in a radar chart.

Citation Information

Patent Citations

  • Historical urban style integrity evaluation method and system based on deep learning

    CN119477804A

  • Historical and cultural block building roof gradient intelligent batch identification method applying YOLO

    CN118608933A

  • Element feature measurement method and device for block style and appearance shaping

    CN120495859A