Urban fifth facade roof space intelligent evaluation system and method and medium
Through the combination of SAM model and RAG technology, the automated evaluation of the fifth facade roof space in the city is achieved, solving the problems of low efficiency and large errors in traditional methods, and providing efficient and accurate evaluation and management solutions.
Patent Information
- Application Number
- CN202510737967.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
AI Technical Summary
The existing technology is difficult to efficiently and accurately perform macro-scale evaluation of the fifth facade roof space in the city, which consumes manpower and has large errors, which cannot meet the needs of large-scale management at the city level.
An image recognition agent based on SAM model is used to perform pixel-level semantic segmentation of roof elements, combined with RAG technology to dynamically call index formulas, and a comprehensive report is generated through Auto-Summary summary model to achieve automated evaluation and management.
It significantly improves the efficiency and accuracy of roof space evaluation, reduces manual errors, supports large-scale assessment and management at the city level, and provides visual reporting and closed-loop feedback mechanisms.
Smart Images

Figure CN120259675A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of big data processing, and particularly relates to an intelligent evaluation system, method and medium for the roof space of the fifth elevation of a city. Background Art
[0002] As a valuable space resource, the roof space of the fifth elevation of a city is of great significance in aspects such as vertical space development, resilient ecological city construction, and the embodiment of urban landscape and form. However, due to the large quantity and complexity of urban buildings, the evaluation of the fifth elevation at the macro scale has always been a management difficulty, lacking a mature technical solution.
[0003] Existing solutions mostly target individual buildings and adopt manual visual interpretation and template-based analysis methods to carry out qualitative evaluations, which not only consume labor costs but also are difficult to support the evaluation requirements of the fifth elevation at a wide range and large scale at the macro scale. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide an intelligent evaluation system, method and medium for the roof space of the fifth elevation of a city in view of the deficiencies in the above-mentioned prior art. It supports three modes of automatic segmentation, interactive segmentation, and prompt-word segmentation to realize the recognition of the roof elements of the fifth elevation, greatly improving the interpretation efficiency. In addition, it supports carrying out index calculations based on the recognition results and outputs an evaluation summary for the recognition results and index situations, reducing manual subjective errors and promoting the intelligent and precise management of the fifth elevation of the city.
[0005] The first aspect of the present invention discloses an intelligent evaluation system for the roof space of the fifth elevation of a city, including: A main control model, which is connected to a large language model and is used to parse the aerial images, texts or Prompt instructions input by users, generate a structured task logic chain and allocate it to the corresponding intelligent agents; An image recognition intelligent agent, which is an intelligent agent generated based on the pre-trained SAM (Segment Anything Model) model. The image recognition intelligent agent includes an image encoder, a prompt encoder and a mask decoder, and is used to perform semantic segmentation on the roof image and output a labeled raster mask as the image recognition result; An index calculation intelligent agent, which is externally connected to an index library and is used to retrieve the scoring formula that matches the requirements in the index library through RAG (Retrieval-Augmented Generation), and drive a spatial operation tool based on the scoring formula to generate the calculation result of the index data of the roof image; An Auto-Summary summary model, which integrates the capabilities of a multi-modal large model and is used to fuse the image recognition result and the calculation result of the index data to generate a comprehensive report on rectification suggestions, and the comprehensive report includes suggestions for roof layout.
[0006] In the above system, the image encoder uses a ViT (Vision Transformer) pre-trained with MAE (Masked Autoencoder) as the backbone network, and constructs a multi-scale feature pyramid in combination with an improved Fpn (Feature Pyramid Network); The prompt encoder supports three interactive Prompt input modes: point, box, and mask, and converts the user's Prompt instructions into sparse or dense embedding vectors; The mask decoder receives the image features and the embedding vectors converted from the Prompt instructions, outputs a binary mask, an IoU confidence level, and the coordinates of the segmentation region, and supports interactive correction based on a confidence threshold.
[0007] In the above system, the exponential calculation agent includes: A RAG retrieval framework, which takes Qwen3-30B-A3B as the base model, fine-tunes the base model through a reinforcement learning method, infers the user's instructions based on the fine-tuned base model, and retrieves matching scoring formulas from the index library; An operation toolchain, which includes an instance segmentation engine, a spatial overlay analyzer, and a formula parser, is used to convert a natural language scoring formula into an executable expression, and perform spatial operations to generate the calculation results of the index data; A spatio-temporal correlation database stores the calculation results of the index data and supports tracing the changes in the roof state according to the time series.
[0008] In the above system, the Auto-Summary summary model further includes: A multi-modal alignment module, which associates text instructions with image features through a cross-modal attention mechanism, and establishes a mapping relationship between spatial coordinates and evaluation metrics. The spatial coordinates are the spatial coordinates where the roof is located corresponding to the roof image, and the evaluation metrics are the evaluation metrics corresponding to the area where the roof is located; A visualization rendering unit that generates a heat map to mark the violation area, a 3D interactive model, and a structured data report; A rectification suggestion generation unit that outputs natural language suggestions based on the standard solutions in the knowledge base.
[0009] In the above system, the Auto-Summary summary model also includes: A closed-loop feedback unit that periodically triggers the re-inspection task of the drone and generates a rectification effect evaluation report by comparing historical data; An early warning push module that automatically generates an early warning notice when it detects new illegal construction or the greening rate drops by more than a preset threshold.
[0010] The second aspect of the present invention discloses an urban fifth facade evaluation method based on the system described in the first aspect, including the following steps: Receive the aerial images and text instructions uploaded by the user, parse the task type through the main control model and allocate it to the corresponding intelligent agent; Call the image recognition intelligent agent to segment the roof elements and generate a raster mask with semantic labels; Through the index calculation intelligent agent, load the scoring formula to calculate the greening coverage rate, space utilization rate and safety hazard level index; Use the Auto-Summary summary model to fuse multi-modal data, generate a visual report and mark the violation areas; Push the analysis results to the management platform and trigger a periodic re-inspection task to form a closed-loop management.
[0011] In the above method, the step of "segmenting roof elements" further includes: Use the SAM model to perform multi-scale feature extraction on the aerial images, and optimize the segmentation accuracy in combination with the user interactive Prompt instructions; Trigger the manual correction process for areas with confidence lower than the threshold to ensure the reliability of the mask output.
[0012] In the above method, the step of "calculating the greening coverage rate, space utilization rate and safety hazard level index" further includes: Dynamically match the index library according to natural language, and retrieve the applicable formula through the RAG technology; First perform id mapping on the roof segmentation mask (roof mask) output by the image recognition intelligent agent and the roof element segmentation mask (mask) required by the user, and then perform overlay operation according to the index formula.
[0013] In the above method, the step of "generating a visual report" further includes: Map the image recognition results and index data to the geographic coordinate system through the GIS engine; Generate an interactive 3D heat map to support users to click and query specific violation items and rectification suggestions.
[0014] The third aspect of the present invention discloses a computer-readable storage medium storing a computer program, which implements the steps of the method described in the second aspect when the program is executed by a processor.
[0015] The present invention has the following advantages compared with the prior art: 1. The image recognition agent based on the SAM model realizes pixel-level semantic segmentation of roof elements (such as rooftop greening, solar panels, equipment rooms, etc.), supports three interactive Prompt input modes: point selection, box selection, and masking, and significantly improves efficiency compared to traditional manual visual interpretation methods, meeting the requirements of large-scale urban areas.
[0016] 2. The exponential calculation agent dynamically calls index formulas (index formulas are pre-stored in the index library) in official documents such as "Guidelines for the Construction of the Fifth Facade of the City" through the RAG technology, automatically matches the calculation formulas of multiple core indicators such as greening coefficient and space utilization rate, eliminates the subjective errors of manual operations, and ensures that the evaluation results meet industry standards.
[0017] 3. The Auto-Summary summary model integrates the image recognition results and index data, generates a structured report including raster masks, heatmap annotations, and 3D interactive models, synchronously outputs the positioning of violation areas and rectification suggestions, and realizes the visual presentation and traceable management of evaluation conclusions.
[0018] The technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0019] Figure 1 It is the system principle block diagram of the present invention.
[0020] Figure 2 It is the principle block diagram of the main control model.
[0021] Figure 3 It is the principle block diagram of the image recognition agent.
[0022] Figure 4 It is the principle block diagram of the image encoder.
[0023] Figure 5 It is the principle block diagram of the prompt encoder.
[0024] Figure 6 It is the principle block diagram of the mask decoder.
[0025] Figure 7 It is the principle block diagram of the exponential calculation agent.
[0026] Figure 8 It is the principle block diagram of the operation generation tool.
[0027] Figure 9 It is the method flow chart for evaluating the greening coverage rate of the fifth facade roof of the city. Detailed Embodiments
[0028] Embodiment 1: As Figure 1As shown in the figure, an intelligent evaluation system for the roof space of the fifth facade in the city includes: a main control model, an image recognition agent, an index calculation agent, and an Auto-Summary summary model.
[0029] The main control model is connected to a large language model and is used to parse the aerial images, texts, or Prompt instructions input by users, generate a structured task logic chain, and allocate it to the corresponding agent. Specifically, as Figure 2 shown, the main control model mainly relies on the LLM ability, realizes the parsing of user input data types and the recognition of user intentions based on prompt engineering, judges the task type (element recognition, index calculation, composite task), and outputs a structured intention label, and calls the relevant workflow based on the label.
[0030] The image recognition agent is an agent generated based on the pre-trained SAM model. The image recognition agent includes an image encoder, a prompt encoder, and a mask decoder, and is used to perform semantic segmentation on the roof image and output a raster mask with labels as the image recognition result.
[0031] Specifically, the image encoder, uses the ViT (Vision Transformer) pre-trained by MAE (Masked Autoencoder) as the backbone network, and combines an improved Fpn (Feature Pyramid Network) to construct a multi-scale feature pyramid. As Figure 3 shown, the image encoder mainly realizes the function of mapping the image to be segmented to the image feature space. The backbone is the ViT pre-trained by MAE, and the neck is the improved Fpn, which realizes the input image and returns three feature vectors: src (the highest-level feature map), pos (the position encoding of each level of feature map), and features (a list of multi-scale feature pyramids).
[0032] The prompt encoder supports three interactive Prompt input modes: point, box, and mask, and converts the user's Prompt instructions into sparse or dense embedding vectors. Specifically, as Figure 4 shown, the prompt encoder mainly realizes the function of converting sparse prompts (points, boxes) and dense prompts (masks) into sparse embeddings and dense embeddings. Input the marker point, box, and mask parameters, and the output is a feature vector.
[0033] The Mask decoder receives the image features and the embedded vector converted from the Prompt instruction, outputs a binary mask, an IoU confidence level, and the coordinates of the segmentation region, and supports interactive correction based on a confidence threshold.
[0034] Specifically, as Figure 5 shown, the Mask decoder is the main module that implements the mask prediction function. Its inputs are the feature vectors generated by the first two modules and the output token (formed by concatenating the learnable parameters iou_token and mask_token), and the outputs are logits, a binary mask, and iou scores.
[0035] The exponential calculation agent is externally connected to an index library, which is used to retrieve the scoring formula that matches the requirements in the index library through RAG, and drives a spatial operation tool to generate the calculation result of the index data of the roof image based on the scoring formula.
[0036] Specifically, as Figure 6 shown, the exponential calculation agent mainly relies on the LLM ability. By fine-tuning the base model and performing RAG retrieval on the externally connected knowledge base, it matches the calculation formula according to the text description and returns it. The exponential calculation agent includes: a RAG retrieval framework that fine-tunes the base model with Qwen3-30B-A3B as the base model through the reinforcement learning method, infers the user instruction based on the fine-tuned base model, and retrieves the matching scoring formula from the index library; an operation tool chain that includes an instance segmentation engine, a spatial overlay analyzer, and a formula parser, which are used to convert the natural language scoring formula into an executable expression and perform spatial operations to generate the calculation result of the index data; a spatio-temporal correlation database that stores the calculation result of the index data and supports tracing the change of the roof state according to the time series.
[0037] It should be noted that, as Figure 7 shown, the operation tool chain mainly based on various built-in tool functions, inputs the roof mask and element masks generated by the image recognition agent, and the calculation formula generated by the exponential calculation agent, and returns the index vector calculation data after instance segmentation, image overlay, and spatial operations.
[0038] The Auto-Summary summary model integrates the capabilities of multi-modal large models and is used to fuse the image recognition results and the calculation results of the index data to generate a comprehensive report on the rectification suggestions, and the comprehensive report includes suggestions for the roof layout.
[0039] Specifically, the Auto-Summary summary model further includes: The multi-modal alignment module associates text instructions with image features through a cross-modal attention mechanism, and establishes a mapping relationship between spatial coordinates and evaluation metrics. The spatial coordinates are the spatial coordinates of the roof corresponding to the aerial image of the roof, and the evaluation metrics are the evaluation metrics corresponding to the area where the roof is located; The visualization rendering unit generates a heat map to mark the violation area, a 3D interactive model, and a structured data report; The rectification suggestion generation unit outputs natural language suggestions based on the standard solutions in the knowledge base.
[0040] It should be noted that the Auto-Summary model further includes: The closed-loop feedback unit periodically triggers the drone re-inspection task, and generates a rectification effect evaluation report by comparing historical data; The early warning push module automatically generates an early warning notice when a new illegal construction is detected or the greening rate drops by more than a preset threshold.
[0041] Embodiment 2: A method for evaluating the fifth elevation of a city based on the above system includes the following steps: Receive the aerial image and text instructions uploaded by the user, parse the task type through the main control model, and allocate it to the corresponding intelligent agent; Call the image recognition intelligent agent to segment the roof elements and generate a raster mask with semantic labels; Through the index calculation intelligent agent, load the scoring formula to calculate the greening coverage rate, space utilization rate, and safety hazard level index; Use the Auto-Summary model to fuse multi-modal data, generate a visualization report, and mark the violation area; Push the analysis results to the management platform, and trigger a periodic re-inspection task to form a closed-loop management.
[0042] It should be noted that the step of "segmenting roof elements" further includes: Use the SAM model to perform multi-scale feature extraction on the aerial image, and optimize the segmentation accuracy in combination with the user interactive Prompt instruction; Trigger the manual correction process for areas with a confidence level lower than the threshold to ensure the reliability of the mask output.
[0043] It should be noted that the step of "calculating the greening coverage rate, space utilization rate, and safety hazard level index" further includes: Dynamically match the index library according to natural language, and retrieve the applicable formula through the RAG technology; The roof segmentation mask (roof mask) output by the image recognition intelligent agent and the mask of the roof element segmentation required by the user are first subjected to id mapping, and then overlay operation is performed according to the index formula.
[0044] It should be noted that the step of "generating a visualization report" further includes: Mapping the image recognition results and index data to a geographic coordinate system through a GIS engine; Generating an interactive 3D heat map that supports users to click and query specific violation items and rectification suggestions.
[0045] As Figure 9 shown below, taking the implementation scenario of the evaluation of the roof greening coverage rate of the fifth facade of a city as an example, the technical solution of the present invention will be further described. The method for evaluating the roof greening coverage rate of the fifth facade of a city includes the following steps: Step 1: Data collection and input: Aerial photographing the target area by a drone equipped with a high-resolution camera (resolution ≥ 0.1 m) to obtain roof RGB image data; The user selects the evaluation task type (such as "greening coverage rate analysis") through the system interaction interface and uploads an image data set containing the roofs of multiple buildings; It should be noted that the input mode supports multi-source data input, including drone aerial photography images, satellite remote sensing data, and local photos manually uploaded by users; The user can attach text instructions (such as "statistical rooftop greening area") or specify the area of interest through an interactive Prompt (points, boxes, masks).
[0046] Step 2: Master model task parsing and allocation: The master model parses the user input based on a large language model (such as using Qwen3-30B-A3B as the base model and fine-tuning the base model through reinforcement learning methods), and identifies the task type as a "composite task" (feature recognition + index calculation); Generating a structured task logic chain: Roof feature segmentation → Greening coverage rate calculation → Report generation; Resource allocation: Invoking an image recognition agent, an index calculation agent, and an Auto-Summary summary model; It should be noted that the master model distributes the aerial photography images to the image recognition agent through the API interface and triggers the knowledge base to load the formula library at the same time.
[0047] Step 3: Operation of the image recognition agent (based on the SAM model): Image Encoder, using the ViT-H / 16 model pre-trained with MAE (Masked Autoencoder) as the image encoder; the input image is segmented into 16×16 pixel blocks by ViT, and a multi-scale feature pyramid (including three groups of vectors: src, pos, and features) is extracted; different-level feature maps are fused through an improved Fpn (Feature Pyramid Network) to enhance the recognition ability of small targets (such as potted greenery). Prompt Encoder, supporting three types of Prompt inputs: point, box, and mask; the system automatically boxes the roof area as the global Prompt; the user clicks on the unrecognized area to add point prompts; the Prompt coordinates or mask are converted into sparse / dense embedding vectors and aligned with the image features. Mask Decoder: Receives the image encoding features and the Prompt embedding vectors, combines the learnable iou_token and mask_token to generate output Tokens; predicts the binary mask through multiple layers of Transformer decoders and calculates the IoU confidence (the threshold is set to 0.85); triggers interactive correction for areas with confidence below the threshold to ensure segmentation accuracy.
[0048] Step 4: Execution of the exponential calculation agent (RAG framework and operation toolchain): Based on the retrieval-augmented generation technology, dynamically call the metric library (such as the metric formula corresponding to the "Roof Greening Management Regulations of XX City"); use the fine-tuned Qwen3-30B-A3B model (the base model has the ability to disassemble into basic formula operations according to the user input requirements) to encode the user instructions into vectors, and retrieve the top-3 relevant formulas through cosine similarity (such as "Greening coverage rate = Green plant projection area / Roof greenable area × 100%"); perform pixel-level classification on the image recognition results and extract the vector boundaries of the greening areas; use the spatial overlay analyzer to calculate the geometric overlay result of the green plant mask and the total roof area, and use the formula parser: convert the natural language scoring formula (such as "Space utilization rate = Equipment room area / Total roof area") into a mathematical expression and substitute the numerical calculation results; the calculation results are stored in the spatio-temporal correlation database, supporting the tracing of roof state changes over time (such as quarterly fluctuations in greening coverage rate).
[0049] Step 5: The Auto-Summary summarization model generates a report: On the text side, extract user requirement keywords (such as "illegal building detection") and the evaluation index system (such as "greening coverage rate ≥ 70% is excellent"); on the image side, identify spatial features such as roof material distribution and equipment room location; through the GIS engine, convert the image pixel coordinates into the WGS-84 geographic coordinate system, and establish a quantitative mapping between spatial features and indicators; the report output form is a structured report: including building number, greening area, coverage rate level (excellent / qualified / failed to meet the standard) and historical comparison data; 3D heat map: mark the illegal areas (such as ungreened areas) in red and the compliant areas in green, supporting interactive zooming; rectification suggestions: generate natural language suggestions based on the knowledge base (such as "It is recommended to demolish the illegal structure in the northwest corner, which is expected to reduce the safety hazard level to B level").
[0050] Step 6: Closed-loop feedback and re-inspection: The system automatically pushes the report to the urban planning management platform and generates a work order (such as number #2024-058); further, it can support API docking with the urban management work order system to trigger manual verification or drone re-inspection tasks; set to automatically call the drone to re-take pictures of the target roof after 30 days, and generate a rectification effect evaluation report by comparing historical data.
[0051] It should be noted that: The training data for SAM model construction includes 100,000 labeled roof images, covering various scenarios such as sunny, rainy, and snowy days to enhance the robustness of the model; the loss function is jointly optimized using cross-entropy loss (classification error) and Dice loss (mask edge smoothness); it supports real-time update of the mask after the user adds point prompts, with a response time ≤ 0.5 seconds. For the knowledge base construction of the exponential calculation agent, convert the building code PDF document into an index formula for storage and use the Faiss vector database for storage (dimension 768); adopt a hybrid strategy during retrieval, first filter by geographical area and then sort by semantic similarity. The parsing logic of the formula is: natural language → mathematical symbol mapping table (such as "proportion" → " / ", "square kilometer" → "km 2 "); support nested calculation of composite formulas (such as "greening coefficient = 0.3 × coverage rate + 0.7 × vegetation diversity index"). When performing multi-modal alignment in Auto-Summary, use the Auto-Summary model constructed by an improved CLIP model, and align text and image features through a cross-attention layer. For heat map generation, map the pixel values of the illegal area to the HSV color space based on the OpenCV library, with red indicating high risk. For the interactive panel, use WebGL technology to achieve dynamic rotation and click query of 3D models.
[0052] Example 3: A computer-readable storage medium stores a computer program, which when executed by a processor implements the steps of the method described in Example 2.
[0053] The storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs.
[0054] As mentioned above, the above are only the preferred embodiments of the present invention and do not impose any limitations on the present invention. Any simple modifications, changes, and equivalent structural changes made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. An intelligent evaluation system for the roof space of the fifth elevation in a city, characterized in that including: The main control model is connected to a large language model, which is used to parse the aerial images, texts or Prompt instructions input by users, generate a structured task logic chain and allocate it to the corresponding agent; The image recognition agent is an agent generated based on the pre-trained SAM model. The image recognition agent includes an image encoder, a prompt encoder and a mask decoder, which are used to perform semantic segmentation on the roof image and output a labeled raster mask as the image recognition result; The index calculation agent is externally connected to an index library, which is used to retrieve the scoring formula that matches the requirements in the index library through RAG, and drive the spatial operation tool based on the scoring formula to generate the calculation result of the index data of the roof image; The Auto-Summary summary model integrates the capabilities of multi-modal large models, which is used to fuse the image recognition result and the calculation result of the index data, generate a comprehensive report on the rectification suggestions, and the comprehensive report includes the roof layout suggestions.
2. The system according to claim 1, wherein The image encoder uses ViT pre-trained by MAE as the backbone network, and combines an improved Fpn to construct a multi-scale feature pyramid; The prompt encoder supports three interactive Prompt input modes: point, box, and mask, and converts the user's Prompt instruction into a sparse or dense embedding vector; The mask decoder receives the image features and the embedding vector converted from the Prompt instruction, outputs a binary mask, an IoU confidence level, and the coordinates of the segmentation region, and supports interactive correction based on the confidence threshold.
3. The system according to claim 1, characterized in that The index calculation agent includes: The RAG retrieval framework, based on Qwen3-30B-A3B as the base model, fine-tunes the base model through the reinforcement learning method, infers the user's instruction based on the fine-tuned base model, and retrieves the matching scoring formula from the index library; The operation tool chain includes an instance segmentation engine, a spatial overlay analyzer and a formula parser, which are used to convert the natural language scoring formula into an executable expression and perform spatial operations to generate the calculation result of the index data; The spatio-temporal association database stores the calculation results of the index data.
4. The system according to claim 1, wherein The Auto-Summary summary model further includes: The multi-modal alignment module associates the text instruction with the image features through the cross-modal attention mechanism, and establishes a mapping relationship between the spatial coordinates and the evaluation indicators. The spatial coordinates are the spatial coordinates of the roof corresponding to the roof image, and the evaluation indicators are the evaluation indicators corresponding to the area where the roof is located; The visualization rendering unit generates a heat map to mark the violation areas, a 3D interactive model and a structured data report; The rectification suggestion generation unit outputs natural language suggestions based on the standard solutions in the knowledge base.
5. The system according to claim 4, wherein The Auto-Summary summary model also includes: The closed-loop feedback unit periodically triggers the drone re-inspection task and generates a rectification effect evaluation report by comparing historical data; The early warning push module automatically generates an early warning notice when it detects new illegal construction or the greening rate drops by more than the preset threshold.
6. A method for evaluating the fifth elevation of a city based on the system according to any one of claims 1-5, characterized in that, including the following steps: Receive the aerial image and text instruction uploaded by the user, parse the task type through the main control model and allocate it to the corresponding agent; Call the image recognition agent to segment the roof elements and generate a raster mask with semantic labels; The index calculation agent loads the scoring formula to calculate the greening coverage rate, space utilization rate, and safety hazard level indicators; Use the Auto-Summary summary model to fuse multi-modal data, generate a visual report, and mark the violation areas; Push the analysis results to the management platform and trigger a periodic re-inspection task to form a closed-loop management.
7. The method according to claim 6, wherein The step of "segmenting roof elements" further includes: Use the SAM model to perform multi-scale feature extraction on the aerial image and optimize the segmentation accuracy in combination with the user interactive Prompt instruction; Trigger the manual correction process for areas with confidence lower than the threshold to ensure the reliability of the mask output.
8. The method according to claim 6, characterized in that, The step of "calculating the greening coverage rate, space utilization rate, and safety hazard level indicators" further includes: Dynamically match the index library according to natural language and retrieve the applicable formula through the RAG technology; The roof segmentation mask output by the image recognition agent and the roof element segmentation mask required by the user are first subjected to id mapping, and then overlay calculation is performed according to the index formula.
9. The method according to claim 6, wherein The step of "generating a visual report" further includes: Map the image recognition results and index data to the geographic coordinate system through the GIS engine; Generate an interactive 3D heat map to support users to click and query specific violation items and rectification suggestions.
10. A computer-readable storage medium storing a computer program, characterized in that, When the program is executed by the processor, the steps of the method described in any one of claims 6-9 are implemented.
Citation Information
Patent Citations
Safety production supervision system and method based on multi-modal large model
CN119274142A
SAM-based self-prompting semantic segmentation method and apparatus, and storage medium
CN119863623A
Pitch determination systems and methods for aerial roof estimation
US20100110074A1
Cited By
Roof greening fairness dynamic evaluation and composite risk identification method
CN121010472A
Underground pipeline pipe gallery resource communication method and device based on large language model
CN122197244A
Underground pipeline pipe gallery resource connection method and device based on large language model
CN122197244B