General building industry polling image analysis and report system based on artificial intelligence
By integrating visual large-scale model technology into an AI-based system, the system achieves automated management and intelligent patrol of cameras, solving the problem of low intelligence in existing monitoring systems, improving monitoring efficiency and anomaly detection accuracy, and providing a brand-new monitoring solution.
Patent Information
- Application Number
- CN202510972345.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-11-11
AI Technical Summary
Existing camera monitoring systems rely heavily on manual operation for the management of preset points and automatic rotation of PTZ cameras. They have a low level of intelligence and are difficult to achieve intelligent analysis of image content and automatic generation of inspection reports, resulting in the difficulty of effectively retaining and analyzing monitoring data.
The system employs an artificial intelligence-based approach, integrating large-scale visual model technology to achieve automated management, intelligent round-robin control, and structured report generation for cameras. Through a front-end and back-end separation architecture and modular design, combined with round-robin scheduling algorithms, intelligent image analysis, and report generation technology, it enables automated monitoring and anomaly detection.
It significantly improves the automation level and anomaly detection efficiency of monitoring operations, reduces manual intervention, enhances the reliability and availability of monitoring data, and provides brand-new security monitoring, industrial inspection and urban management solutions.
Smart Images

Figure CN120929052A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent monitoring and image analysis technology, and in particular to an artificial intelligence-based system for rotating image analysis and reporting in the construction industry. Background Technology
[0002] In modern security monitoring, industrial inspection, and urban management, camera surveillance systems serve as crucial infrastructure, widely used in industries such as engineering construction, park operation and maintenance, and transportation. However, existing camera surveillance systems generally suffer from low levels of intelligence, particularly in the management of preset points and automatic rotation of PTZ cameras, which still heavily rely on manual operation. Monitoring personnel need to manually switch camera views and continuously observe the screen with their naked eyes to identify anomalies. This method is not only labor-intensive but also prone to missing critical events due to subjective factors or lack of attention. Furthermore, traditional surveillance systems lack intelligent analysis capabilities for image content, cannot perform structured image processing, and lack mechanisms for automatically generating inspection reports. This makes it difficult to effectively retain and subsequently analyze monitoring data, thus limiting their application value in areas such as intelligent security and decision support.
[0003] While some current technical solutions attempt to introduce automation, their level of intelligence remains insufficient, particularly in image analysis and anomaly detection. Traditional methods, often based on simple rule matching or fixed templates, struggle to address the diverse needs of complex scenarios. Furthermore, these systems suffer from inefficiency and lack of flexibility in scheduling and executing patrol tasks, failing to meet users' demands for refined time control and efficient task management. Therefore, developing a system capable of intelligent camera management, automatic patrolling, real-time image analysis, and structured report generation is crucial for improving monitoring efficiency and accuracy.
[0004] This invention aims to solve the aforementioned problems in existing technologies by introducing artificial intelligence, image recognition, workflow, and automation technologies. Specifically, this invention provides an AI-based system for analyzing and reporting camera patrol images in the broader construction industry. This system integrates large-scale visual model technology to achieve automated management, intelligent patrol control, real-time image analysis, and structured report generation for cameras. This innovation not only significantly improves the automation level and anomaly detection efficiency of daily monitoring operations but also provides a novel solution for fields such as security monitoring, industrial inspection, and urban management, possessing significant technical importance and application prospects. Summary of the Invention
[0005] The purpose of this invention is to provide an AI-based system for analyzing and reporting images in the broader construction industry, thereby solving the problems mentioned in the background section.
[0006] This invention is implemented as follows: an AI-based image analysis and reporting system for the broader construction industry, comprising a front-end interface layer, a back-end service layer, and an intelligent analysis layer. The front-end interface layer, developed using the Nuxt 3 and Vue 3 frameworks, implements user interaction functions. The back-end service layer, built on Express and MongoDB, handles core logic processing. The intelligent analysis layer integrates the Qwen VL Plus visual model for intelligent image analysis, anomaly detection, and structured report generation. The front-end interface layer provides modules for camera management, preset point configuration, patrol plan settings, and report viewing, allowing users to add, delete, modify, and query cameras, set preset points, and configure patrol tasks. Furthermore, the back-end service layer provides a RESTful API interface via Express to receive front-end requests and perform operations such as data management, patrol scheduling, image acquisition, and storage. Specifically, the intelligent analysis layer utilizes the Qwen VL Plus visual model to perform object recognition, scene understanding, anomaly detection, and risk assessment on the acquired images, ultimately generating structured analysis results.
[0007] Furthermore, the core functional modules of this invention include a camera and preset point management module, a patrol plan management module, a patrol execution module, and an intelligent analysis and report generation module. The camera and preset point management module supports adding, editing, deleting, and monitoring the status of cameras, and can preview camera images in real time. Specifically, this module maps camera positions onto a project floor plan or aerial photograph, displaying the spatial distribution of cameras in icon form and saving their coordinates. Furthermore, the preset point management function allows users to set preset points by moving the PTZ camera view to a designated location and taking a screenshot, while also supporting the editing, deletion, and image management of preset points. The patrol plan management module supports the creation, editing, and deletion of patrol plans. Users can name patrol plans, add specific points, and describe the plan content. Specifically, this module provides a time window setting function, allowing users to limit the execution scope of patrol tasks by inputting the patrol start and end times. Furthermore, users can adjust the order of preset points in the patrol plan by dragging and dropping, and configure the dwell time at each point.
[0008] The polling execution module automatically initiates polling tasks based on user-defined time windows and execution counts. This module uses a timer mechanism to execute polling tasks at calculated time points and updates the execution status to schedule the next execution. Specifically, during polling, the system automatically switches cameras to preset locations and stays at each location for a specified time to capture images. Furthermore, the image acquisition and storage function automatically saves images acquired during polling and records relevant metadata, including capture time, camera number, and preset point information. The intelligent analysis and report generation module integrates the Qwen VL Plus visual model to intelligently analyze the acquired images. This module performs image preprocessing, feature extraction, object recognition, and classification to achieve scene understanding and contextual analysis. Specifically, the module combines visual and contextual information for comprehensive anomaly detection and generates structured analysis results. Furthermore, the report generation process includes collecting polling metadata and image analysis results, extracting key findings and recommendations, and generating HTML and JSON format reports based on templates. Specifically, the generated report content includes findings, recommendations, and an overall assessment, and persistent storage ensures long-term data availability.
[0009] The key technologies and algorithms of this invention include a round-robin scheduling algorithm, an intelligent image analysis algorithm, and a report generation technology. The round-robin scheduling algorithm employs an intelligent time window scheduling strategy to ensure efficient execution of round-robin tasks. S1: Calculate the round-robin time window and determine the start and end times. S2: Determine the execution strategy based on the execution type. S3: For a finite number of executions, calculate the optimal execution interval within the time window. S4: Set a timer to execute the round-robin task at the calculated time points. S5: After the task execution is complete, update the execution status and schedule the next execution. Specifically, the algorithm considers factors such as system load, execution history, and priority to ensure even distribution and efficient execution of round-robin tasks.
[0010] Image intelligent analysis algorithms achieve object recognition, scene understanding, and anomaly detection through multi-step image processing. S1: Preprocess and extract features from the acquired image. S2: Identify and classify objects in the image using a deep learning model. S3: Understand the scene and perform contextual analysis based on contextual information. S4: Identify anomalies in the image using an anomaly detection model and conduct risk assessment. S5: Generate structured analysis results, including an object list, scene description, and anomaly information. In particular, the algorithm combines visual and contextual information, improving the accuracy and reliability of anomaly detection.
[0011] The report generation technology employs a template-driven approach, combined with the text generation capabilities of a large model, to automatically generate structured patrol reports. S1: Collect patrol metadata and image analysis results. S2: Extract key findings and recommendations, including anomalies, risk assessments, and improvement measures. S3: Generate report content based on a predefined template, ensuring a consistent and easy-to-read report format. S4: Construct HTML and JSON format reports for easy user viewing and subsequent processing. S5: Persistently store the generated reports and related resources, ensuring data security and traceability. In particular, this technology, through structured storage and analysis, uncovers the deeper value of monitoring data.
[0012] Furthermore, this invention achieves high scalability and flexibility through modular design and a front-end / back-end separation architecture. The front-end interface layer uses the Nuxt 3 and Vue 3 frameworks, supporting rapid iteration and efficient development; the back-end service layer is built on Express and MongoDB, possessing high performance and high reliability. In particular, the intelligent analysis layer, through the integration of the QwenVL Plus visual model, achieves deep understanding of image content and comprehensive anomaly detection. Furthermore, the system employs a retry mechanism and error handling strategy to ensure stable operation in complex environments.
[0013] The innovations of this invention include an intelligent cruise scheduling algorithm, AI-driven image analysis, automated report generation, multimodal anomaly detection, visualized map management, and flexible execution control. The intelligent cruise scheduling algorithm automatically arranges cruise tasks based on time windows and execution counts, reducing manual intervention. Furthermore, AI-driven image analysis achieves a deep understanding of image content by integrating a large visual model. In particular, the automated report generation function automatically generates structured cruise reports based on analysis results, improving work efficiency. Multimodal anomaly detection combines visual and contextual information for comprehensive judgment, enhancing the accuracy of anomaly detection. Visualized map management intuitively displays the spatial relationship between cameras and preset points, improving the user experience. Flexible execution control supports multiple execution modes and fine-grained time control to meet the needs of different scenarios.
[0014] The technical effects of this invention are achieved through the above-described technical solutions. First, the system significantly reduces manual intervention and improves monitoring efficiency through intelligent cruise scheduling algorithms and automated report generation. Second, AI-driven image analysis and multimodal anomaly detection enhance the accuracy and reliability of anomaly detection. Furthermore, visualized map management and flexible execution control enhance user experience and system adaptability. In particular, the modular design and front-end / back-end separation architecture ensure high scalability and flexibility of the system. In summary, this invention provides a novel solution for security monitoring, industrial inspection, urban management, and other fields, possessing broad application prospects and technical value. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the overall system structure of the present invention, showing the front-end and back-end separation architecture of the system, including three main parts: the front-end interface layer, the back-end service layer, and the intelligent analysis layer;
[0016] Figure 2 This is a flowchart of the polling plan and execution process of the present invention, which describes the complete process from the user creating a polling plan to the system automatically executing the polling task and generating a polling report;
[0017] Figure 3 This is a timing interaction diagram of the camera patrol control system of the present invention, which shows in detail the interaction process between user operation and system response, including steps such as login, configuring task parameters, performing shooting and status update. Detailed Implementation
[0018] This invention provides an artificial intelligence-based system for analyzing and reporting images from camera patrols in the broader construction industry. The specific implementation is described in conjunction with the accompanying drawings. Figure 1 , Figure 2 and Figure 3 A detailed description is provided below. The system adopts a front-end and back-end separation architecture, comprising three main parts: the front-end interface layer, the back-end service layer, and the intelligent analysis layer. The front-end interface layer is developed based on the Nuxt 3 and Vue 3 frameworks, the back-end service layer is built on Express and MongoDB, and the intelligent analysis layer integrates the Qwen VL Plus visual model. The following section, with reference to the accompanying diagrams and specific application scenarios, comprehensively explains the system's operating principles and implementation process.
[0019] The overall system structure is as follows Figure 1 As shown, the front-end interface layer is responsible for user interaction functions, including camera management, preset point configuration, polling plan settings, and report viewing. The back-end service layer provides a RESTful API interface through Express to handle core logic such as data management, polling scheduling, image acquisition, and storage. The intelligent analysis layer utilizes the Qwen VLPlus visual big data model to achieve intelligent image analysis, anomaly detection, and structured report generation. The layers communicate with each other through API interfaces to ensure efficient and stable data flow. In practical applications, users complete the CRUD operations for cameras and configure polling plans through the front-end interface layer. The back-end service layer receives requests from the front-end, breaks down the tasks into specific execution steps, and calls the intelligent analysis layer to perform in-depth analysis of the acquired images. Finally, the system generates a structured polling report and stores it in a MongoDB database.
[0020] The rotation plan and execution process are as follows: Figure 2As shown, the process from user-created patrol plan to system-automated execution of patrol tasks and generation of patrol reports consists of multiple stages. First, the user creates the patrol plan through the front-end interface, including naming the plan, adding specific locations, describing the plan content, and setting a time window. After receiving the patrol plan, the back-end service layer calculates the optimal execution interval based on the user-defined time range and number of executions, and starts the scheduling task through a timer mechanism. Once the patrol task starts, the system automatically switches the camera to the preset location and stays at each location for a specified time to capture images. During the capture process, the system saves images and related metadata in real time, including the capture time, camera number, and preset location information. The intelligent analysis layer performs object recognition, scene understanding, and anomaly detection on the collected images, extracting key findings and suggestions. Finally, the system generates an HTML-formatted patrol report based on a predefined template and persistently stores the report and related resources.
[0021] The timing interaction of the camera polling control system is as follows: Figure 3 The diagram illustrates the interaction process between user operations and system responses. After user login, the system loads homepage data and displays the current camera status, patrol plan execution status, and historical report information on the dashboard. Users can configure patrol task parameters through the front-end interface layer, such as selecting cameras, setting the preset point order, and dwell time. After configuration, the user submits a patrol task creation request, which the back-end service layer receives and verifies the validity of the parameters. Upon successful verification, the back-end service layer sets a timer task to trigger the patrol task at the calculated time points. During patrol task execution, the system calls the camera API interface to control the PTZ camera to move to the preset point position and take a picture. After the picture is taken, the system stores the image and metadata in a MongoDB database and updates the execution status to schedule the next execution.
[0022] The system's core functional modules include a camera and preset point management module, a patrol plan management module, a patrol execution module, and an intelligent analysis and report generation module. The camera and preset point management module supports adding, editing, deleting, and monitoring the status of cameras. Users can register cameras through the front-end interface and map their locations onto the project's floor plan or aerial image. After mapping, the system displays the spatial distribution of cameras as icons and saves their coordinates. The preset point management function allows users to set preset points by moving the PTZ camera view to a designated location and taking a screenshot. It also supports editing, deleting, and managing images of preset points. The patrol plan management module supports creating, editing, and deleting patrol plans. Users can name patrol plans, add specific points, and describe the plan content. The module provides a time window setting function, allowing users to limit the execution scope of patrol tasks by inputting the patrol start and end times. Users can also adjust the order of preset points in the patrol plan by dragging and dropping, and configure the dwell time for each point.
[0023] The polling execution module automatically initiates polling tasks based on the user-defined time window and execution count. The module employs an intelligent time window scheduling algorithm to ensure efficient execution of polling tasks. First, the system calculates the polling time window, determining the start and end times. Next, it determines the execution strategy based on the execution type. For a limited number of executions, the system calculates the optimal execution interval within the time window and sets a timer to execute the polling task at the calculated time points. After task execution is complete, the system updates the execution status and schedules the next execution. During polling, the system automatically switches the camera to preset locations and stays at each location for a specified time to capture images. After capturing images, the system automatically saves the acquired images and records relevant metadata.
[0024] The intelligent analysis and report generation module integrates the Qwen VL Plus visual model to perform intelligent analysis on the acquired images. The module achieves scene understanding and contextual analysis through image preprocessing, feature extraction, object recognition, and classification. Specifically, the system first preprocesses and extracts features from the acquired images, removing noise and enhancing key features. Next, it uses a deep learning model to identify and classify objects in the images, generating an object list. Then, the system combines contextual information to understand the scene and perform contextual analysis, determining if any anomalies exist. The anomaly detection model uses multimodal analysis technology, combining visual and contextual information to make a comprehensive anomaly judgment and generate a risk assessment report. Finally, the system generates an HTML-formatted polling report based on a predefined template and persistently stores the report and related resources.
[0025] Key technologies and algorithms include a round-robin scheduling algorithm, an intelligent image analysis algorithm, and a report generation technology. The round-robin scheduling algorithm employs an intelligent time window scheduling strategy to ensure even distribution and efficient execution of round-robin tasks. The algorithm dynamically adjusts the execution time of round-robin tasks based on factors such as system load, execution history, and priority. The intelligent image analysis algorithm performs multi-step image processing to achieve object recognition, scene understanding, and anomaly detection. The system first preprocesses and extracts features from the acquired images, then uses a deep learning model to identify and classify objects in the images. Next, the system combines contextual information to understand the scene and perform situational analysis, and identifies anomalies in the images using an anomaly detection model. The report generation technology uses a template-driven approach, combined with the text generation capabilities of a large model, to automatically generate structured round-robin reports. The system first collects round-robin metadata and image analysis results, extracting key findings and recommendations. Then, it generates report content based on a predefined template and constructs HTML and JSON format reports. Finally, the system persistently stores the generated reports and related resources to ensure data security and traceability.
[0026] The system achieves high scalability and flexibility through modular design and a front-end / back-end separation architecture. The front-end interface layer uses the Nuxt 3 and Vue 3 frameworks, supporting rapid iteration and efficient development. The back-end service layer is built on Express and MongoDB, ensuring high performance and reliability. The intelligent analysis layer integrates the Qwen VL Plus visual model to achieve deep understanding of image content and comprehensive anomaly detection. The system employs retry mechanisms and error handling strategies to ensure stable operation in complex environments. For example, in the event of network interruption or camera malfunction, the system will automatically attempt to reconnect or skip the current task, and log errors for subsequent troubleshooting.
[0027] The innovations of this invention include an intelligent cruise scheduling algorithm, AI-driven image analysis, automated report generation, multimodal anomaly detection, visualized map management, and flexible execution control. The intelligent cruise scheduling algorithm automatically arranges cruise tasks based on time windows and execution counts, reducing manual intervention. AI-driven image analysis achieves deep understanding of image content by integrating large-scale visual models. The automated report generation function automatically generates structured cruise reports based on analysis results, improving work efficiency. Multimodal anomaly detection combines visual and contextual information for comprehensive judgment, enhancing the accuracy and reliability of anomaly detection. Visualized map management intuitively displays the spatial relationship between cameras and preset points, improving user experience. Flexible execution control supports multiple execution modes and fine-grained time control to meet the needs of different scenarios.
[0028] The technical effects of this invention are achieved through the above-described technical solutions. The system significantly reduces manual intervention and improves monitoring efficiency through intelligent cruise scheduling algorithms and automated report generation. AI-driven image analysis and multimodal anomaly detection enhance the accuracy and reliability of anomaly detection. Visualized map management and flexible execution control enhance user experience and system adaptability. Modular design and a front-end / back-end separation architecture ensure high scalability and flexibility. In summary, this invention provides a novel solution for security monitoring, industrial inspection, urban management, and other fields, possessing broad application prospects and technological value.
[0029] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An AI-based system for analyzing and reporting images in the broader construction industry, characterized in that: It includes a front-end interface layer, a back-end service layer, and an intelligent analysis layer. The front-end interface layer is developed based on the Nuxt 3 and Vue 3 frameworks to implement user interaction functions. The back-end service layer is built based on Express and MongoDB to handle data management, polling scheduling, image acquisition and storage tasks. The intelligent analysis layer integrates the Qwen VL Plus visual big model for intelligent image analysis, anomaly detection, and structured report generation.
2. The AI-based panoramic image analysis and reporting system for the construction industry as described in claim 1, characterized in that, The front-end interface layer provides modules for camera management, preset point configuration, patrol plan setting, and report viewing, supporting users to perform operations such as adding, deleting, modifying, and querying cameras and setting preset points.
3. The AI-based panoramic image analysis and reporting system for the construction industry as described in claim 2, characterized in that, Furthermore, the spatial distribution of cameras is displayed as icons by mapping their locations onto the project floor plan or aerial photograph, and the location coordinates are saved.
4. The AI-based panoramic image analysis and reporting system for the construction industry as described in claim 1, characterized in that, The backend service layer provides a RESTful API interface through Express to receive frontend requests and perform data management, round-robin scheduling, image acquisition and storage operations.
5. The AI-based panoramic image analysis and reporting system for the construction industry as described in claim 4, characterized in that, The backend service layer uses an intelligent time window scheduling algorithm to calculate the optimal execution interval for the polling task, and starts the polling task at a specified time point through a timer mechanism.
6. The AI-based panoramic image analysis and reporting system for the construction industry as described in claim 1, characterized in that, The intelligent analysis layer performs preprocessing, feature extraction, object recognition, and classification on the acquired images to understand the scene and analyze the context, and combines visual and contextual information to make comprehensive anomaly judgments.
7. The AI-based panoramic image analysis and reporting system for the construction industry as described in claim 6, characterized in that, The intelligent analysis layer generates round-robin reports in HTML and JSON formats based on predefined templates, and persists the generated reports and related resources.