Construction site supervision system and method based on edge-cloud collaboration and long-term memory enhancement

By combining edge-cloud collaboration with long-term memory enhancement, the construction site supervision system solves the problems of high false alarm rate and low management efficiency of existing monitoring systems at construction sites, and realizes intelligent management and safety assurance.

CN122492115APending Publication Date: 2026-07-31STATE GRID HUNAN ELECTRIC POWER CO +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID HUNAN ELECTRIC POWER CO
Filing Date
2026-05-07
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing video surveillance systems are unable to meet the needs of refined and cognitive management at construction sites. They have high false alarm rates, lack business context awareness, have rigid logic, lack continuous learning and adaptive capabilities, have low PTZ control precision, and lack closed-loop feedback mechanisms.

Method used

The construction site supervision system adopts edge-cloud collaboration and long-term memory enhancement. Through the collaborative work of the edge terminal and cloud processing terminal, combined with the multi-agent decision-making subsystem and the hybrid memory storage subsystem, it realizes real-time processing of video data and complex decision-making. It uses natural language processing and multimodal analysis for intelligent judgment, reduces false alarm rate and improves management efficiency.

Benefits of technology

It reduces false alarm rates, improves management efficiency, ensures on-site safety, lowers the barrier to entry, enables continuous system evolution and cost-effectiveness, and is suitable for large-scale deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492115A_ABST
    Figure CN122492115A_ABST
Patent Text Reader

Abstract

This invention discloses a construction site supervision system and method with edge-cloud collaboration and long-term memory enhancement; it relates to the intersection of architectural engineering and artificial intelligence; the system includes an edge terminal and a cloud processing terminal, which communicate with the edge terminal via a network; the edge terminal includes a video acquisition module, a target detection module, a sentinel management module, and a tracking control module; the cloud processing terminal includes a multi-agent decision-making subsystem and a hybrid memory storage subsystem; the multi-agent decision-making subsystem runs on a large model orchestration platform, which includes an intent routing agent, a visual intelligence agent, an on-site action agent, and a safety compliance agent; the hybrid memory storage subsystem includes: an SQL database and a vector database; this application can make accurate decisions based on comprehensive business context and visual information, distinguish between compliant and non-compliant operations, log only compliant operations without triggering alarms, and trigger alarms and specify the reasons for non-compliant operations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the intersection of architectural engineering and artificial intelligence, and in particular to a construction site supervision system and method with edge-cloud collaboration and long-term memory enhancement. Background Technology

[0002] With the advancement of smart grid construction, the requirements for safety supervision standards at power construction sites are increasingly stringent. However, most existing video surveillance systems remain at the "seeing" or simple "perception" stage, failing to meet the needs of refined and cognitive management. Existing technologies suffer from the following deep-seated pain points and defects in practical engineering applications: high false alarm rates; a lack of "pseudo-intelligence" with business context awareness, unable to combine business contexts such as construction plans, electronic work orders, personnel qualification databases, or on-site operation status to determine the legality of actions; existing monitoring systems are typically based on predefined rule engines, resulting in rigid logic and an inability to understand complex on-site semantics and dynamic risks; a lack of continuous learning and adaptive capabilities (long-term memory loss), preventing users from correcting the system's judgment logic through simple feedback mechanisms; and low PTZ control precision and a lack of closed-loop feedback mechanisms. Summary of the Invention

[0003] To construct a digital supervisor with perception, cognition, decision-making, execution, and evolution capabilities, this application provides a construction site supervision system and method with edge-cloud collaboration and long-term memory enhancement.

[0004] Firstly, this application provides a construction site supervision system with edge-cloud collaboration and long-term memory enhancement, employing the following technical solution:

[0005] A construction site supervision system with edge-cloud collaboration and long-term memory enhancement includes an edge terminal deployed on-site and a cloud processing terminal deployed on a cloud server, with the cloud processing terminal and the edge terminal communicating via a network.

[0006] The edge device includes a video acquisition module, a target detection module, a sentry management module, and a tracking control module: the video acquisition module is used to acquire video data streams and transmit them to the target detection module; the target detection module is used to identify targets and output bounding boxes to the sentry management module; the sentry management module is used to calculate the video data streams and determine whether any target bounding boxes have entered the restricted area;

[0007] The cloud processing unit includes a multi-agent decision-making subsystem and a hybrid memory storage subsystem;

[0008] The multi-agent decision-making subsystem runs on a large model orchestration platform and includes an intent routing agent, a visual intelligence agent, a field action agent, and a safety and compliance agent. The intent routing agent serves as the user interaction entry point, identifying intents and distributing them to other agents. The visual intelligence agent calls the multimodal model to analyze field snapshots and answer complex questions about image details. The field action agent is responsible for translating abstract natural language instructions into specific hardware control parameter sequences. The safety and compliance agent is responsible for querying structured databases, comparing regulations with work plans, managing long-term memory, and generating compliance reports.

[0009] The hybrid memory storage subsystem includes:

[0010] SQL database: Stores job plan tables, original intrusion event tables, and analysis result record tables.

[0011] Vector database: Stores expert experience memos and safe operating procedures.

[0012] Optionally, the video acquisition module connects the fixed camera and the PTZ camera via the RTSP / ONVIF protocol to acquire real-time video data streams and maintain a circular frame buffer, supporting the playback of footage before the event occurred.

[0013] Optionally, the target detection module runs a lightweight deep learning model for target detection that has been pruned and quantized to detect targets such as people, safety helmets, reflective vests, engineering vehicles, and fireworks in real time, and outputs bounding boxes and category confidence scores.

[0014] Optionally, the sentry management module provides a visual area drawing tool that performs geometric calculations on the fixed camera footage based on preset electronic fence coordinates to determine whether anyone has entered the restricted area.

[0015] Optionally, the tracking control module implements a visual feedback-based PID servo control logic for the gimbal, which is responsible for calculating the target centroid offset and controlling the gimbal movement through the PTZ protocol to keep the target in the center of the screen.

[0016] Secondly, this application provides a construction site supervision method based on edge-cloud collaboration and long-term memory enhancement, employing the following technical solution:

[0017] A construction site supervision method combining edge-cloud collaboration and long-term memory enhancement includes the following:

[0018] Plan entry stage: User inputs plan;

[0019] The security and compliance intelligent agent extracts key entities through natural language processing and transforms unstructured speech into structured data;

[0020] Write the structured data into the work plan table of the SQL database and create a business whitelist;

[0021] Abnormal triggering phase:

[0022] The sentry management module at the edge analyzes the video stream from the fixed camera in real time to detect whether anyone has entered the area;

[0023] The system immediately captures a snapshot of the scene and records the timestamp at that moment. Then, it proactively calls back to the cloud-based decision-making subsystem via API to upload the event metadata.

[0024] Intelligent analysis phase:

[0025] Upon receiving the callback, the security compliance agent first queries the work plan table based on the region and the current alarm time.

[0026] If the query result is empty, the intelligent agent directly determines it as "illegal intrusion", triggers a high-level alarm, and pushes a notification to the security personnel;

[0027] If a work plan is found, the agent calls the visual intelligence agent, passing in the on-site snapshot and a description of the work plan;

[0028] Visual intelligence agents utilize multimodal capabilities to analyze snapshots and conduct compliance checks;

[0029] Final ruling: If the personnel characteristics meet the safety regulations and the operation is within the planned time, the system will determine it as compliant operation, record the information in the log, and not trigger an alarm; otherwise, an alarm for non-compliant operation will be triggered, and the reason for the violation will be noted.

[0030] In summary, this application includes the following beneficial technical effects:

[0031] 1. Reduce false alarm rate and improve management efficiency: By introducing a work plan database and long-term memory memo, the system can understand business logic and special rules, effectively filtering out non-compliant "abnormal" events (such as compliant maintenance work), allowing managers to focus on real risks;

[0032] 2. High real-time performance, ensuring on-site safety: By pushing computationally intensive video stream analysis, target detection, and motion control to the edge, millisecond-level response speeds are ensured, enabling reliable operation even in construction sites with unstable networks, without missing any momentary risks;

[0033] 3. Natural interaction, lowering the barrier to entry: Through multi-agent orchestration, users can use natural language to complete all operations from viewing to configuration, greatly reducing the barrier to entry and enabling non-technical personnel to easily manage complex security systems;

[0034] 4. Continuous Evolution, Getting Smarter with Use: The system has a learning ability similar to that of a human apprentice. Through daily use and feedback, the system's judgment standards will become more and more in line with the actual needs on site. As the usage time increases, the system becomes smarter with use, and the maintenance cost is significantly reduced.

[0035] 5. High cost-effectiveness and easy to promote: Software algorithms make up for the shortcomings of hardware capabilities (such as using visual closed loop to replace expensive closed loop gimbals), and the API call cost is greatly reduced by "on-demand calling" large models, making it suitable for large-scale deployment. Attached Figure Description

[0036] Figure 1 This is a schematic diagram of the system architecture of this application;

[0037] Figure 2 This is a flowchart of the methodology evaluation process for this application;

[0038] Figure 3 This is the flowchart of the long-term memory adaptive learning process in this application. Detailed Implementation

[0039] The following is in conjunction with the appendix Figure 1-3 This application will be described in further detail.

[0040] This application discloses a construction site supervision system with edge-cloud collaboration and long-term memory enhancement, including an edge terminal deployed on site and a cloud processing terminal deployed on a cloud server, wherein the cloud processing terminal and the edge terminal communicate through a network.

[0041] The edge device includes a video acquisition module, a target detection module, a sentry management module, and a tracking and control module.

[0042] The video acquisition module connects the fixed camera and the pan-tilt camera via the RTSP / ONVIF protocol to acquire real-time video data streams and transmit them to the target detection module. Simultaneously, it maintains a circular frame buffer; this buffer is an efficient data storage mechanism that allocates a fixed-size area in memory to cyclically store video frame data from the most recent period. When the buffer is full, new video frames automatically overwrite the oldest, thus continuously updating and retaining the latest historical footage within limited storage resources.

[0043] The object detection module runs a lightweight object detection deep learning model that has undergone pruning and quantization to detect targets such as people, safety helmets, reflective vests, construction vehicles, and fireworks in real time, outputting bounding boxes and class confidence scores; it also outputs the bounding boxes of the targets to the sentinel management module. The lightweight object detection deep learning model refers to a neural network architecture designed with computational efficiency and model size in mind, such as the MobileNet series, YOLO-Tiny versions, or EfficientDet-Lite; it employs L1 / L2 norm pruning to identify and remove connections with small weights; or it uses structured pruning to directly remove entire convolutional kernels or layers that have little impact on model performance. Simultaneously, the "quantization" technique converts the model's parameters and activation values ​​from high-precision floating-point numbers (such as FP32) to low-precision fixed-point numbers (such as INT8), further reducing the model's storage requirements and computational load, while accelerating the inference process.

[0044] The sentry management module provides a visual area drawing tool that performs geometric calculations on the fixed camera footage based on preset electronic fence coordinates to determine whether anyone has entered the restricted area.

[0045] The Sentinel Management module is a core component deployed at the edge. It is responsible for real-time monitoring of the video stream and identifying potential intrusion behaviors based on the output of the target detection module. This module can be a standalone software service running on an embedded system or industrial PC at the edge, receiving the target detection module's output via an API interface and interacting with the tracking and control module. Alternatively, this module can be integrated into the same process as the target detection module as a sub-functional unit, sharing computing resources to improve processing efficiency and reduce communication latency.

[0046] For example, the sentinel management module can preset multiple geometric regions as restricted areas. When the bounding box output by the target detection module geometrically overlaps with these preset restricted areas, the sentinel management module determines that a target has entered the restricted area.

[0047] The tracking control module implements visual feedback-based PID servo control logic for the gimbal, which is responsible for calculating the target's centroid offset and controlling the gimbal movement through the PTZ protocol to keep the target in the center of the screen.

[0048] For example, the tracking control module can receive target position information from the sentry management module and generate directional commands based on the target's relative position in the frame. These commands are then sent to a motion-capable camera device to adjust its viewing angle to keep the target within its field of view. Alternatively, the tracking control module can calculate the target's motion trend based on its displacement across consecutive frames. Based on the calculation results, the module can output adjustment signals, such as adjusting the camera's horizontal or vertical angle, to attempt to follow the target.

[0049] The cloud processing unit includes a multi-agent decision-making subsystem and a hybrid memory storage subsystem.

[0050] The multi-agent decision-making subsystem runs on a large model orchestration platform and includes an intent routing agent, a visual intelligence agent, a field action agent, and a security compliance agent.

[0051] The intent routing agent, serving as the user interaction entry point, is configured to identify user intents and distribute them to other agents. For example, the intent routing agent can receive text or voice commands input by the user and identify the user's basic intent through keyword matching or rule parsing. The identified intent is then mapped to a pre-defined agent function, and the request is forwarded to the appropriate agent for processing.

[0052] The visual intelligence agent is configured to invoke a multimodal model to analyze a snapshot of the scene in order to answer complex questions about the details of the image. For example, the visual intelligence agent can invoke a pre-trained image recognition model to classify and identify objects and scenes in the snapshot. By combining and reasoning about the recognition results, the agent can answer simple questions about the visible elements in the image.

[0053] The field agent is configured to translate abstract natural language instructions into specific sequences of hardware control parameters. For example, the field agent can maintain a mapping table between instructions and parameters. When it receives a natural language instruction such as "move camera," the agent queries this table and translates it into predefined hardware control commands and parameter values.

[0054] The safety compliance agent is configured to query structured databases, compare regulations with work plans, manage long-term memory, and generate compliance reports. For example, the agent can execute SQL queries to retrieve work plans and regulatory provisions from the structured database. Through string matching or logical rules, the agent compares the on-site situation with this information to determine if any non-compliant behavior exists and generates a text-based report. Regarding long-term memory management, the agent can be responsible for writing new event data or analysis results to designated storage areas.

[0055] The hybrid memory storage subsystem comprises an SQL database and a vector database. The SQL database is configured to store job schedule tables, raw intrusion event tables, and analysis result records. For example, the SQL database can employ a common relational database management system, storing this structured data through predefined table structures and supporting transaction processing and complex queries. The vector database is configured to store expert experience memos and security operating procedures. For instance, unstructured text content such as expert experience memos and security operating procedures is processed using an embedding model and transformed into high-dimensional vectors, then stored in the vector database for semantic similarity searching.

[0056] Through an edge-cloud collaborative architecture, real-time processing of construction site data and effective complex decision-making are combined, thereby reducing the false alarm rate. The multi-agent decision-making subsystem can understand on-site semantics and dynamic risks. Combined with the long-term memory capabilities provided by the hybrid memory storage subsystem, the system can continuously learn and adaptively correct its judgment logic, thus improving the accuracy and flexibility of decision-making. In addition, the collaborative work between the tracking and control module and the on-site action agents improves the tracking accuracy of on-site targets, comprehensively optimizing the intelligent supervision efficiency of the construction site.

[0057] The monitoring method of the system of this application will be described below through three embodiments;

[0058] Example 1: Intelligent assessment process for operational compliance

[0059] This embodiment demonstrates how the system avoids false alarms and achieves a closed-loop business process by integrating business data and visual data.

[0060] Planned data entry phase:

[0061] The user inputs via voice: "Tomorrow afternoon from 2 pm to 4 pm, Engineer Zhang will be conducting transformer maintenance in Area A."

[0062] The security and compliance intelligent agent extracts key entities through natural language processing and transforms unstructured speech into structured data: {Start_Time: "202X-XX-XX 14:00:00", End_Time: "16:00:00", Zone_ID: "Zone_A", Personnel: "Zhang Gong", Task_Type: "Inspection"}.

[0063] The system writes the structured data into the work plan table of the SQL database and establishes a business whitelist.

[0064] Abnormal triggering phase:

[0065] The sentry management module at the edge analyzes the video stream from the fixed camera in real time and detects that someone has entered Zone_A.

[0066] The system immediately captures a snapshot of the scene, generates a unique event_id, records the timestamp at that moment, and then proactively calls back to the cloud-based decision-making subsystem via API to upload the event metadata.

[0067] Intelligent analysis phase:

[0068] Upon receiving the callback, the security compliance agent first queries the work plan table based on area A and the current alarm time.

[0069] Scenario A (Unplanned): If the query result is empty, the agent directly determines it as "illegal intrusion", triggers a high-level alarm, and pushes a notification to the security personnel.

[0070] Scenario B (with a plan): If a work plan is found, the agent calls the visual intelligence agent, passing in the on-site snapshot and the work plan description.

[0071] The visual intelligence agent uses multimodal capabilities to analyze snapshots, focusing on checking: Are personnel wearing the prescribed work clothes? Are they wearing safety helmets? Are there any violations such as smoking or using mobile phones? Is it consistent with the planned task description?

[0072] Final ruling: If the personnel characteristics meet the safety regulations and are within the planned time, the system will determine it as "compliant operation", only log the information, and will not trigger an alarm; otherwise, it will trigger an "illegal operation" alarm and indicate the reason for the violation, such as "planned but not wearing a safety helmet".

[0073] Example 2: Adaptive Learning Process Based on Long-Term Memory

[0074] This embodiment demonstrates how the system evolves through a "memory" mechanism to solve the problem of "false alarms in specific scenarios".

[0075] Scenario trigger: The system previously issued a violation alert to a manager wearing a yellow vest for "not wearing standard blue work clothes".

[0076] Human feedback: The user enters a correction command on the interactive interface: "Note that the person wearing the yellow vest is the on-site safety officer, who is a compliance personnel. Do not call the police in the future."

[0077] Knowledge Extraction:

[0078] The intent routing agent recognizes that this is a "teaching / correction" intent and forwards the instruction to the security compliance agent.

[0079] Intelligent agent extracts rule features: {Object features: "yellow vest", Scenario: "universally applicable", Judgment: "compliant / exempt"}.

[0080] The system converts the rule into a natural language description; for example, "Rule #101: When judging a dress code violation, if a person is identified as wearing a yellow vest, they are considered a compliant security manager and are exempted," and calls the Embedding model to generate a vector, which is then stored in the "Expert Experience Base."

[0081] Rule application:

[0082] In the next similar intrusion or clothing detection incident, the security compliance agent is configured to force a search of the knowledge base.

[0083] The RAG engine calculates vector similarity based on the description of the current screen, such as "A person wearing a yellow vest has been detected entering...", and retrieves the aforementioned memo.

[0084] The large model receives the memo information in the "context" part of the prompt word, and accordingly overrides the default judgment logic to output the "compliant" conclusion, thereby avoiding repeated false alarms and realizing the self-evolution of the system.

[0085] Example 3: Binocular Coordination and Visual Guidance Tracking

[0086] This example demonstrates how to achieve high-precision tracking using low-cost hardware to solve the problem of target loss.

[0087] Wide-area detection: Fixed cameras provide real-time panoramic monitoring. When the edge detection algorithm detects a target (such as a "person") entering the detection zone, it calculates the normalized center point (x_fixed, y_fixed) of the target in the fixed-view coordinate system.

[0088] Strategy guidance: The edge device searches a preset mapping table, which records the correspondence between fixed screen areas and preset points of the PTZ, and finds the preset point ID of the PTZ camera corresponding to the warning area.

[0089] Rapid response: The PTZ camera receives the command and turns to the preset point at maximum speed to quickly bring the target into the field of view.

[0090] Visual locking and correction:

[0091] Once the gimbal is in position, the edge device uses a target detection algorithm to search for the target again in the current field of view. To prevent target confusion, the system can further extract the target's Re-ID (person re-identification) feature vector for comparison.

[0092] Closed-loop control: If the target is at the edge of the screen, the system calculates the deviation vector (dx, dy) between the target center and the screen center.

[0093] The system dynamically adjusts the rotation speed of the gimbal based on the magnitude of the deviation using a proportional control algorithm: the larger the deviation, the faster the speed, and the smaller the deviation, the slower the speed.

[0094] In each frame processing, the system continuously adjusts the gimbal speed through visual feedback, so that the target quickly returns to the center of the image and remains locked, and can closely follow the target even if it moves irregularly.

[0095] Multimodal confirmation: After tracking stabilizes (e.g., the target is located in the central region for 10 consecutive frames), the system automatically captures close-up images and sends them to the cloud-based large model to confirm the target's detailed identity or behavioral status (e.g., "whether it is smoking" or "whether it has operated a switch").

[0096] Through the above technical solution, this application provides a construction site supervision method with edge-cloud collaboration and long-term memory enhancement, which effectively solves the problem of high false alarm rate caused by lack of business context awareness in the existing system.

[0097] During the plan entry phase, the security and compliance intelligent agent uses natural language processing technology to transform the unstructured plans input by users into structured data and store them in the work plan table of the SQL database, establishing a reliable business whitelist and providing accurate contextual basis for subsequent intelligent analysis.

[0098] During the anomaly triggering phase, the sentinel management module at the edge can detect intrusion events in the area in real time, capture snapshots of the scene and record timestamps in a timely manner, and actively call back to the cloud decision-making subsystem via API to ensure the rapid and complete uploading of event information.

[0099] During the intelligent analysis phase, the security and compliance intelligence agent first queries the work plan table based on the event area and time, quickly filtering out unplanned "illegal intrusions" and triggering high-level alarms, thereby avoiding delayed responses to non-compliant behaviors. For events with work plans, the system further invokes the visual intelligence intelligence agent, utilizing its multimodal capabilities to conduct in-depth analysis of on-site snapshots and work plan descriptions, performing refined compliance checks, such as verifying whether personnel are wearing safety equipment and whether work behaviors comply with regulations.

[0100] Ultimately, the system can make accurate judgments based on comprehensive business context and visual information, distinguishing between compliant and non-compliant operations. It logs compliant operations without triggering alarms, while triggering alarms and specifying the reasons for non-compliant operations. This intelligent judgment mechanism, which combines business context, significantly reduces the false alarm rate and improves the intelligence level and management efficiency of the supervision system.

[0101] Meanwhile, this method works in conjunction with the video acquisition module, target detection module, sentinel management module, and tracking control module in the aforementioned system, enabling the edge to perform preliminary detection and data acquisition efficiently, while the cloud focuses on complex intelligent analysis and decision-making. This achieves complementary advantages of edge-cloud collaboration and further enhances the overall performance and accuracy of the system.

[0102] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. A construction site supervision system with edge-cloud collaboration and long-term memory enhancement, characterized in that: This includes edge devices deployed on-site and cloud processing devices deployed on cloud servers, with the cloud processing devices and edge devices communicating via a network; The edge device includes a video acquisition module, a target detection module, a sentry management module, and a tracking control module: the video acquisition module is used to acquire video data streams and transmit them to the target detection module; the target detection module is used to identify targets and output bounding boxes to the sentry management module; the sentry management module is used to calculate the video data streams and determine whether any target bounding boxes have entered the restricted area; The cloud processing unit includes a multi-agent decision-making subsystem and a hybrid memory storage subsystem; The multi-agent decision-making subsystem runs on a large model orchestration platform and includes an intent routing agent, a visual intelligence agent, a field action agent, and a safety and compliance agent. The intent routing agent serves as the user interaction entry point, identifying intents and distributing them to other agents. The visual intelligence agent calls the multimodal model to analyze field snapshots and answer questions about details in the images. The field action agent is responsible for converting abstract natural language instructions into specific sequences of hardware control parameters. The safety and compliance agent is responsible for querying structured databases, comparing regulations with work plans, managing long-term memory, and generating compliance reports. The hybrid memory storage subsystem includes: SQL database: Stores job plan tables, original intrusion event tables, and analysis result record tables; Vector database: Stores expert experience memos and safe operating procedures.

2. The construction site supervision system with edge-cloud collaboration and long-term memory enhancement as described in claim 1, characterized in that: The video acquisition module connects the fixed camera and the pan-tilt camera via the RTSP / ONVIF protocol to acquire real-time video data streams and maintain a circular frame buffer, supporting the playback of footage before an event occurs.

3. The construction site supervision system with edge-cloud collaboration and long-term memory enhancement according to claim 2, characterized in that: The target detection module runs a lightweight deep learning model for target detection that has been pruned and quantized, and detects people, safety helmets, reflective vests, engineering vehicles, and fireworks targets in real time, outputting bounding boxes and category confidence scores.

4. The construction site supervision system with edge-cloud collaboration and long-term memory enhancement as described in claim 3, characterized in that: The sentry management module provides a visual area drawing tool that performs geometric calculations on the fixed camera footage based on preset electronic fence coordinates to determine whether anyone has entered the restricted area.

5. The construction site supervision system with edge-cloud collaboration and long-term memory enhancement according to claim 4, characterized in that: The tracking control module implements visual feedback-based PID servo control logic for the gimbal, which is responsible for calculating the target's centroid offset and controlling the gimbal movement through the PTZ protocol to keep the target in the center of the screen.

6. A construction site supervision method with edge-cloud collaboration and long-term memory enhancement, employing the construction site supervision system with edge-cloud collaboration and long-term memory enhancement as described in claim 5, characterized in that: Includes the following: Plan entry stage: User inputs plan; The security and compliance intelligent agent extracts key entities through natural language processing and transforms unstructured speech into structured data; Write the structured data into the work plan table of the SQL database and create a business whitelist; Abnormal triggering phase: The sentry management module at the edge analyzes the video stream from the fixed camera in real time to detect whether anyone has entered the area; The system immediately captures a snapshot of the scene and records the timestamp at that moment. Then, it proactively calls back to the cloud-based decision-making subsystem via API to upload the event metadata. Intelligent analysis phase: Upon receiving the callback, the security compliance agent first queries the work plan table based on the region and the current alarm time. If the query result is empty, the intelligent agent directly determines it as an unauthorized intrusion, triggers a high-level alarm, and pushes a notification to the security personnel; If a work plan is found, the agent calls the visual intelligence agent, passing in the on-site snapshot and a description of the work plan; Visual intelligence agents utilize multimodal capabilities to analyze snapshots and conduct compliance checks; Final ruling: If the personnel characteristics meet the safety regulations and the operation is within the planned time, the system will determine it as compliant operation, record the information in the log, and not trigger an alarm; otherwise, an alarm for non-compliant operation will be triggered, and the reason for the violation will be noted.