Construction operation safety intelligent supervision system and method based on AI identification algorithm

The construction operation safety intelligent supervision system based on AI recognition algorithms, combined with computer vision and deep learning technologies, solves the problems of weak recognition capabilities and insufficient edge computing in existing intelligent monitoring systems for construction operation safety supervision. It realizes multi-target real-time intelligent analysis and efficient supervision, and improves the intelligence and adaptability of the system.

CN120913159AActive Publication Date: 2025-11-07CHINA TOWER CO LTD XIANGTAN BRANCH +1

Patent Information

Application Number
CN202511450925.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2025-11-07
Estimated Expiration
2045-10-11

AI Technical Summary

Technical Problem

Existing intelligent monitoring systems for construction safety supervision suffer from several problems, including weak identification capabilities, lack of real-time and intelligent image recognition technology, insufficient integration of edge computing and the Internet of Things, susceptibility of worker identification to environmental interference, weak adaptability of behavior supervision rules to different scenarios, and difficulty in focusing on risk events.

Method used

The construction operation safety intelligent supervision system adopts AI recognition algorithm, which combines computer vision and deep learning technology to achieve multi-objective real-time intelligent analysis. It optimizes the data processing process by integrating edge computing and IoT, improves the success rate of identity recognition by cross-authentication, and enhances the adaptability of supervision through risk prediction and rule optimization.

Benefits of technology

It enables real-time intelligent analysis of multiple features of operators and equipment, reduces reliance on manual labor, enhances the system's intelligent recognition capabilities and adaptability, meets the in-depth application needs of different industry business scenarios, reduces data transmission pressure and costs, and improves supervision efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913159A_ABST
    Figure CN120913159A_ABST
Patent Text Reader

Abstract

The invention discloses a construction operation safety intelligent supervision system and method based on an AI identification algorithm, and relates to the technical field of operation supervision. The method comprises the following steps: acquiring a defect image sample and a non-defect image sample, performing defect classification and target feature library training, acquiring an original image, performing image pre-processing, feature library comparison and processed image output, framing a detected target by using an outer line rectangle, and obtaining a non-defect image sample; the verification terminal and a work card of an operator are networked and identity information is verified, the camera analyzes the identity information of the operator based on AI identification, and the verification terminal and the camera perform cross authentication to confirm the operator. Multi-target real-time intelligent analysis and identification of a non-digital display vernier caliper, a grounding resistance megger, a test pen, an anti-falling device and a safety cone barrel are realized, the requirements of three-dimensional supervision, operation and service in the industrial field are met, and the intelligent identification capability of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of construction supervision, in particular to a construction operation safety intelligent supervision system and method based on an AI recognition algorithm. BACKGROUND

[0002] With the continuous progress of image and video processing, analysis, transmission technology and artificial intelligence technology, the monitoring system develops from pure analog system to analog-digital combination, pure IP monitoring mode, and constantly moves towards high definition and networking. The intelligent analysis demand of the monitoring system also emerges as the times require, that is, intelligence. The demand of the intelligent recognition market comes from various actual needs of specific industry characteristic monitoring. Each industry monitoring may need several kinds of intelligent monitoring technology. The personalized demand of the subdivided market determines that we must pay close attention to the particularity of different systems, take marketization as the orientation of the whole company, and closely cooperate and cooperate between the sales, market, technology, research and development, product, etc. departments to provide customized products and services that meet the unique needs of each industry customer, provide full high-definition video stream and intelligent intelligent recognition system and platform.

[0003] The existing intelligent monitoring system mainly has the following two technical schemes: 1. Target extraction and event analysis technology based on rules. This scheme extracts and detects the target in the video picture through picture segmentation, foreground extraction and other methods, and then distinguishes different events based on preset rules to realize judgment and alarm linkage. Its principle is to separate the foreground and the background through image segmentation algorithm, the foreground is the moving target, and specific rules such as "crossing" and "area intrusion" are set as the judgment conditions. When the target triggers the rules, the alarm is triggered. It can be applied to early behavior analysis or traffic event detection. This scheme depends on fixed rules, is effective for single event recognition of simple scene, but has poor flexibility and is difficult to cope with complex environment or multi-feature target analysis. 2. Specific object detection technology based on pattern recognition. This scheme uses pattern recognition technology to model specific objects in the picture, trains the model through a large number of samples, and realizes the detection and recognition of specific objects. Its principle is to construct a feature model of a specific object based on machine learning, train the model through a large number of labeled samples, so that the model has the ability to locate and recognize the target from the video or image. It can be applied to personnel detection and specific object defect recognition. This scheme can realize high-precision recognition of specific objects, but relies on a large number of samples for training, has limited generalization ability, and focuses on single object recognition, lacking comprehensive analysis ability for multiple targets and multiple features.

[0004] The prior art scheme has the following disadvantages: 1. mainly using a camera as a collection source, lacking real-time intelligent analysis capability for target features, and lacking real-time and intelligent recognition capability; 2. image recognition technology is usually based on basic management applications, lacking application technology for different industry business scenarios; 3. most of the snake-shaped head edge computing capabilities are simple and single, and need to transmit video streams to a cloud server for centralized processing through a network, lacking division and optimization of data transmission and calculation; 4. in construction operation safety supervision, the identity of the operation personnel is easily disturbed by the environment, single authentication is unreliable, the behavior supervision rule adaptation scene capability is weak and the illegal response is lagging, and the risk events are difficult to focus, and the supervision rule is difficult to dynamically optimize. Therefore, in view of the above defects, designing a construction operation safety intelligent supervision system and application method based on an AI recognition algorithm is an urgent problem to be solved by the person skilled in the art. SUMMARY

[0005] In view of the problems of weak intelligent recognition capability, low image detection application level, insufficient fusion of edge computing and Internet of Things in the prior art, the present application aims to provide a construction operation safety intelligent supervision system and method based on an AI recognition algorithm, and the specific purposes include breaking through the limitation of single intelligent recognition type, realizing real-time intelligent analysis of multiple features of people and objects, reducing the dependence on artificial labor, improving the professionalization and scale of image detection, strengthening the deep application for different industry business scenarios, fully tapping the potential of image resources, fusing edge computing and Internet of Things technology, optimizing data processing flow, reducing transmission pressure, cost and bandwidth demand, and improving the overall efficiency of the system.

[0006] To achieve the above purposes, the present application is implemented by the following technical scheme: In a first aspect, the present application provides a construction operation safety intelligent supervision method based on an AI recognition algorithm. In a complex scene, target automatic extraction and recognition are particularly important when multiple targets need to be processed in real time. Based on the application of computer vision principles, computer image processing technology is used to track the target in real time, combining high-definition video acquisition, image processing and deep learning to realize the detection and recognition of general targets and customized targets in complex environments. The method comprises the following steps: Step 1: realized based on a feature training sub-module, the feature training sub-module is integrated in a cloud server, the module acquires defect image samples and non-defect image samples, and is responsible for defect classification and target feature library training; Step 2: Based on the real-time analysis sub-module implementation, the feature training sub-module is integrated in the edge computing server. The module obtains the original image, performs image preprocessing, compares the feature library, i.e., the deep learning target detection algorithm, outputs the processed image, and frames the detected target with an external rectangular box. Users can annotate the reference size of the original image. According to the image gray point calculation, the position and size information of the external rectangular box are calculated and displayed in the processed image. Step 3: The verification terminal and the worker's card are networked and the identity information is verified. The camera recognizes and analyzes the identity information of the worker based on AI, and the verification terminal and the camera cross-authenticate to confirm the worker; Step 4: The supervisor sets the total rule, the cloud server automatically generates sub-rules based on the processed image and sends them to the edge computing server for execution. The worker triggers the total rule or sub-rule to push the alarm and record the event to the client for the supervisor to check; Step 5: The edge computing server predicts the risk of the worker's behavior characteristics based on the total rule and sub-rule.

[0007] Further, the defect classification and target feature library training specifically includes the following steps: Step 101: The cloud server uses AI intelligent algorithms in the AI algorithm capability core engine to classify defect targets for defect-free image samples and defect image samples. After classification, the next step is executed; Step 102: The cloud server uses AI models in the AI algorithm capability core engine to train the curve classification target to obtain a defect classification target feature library. Save the defect classification target feature library for subsequent repeated retrieval. Execute the following step; Step 103: Transfer the saved defect classification target feature library to the edge computing server to perform feature library comparison in step 203. Repeat step 103 until the cloud server obtains new defect image samples and non-defect image samples.

[0008] Further, the local area network camera collects original images directly or indirectly transmitted to the edge computing server. When directly transmitted, the remote device is directly transmitted to the edge computing server through the remote network. The remote device is the camera. When indirectly transmitted, the storage and calculation separation technology is adopted, i.e., data storage and image processing are performed by different devices. The original image obtained by the camera is saved in the network attached storage. The hard disk file system built-in the network attached storage transmits the original image to the edge computing server through file IO. The network attached storage is the storage end, and the edge computing server is the computing end. The system IO built-in the verification terminal transmits the verification data to the edge computing server through the system serial port or other interfaces. The verification data is also obtained by the image acquisition module of the edge computing server. The verification data and the original image are bound together for subsequent processing. The processing of the original image specifically includes the following steps: Step 201: The image acquisition module of the real-time analysis submodule obtains the original image from the remote device, the hard disk file system and the system IO respectively, and transmits the original image to the image pre-processing module for the next step; Step 202: The image pre-processing module performs noise reduction and white balance processing on the original image to obtain a pre-processed image, and executes the next step; Step 203a: Compare the pre-processed image with the defect classification target feature library output by the feature training submodule to obtain target recognition, and frame the detected target with an outer rectangle; Step 203b: Perform reference size labeling on the pre-processed image to obtain target size labeling; Step 203c: Perform clarity judgment on the pre-processed image, and output image clarity evaluation; Among them, steps 203a, 203b and 203c are executed in parallel, and after execution is completed, they are merged into a processed image and output uniformly, and the next step is executed; Step 204: The real-time analysis submodule outputs the processed image, and sequentially labels the image clarity, defect target size and defect target area, and jumps to step 201 for repeated execution.

[0009] Further, the cross-authentication confirmation of the verification terminal and the camera to the job worker specifically includes the following steps: The pre-deployment process is as follows: Step 301: Provide each job worker with a work card with a built-in Bluetooth module, and record the work card ID in the Bluetooth module. The work card ID is encrypted using an RSA key. The verification terminal is built-in with an RSA key for decryption. The identity information of each job worker is pre-set in the edge computing server and a mapping table of identity information and work card ID is established. The next step is executed; Step 302: Configure the verification terminal as a Bluetooth Mesh gateway and complete the initialization. Connect each job worker's work card Bluetooth module to the verification terminal to complete the pre-deployment process. The Bluetooth modules between the work cards can be used as relays to transmit encrypted data to the verification terminal. Each camera is also configured with the same Bluetooth module as the work card to participate in networking. The Bluetooth module configured in the camera records the device number of each camera, which is used by the verification terminal to determine the specific camera; When the job worker appears in the camera monitoring screen, the verification operation starts, and the specific steps are as follows: Step 303: The camera collects the original image and transmits it to the edge computing server. The edge computing server processes the original image and outputs a processed image, and compares it with the built-in identity information, including the name, personal number and portrait of the worker. The edge computing server uses MSE (Mean Square Error) and PSNR (Peak Signal to Noise Ratio) to calculate the similarity between the processed image and the personal portrait. The edge computing server outputs the identity information with the highest similarity, and proceeds to the next step. Step 304: The verification terminal obtains the device number from the Bluetooth module of the camera that triggered the verification operation. It also obtains the encrypted data of the Bluetooth module with the strongest connection signal near the camera. The worker who triggered the verification is closest to the camera, and the Bluetooth connection signal is the strongest. The worker uses the built-in RSA key to decrypt the encrypted data to obtain the ID of the worker card, and proceeds to the next step. Step 305: The verification terminal transmits the ID of the worker card to the edge computing server. The edge computing server finds the corresponding identity information from the mapping table based on the ID of the worker card, and proceeds to the next step. Step 306: The edge computing server compares the identity information obtained in steps 303 and 305. If they match, the worker's identity is confirmed. The edge computing server uploads the matching result to the cloud server. If they do not match, the edge computing server retrieves the historical records of successful matching of identity information for the corresponding verification terminal and the corresponding camera from the cloud server, and outputs the corresponding identity information based on the number of successful times.

[0010] Further, the determination of the total rule and the sub-rule specifically includes the following steps: Step 401: The supervisor logs in to the system through the client and enters the camera management interface. The supervisor sets a special tag for each camera according to the monitoring area and the type of work. The edge computing server classifies all cameras based on the tags, and proceeds to the next step. Step 402: The supervisor sets the total rule on the client. The total rule is uploaded to the cloud server through the client, and then synchronized to the edge computing server by the cloud server, and proceeds to the next step. Step 403: The edge computing server uploads the processed image of each camera to the cloud server. The cloud server analyzes the behavior characteristics of the workers through the AI intelligent algorithm of the AI algorithm capability core engine. The behavior characteristics include the actions and operation processes of the workers, and proceeds to the next step. Step 404: The cloud server calls the AI model to analyze the behavior characteristics of the workers, and automatically generates sub-rules. The sub-rules are further constraint conditions under the total rule. The cloud server transmits the sub-rules to the edge computing server through the distribution device for saving and enabling, and proceeds to the next step. Step 405: When the original image captured by the camera appears a picture change, the edge computing server extracts the processed image from the original image in real time and matches it with the enabled total rules and sub-rules. If the processed image triggers the total rules or sub-rules, proceed to step 406. If the total rules or sub-rules are not triggered, do not take any operation, and repeat step 405. Step 406: When the total rules are triggered, the edge computing server automatically generates alarm information and pushes it to the client through the distribution device for real-time viewing by the supervisor. When the sub-rules are triggered, automatically generate a record event and push it to the client in chronological order for subsequent browsing by the supervisor. Go to step 405.

[0011] Further, the risk prediction specifically includes the following steps: Step 501: The edge computing server calculates the risk value of each camera monitoring area based on the AI intelligent algorithm of the AI algorithm capability core engine according to the enabled sub-rules and the real-time collected worker behavior data, generates a risk prediction event, and proceeds to the next step. Step 502: The edge computing server automatically classifies and sorts all cameras according to the risk level of the risk prediction event, which is divided into high, medium, and low. The prediction event with a high risk level is pushed to the client of the supervisor through the distribution device along with the corresponding camera screenshot. The client interface displays high-risk events first, making it easy for supervisors to understand key risk points in a timely manner, and proceeds to the next step. Step 503: The supervisor views the pushed risk prediction event on the client and labels the processing result according to the actual situation. The processing result is divided into handled, unhandled, and ignored. Handled means the worker has been contacted for rectification, unhandled means it is temporarily impossible to contact for follow-up, and ignored means it is determined to be a false positive or does not need to be handled. The processing result is uploaded to the edge computing server through the client. Set a cumulative threshold number and determine whether the number of processing result uploads exceeds the cumulative threshold number. If it does not exceed, go to step 501 and repeat. If it exceeds, proceed to the next step. Step 504: The edge computing server counts the number of labels for each processing result. If the most frequent label is handled, the current settings of the sub-rule are retained. If the most frequent label is unhandled, the edge computing server will increase the risk level weight of the corresponding sub-rule and subsequently prioritize the push of this type of risk prediction event. If the most frequent label is ignored, the edge computing server uploads the ignored processing result to the cloud server and executes step 404 to regenerate the sub-rule.

[0012] In a second aspect, the present application provides a construction work safety intelligent supervision system based on an AI recognition algorithm, comprising a client, a cloud server, a distribution device, an edge computing server and a plurality of cameras, the distribution device comprising a router and a switch, the cloud server establishing bidirectional communication with the client and the distribution device through the Internet, the distribution device establishing bidirectional communication with the edge computing server and the plurality of cameras through an internal local area network, and the verification terminal establishing bidirectional communication with the distribution device through wireless connection or wired connection; The cloud server internally builds an AI intelligent recognition platform, which comprises a basic information management module, an AI video supervision center, a streaming media service platform and an AI service platform, the basic information management module comprises perception device basic information management, first log management, camera management and AI intelligent algorithm setting, the AI video supervision center comprises AI monitoring and alarming, video carousel supervision, offline monitoring and analysis and AI statistics, and the streaming media service platform comprises streaming media basic service, alarm management, camera management, subscription time correction notification, RTSP server management, video acquisition, video distribution, video storage, historical video management, audio and video codec, download management, voice broadcast and voice intercom; The AI service platform comprises an AI algorithm capability core engine and AI intelligent services, the AI algorithm capability core engine comprises an AI algorithm model optimization, marking autonomous learning engine, AI algorithm basic framework service, RTSP video acquisition management, AI model, AI intelligent algorithm and AI core algorithm power acceleration, and the AI intelligent services comprise perception device basic information management, offline data acquisition, second log management, AI recognition and retrieval service, business scenario AI service application, message notification, AI video service and AI picture service.

[0013] Further, in the basic information management module, the perception device basic information management is used for inputting and updating the basic data of the model, position and access mode of the perception device camera, ensuring device compliance access and state traceability, ensuring unified control of the system on the device, the first log management is used for recording system operation behavior, device running state and abnormal information, providing data basis for fault troubleshooting and responsibility tracing, the camera management is used for configuring the resolution, frame rate and picture angle of the camera, and the AI intelligent algorithm setting is used for adjusting the detection threshold and recognition category of the AI algorithm, adapting to different recognition scenarios of construction workers, devices and defects, and the significance is to improve the accuracy of the algorithm in a specific scene. In the AI video supervision center, the AI monitoring alarm detects irregular behavior by analyzing video streams in real time with the help of the YOLOv12 open source model and triggers an alarm. The video carousel supervision is used to cycle through the real-time images of multiple cameras. The offline monitoring analysis is used to monitor the online status of the camera and edge computing server in real time, analyze the offline reasons, ensure the continuity of data collection, avoid supervision vacuum, and the AI statistics are used to summarize the number of alarms, recognition accuracy and device online rate.

[0014] Further, in the streaming media service platform, the streaming media basic service is used to process the RTSP video transmission protocol to ensure stable transmission of video streams, connect video collection with subsequent analysis and playback links, the alarm management is used to store, classify and associate alarm information with corresponding video segments to facilitate traceability and batch processing of alarm information, the camera management and basic information management module are coordinated to supplement the camera coding format and stream transmission parameter configuration at the streaming media level to ensure compatibility of the video stream output by the camera with the streaming media system, the subscription time correction notification is used to push time correction information to each device to ensure time synchronization of the camera and edge server to avoid video data timestamp deviation, the RTSP server management is used to deploy and maintain the open source RTSP server to control the start and stop of the server, limit the number of connections, and ensure efficient transmission of video streams to the AI model and client through the RTSP protocol, the video collection is used to obtain the original video stream by calling the camera interface, perform preliminary frame extraction, and provide the original data source for AI analysis, the video distribution is used to distribute the original video stream or AI processed video stream to the client and edge computing server to meet the needs of multiple terminals to view simultaneously, reduce the pressure of repeated transmission in the cloud, the video storage is used to store video data in a network attached storage to realize long-term storage of historical video, facilitate post-accident review, the historical video management is used to provide search and playback functions for historical video, support tracing of past construction scenes, and troubleshoot potential problems, the audio and video codec is used for video compression and audio and video decoding to reduce video storage space and transmission bandwidth, and ensure data transmission efficiency, the download management is used to support users to download video segments to facilitate offline analysis and evidence preservation, the voice broadcast and voice intercom are used for real-time voice interaction between the supervision end and the construction end to timely convey safety instructions and respond to workers' inquiries to improve communication efficiency.

[0015] Further, in the AI service platform, the AI algorithm model optimization marking autonomous learning engine of the AI algorithm capability core engine re-labels the mis-identified and missed-identified samples, iteratively optimizes the AI model, continuously improves the model identification accuracy, adapts to the changes of the construction scene, the AI algorithm basic framework service provides sample preprocessing, concurrent processing, communication guarantee and video stream and picture identification, provides a basic environment for the AI algorithm operation, guarantees the efficient and stable work of the algorithm, the RTSP video acquisition management obtains the RTSP video stream and analyzes it into frame data transmission to the AI model, connects the streaming media acquisition and AI analysis, ensures the real-time data transmission, the AI model is used for multi-target detection and identification of workers, equipment and defects in the construction scene, replaces manual real-time intelligent analysis, improves the supervision efficiency, the AI intelligent algorithm extracts target features, classifies defects, improves the accuracy of target classification and defect identification, reduces misjudgment, the AI core algorithm accelerates the inference speed of the AI model, reduces the calculation delay of the edge and cloud, adapts to the ARM architecture embedded device; The perception device basic information management and the basic information management module of the AI intelligent service are coordinated, the latest state of the device is synchronized to the AI video supervision center, it is ensured that the identification strategy can be adjusted in combination with the device state during AI analysis, the state data of the device offline is acquired for collecting the device context information for the AI model, misanalysis caused by device offline is avoided, the second log management records the calling record, the identification result and the model running log of the AI service, which is convenient for AI service troubleshooting and performance optimization, the AI identification and retrieval service is used for fast retrieval of target identification results and historical identification data, improves the query efficiency of the supervision personnel, quickly locates key information, the business scene AI service application adapts the AI capability to specific construction scenes, realizes the transformation of AI technology from general identification to industry landing, solves the actual construction supervision problem, the message notification is used for pushing the alarm information and the identification result to the supervision personnel through the short message or the system message, ensures that the key information is timely reached, avoids delay disposal, the AI video service provides the enhanced video stream after AI analysis, intuitively displays the identification result, and the supervision personnel can quickly identify the problem, the AI picture service provides the labeled picture after AI processing, which is convenient for detail checking, archiving and subsequent analysis, and improves the application efficiency of the picture detection.

[0016] The application has the following beneficial effects: 1. By fusing computer vision, artificial intelligence, pattern recognition and other technologies, real-time intelligent analysis and identification of multiple targets such as non-digital vernier caliper, grounding resistance swing meter, electroscope, anti-falling device and safety cone barrel are realized, manual video content identification and judgment are replaced, industrial field stereoscopic supervision, operation and service needs are met, and the intelligent identification capability of the system is improved.

[0017] 2. Establish a customized intelligent application mechanism for industry business scenarios, support quick extraction of key information from massive videos, promote the transformation of image resources from basic general-purpose to deep application, improve system comprehensive efficiency, meet the needs of intelligent and socialized management, and improve the efficiency of image detection application.

[0018] 3. Through the integration of edge computing and the Internet of Things, the edge end of the camera is used to realize real-time data processing, reduce the data transmission amount to the cloud, reduce the network bandwidth demand and cloud computing cost, and improve the data processing speed and application efficiency, realize low-delay, low-cost and high real-time intelligent identification and management, and improve resource utilization.

[0019] 4. Cross-certification is used to improve the success rate of identity recognition, scene-based rules are added to manual adjustment to ensure regulatory adaptability, risk prediction grading is used to push and sub-rule iteration optimization is used to improve disposal efficiency, and deep adaptation to different working environments and business scenarios.

[0020] Of course, implementing any product of the present application does not necessarily require all the advantages described above. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0022] Figure 1 The block diagram of the construction work safety intelligent monitoring system based on AI recognition algorithm of the present application; Figure 2 The execution flow diagram of the feature training sub-module and the real-time analysis sub-module of the present application; Figure 3 The framework diagram of the AI intelligent recognition platform of the cloud server of the present application. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present application will be described in detail below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0024] The technical problems to be solved by the present application are as follows: 1. Pre-job safety material carrying missing problem, workers often do not carry key safety materials when entering the work area, and the lack of on-site supervision personnel leads to high missed detection rate, the traditional manual point inspection mode cannot be fully covered and real-time verification, and accident hidden dangers such as electric shock and falling are buried; 2. Safety cone barrel placement deviant problem in operation, relying on on-site supervision patrol is low in efficiency, safety cone barrels are not placed, leading to the mistake of other personnel, and safety risks are generated; 3. Precise instrument manual reading error problem, non-digital display tools rely on manual reading, for example, vernier caliper and ground resistance swing watch, there are defects of low efficiency and high error rate, efficiency is more than 30 seconds per time, error rate is more than 18%, and chain accidents such as part batch scrap and equipment grounding failure are easily caused; 4. Supervision vacuum problem caused by lack of supervision personnel, supervision of high-risk operation surface cannot be fully covered, manual patrol has blind area and time delay, leading to the fact that illegal operation cannot be intercepted in real time, for example, not using a test pen and a virtual hanging of a falling protector, and an accident prevention mechanism is just a formality; 5. Limitation problem of intelligent identification capability, existing intelligent identification is single, mainly relies on ordinary monitoring cameras, and lacks real-time intelligent analysis capability for multiple characteristics of people and objects, for example, face, biology, physics and defects, intelligent algorithms are concentrated in basic target extraction or specific object detection, and are insufficient in robustness, are limited by use environment, mainly rely on manual post-analysis, and cannot meet the needs of stereoscopic supervision, operation and service in the communication field; 6. Low efficiency problem of image detection application, intelligent identification systems are mainly based on basic management applications, lack of professional application mechanism for industry business scenes, lack of comprehensive intelligent technology support, for example, quickly positioning a key image in a large amount of video, the professionalization and scale of image detection are low, image resources only stay at a basic browsing level, potential is not fully tapped, the system comprehensive efficiency cannot be improved, and the intelligent and socialized application management needs cannot be met; 7. Insufficient fusion problem of edge computing and Internet of Things, most cameras on the market are simple and single in edge computing capability, need to transmit data to a cloud server for processing through a network, and the deep fusion of edge computing and Internet of Things technology is not realized; data rely on cloud centralized processing, leading to large transmission pressure, high delay and high cost, real-time and low-cost processing needs cannot be met, and the overall system efficiency is limited.

[0025] Embodiment one Please refer to Figures 1-3 The present application provides a technical solution: an operation safety intelligent supervision method based on an AI recognition algorithm, in a complex scene, target automatic extraction and recognition are particularly important when multiple targets need to be processed in real time, based on the application of computer vision principles, real-time tracking of targets is realized by using computer image processing technology, high-definition video acquisition, image processing and deep learning are combined, and detection and identification of general targets and customized targets in a complex environment are realized. The method comprises the following steps: Step 1: based on the feature training sub-module implementation, the feature training sub-module is integrated in the cloud server, the module obtains defect image samples and non-defect image samples from the industry general training data set, and is responsible for defect classification and target feature library training; Step 2: based on the real-time analysis sub-module, the feature training sub-module is integrated in the edge computing server, the module obtains the original image from the image acquisition module, the user can perform reference size annotation on the original image, train the model, and after image preprocessing, feature library comparison, that is, through the deep learning target detection algorithm processing, output the processed image, the detected target is boxed with an outer rectangle, according to the image gray point calculation, the position and size information of the outer rectangle are calculated, and are displayed in the processed image, according to the image definition algorithm, whether the original image is clear is automatically judged, so that the user adjusts the shooting parameters, the definition has a certain relationship with the image target detection, and the provided image meets the target detection requirement; It should be noted that: the user uses the labelImg tool to mark the position of the target to be recognized in the original image, generates a.txt annotation file corresponding to each picture, after the annotation is completed, the entire image data set is randomly divided into a training set and a validation set two parts, wherein the training set is used as the main data for model weight update; the validation set is used to monitor the model performance during the training process to prevent overfitting; a YOLO model configuration file is selected, and the number of classes is changed to the number of classes to be recognized according to the modification of the key parameters therein; finally, the model is trained, including: selecting a pre-trained weight file as an initial point for training, wherein the training process is an iterative loop, and the following operations are performed in each traversal of the entire training set: forward propagation: input a batch of training images into the model. The model outputs the prediction for each grid according to the current weight parameters, including the bounding box coordinates, confidence and class probability. Loss calculation: compare the model's prediction results with the true labels generated by labelImg. The loss function calculates a total error value. Backpropagation and optimization: through the backpropagation algorithm, the gradient of the loss function to each weight of the model (that is, each weight needs to bear how much responsibility for the total error) is calculated. The optimizer updates all weights of the model according to the gradient and a certain learning rate, aiming to reduce the loss value in the next iteration. Verification and monitoring: after each traversal of the entire training set is completed, the model will run on the validation set once to calculate various performance indicators.

[0026] Step 3: the verification terminal and the work card of the operator are networked and the identity information is verified, the camera identifies and analyzes the identity information of the operator based on AI, and the verification terminal and the camera cross-authenticate the work personnel; Step 4: The supervisor sets the total rule, the cloud server automatically generates the sub-rule according to the processed image and delivers it to the edge computing server for execution, and the worker triggers the total rule or the sub-rule to push the alarm and record the event to the client for the supervisor to view; Step 5: The edge computing server predicts the risk of the worker's behavior characteristics according to the total rule and the sub-rule.

[0027] As shown in Figure 2 , the defect classification and target feature library training specifically includes the following steps: Step 101: The cloud server uses AI intelligent algorithms in the AI algorithm capability core engine to classify defect targets for defect-free image samples and defect image samples, and after classification, the next step is executed; Step 102: The cloud server uses AI models in the AI algorithm capability core engine to train the curve classification target, obtains a defect classification target feature library, saves the defect classification target feature library, saves the defect classification target feature library for subsequent repeated retrieval, and executes the following step; Step 103: The saved defect classification target feature library is transmitted to the edge computing server, and step 203 is executed for feature library comparison, and step 103 is repeatedly executed until the cloud server obtains new defect image samples and non-defect image samples.

[0028] Among them, the original image collected by the camera in the local area network is directly or indirectly transmitted to the edge computing server. When directly transmitted, the remote device is directly transmitted to the edge computing server through the remote network. The remote device is the camera. When indirectly transmitted, the storage and calculation separation technology is adopted, that is, data storage and image processing are performed by different devices. The original image obtained by the camera is saved in the network attached storage, and the hard disk file system built-in the network attached storage transmits the original image to the edge computing server through file IO. The network attached storage is the storage end, and the edge computing server is the computing end. The system IO built-in the verification terminal transmits verification data to the edge computing server through the system serial port or other interfaces. The verification data is also obtained by the image acquisition module of the edge computing server. The verification data and the original image are bound together for subsequent processing; As shown in Figure 2 , the processed original image specifically includes the following steps: Step 201: The image acquisition module of the real-time analysis submodule obtains the original image from the remote device, the hard disk file system and the system IO respectively, and transmits the original image to the image preprocessing module for the next step; Step 202: The image preprocessing module performs noise reduction and white balance processing on the original image to obtain a preprocessed image, and executes the next step; Step 203a: compare the pre-processed image with the defect classification target feature library output by the feature training submodule to obtain target recognition, and use an outer rectangle to frame the detected target; Step 203b: perform reference size labeling on the pre-processed image to obtain target size labeling; Step 203c: perform clarity judgment on the pre-processed image to output image clarity evaluation; Steps 203a, 203b, and 203c are executed in parallel, and after execution is completed, they are merged into a processed image and output uniformly, and the next step is executed; Step 204: the real-time analysis submodule outputs the processed image, and sequentially labels the image clarity, defect target size, and defect target area, and jumps to step 201 for repeated execution.

[0029] The cross-authentication confirmation of the verification terminal and the camera to the specific work personnel includes the following steps: The pre-deployment process is as follows: Step 301: provide each work personnel with a work card with a built-in Bluetooth module, record the work card ID in the Bluetooth module, encrypt the work card ID using RSA keys, the verification terminal is built-in with RSA keys for decryption, the identity information of each work personnel is pre-set in the edge computing server and a mapping table of identity information and work card ID is established, and the next step is executed; Step 302: configure the verification terminal as a Bluetooth Mesh gateway and complete initialization, connect each work personnel's work card Bluetooth module to the verification terminal, complete the pre-deployment process, the Bluetooth modules between the work cards can be used as relays to transmit encrypted data to the verification terminal, and each camera is also configured with the same Bluetooth module as the work card to participate in networking, the Bluetooth module configured in the camera records the device number of each camera, which is used by the verification terminal to determine the specific camera; When the work personnel appears in the camera monitoring picture, the verification operation starts, and the specific steps are as follows: Step 303: the camera collects the original image and transmits it to the edge computing server, the edge computing server processes the original image to output a processed image and compares it with the built-in identity information, the identity information includes the name, personal number, and personal portrait of the work personnel, when comparing, the edge computing server uses MSE mean square error and PSNR peak signal-to-noise ratio to calculate the similarity between the processed image and the personal portrait, the similarity weight calculated by MSE mean square error and PSNR peak signal-to-noise ratio is 0.5:0.5, the edge computing server outputs the identity information with the highest similarity, and the next step is entered; Step 304: The terminal verifies that the device number is obtained from the Bluetooth module of the camera triggering the verification operation, the encrypted data of the Bluetooth module with the strongest connection signal near the camera Bluetooth module is obtained, the verification operator is closest to the camera, the Bluetooth connection signal is the strongest, the RSA key is used to decrypt the encrypted data to obtain the ID of the work card, and the next step is entered; Step 305: The terminal transmits the ID of the work card to the edge computing server, and the edge computing server finds the corresponding identity information from the mapping table according to the ID of the work card, and the next step is entered; Step 306: The edge computing server compares the identity information obtained in steps 303 and 305, and if they match, the identity of the operator is confirmed, the edge computing server uploads the matching result to the cloud server, and if they do not match, the edge computing server retrieves the historical records of the corresponding verification terminal and the corresponding camera from the cloud service, and the one with more successful times is used as the identity information.

[0030] The determination of the total rule and the sub-rule specifically includes the following steps: Step 401: The supervisor logs in to the system through the client and enters the camera management interface, sets a special tag for each camera according to the monitoring area and the operation type, and the edge computing server classifies all cameras according to the tags, and the next step is entered. Step 402: The supervisor sets the total rule on the client, such as "high-altitude operation requires wearing a fall protection device" and "electrical operation requires using an electrical tester", the total rule is uploaded to the cloud server through the client, and then synchronized to the edge computing server by the cloud server, and the next step is entered. Step 403: The edge computing server uploads the original image of each camera to the cloud server after processing, and the cloud server analyzes the behavior characteristics of the operator through the AI intelligent algorithm of the AI algorithm capability core engine, including the action and operation process of the personnel, such as "not wearing a fall protection device" and "not using an electrical tester", and the next step is entered. Step 404: The cloud server calls the AI model based on the YOLOv12 model and the supporting behavior analysis model to analyze the behavior characteristics of the operator, and automatically generates a sub-rule, which is a further constraint condition under the total rule, such as "virtual hanging of the fall protection device triggers an alarm", which makes more detailed operation requirements compared to the total rule "not wearing a fall protection device", and the cloud server transmits the sub-rule to the edge computing server through the distribution device for saving and enabling, and the next step is entered. Step 405: When the original image captured by the camera appears to have a picture change, the edge computing server extracts the processed image from the original image in real time and matches it with the enabled total rules and sub-rules. If the processed image triggers a total rule or a sub-rule, proceed to step 406. If no total rule or sub-rule is triggered, do not take any action and repeat step 405. Step 406: When a total rule is triggered, the edge computing server automatically generates an alarm message and pushes it to the client through the distribution device for real-time viewing by supervisors. When a sub-rule is triggered, an event is automatically recorded and pushed to the client in chronological order for subsequent browsing by supervisors. Go to step 405.

[0031] The risk prediction specifically includes the following steps: Step 501: The edge computing server calculates the risk value of each camera monitoring area based on the AI intelligent algorithm of the AI algorithm capability core engine, generates a risk prediction event, such as "High-altitude-01 camera area: 2 times of not wearing a fall protector in 10 minutes, high risk level", and proceeds to the next step, according to the enabled sub-rules and combined with the real-time collected behavior data of the workers. Step 502: The edge computing server automatically classifies and sorts all cameras according to the risk level of the risk prediction event, which is divided into high, medium, and low. The prediction event with a high risk level is pushed to the client of the supervisor through the distribution device, along with the corresponding camera screenshot. The client interface displays high-risk events first, making it easy for supervisors to understand key risk points in a timely manner, and proceeds to the next step. Step 503: The supervisor views the pushed risk prediction event on the client and labels the processing result according to the actual situation. The processing result is divided into handled, unhandled, and ignored. Handled means the worker has been contacted for rectification, unhandled means it is temporarily impossible to contact for follow-up, and ignored means it is determined to be a false positive or does not need to be handled. The processing result is uploaded to the edge computing server through the client. Set the cumulative threshold number to 15. Determine whether the number of processing result uploads exceeds the cumulative threshold number. If it does not exceed, go to step 501 and repeat. If it exceeds, proceed to the next step. Step 504: The edge computing server counts the number of labels for each processing result. If the most frequent label is handled, the current settings of the sub-rule are retained. If the most frequent label is unhandled, the edge computing server will increase the risk level weight of the corresponding sub-rule and subsequently prioritize the push of this type of risk prediction event. If the most frequent label is ignored, the edge computing server uploads the ignored processing result to the cloud server and executes step 404 to regenerate the sub-rule.

[0032] An AI recognition algorithm-based construction work safety intelligent supervision system, as shown in Figure 1 includes a client, a cloud server, a distribution device, an edge computing server and n cameras, the distribution device includes a router and a switch, the cloud server establishes bidirectional communication with the client and the distribution device through the Internet respectively, the distribution device establishes bidirectional communication with the edge computing server and a plurality of n cameras through an internal local area network respectively, and a verification terminal establishes bidirectional communication with the distribution device in a wireless connection or wired connection manner; wherein n is a positive integer; As shown in Figure 3 , an AI intelligent recognition platform is built inside the cloud server, the AI intelligent recognition platform includes a basic information management module, an AI video supervision center, a streaming media service platform and an AI service platform, the basic information management module includes perception device basic information management, first log management, camera management and AI intelligent algorithm setting, the AI video supervision center includes AI monitoring alarm, video carousel supervision, offline monitoring analysis and AI statistics, the streaming media service platform includes streaming media basic service, alarm management, camera management, subscription time notification, RTSP server management, video acquisition, video distribution, video storage, historical video management, audio and video codec, download management, voice broadcast and voice intercom; The AI service platform includes an AI algorithm capability core engine and an AI intelligent service, the AI algorithm capability core engine includes an AI algorithm model optimization, a self-learning engine, an AI algorithm basic framework service, an RTSP video acquisition management, an AI model, an AI intelligent algorithm and an AI core algorithm acceleration, and the AI intelligent service includes perception device basic information management, offline data acquisition, second log management, AI recognition and retrieval service, business scenario AI service application, message notification, AI video service and AI picture service.

[0033] Among them, in the basic information management module, the perception device basic information management is used for entering and updating the basic data of the model, position and access mode of the perception device camera, ensuring device compliance access and state traceability, ensuring unified control of the system on the device, the first log management is used for recording system operation behavior, device running state and abnormal information, providing data basis for fault troubleshooting and responsibility tracing, which can cooperate with the open source Pandas library for log data statistical analysis, the camera management is used for configuring the resolution, frame rate and picture angle of the camera, ensuring that the camera stably outputs the original image or video stream meeting the AI analysis demand, and the AI intelligent algorithm setting is used for adjusting the detection threshold and recognition category of the AI algorithm, adapting to different recognition scenarios of construction workers, equipment and defects, which means to improve the accuracy of the algorithm in a specific scene, and the open source YOLOv12 algorithm can be used for multi-target detection parameter optimization; In the AI video supervision center, the AI monitoring alarm detects irregular behavior through real-time analysis of video streams and triggers an alarm by using the YOLOv12 open source model, which can timely warn of potential safety hazards and prevent accidents. The video carousel supervision is used to switch the real-time images of multiple cameras in a loop, achieving synchronized monitoring of multiple areas of construction work and reducing blind spots in supervision. The offline monitoring and analysis is used to monitor the online status of cameras and edge computing servers in real time, analyze the reasons for offline, and ensure the continuity of data collection to avoid a regulatory vacuum. The AI statistics are used to aggregate the number of alarms, recognition accuracy, and device online rate to provide data support for regulatory strategy optimization and system performance improvement.

[0034] In the stream media service platform, the stream media basic service is used to process the RTSP video transmission protocol to ensure stable transmission of video streams and connect video collection with subsequent analysis and playback. The alarm management is used to store and classify alarm information and associate corresponding video clips to facilitate traceability and batch processing of alarm information. The camera management and basic information management module work together to supplement the camera coding format and stream transmission parameter configuration at the stream media level to ensure compatibility of the video stream output by the camera with the stream media system. The subscription time correction notification is used to push time correction information to devices to ensure synchronization of device time for cameras and edge servers and avoid timestamp deviation of video data. The RTSP server management is used to deploy and maintain the open source live555 RTSP server to control the start and stop of the server, limit the number of connections, and ensure efficient transmission of video streams to AI models and clients through the RTSP protocol. The video collection obtains raw video streams by calling camera interfaces and uses the open source OpenCV library for preliminary frame extraction to provide raw data sources for AI analysis. The video distribution is used to distribute raw video streams or AI-processed video streams to clients and edge computing servers to meet the demand for simultaneous viewing by multiple terminals and reduce the pressure of repeated transmission in the cloud. The video storage is used to store video data in the open source MinIO object storage or network attached storage to achieve long-term retention of historical video and facilitate post-accident review. The historical video management provides search and playback functions for historical video to support tracing of past construction scenes and identifying potential problems. The audio and video codec uses the open source x264 algorithm for video compression and FFmpeg for audio and video decoding to reduce video storage space and transmission bandwidth and ensure data transmission efficiency. The download management supports user download of video clips to facilitate offline analysis and evidence preservation. The voice broadcast and voice intercom based on the open source WebRTC protocol are used to realize real-time voice interaction between the supervision end and the construction end, timely convey safety instructions, and respond to inquiries from workers to improve communication efficiency.

[0035] In the AI service platform, the AI algorithm model optimization mark autonomous learning engine of the AI algorithm capability core engine re-labels the mis-identified and missed-identified samples through the open source LabelStudio tool, iteratively optimizes the AI model in combination with the Adam optimizer, continuously improves the model identification accuracy, adapts to the changes of the construction scene, the AI algorithm basic framework service provides sample preprocessing, concurrent processing, communication guarantee and video stream and picture identification, the sample preprocessing is based on OpenCV for image rotation and cropping data enhancement, the concurrent processing is based on the open source TensorFlowServing to realize multi-request scheduling, provides a basic environment for the AI algorithm operation, guarantees the efficient and stable work of the algorithm, the RTSP video acquisition management acquires the RTSP video stream through the live555 server and parses it into frame data transmitted to the AI model, connects the streaming media acquisition and AI analysis, ensures the real-time data transmission, the AI model adapts to the open source YOLOv12 model, is used for multi-target detection and identification of workers, equipment and defects in the construction scene, replaces manual work to complete real-time intelligent analysis, improves the supervision efficiency, the AI intelligent algorithm extracts target features by using the open source ResNet algorithm, classifies defects by using the support vector machine (SVM) algorithm, improves the accuracy of target classification and defect identification, reduces the misjudgment, the AI core computing power acceleration is based on the OpenCL convolution software acceleration scheme, uses the GPU computing core and LocalMemory to improve the AI model inference speed, reduces the calculation delay of the edge and cloud, adapts to the ARM architecture embedded device; The perception device basic information management of the AI intelligent service cooperates with the basic information management module, synchronizes the latest state of the device to the AI video supervision center, ensures that the AI analysis can adjust the identification strategy in combination with the device state, acquires the state data of the device offline for collecting the device context information for the AI model, avoids the misanalysis caused by the device offline, the second log management records the calling record, the identification result and the model running log of the AI service, facilitates the AI service troubleshooting and performance optimization, the AI identification and retrieval service is based on the open source Elasticsearch and is used for realizing the quick retrieval of the target identification result and historical identification data, improves the query efficiency of the supervision personnel, quickly locates the key information, the AI service application of the business scene adapts the AI capability to the specific construction scene, for example, the state detection of the anti-falling device for high-altitude operation, the compliance inspection of the safety cone barrel laying, the YOLOv12 model is based on the scene optimization, realizes the transformation of the AI technology from general identification to industry landing, solves the actual construction supervision problem, the message notification is used for pushing the alarm information and the identification result to the supervision personnel in the form of a short message or a system message, ensures that the key information is timely reached, avoids the delay of disposal, the AI video service provides the enhanced video stream after the AI analysis, for example, the target is labeled with a rectangular box and the defect information is superimposed, the identification result is directly displayed, and the supervision personnel can quickly identify the problem, the AI picture service provides the labeled picture after the AI processing, for example, the defect area and size information are labeled, and the details are viewed, archived and analyzed, and the application efficiency of the picture detection is improved.

[0036] Embodiment two The application also provides a technical solution: a construction operation safety intelligent supervision system based on an AI identification algorithm, including neural network computing power acceleration and optimization, AI and related modules, non-digital vernier caliper reading AI identification, grounding resistance shaking table reading AI identification, test pen AI identification, anti-falling device AI identification and safety cone barrel AI identification. The neural network computing power acceleration and optimization is based on the convolution software acceleration optimization scheme of OpenCL. OpenCL performs specific convolution optimization on specific network models, including fully utilizing GPU LocalMemory, fully utilizing GPU computing cores, GPU having multiple computing cores, each computing core having many work units, each work unit being equivalent to a thread, fully utilizing data carried each time: as many points as possible are calculated each time, and loops are unfolded, such as The for loop is expanded into 36 serial single instructions; GPU SIMD GPU level multi-stage pipeline overlap and optimization technology, the scheme is based on OpenCL and specific convolution optimization for specific network models, has strong applicability, can be widely used in PC and embedded devices supporting OpenCL, especially in embedded devices that do not support Nvidia graphics cards, the scheme has strong practical significance, even if the device supports Nvidia graphics card, considering the price of Nvidia graphics card, this scheme can be used to greatly improve the system computing performance without increasing the hardware cost, at present, terminal widely adopts ARM architecture, and the operating system supports OpenCL, the scheme can be widely applied to the calculation acceleration of terminal device neural network application scene; Among them, the AI and related modules include AI algorithm model optimization mark autonomous learning engine and AI algorithm basic framework service, the AI algorithm model optimization mark autonomous learning engine is used for misidentification and missed identification, and the AI algorithm basic framework service is used for sample table preprocessing, concurrent processing, communication guarantee, video stream identification and picture identification; Among them, the non-digital vernier caliper is usually used for precise measurement, and the non-digital vernier caliper reading AI recognition automatically identifies the scale of the vernier caliper and reads the accurate value by AI technology, therefore, we need to first collect images containing vernier calipers of different brands and models, and mark the positions of the vernier, the scale and the scale line in each image, these marked data will be used to train the YoloV12 model to enable it to detect and locate each part of the vernier caliper, in the aspect of image processing, we use data enhancement techniques such as rotation, scaling, cropping, etc. to enhance the robustness of the model, in the model training stage, we use the target detection capability of YoloV12 to detect the outline of the vernier caliper, the specific position of the vernier and the scale line, and through the regression network to predict the distance between the vernier and the reference scale line, and combine these information to calculate the reading, finally, the model will output the accurate reading value, reducing the manual measurement error, improving the work efficiency and accuracy, in order to improve the recognition accuracy of the model, in the future, we can continue to expand the data set and introduce high-resolution training data to further optimize the generalization ability of the model, especially in different light and angle; Among them, the ground resistance dial gauge is usually used to measure the ground resistance value of electrical equipment, and the AI recognition core of the ground resistance dial gauge reading is to automatically read the pointer position and its digital display on the dial gauge to obtain the resistance value. First of all, we need to collect images of ground resistance dial gauges of different models and brands, and label the dial, pointer and digital area in each image. Image enhancement techniques will also be applied in this stage to adapt to changes in different angles, lighting and dial reflection, etc. In the model training stage, we still use YoloV12 for target detection to identify the positions of the dial, pointer and digital area. Through the regression network, the pointer angle is predicted, and combined with the scale of the dial, the accurate resistance value is finally calculated. At the same time, the digital area will be identified by a classification algorithm to ensure the accuracy of the reading. This model will output the value of the ground resistance in real time. Through model optimization, such as improving data preprocessing and image enhancement techniques, its stability under complex lighting conditions can be further improved, ensuring that it can quickly and accurately complete the measurement in different environments. Among them, the ground resistance dial gauge is usually used to measure the ground resistance value of electrical equipment, and the AI recognition core of the ground resistance dial gauge reading is to automatically read the pointer position and its digital display on the dial gauge to obtain the resistance value. First of all, we need to collect images of ground resistance dial gauges of different models and brands, and label the dial, pointer and digital area in each image. Image enhancement techniques will also be applied in this stage to adapt to changes in different angles, lighting and dial reflection, etc. In the model training stage, we still use YoloV12 for target detection to identify the positions of the dial, pointer and digital area. Through the regression network, the pointer angle is predicted, and combined with the scale of the dial, the accurate resistance value is finally calculated. At the same time, the digital area will be identified by a classification algorithm to ensure the accuracy of the reading. This model will output the value of the ground resistance in real time. Through model optimization, such as improving data preprocessing and image enhancement techniques, its stability under complex lighting conditions can be further improved, ensuring that it can quickly and accurately complete the measurement in different environments. Among them, the fall arrestor is usually used in high-altitude operation, and its main function is to prevent personnel from falling. The fall arrestor AI recognition is used to judge the working state of the fall arrestor, such as whether it is locked or in an active state. The image data of the fall arrestor is collected, and the image contains the fall arrestor in different states to ensure that the data set covers a variety of equipment and scenes. The image annotation will cover the shape characteristics of the fall arrestor and its working state, such as: locking, loosening, etc. In the training process, YoloV12 will be used to detect the position and shape characteristics of the fall arrestor. Next, combined with the appearance changes of the fall arrestor, the model will judge whether it is in normal working state or whether there is a fault. This task requires accurate judgment of the changes of the fall arrestor, especially when the fall arrestor is blocked or in poor lighting conditions. Through optimization of the data set and training strategy, the model's adaptability to these environmental changes can be further improved in the future. The system will provide real-time feedback on the working state of the fall arrestor on the interface, helping workers to discover potential safety hazards in time and avoid accidents. Among them, the safety cone barrel is usually used in road construction or safety warning. The safety cone barrel AI recognition judges whether there is a safety cone barrel in the image, and further judges the integrity of the cone barrel, such as whether it is collapsed or damaged. In order to train the model, we need to collect images containing safety cone barrels to ensure that the data set covers safety cone barrels in various environmental conditions, such as daytime, nighttime, different weather, etc. During annotation, in addition to the positioning of the cone barrel, the state information of the cone barrel needs to be annotated, such as intact, inclined or collapsed. In the model training stage, YoloV12 will be used to detect the cone barrel in the image and judge the state of the cone barrel. Since the color of the cone barrel may be affected by the environmental light, we need to introduce diverse lighting conditions in the data set annotation. Through this strategy, the model will be able to accurately identify the state of the cone barrel in different backgrounds and environments. The goal in the training process is to judge the state of the cone barrel through multi-classification, such as intact, collapsed, missing, etc. By strengthening the learning of the collapsed and blocked scenes of the cone barrel, the robustness of the model is further improved to ensure its stability in complex environments. By combining target detection and segmentation technology, the detail recognition ability of the cone barrel can be further improved to provide more accurate working state feedback. The key points that need to be explained in this embodiment are as follows: 1. Technological architecture innovation, building a typical architecture of deep integration of edge computing and cloud computing, realizing real-time identity detection, recognition and authentication in ubiquitous access scenarios, solving the problem of insufficient cooperation between edge and cloud in traditional monitoring systems; ubiquitous monitoring and industry adaptation, realizing the customization of "person" ubiquitous monitoring, recognition and access in specific scenarios of communication industry maintenance work, meeting the personalized needs of whole life cycle monitoring of construction work; intelligent technology fusion application, integrating artificial intelligence, big data, Internet of Things, mobile Internet and other technologies, improving the robustness of intelligent recognition, such as target detection and behavior analysis in complex environment, breaking through the bottleneck of existing algorithm environment limitation; whole link security and credibility, establishing the trusted access, interaction and access mechanism of equipment and target, forming an integrated identity authentication system of edge and cloud, ensuring the security and legality of Internet of Things access; through the construction of high-precision intelligent service ecosystem, supporting whole life cycle monitoring and adaptability management of construction work, improving the level of comprehensive intelligent supervision, operation and service; whole process safety supervision, from image acquisition, feature extraction, model training to application scenario landing, forming a complete construction work safety supervision closed loop, through real-time monitoring, automatic recording, intelligent early warning and other functions, realizing all-round safety control of construction process, effectively avoiding safety problems caused by improper use of equipment or data measurement error; 2. Identity authentication architecture of edge and cloud cooperation, protecting the cooperation mechanism of real-time detection of edge node and deep analysis of cloud, including data interaction rules, computing power allocation strategy and authentication process, solving the problem of high delay and large bandwidth pressure of traditional single cloud processing, real-time recognition technology in ubiquitous access scenarios, protecting multi-modal recognition algorithm for "person" ubiquitous monitoring, such as fusion recognition of personnel biological characteristics and object physical characteristics, and adaptive optimization scheme in complex environment of construction work, industry customized intelligent analysis system, protecting customized function modules for whole life cycle monitoring of construction work, such as real-time processing of full HD video stream, intelligent research and judgment of specific behavior, whole link data traceability, etc., device trusted access and safe interaction mechanism, protecting the identity authentication protocol, data encryption transmission method and access permission dynamic management strategy of Internet of Things device access, ensuring legal access and trusted interaction of devices, multi-technology fusion demonstration application mode, protecting the integrated application scheme of cloud computing, big data, artificial intelligence and edge computing in construction work monitoring scene, including whole life cycle data linkage and adaptive management decision support model, intelligent algorithm robustness improvement scheme, protecting algorithm optimization technology for complex environment, such as sample enhancement and scene adaptive training model, breaking through the problem of existing algorithm environment limitation, safety supervision system architecture and process, covering the whole architecture and running process of construction work safety supervision system from image acquisition device layout, data transmission processing, model deployment application to safety warning mechanism, which is a system scheme for comprehensive, real-time and intelligent supervision of construction site equipment.

[0037] The above merely provides the preferred embodiment of the present application, and the protection scope of the present application is not limited thereto. Any person skilled in the art, according to the technical range disclosed by the present application and the inventive concept, can make equivalent replacements or changes, and all of them should be covered in the protection scope of the present application.

Claims

1. A construction operation safety intelligent monitoring method based on an AI recognition algorithm, characterized by comprising the following steps: Step 1: Obtain defect image samples and non-defect image samples, responsible for defect classification and target feature library training; Step 2: Obtain the original image, perform image preprocessing, feature library comparison, output the processed image, and frame the detected target with an outer rectangle; Step 3: Verify the terminal and the worker's card to form a network and verify the identity information, the camera analyzes the identity information of the worker based on AI recognition, and the terminal and the camera cross-authenticate to confirm the worker; Step 4: The supervisor sets the general rules, the cloud server automatically generates sub-rules based on the processed image and sends them to the edge computing server for execution, and the worker triggers the general rules or sub-rules to push alarms and record events; Step 5: The edge computing server predicts the risk of the worker's behavior characteristics based on the general rules and sub-rules. The defect classification and target feature library training specifically includes the following steps: 2.The AI identification algorithm-based construction operation safety intelligent supervision method according to claim 1, characterized in that, Step 101: The cloud server uses AI intelligent algorithms to classify defect targets for non-defect image samples and defect image samples, and performs the next step after classification is complete; Step 102: The cloud server uses AI model in the AI algorithm capability core engine to train the curve classification target, obtains the defect classification target feature library, saves the defect classification target feature library, and executes the next step; Step 103: The saved defect classification target feature library is transmitted to the edge computing server for feature library comparison, and step 103 is repeatedly executed until the cloud server obtains new defect image samples and non-defect image samples. The local area network camera collects original images directly or indirectly transmitted to the edge computing server, when directly transmitted, the remote device is directly transmitted to the edge computing server through the remote network, the remote device is the camera, when indirectly transmitted, the original image obtained by the camera is saved in the network attached storage, and the hard disk file system built in the network attached storage is transmitted to the edge computing server through file IO, and the system IO built in the verification terminal is transmitted to the edge computing server through the system serial port or other interface; 3.The AI identification algorithm-based construction operation safety intelligent supervision method of claim 1, wherein, Processing the original image specifically includes the following steps: Step 201: The image acquisition module of the real-time analysis submodule obtains the original image from the remote device, the hard disk file system, and the system IO, respectively, and transmits the original image to the image preprocessing module for the next step; Step 202: The image preprocessing module performs noise reduction and white balance processing on the original image to obtain a preprocessed image, and executes the next step; Step 203a: Compare the preprocessed image with the defect classification target feature library output by the feature training submodule to obtain target recognition, and frame the detected target with an outer rectangle; Step 203b: Perform reference size labeling on the preprocessed image to obtain target size labeling; Step 203c: Perform clarity judgment on the preprocessed image, and output the image clarity evaluation; Wherein, steps 203a, 203b and 203c are executed in parallel, and are merged into a processed image for unified output after execution, and the next step is executed; ​ Step 204: The real-time analysis submodule outputs the processed image and sequentially labels the image clarity, defect target size, and defect target area, and jumps back to step 201 for repeated execution. 4.The AI identification algorithm-based construction operation safety intelligent supervision method according to claim 1, characterized in that, The cross-verification and confirmation of the terminal and the camera with the work personnel specifically includes the following steps: The early deployment process is as follows: Step 301: Each work personnel is equipped with a work card with a built-in Bluetooth module, the Bluetooth module has the work card ID recorded therein, the work card ID is encrypted using an RSA key, the verification terminal has a built-in RSA key for decryption, the edge computing server has pre-set identity information of each work personnel and has established a mapping table of the identity information and the work card ID, and the next step is executed; Step 302: The verification terminal is configured as a gateway and is initialized, the Bluetooth module of each work card is connected to the verification terminal, the early deployment process is completed, the Bluetooth modules of the work cards can be used as relays to transmit encrypted data to the verification terminal, and each camera is also configured with the same Bluetooth module as the work card to participate in networking, and the Bluetooth module of the camera has the device number of each camera recorded therein; When the work personnel appears in the camera monitoring picture, the verification operation is started, and the specific steps are as follows: Step 303: The camera collects the original image and transmits it to the edge computing server, the edge computing server processes the original image to output a processed image and compares it with the built-in identity information, the identity information includes the name, personal number, and portrait of the work personnel, the edge computing server uses MSE mean square error and PSNR peak signal-to-noise ratio to calculate the similarity of the processed image and the portrait, the edge computing server outputs the identity information with the highest similarity, and the next step is entered; Step 304: The verification terminal obtains the device number from the Bluetooth module of the camera that triggers the verification operation, obtains the encrypted data of the Bluetooth module with the strongest connection signal near the camera Bluetooth module, decrypts the encrypted data using the built-in RSA key to obtain the work card ID, and enters the next step; Step 305: The verification terminal transmits the work card ID to the edge computing server, and the edge computing server finds the corresponding identity information from the mapping table according to the work card ID, and enters the next step; Step 306: The edge computing server compares the identity information obtained in steps 303 and 305, and if they match, the identity of the work personnel is confirmed, the edge computing server uploads the matching result to the cloud server, and if they do not match, the edge computing server retrieves the historical records of the successful matching of the identity information of the verification terminal and the camera from the cloud server, and takes the one with more successful times as the standard, and outputs the corresponding identity information. 5.The AI identification algorithm-based construction operation safety intelligent supervision method according to claim 1, characterized in that, The determination of the total rule and the sub-rule specifically includes the following steps: Step 401: The supervisor logs in to the system through the client and enters the camera management interface, sets a special tag for each camera according to the monitoring area and the work type, and the edge computing server classifies all the cameras according to the tags, and enters the next step; Step 402: The supervisor sets the total rule on the client, the total rule is uploaded to the cloud server through the client, and then the cloud server synchronizes the total rule to the edge computing server, and enters the next step; Step 403: The edge computing server uploads the processed original image of each camera to the cloud server, which analyzes the behavior characteristics of the workers through the AI intelligent algorithm of the AI algorithm capability core engine. The behavior characteristics include personnel actions and operation processes, and the next step is entered. Step 404: The cloud server calls the AI model to analyze the behavior characteristics of the workers and automatically generates sub-rules. The sub-rules are further constraint conditions under the total rule. The cloud server transmits the sub-rules to the edge computing server through the distribution device for storage and activation, and the next step is entered. Step 405: When the original image captured by the camera shows a picture change, the edge computing server extracts the processed image from the original image in real time and matches it with the enabled total rule and sub-rule. If the processed image triggers the total rule or sub-rule, the next step 406 is performed. If the total rule or sub-rule is not triggered, no action is taken, and step 405 is repeated. Step 406: When the total rule is triggered, the edge computing server automatically generates an alarm message and pushes it to the client through the distribution device for real-time viewing by supervisors. When the sub-rule is triggered, an event is automatically recorded and pushed to the client in chronological order for subsequent browsing by supervisors. Go to step 405. 6.The AI identification algorithm-based construction operation safety intelligent supervision method according to claim 1, characterized in that, Risk prediction specifically includes the following steps: Step 501: The edge computing server calculates the risk value of each camera monitoring area based on the AI intelligent algorithm of the AI algorithm capability core engine, generates a risk prediction event, and enters the next step according to the enabled sub-rules and real-time collected worker behavior data. Step 502: The edge computing server automatically classifies and sorts all cameras according to the risk level of the risk prediction event, which is divided into high, medium, and low. It prioritizes pushing the prediction events with high risk levels, along with the corresponding camera screenshots, to the client of the supervisor through the distribution device. The client interface displays high-risk events first, and the next step is entered. Step 503: The supervisor views the pushed risk prediction events on the client and labels the processing results according to the actual situation. The processing results are divided into handled, unhandled, and ignored. The supervisor uploads the processing results to the edge computing server through the client. A cumulative threshold number is set to determine whether the number of processing result uploads exceeds the cumulative threshold number. If it does not exceed, go to step 501 and repeat the execution. If it exceeds, perform the next step. Step 504: The edge computing server counts the number of labels for each processing result. If the most frequent label is handled, the current settings of the sub-rule are retained. If the most frequent label is unhandled, the edge computing server increases the risk level weight of the corresponding sub-rule and prioritizes pushing this type of risk prediction event in the future. If the most frequent label is ignored, the edge computing server uploads the ignored processing results to the cloud server to regenerate the sub-rule.

7. An intelligent supervision system for construction safety based on an AI recognition algorithm, characterized in that The AI recognition algorithm-based construction work safety intelligent supervision method according to any one of claims 1-6 is implemented, including a client, a cloud server, a distribution device, an edge computing server and a plurality of cameras, the distribution device includes a router and a switch, the cloud server establishes bidirectional communication with the client and the distribution device through the Internet respectively, the distribution device establishes bidirectional communication with the edge computing server and the n cameras through an internal local area network respectively, and the verification terminal establishes bidirectional communication with the distribution device in a wireless connection or wired connection manner; An AI intelligent identification platform is built inside the cloud server, the AI intelligent identification platform includes a basic information management module, an AI video supervision center, a streaming media service platform and an AI service platform, the basic information management module includes perception device basic information management, first log management, camera management and AI intelligent algorithm setting, the AI video supervision center includes AI monitoring alarm, video carousel supervision, offline monitoring analysis and AI statistics, and the streaming media service platform includes streaming media basic service, alarm management, camera management, subscription time correction notification, RTSP server management, video acquisition, video distribution, video storage, historical video management, audio and video codec, download management, voice broadcast and voice intercom; The AI service platform includes an AI algorithm capability core engine and AI intelligent services, the AI algorithm capability core engine includes an AI algorithm model optimization, marking autonomous learning engine, an AI algorithm basic framework service, RTSP video acquisition management, an AI model, an AI intelligent algorithm and AI core algorithm power acceleration, and the AI intelligent services include perception device basic information management, offline data acquisition, second log management, AI identification and retrieval services, business scenario AI service application, message notification, AI video services and AI picture services. 8.The AI identification algorithm-based construction operation safety intelligent monitoring system according to claim 7, characterized in that, In the basic information management module, the perception device basic information management is used for inputting and updating the basic data of the model, position and access mode of the perception device camera, the first log management is used for recording system operation behavior, device running state and abnormal information, the camera management is used for configuring the resolution, frame rate and picture angle of the camera, and the AI intelligent algorithm setting is used for adjusting the detection threshold and identification category of the AI algorithm to adapt to different identification scenes of construction workers, equipment and defects; In the AI video supervision center, the AI monitoring alarm analyzes the video stream in real time, detects illegal behavior by means of a YOLOv12 open source model and triggers an alarm, the video carousel supervision is used for cyclically switching the real-time pictures of a plurality of cameras, the offline monitoring analysis is used for monitoring the online state of the cameras and the edge computing server in real time, analyzing the offline reasons, and the AI statistics is used for summarizing the alarm times, identification accuracy and device online rate. 9.The AI identification algorithm-based construction operation safety intelligent monitoring system according to claim 7, wherein, In the streaming media service platform, the streaming media basic service is used for processing the RTSP video transmission protocol, the alarm management is used for storing, classifying alarm information and associating corresponding video clips, the camera management cooperates with the basic information management module, supplements the camera coding format and stream transmission parameter configuration of the streaming media level, the subscription time correction notification is used for pushing time correction information to each device, the RTSP server management is used for deploying and maintaining the open source live555 RTSP server, controlling the start and stop of the server, limiting the connection number, the video acquisition is used for acquiring the original video stream by calling the camera interface, performing preliminary frame extraction, the video distribution is used for distributing the original video stream or the AI processed video stream to the client and the edge computing server, and is used for storing the video data to the network attached storage, the historical video management is used for providing the search and playback functions of the historical video, the audio and video coding and decoding are used for video compression and audio and video decoding, the download management is used for supporting the user to download the video clip, and the voice broadcast and voice talkback are used for the real-time voice interaction between the supervision end and the construction end. 10.The AI identification algorithm-based construction operation safety intelligent monitoring system of claim 7, wherein, In the AI service platform, the AI algorithm model optimization label autonomous learning engine of the AI algorithm capability core engine re-labels the misidentification and missed identification samples, iteratively optimizes the AI model, the AI algorithm basic framework service provides sample preprocessing, concurrent processing, communication guarantee and video stream and picture identification, the RTSP video acquisition management acquires the RTSP video stream and analyzes it into frame data transmission to the AI model, the AI model is used for multi-target detection and identification of workers, equipment and defects in the construction scene, the AI intelligent algorithm extracts target features, classifies defects, and the AI core power accelerates the inference speed of the AI model; The perception device basic information management of the AI intelligent service cooperates with the basic information management module, synchronizes the latest state of the device to the AI video supervision center, acquires the state data of the device offline, the second log management records the calling record, identification result and model running log of the AI service, the AI identification and retrieval service is used for quick retrieval of target identification results and historical identification data, the business scene AI service application adapts the AI capability to the specific construction scene, the message notification is used for pushing the alarm information and identification result to the supervision personnel through the short message or system message, the AI video service provides the enhanced video stream after AI analysis, and the AI picture service provides the labeled picture after AI processing.

Citation Information

Patent Citations

  • Video-based non-stop construction unsafe behavior identification system and method

    CN113392760A

  • Food processing safety management method and system based on machine vision

    CN115439935A

  • Video analysis method, device and equipment

    CN116645631A

  • Edge detection system for power grid inspection and monitoring

    CN116846059A

  • Alarm system for automatically identifying nonstandard lightering loading operation based on AI image

    CN119252000A

Cited By

  • Intelligent image recognition and judgment system

    CN121259510A

  • An image intelligent recognition and determination system

    CN121259510B

  • Construction risk intelligent monitoring and early warning method and system based on behavior recognition and operation permission verification

    CN121747033A