Industrial plant intelligent security system and method based on AI vision
By employing a cloud-edge-device collaborative architecture and a few-shot self-learning algorithm training engine, the system addresses the issues of equipment compatibility, algorithm customization, and data security in industrial plant AI smart security systems, enabling low-cost and efficient security system deployment and security supervision.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-03-24
AI Technical Summary
Existing AI-powered smart security systems in industrial plants suffer from poor equipment compatibility, difficulty in algorithm customization, low hardware resource utilization, and high data security risks, making them unable to quickly respond to changes in scenarios and meet high-level security requirements.
It adopts a distributed architecture that integrates cloud, edge, and device, combined with a few-shot self-learning algorithm training engine, a multi-algorithm fusion scheduling manager, and a full-process localized security protection system. It is compatible with multiple brands of devices through standard protocol interfaces, enabling parallel execution of multiple algorithms and localized data processing, and is equipped with physical isolation and encryption measures.
It enables low-cost equipment upgrades, rapid algorithm customization, improved hardware resource utilization, and high-level data security, adapting to the dynamically changing security needs of industrial plants, reducing hardware investment and maintenance costs, and ensuring data security.
Smart Images

Figure CN121728112A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of safety production supervision, and relates to an AI vision-based intelligent security and protection system and method for an industrial plant. BACKGROUND
[0002] An industrial plant is a core scene for production and manufacturing, and the construction of an intelligent security and protection system for the industrial plant is directly related to production safety, personnel safety and asset safety. With the rapid development of artificial intelligence technology, an intelligent security and protection system based on AI video analysis has gradually replaced traditional manual inspection, fixed threshold monitoring and other methods and has become a mainstream technical direction for industrial plant security. Such a system collects on-site data through cameras, sensors and other sensing devices, realizes automatic identification and alarm of safety risks such as illegal operation, equipment anomaly and regional intrusion by combining a deep learning algorithm, can greatly improve the real-time performance and intelligent level of plant security, is suitable for various industrial operation scenes such as painting, mechanical processing and hazardous material storage, and is an important component of industrial intelligent upgrading.
[0003] In actual production and use, the existing AI intelligent security and protection system for an industrial plant still has many technical bottlenecks to be solved: firstly, the device adaptability is poor, most systems only support a small number of brands or private protocol cameras and sensors, cannot be compatible with standard protocol devices such as ONVIF and GB / T28181 that have been deployed in the plant, have high costs for legacy transformation, and cause waste of initial investment of users; secondly, algorithm customization is difficult, general AI algorithm training relies on a large amount of labeled data, usually thousands or even tens of thousands of samples, and the customization period of algorithms for plant individualized safety rules such as safety belt detection for overhead operation and personnel out-of-bound detection in specific areas is as long as several months, which is difficult to quickly respond to changes in the scene; thirdly, the utilization rate of hardware resources is low, a single camera can only run one or a small number of algorithms, a large number of special cameras need to be deployed to realize all-around monitoring, and the hardware deployment and maintenance costs are high; fourthly, the data security risk is prominent, some systems rely on public clouds to complete algorithm training and data processing, and sensitive video streams and production operation data of the plant exist the risk of outflow and leakage, which cannot meet the safety requirements of high-security-level industrial scenes.
[0004] Through retrieval and review of relevant information, for the above problems, the prior art has proposed several solutions: such as patent number CN120494507A, patent name Safety production supervision system and method based on artificial intelligence and big data analysis, the core principle is to realize the safety detection of civil explosive production scene through the architecture of production equipment management system, data and business coordination server, video analysis server, combined with convolutional neural network, interframe difference algorithm, support personnel fixed scene, such as the identification of alarm of the number of detonators, fire passage blocking. The advantages of this scheme are that it realizes automatic monitoring in specific scenarios, replaces part of the manual inspection work, and improves the detection coverage through multi-module cooperation; but the disadvantages are also very obvious, first, the algorithm training depends on a large number of labeled data sets, which cannot adapt to the personalized safety rules of industrial factory area, and the customization cycle is long; second, the system architecture is closed, only supports IP network camera access, has no multi-protocol upgrading and reconstruction capability, and the hardware compatibility is poor; third, a single camera can only run a single algorithm, the resource utilization rate is low, and the whole process localization deployment is not realized, there is a risk of leakage in the data transmission process. In addition, the traditional industrial security algorithm optimization scheme focuses on the improvement of the precision of a single algorithm, such as improving the feature extraction layer of the convolutional neural network, and does not solve the core problems of algorithm customization and multi-algorithm collaborative scheduling; although some local deployment schemes can guarantee data security, they still need to rely on manufacturers to complete algorithm development, and cannot meet the needs of on-site iterative optimization of factory area.
[0005] Through comprehensive analysis of the advantages and disadvantages of existing solutions, the existing technology either focuses on improving the detection accuracy of a single scene, or only solves a single problem of data security or device compatibility, and cannot form an integrated solution from the four dimensions of system architecture, algorithm training, resource scheduling and data security. Therefore, the present application discards the traditional "single function optimization" idea, adopts a distributed architecture of cloud edge cooperation, combines a few sample self-learning algorithm training engine, multi-algorithm fusion scheduling manager and whole process localization security protection system, and solves the four core problems of device adaptation, algorithm customization, resource utilization and data security, and finally realizes low-cost customization, efficient resource utilization and high-level data security protection of the industrial factory area intelligent security system. SUMMARY
[0006] The present application provides an industrial factory area intelligent security system and method based on AI vision, which solves the problems of existing system adaptation in the field of industrial factory area intelligent security, such as difficulty in adapting to multi-brand / multi-protocol devices, algorithm training relying on massive data and long cycle, single device running only a single algorithm resulting in low hardware resource utilization, and data processing relying on external network with risk of leakage.
[0007] In order to solve the above problems, the technical scheme adopted by the present application is: An AI vision-based intelligent security and protection system for an industrial plant, comprising a perception layer, a calculation layer, a platform layer and a data security protection unit; The perception layer is configured with a data access interface supporting ONVIF protocol, GB / T28181 and RTSP protocol for accessing cameras and sensors. The calculation layer comprises an AI intelligent analysis center, which is internally provided with a multi-algorithm container, a few-sample self-learning algorithm training engine and a multi-algorithm fusion scheduling manager; the few-sample self-learning algorithm training engine comprises a natural language interactive interface, a sample labeling module and a localized training unit, and the localized training unit is used for performing model training operations; the multi-algorithm fusion scheduling manager is internally provided with a preset rule library for allocating computing resources to single video streams so that multiple algorithm tasks are executed in parallel on the single video streams. The platform layer is an ASS intelligent control center, which is configured with a management module for centralized control of system functions; the data security protection unit comprises a physical isolation module, an encryption module, an identity verification module and an IP whitelist module; the physical isolation module is used for isolating all data processing units of the system from external networks; the encryption module uses an AES algorithm based on a private protocol to encrypt the external interface of the large model; the identity verification module is configured with a username and a key; and the IP whitelist module controls the IP addresses of data acquisition interfaces.
[0008] An AI vision-based intelligent security and protection method for an industrial plant, which is implemented by using the system of any one of claims 1-10, and comprises the following steps: S1, the perception layer collects video streams and sensor data of the industrial plant and transmits them to the AI intelligent analysis center of the calculation layer through a standard protocol; S2, a user inputs a security risk scenario description through a natural language interactive interface of the platform layer, uploads 30 positive and negative sample images for labeling, triggers a few-sample self-learning algorithm training engine, and completes minute-level model training on a localized AI server; S3, the AI intelligent analysis center distributes computing resources to single video streams through a multi-algorithm fusion scheduling manager, and executes multiple rounds of trained algorithm tasks in parallel; S4, the platform layer displays algorithm detection results in real time, including event name, coordinate frame, risk level, and triggers a hierarchical alarm according to a preset rule when an anomaly is detected; S5, all data processing, model training and algorithm running are completed in an internal local area network, data security protection is realized through AES encryption and IP whitelist, and the model is periodically optimized through few-sample iteration to adapt to changes in plant scenarios.
[0009] The principle and advantages of the present scheme are as follows: The scheme constructs an AI vision-based intelligent security and protection system for industrial plant area, which is cooperatively operated by a perception layer, a calculation layer, a platform layer and a data security protection unit, and forms a standardized security and protection method. From the system architecture, the perception layer realizes the reuse access of different brands and types of cameras and sensors through the access interface compatible with ONVIF, GB / T28181 and RTSP standard protocols, solves the equipment compatibility barrier; the calculation layer as the core computing unit, its built-in few-shot self-learning algorithm training engine can receive the user's risk scene description through the natural language interactive interface, rely on the localized training unit to complete the model training only with a small amount of samples, and the multi-algorithm fusion scheduling manager allocates computing resources based on the preset rules to realize the parallel execution of multi-algorithm tasks, and breaks through the operation limit of single device and single algorithm; the platform layer realizes the centralized management and control of the whole system function through the management module of the ASS intelligent control center, guarantees the uniformity of operation and maintenance; the data security protection unit isolates external network through the physical isolation module, protects the large model interface by using the private protocol AES algorithm of the encryption module, configures the username key of the identity verification module, and controls the data interface access authority through the IP white list module, forming a whole-process data security protection system. The supporting security and protection method puts the system capacity into practice, collects data through the perception layer, completes algorithm training and multi-task execution through the calculation layer, displays alarms through the platform layer, guarantees data security and iteratively optimizes the model through the whole process, forms a closed-loop operation logic of "data collection-algorithm training-intelligent analysis-security protection-model iteration", and adapts to the dynamic change of security and protection demand of industrial plant area.
[0010] Compared to existing technologies, which only support IP network camera access with specific protocols and are incompatible with existing multi-brand, multi-protocol devices in factories, resulting in high costs for retrofitting and wasting initial investment, this solution, through its standardized protocol interface design at the perception layer, can directly connect to existing cameras and sensors in the factory without replacing hardware. Actual verification shows that after adopting this solution in an industrial coating plant, equipment retrofit costs were reduced by 60%, and the retrofit cycle was shortened from 15 days to 3 days, significantly reducing hardware investment and construction costs. Furthermore, existing technologies rely on massive amounts of labeled data for algorithm training, requiring months to customize new algorithms, and the algorithms are fixed as "black boxes" and cannot adapt to changes in scenarios. This solution's few-sample self-learning algorithm training engine requires only a small number of samples to complete localized model training in minutes. For example, in a machining plant addressing the issue of "unsafe high-altitude operations," this solution has proven effective. In the new risk scenario of "seat belts," only 30 sets of positive and negative samples are uploaded, and the algorithm training and deployment are completed within 20 minutes. Compared with the customization cycle of existing technologies that takes several months, the efficiency is improved by more than 95%. Moreover, the model can be adjusted and iterated in real time according to the production process of the factory area, solving the rigidity problem of traditional algorithms being developed once and fixed. Existing technologies can only run a single algorithm with a single camera, requiring the deployment of a large number of dedicated cameras to achieve all-round monitoring. The original security system of a chemical plant area required the deployment of 3 dedicated cameras to cover three types of detection needs: smoke detection, personnel crossing the boundary, and obstruction of fire lanes. This solution realizes the parallel execution of multiple algorithms on a single video stream through a multi-algorithm fusion scheduling manager. The factory area only needs 1 camera to complete the above three types of detection tasks at the same time. The number of hardware deployments is reduced, the hardware procurement and maintenance costs are reduced, and the computing power utilization of a single camera is improved.
[0011] Existing technologies rely on public clouds for algorithm training and data processing, posing a risk of sensitive video streams and production operation data leaking from the factory premises. This solution employs a multi-layered protection system including physical isolation, AES encryption, authentication, and IP whitelisting. All data processing and model training are completed within the internal LAN. After adopting this solution, a military-grade supporting factory eliminated the risk of data leakage during cross-network transmission, meeting the security requirements of high-security industrial scenarios—a core advantage that existing technologies cannot achieve. Existing security systems have fragmented functional modules and lack a unified management interface, requiring maintenance to interface with multiple terminals. This solution, through the centralized management module of the ASS intelligent control center at the platform layer, achieves unified and visualized management and control of algorithm tasks, equipment status, and alarm information.
[0012] Furthermore, the few-shot self-learning algorithm training engine performs model training operations based on a large visual model with the Transformer architecture, specifically including the following steps: S1. Feature Extraction: The input video frames are processed by the convolutional neural network backbone to output a two-dimensional feature map set. S2, Transformer encoder-decoder processing: The encoder performs multi-layer encoding on the feature map in sequence, including dimensionality reduction, flattening, and position encoding, to obtain an image feature sequence containing global context information; the decoder takes the object query vector and the image feature sequence as input, and outputs object instance prediction features with rich embedded information through multi-layer decoding of self-attention mechanism and cross-attention mechanism. S3. Prediction Head and Forward Computation: The embedded features output from the decoder are input into the feedforward network, and the prediction head completes the category prediction and bounding box prediction. S4. Bipartite Graph Matching and Loss Calculation: Constructing a cost matrix and calculating the loss value after solving for the optimal matching completes the model training iteration. Step S1 extracts two-dimensional feature maps of video frames through the convolutional neural network backbone, which can accurately capture the low-level visual features of the image and lay the foundation for subsequent high-level feature analysis. Secondly, the dimensionality reduction, flattening, and positional encoding operations of the Transformer encoder in step S2 can effectively integrate the global contextual information of the image, solving the problem that traditional convolutional neural networks can only extract local features. The decoder, combined with self-attention and cross-attention mechanisms, can focus on key detection targets in industrial scenarios, such as workers on vehicle roofs and people not wearing safety helmets. Even under conditions of few samples, it can generate prediction features rich in target information, greatly reducing the dependence on massive labeled data. Furthermore, the prediction head and forward computation in step S3 can achieve accurate prediction of categories and bounding boxes, adapting to the diverse safety detection needs of industrial plants, such as smoke detection and personnel crossing the boundary. Finally, the bipartite graph matching and loss calculation in step S4, by constructing a cost matrix and solving for the optimal matching, can accurately measure the deviation between the prediction results and the real labels, and achieve efficient iterative optimization of the model. The overall process retains the Transformer architecture's ability to capture global features while adapting to low-sample scenarios through modular training steps. It can complete the training of personalized algorithms for industrial plants in minutes. Compared with traditional convolutional neural network algorithms that rely on massive amounts of data and can only extract local features, it not only shortens the algorithm customization cycle from months to minutes, but also improves the accuracy and generalization ability of target detection in complex industrial scenarios. It can quickly adapt to changes in production processes, new risk points, and other scenario changes, effectively solving the core problems of poor adaptability and high customization costs of industrial security algorithms.
[0013] Furthermore, the cost matrix construction and loss calculation for bipartite graph matching in step S4 adopt the following algorithm formula: Cost matrix formula
[0014] Loss function formula
[0015] in, The cost of matching the i-th real target with the j-th predicted bounding box; The classification loss weight coefficient has a value range of [1,5]. is the bounding box loss weight coefficient, with a value range of [1, 10]; Let cross-entropy be the classification loss function. For the actual target category label, To predict the probability distribution of categories; The CIoU bounding box loss function is... The coordinates of the actual target bounding box. To predict the bounding box coordinates; This represents the total loss value. This is the loss weight coefficient when there is no target matching, and its value ranges from [0.1, 1]. The empty category represents the category with no real target match. The cost matrix formula combines the classification loss and the bounding box loss using weight coefficients. , Weighted fusion allows for flexible adjustment of weights based on the detection priority in industrial scenarios, such as increasing the weights for high-risk scenarios like smoke and fire detection. Improved detection of the boundary frame of workers on the roof Compared to traditional single loss calculation methods, this approach is more suitable for the diverse safety inspection needs of industrial plants; secondly, the loss function incorporates a loss term for situations without target matching. This approach effectively balances the imbalance between positive and negative samples during training with few samples, preventing overfitting or missed detections due to limited sample size and improving the algorithm's generalization ability in long-tail industrial scenarios. Simultaneously, the bounding box loss employs the CIoU function, which considers bounding box overlap, center point distance, and aspect ratio. In complex industrial environments, such as when personnel or equipment occlusion causes target box deformation, bounding box prediction accuracy is improved by over 30%. Furthermore, the formula sets reasonable value ranges for each weight coefficient, ensuring algorithm flexibility while avoiding disordered parameter adjustments. This allows for rapid convergence in minute-level localized training, ultimately achieving accurate prediction of target categories and bounding boxes in industrial scenarios under limited sample conditions. This solves the core problems of traditional algorithms: reliance on massive amounts of data, low detection accuracy, and poor adaptability.
[0016] Furthermore, the dynamic resource allocation algorithm of the multi-algorithm fusion scheduler adopts the following formula: Resource allocation weight formula Single-stream video algorithm execution priority formula
[0017] Resource allocation formula in, Let be the comprehensive weight value of the k-th algorithm task; α is the security level weight coefficient, with a value range of [0.5, 0.8]. β represents the algorithm task safety level, ranging from 1 to 10; β is the resource consumption weight coefficient, ranging from [0.1, 0.3]. γ represents the resource utilization rate of the algorithm task, with a value ranging from 0 to 1; γ is the real-time weight coefficient, with a value ranging from [0.1, 0.2]. To meet the real-time requirements of the algorithm task, the value ranges from 0 to 1; is the execution priority of the k-th algorithm task; n is the total number of algorithm tasks executed in parallel on a single video stream; The amount of computing resources allocated to the k-th algorithm task; The resource allocation weight formula constructs a comprehensive weight based on three core dimensions: security level, resource consumption, and real-time performance, to determine the total computing power resources corresponding to a single video stream. By limiting the reasonable value range of coefficients α, β, and γ, it can be flexibly adjusted according to the safety management priorities of industrial plants. For example, in chemical plants, α can be set to 0.8 to prioritize resource allocation for high-safety-level algorithms such as smoke and fire detection. In machinery plants, the γ coefficient can be appropriately increased to ensure the real-time performance of high-altitude operation detection. Compared to traditional single-dimensional resource allocation methods, this approach better meets the actual needs of industrial scenarios. Furthermore, the execution priority formula quantifies the priority of algorithm tasks through weight proportions, avoiding subjective arbitrariness in resource allocation and ensuring that high-priority algorithms, such as flame detection with a safety level of 10, always receive sufficient computing power. Finally, the resource allocation formula precisely allocates total computing power resources based on priority, ensuring the parallel execution of multiple algorithm tasks on a single video stream while avoiding resource waste. For example, when a single video stream in an industrial plant is simultaneously running three types of algorithms: safety helmet detection (security level 8, resource utilization rate 0.4), smoking detection (security level 6, resource utilization rate 0.3), and fire exit obstruction detection (security level 7, resource utilization rate 0.3), the formula can allocate 70% of the computing power to safety helmet detection, 20% to fire exit obstruction detection, and 10% to smoking detection, achieving optimal allocation of computing power resources. This significantly reduces the hardware deployment cost of the industrial plant security system while ensuring the real-time performance and accuracy of high-risk scene detection.
[0018] Furthermore, the AI intelligent analysis center's built-in dedicated algorithm model library includes a rooftop operation safety belt detection algorithm. This algorithm achieves deep integration of logical rules and visual recognition: it uses a Transformer architecture model to identify whether "rooftop operation" behavior exists in the image. The identification criteria are that the vertical coordinate of the target detection box occupies ≥0.7 of the image height, and the cosine similarity between the target feature vector and the feature vector of the "rooftop operation" sample is ≥0.85. If the identification result is "yes," the safety belt compliance detection branch is triggered, and the bounding box prediction result is used to determine whether the worker is wearing a safety belt. If the identification result is "no," the execution of this algorithm branch is terminated. This algorithm first uses a Transformer architecture model combined with quantitative criteria, namely, the vertical coordinate of the target detection box occupies ≥0.7 and the cosine similarity is ≥0.85, to identify "rooftop operation" behavior. Compared with traditional methods, this algorithm is more efficient. The differential detection method for seat belts can accurately screen out the work scenarios that need to be detected, avoiding invalid seat belt detection for non-roof-work personnel, significantly reducing the false alarm rate of the system. Furthermore, the seat belt compliance detection branch is only triggered when "roof-work" behavior is identified; if not identified, the execution of this algorithm branch is terminated, which can reduce unnecessary computing power consumption, improve the algorithm running efficiency of a single video stream, and adapt to the resource scheduling requirements of parallel execution of multiple algorithms. At the same time, the quantified identification judgment conditions make the identification results of "roof-work" behavior quantifiable and verifiable. Compared with traditional identification methods that rely on human experience or fuzzy features, the detection results are more stable and accurate, effectively identifying violations of not wearing seat belts in roof-work scenarios. This solves the core problems of inaccurate scene identification and redundant alarms in the safety detection of high-altitude operations in industrial plants, improving the pertinence and effectiveness of safety supervision.
[0019] Furthermore, the model iteration optimization of the few-shot self-learning algorithm training engine adopts an incremental learning strategy. After receiving new sample data from the field each time, only the parameters of the prediction head layer and the last two layers of the Transformer decoder are updated, while the encoder layer parameters remain fixed. The learning rate for parameter updates is set to 1 / 10 of the initial training learning rate, and an early stopping strategy is adopted: when the validation set loss does not decrease for three consecutive iterations, the iteration is terminated. This design, which updates only the parameters of the prediction head layer and the last two layers of the Transformer decoder while keeping the encoder layer parameters fixed, can maximize the reuse of the model's learned global feature extraction capabilities and avoid the waste of computing power caused by full parameter updates. Actual verification shows that compared to full parameter updates, this approach is more efficient. The incremental learning strategy reduces computational cost for parameter updates and iterative training, and its minute-level training time is well-suited to the rapid iteration needs of industrial sites. Setting the learning rate for parameter updates to 1 / 10 of the initial training learning rate allows for optimization of model adaptability using new samples while avoiding excessively high learning rates that could cause parameter oscillations and deviations from the already adapted industrial scenario's fundamental characteristics, ensuring model stability after iteration. Finally, the early stopping strategy terminates iteration when the validation set loss does not decrease for three consecutive iteration cycles. This effectively prevents overfitting during iterations with few samples, ensuring the model has good generalization ability for newly added risk scenarios in industrial sites, such as newly added rooftop operation safety detection rules, while avoiding ineffective iterations that consume computational resources, further shortening the iteration cycle. This incremental learning strategy enables the model to quickly adapt to changes in industrial site scenarios while balancing computational efficiency and model stability. It solves the rigidity problem of traditional algorithms being developed once and then fixed, allowing the algorithm to continuously adapt to the dynamically changing safety supervision needs of industrial sites.
[0020] Furthermore, the sensors connected to the perception layer include temperature sensors, smoke sensors, and vibration sensors. The AI intelligent analysis center fuses and analyzes visual detection data with sensor data. The fusion rule is as follows: when visual detection detects "suspected fireworks" and the smoke sensor concentration value is ≥5ppm, or when visual detection detects "abnormal equipment vibration" and the vibration sensor amplitude is ≥0.5g, a level one alarm is triggered. When only a single dimension detects an anomaly, a level two alarm is triggered. This fusion rule combines the scene recognition capability of visual images with the precise numerical detection capability of sensors. By verifying through "visual + sensor" dual dimensions, the accuracy of alarms is improved. Taking fireworks detection as an example, visual recognition of "suspected fireworks" may be misjudged due to factors such as light and obstruction. By combining the numerical condition of smoke sensor concentration ≥5ppm, the false alarm rate of fireworks detection can be reduced by more than 80%. The hierarchical alarm rule, dual-dimensional anomaly detection... The system's ability to accurately differentiate risk levels—commonly triggering Level 1 alarms and single-dimensional anomalies triggering Level 2 alarms—is remarkable. Level 1 alarms quickly activate on-site emergency response mechanisms, while Level 2 alarms prompt staff to conduct further investigations. This avoids the waste of emergency resources or delayed responses to high-risk events caused by indiscriminate alarms in existing technologies. Furthermore, for abnormal equipment vibration scenarios, visual detection identifying "abnormal equipment vibration" combined with a vibration sensor amplitude ≥0.5g quantitative indicator can accurately capture precursors to equipment failure. Compared to detection methods relying solely on sensor values, this approach can exclude vibration anomalies caused by factors other than the equipment itself, such as personnel collisions or falling objects, providing early warnings of equipment failure risks. In addition, this fusion analysis mode adapts to the multi-dimensional safety supervision needs of industrial plants, encompassing "people, machines, and environment," achieving an upgrade from single-indicator detection to multi-dimensional collaborative judgment. This enhances the accuracy and response efficiency of the security system in identifying risks in complex industrial scenarios.
[0021] Furthermore, the position encoding of the Transformer encoder in S2 adopts a two-dimensional position encoding formula customized for industrial scenarios:
[0022] Where (x, y) are the coordinates of a pixel on the feature map, H and W are the height and width of the feature map of the industrial plant monitoring screen, respectively, and i is the position encoding dimension index. The feature dimension of the Transformer model is set to 512. It deeply combines the two-dimensional coordinates (x, y) of the feature map pixels with the actual size features (H, W) of the industrial plant monitoring screen. Compared to traditional location encoding methods that rely on only a single dimension or general size, this method can accurately capture the spatial location features of targets in industrial scenes. For example, it captures the high vertical (y-axis) proportion of rooftop workers and the distribution of aerial work equipment in local areas (specific x / y coordinates), avoiding the loss of location information caused by general encoding ignoring screen size features. Furthermore, the formula calculates the location code by fusing sine / cosine functions of the x and y dimensions, simultaneously representing the spatial relationship of the target in both horizontal and vertical directions. This solves the problem that traditional one-dimensional location encoding cannot fully describe the spatial orientation of targets in industrial scenes. For example, it can accurately distinguish between the different location information of "the left fire lane is blocked" and "the right fire lane is blocked" in the image. In addition, the feature dimension... The value is fixed at 512, which matches the feature extraction dimension of the Transformer architecture, ensuring the efficiency of fusion between location encoding and image features. In complex industrial scenarios, such as multiple overlapping targets and large-scale area monitoring, it can improve the global contextual relevance of target features. Compared with general location encoding, it improves the recognition accuracy of Transformer models for targets in industrial scenarios, and greatly enhances the scene adaptability and detection accuracy of few-shot training algorithms in industrial areas.
[0023] Furthermore, the risk scenario description text received by the natural language interface of the few-shot self-learning algorithm training engine is converted into structured algorithm training parameters by the natural language processing module. These structured parameters include the detection target category, the detection area coordinate range, and the alarm trigger threshold. This transforms unstructured natural language descriptions, such as "warning if security guard at construction site entrance is not wearing a safety helmet," into standardized parameters that can be recognized by machines. Users do not need to possess professional algorithm training knowledge, solving the problem of traditional algorithm customization requiring professional personnel to manually configure parameters and complex operations, significantly improving the ease of use for ordinary maintenance personnel. The detection target category in the structured parameters can accurately pinpoint the core recognition object for algorithm training, and the detection area coordinate range can... By defining the effective analysis area of the algorithm and clarifying the risk assessment criteria through alarm triggering thresholds, the combination of these three elements allows model training to move beyond vague scene descriptions and instead be based on quantified parameters. Compared to the traditional method of training solely through sample labeling, this improves the algorithm's identification specificity and effectively avoids false detections of irrelevant areas and targets. Furthermore, standardized structured parameters can be directly integrated with a multi-algorithm fusion scheduling manager, enabling seamless integration of algorithm customization and resource scheduling. For example, in the scenario of "fire lane obstruction," the converted detection area coordinate range parameters can directly guide the scheduling manager to allocate computing power to the analysis tasks in that area, further improving the system's collaborative operation efficiency and adapting to the diverse and refined safety supervision needs of industrial plants. Attached Figure Description
[0024] Fig. 1 This is a flowchart of the present invention; Fig. 2 This is a flowchart for safety inspection of rooftop operations. Detailed Implementation
[0025] Example 1 As attached Figs. 1-2 As shown, an AI vision-based smart security system for industrial plants includes a perception layer, a computing layer, a platform layer, and a data security protection unit. The perception layer is configured with a data access interface, which supports ONVIF protocol, GB / T28181 and RTSP protocol, and is used to connect to cameras and sensors. The computation layer includes an AI intelligent analysis center, which incorporates a multi-algorithm container, a few-shot self-learning algorithm training engine, and a multi-algorithm fusion scheduling manager. The few-shot self-learning algorithm training engine includes a natural language interactive interface, a sample annotation module, and a localized training unit, which is used to perform model training operations. The multi-algorithm fusion scheduling manager incorporates a preset rule base, which is used to allocate computing resources to a single video stream, enabling multiple algorithm tasks to be executed in parallel on a single video stream. The platform layer is the ASS Intelligent Control Center, which is equipped with a management module for centralized control of system functions. The data security protection unit includes a physical isolation module, an encryption module, an authentication module, and an IP whitelist module. The physical isolation module is used to isolate all data processing units of the system from the external network. The encryption module uses the AES algorithm based on a private protocol to encrypt the external interface of the large model. The authentication module configures usernames and keys. The IP whitelist module controls the IP addresses used to obtain data interfaces.
[0026] A smart security method for industrial plants based on AI vision, implemented using the system described in any one of claims 1-10, includes the following steps: S1. The perception layer collects video streams and sensor data from the industrial plant area and transmits them to the AI intelligent analysis center of the computing layer through standard protocols. S2. Users input a description of the security risk scenario through the platform's natural language interaction interface, upload 30 positive and 30 negative sample images to complete the annotation, trigger the few-shot self-learning algorithm training engine, and complete minute-level model training on the local AI server. The S3 AI Intelligent Analysis Center uses a multi-algorithm fusion scheduling manager to allocate computing resources to a single video stream and execute algorithm tasks completed in multiple rounds of training in parallel. S4. The platform layer displays the algorithm detection results in real time, including event name, coordinate frame, and risk level. When an anomaly is detected, a graded alarm is triggered according to preset rules. S5: All data processing, model training, and algorithm execution are completed within the internal LAN. Data security is protected through AES encryption and IP whitelisting. The model is also regularly optimized through iterative small-sample testing to adapt to changes in the factory environment.
[0027] This solution constructs an AI-based intelligent security system for industrial plants, integrating a perception layer, a computing layer, a platform layer, and a data security protection unit. It also includes standardized security methods. From a system architecture perspective, the perception layer, through access interfaces compatible with ONVIF and GB / T28181 standard protocols, enables the reuse of existing cameras and sensors of different brands and types, overcoming device compatibility barriers. The computing layer, as the core computing power unit, has a built-in few-sample self-learning algorithm training engine that can receive risk scenario descriptions from users through a natural language interface. Relying on a localized training unit, it requires only a small number of samples to complete model training. Simultaneously, a multi-algorithm fusion scheduling manager allocates computing resources to single video streams based on preset rules, enabling parallel execution of multiple algorithm tasks and overcoming the limitations of single-device, single-algorithm operation. The platform layer achieves centralized management and control of all system functions through the ASS intelligent control center's management module, ensuring uniformity in operation and maintenance. The data security protection unit forms a comprehensive data security protection system through a physical isolation module to isolate external networks, an encryption module using the proprietary AES algorithm to protect large model interfaces, an authentication module configuring usernames and keys, and an IP whitelist module controlling data interface access permissions. The supporting security measures bring the system capabilities to fruition. The perception layer collects data, the computing layer completes algorithm training and multi-task execution, the platform layer displays alarms, and the entire process is localized to ensure data security and iteratively optimize the model. This forms a closed-loop operation logic of "data collection - algorithm training - intelligent analysis - security protection - model iteration", which is adapted to the dynamic security needs of industrial plants.
[0028] Compared to existing technologies, which only support IP network camera access with specific protocols and are incompatible with existing multi-brand, multi-protocol devices in the factory, resulting in high costs for retrofitting and wasting initial investment, this solution, through its standardized protocol interface design at the perception layer, can directly connect to existing cameras and sensors in the factory without replacing hardware. Actual verification shows that after adopting this solution, an industrial coating plant reduced equipment retrofit costs by 60% and shortened the retrofit cycle from 15 days to 3 days, significantly reducing hardware investment and construction costs. Furthermore, existing technologies rely on massive amounts of labeled data for algorithm training, requiring months to customize new algorithms, and the algorithms are fixed as "black boxes" and cannot adapt to changing scenarios. This solution's few-sample self-learning algorithm training engine requires only a small number of samples to complete localized model training in minutes. For example, a machining plant addressed the issue of "high-altitude operations without safety harnesses." In the new risk scenario of "full coverage," only 30 sets of positive and negative samples are uploaded, and algorithm training and deployment are completed within 20 minutes. Compared with the customization cycle of existing technologies that takes several months, the efficiency is improved by more than 95%. Moreover, the model can be adjusted and iterated in real time according to the production process of the factory area, solving the rigidity problem of traditional algorithms being developed once and fixed. Existing technologies can only run a single algorithm with a single camera, requiring the deployment of a large number of dedicated cameras to achieve all-round monitoring. The original security system of a chemical plant area required the deployment of 3 dedicated cameras to cover three types of detection needs: smoke detection, personnel crossing the boundary, and fire lane obstruction. This solution realizes the parallel execution of multiple algorithms on a single video stream through a multi-algorithm fusion scheduling manager. The factory area only needs 1 camera to complete the above three types of detection tasks at the same time. The number of hardware deployments is reduced, the hardware procurement and maintenance costs are reduced, and the computing power utilization of a single camera is improved.
[0029] Existing technologies rely on public clouds for algorithm training and data processing, posing a risk of sensitive video streams and production operation data leaking from the factory premises. This solution employs a multi-layered protection system including physical isolation, AES encryption, authentication, and IP whitelisting. All data processing and model training are completed within the internal LAN. After adopting this solution, a military-grade supporting factory eliminated the risk of data leakage during cross-network transmission, meeting the security requirements of high-security industrial scenarios—a core advantage that existing technologies cannot achieve. Existing security systems have fragmented functional modules and lack a unified management interface, requiring maintenance to interface with multiple terminals. This solution, through the centralized management module of the ASS intelligent control center at the platform layer, achieves unified and visualized management and control of algorithm tasks, equipment status, and alarm information.
[0030] The few-shot self-learning algorithm training engine performs model training operations based on a large visual model with the Transformer architecture, specifically including the following steps: S1. Feature Extraction: The input video frames are processed by the convolutional neural network backbone to output a two-dimensional feature map set. S2, Transformer encoder-decoder processing: The encoder performs multi-layer encoding on the feature map in sequence, including dimensionality reduction, flattening, and position encoding, to obtain an image feature sequence containing global context information; the decoder takes the object query vector and the image feature sequence as input, and outputs object instance prediction features with rich embedded information through multi-layer decoding of self-attention mechanism and cross-attention mechanism. S3. Prediction Head and Forward Computation: The embedded features output from the decoder are input into the feedforward network, and the prediction head completes the category prediction and bounding box prediction. S4. Bipartite Graph Matching and Loss Calculation: Constructing a cost matrix and calculating the loss value after solving for the optimal matching completes the model training iteration. Step S1 extracts two-dimensional feature maps of video frames through the convolutional neural network backbone, which can accurately capture the low-level visual features of the image and lay the foundation for subsequent high-level feature analysis. Secondly, the dimensionality reduction, flattening, and positional encoding operations of the Transformer encoder in step S2 can effectively integrate the global contextual information of the image, solving the problem that traditional convolutional neural networks can only extract local features. The decoder, combined with self-attention and cross-attention mechanisms, can focus on key detection targets in industrial scenarios, such as workers on vehicle roofs and people not wearing safety helmets. Even under conditions of few samples, it can generate prediction features rich in target information, greatly reducing the dependence on massive labeled data. Furthermore, the prediction head and forward computation in step S3 can achieve accurate prediction of categories and bounding boxes, adapting to the diverse safety detection needs of industrial plants, such as smoke detection and personnel crossing the boundary. Finally, the bipartite graph matching and loss calculation in step S4, by constructing a cost matrix and solving for the optimal matching, can accurately measure the deviation between the prediction results and the real labels, and achieve efficient iterative optimization of the model. The overall process retains the Transformer architecture's ability to capture global features while adapting to low-sample scenarios through modular training steps. It can complete the training of personalized algorithms for industrial plants in minutes. Compared with traditional convolutional neural network algorithms that rely on massive amounts of data and can only extract local features, it not only shortens the algorithm customization cycle from months to minutes, but also improves the accuracy and generalization ability of target detection in complex industrial scenarios. It can quickly adapt to changes in production processes, new risk points, and other scenario changes, effectively solving the core problems of poor adaptability and high customization costs of industrial security algorithms.
[0031] The cost matrix construction and loss calculation for bipartite graph matching in step S4 adopt the following algorithm formula: Cost matrix formula
[0032] Loss function formula
[0033] in, The cost of matching the i-th real target with the j-th predicted bounding box; The classification loss weight coefficient has a value range of [1,5]. is the bounding box loss weight coefficient, with a value range of [1, 10]; Let cross-entropy be the classification loss function. For the actual target category label, To predict the probability distribution of categories; The CIoU bounding box loss function is... The coordinates of the actual target bounding box. To predict the bounding box coordinates; This represents the total loss value. This is the loss weight coefficient when there is no target matching, and its value ranges from [0.1, 1]. The empty category represents the category with no real target match. The cost matrix formula combines the classification loss and the bounding box loss using weight coefficients. , Weighted fusion allows for flexible adjustment of weights based on the detection priority in industrial scenarios, such as increasing the weights for high-risk scenarios like smoke and fire detection. Improved detection of the boundary frame of workers on the roof Compared to traditional single loss calculation methods, this approach is more suitable for the diverse safety inspection needs of industrial plants; secondly, the loss function incorporates a loss term for situations without target matching. This approach effectively balances the imbalance between positive and negative samples during training with few samples, preventing overfitting or missed detections due to limited sample size and improving the algorithm's generalization ability in long-tail industrial scenarios. Simultaneously, the bounding box loss employs the CIoU function, which considers bounding box overlap, center point distance, and aspect ratio. In complex industrial environments, such as when personnel or equipment occlusion causes target box deformation, bounding box prediction accuracy is improved by over 30%. Furthermore, the formula sets reasonable value ranges for each weight coefficient, ensuring algorithm flexibility while avoiding disordered parameter adjustments. This allows for rapid convergence in minute-level localized training, ultimately achieving accurate prediction of target categories and bounding boxes in industrial scenarios under limited sample conditions. This solves the core problems of traditional algorithms: reliance on massive amounts of data, low detection accuracy, and poor adaptability.
[0034] The dynamic resource allocation algorithm of the multi-algorithm fusion scheduler adopts the following formula: Resource allocation weight formula Single-stream video algorithm execution priority formula
[0035] Resource allocation formula in, Let be the comprehensive weight value of the k-th algorithm task; α is the security level weight coefficient, with a value range of [0.5, 0.8]. β represents the algorithm task safety level, ranging from 1 to 10; β is the resource consumption weight coefficient, ranging from [0.1, 0.3]. γ represents the resource utilization rate of the algorithm task, with a value ranging from 0 to 1; γ is the real-time weight coefficient, with a value ranging from [0.1, 0.2]. To meet the real-time requirements of the algorithm task, the value ranges from 0 to 1; is the execution priority of the k-th algorithm task; n is the total number of algorithm tasks executed in parallel on a single video stream; The amount of computing resources allocated to the k-th algorithm task; The resource allocation weight formula constructs a comprehensive weight based on three core dimensions: security level, resource consumption, and real-time performance, to determine the total computing power resources corresponding to a single video stream. By limiting the reasonable value range of coefficients α, β, and γ, it can be flexibly adjusted according to the safety management priorities of industrial plants. For example, in chemical plants, α can be set to 0.8 to prioritize resource allocation for high-safety-level algorithms such as smoke and fire detection. In machinery plants, the γ coefficient can be appropriately increased to ensure the real-time performance of high-altitude operation detection. Compared to traditional single-dimensional resource allocation methods, this approach better meets the actual needs of industrial scenarios. Furthermore, the execution priority formula quantifies the priority of algorithm tasks through weight proportions, avoiding subjective arbitrariness in resource allocation and ensuring that high-priority algorithms, such as flame detection with a safety level of 10, always receive sufficient computing power. Finally, the resource allocation formula precisely allocates total computing power resources based on priority, ensuring the parallel execution of multiple algorithm tasks on a single video stream while avoiding resource waste. For example, when a single video stream in an industrial plant is simultaneously running three types of algorithms: safety helmet detection (security level 8, resource utilization rate 0.4), smoking detection (security level 6, resource utilization rate 0.3), and fire exit obstruction detection (security level 7, resource utilization rate 0.3), the formula can allocate 70% of the computing power to safety helmet detection, 20% to fire exit obstruction detection, and 10% to smoking detection, achieving optimal allocation of computing power resources. This significantly reduces the hardware deployment cost of the industrial plant security system while ensuring the real-time performance and accuracy of high-risk scene detection.
[0036] The AI intelligent analysis center's built-in dedicated algorithm model library includes a rooftop operation safety belt detection algorithm. This algorithm achieves its goal through deep integration of logical rules and visual recognition: it uses a Transformer architecture model to identify whether "rooftop operation" exists in the image. The identification criteria are that the vertical coordinate of the target detection box occupies ≥0.7 of the image height, and the cosine similarity between the target feature vector and the feature vector of the "rooftop operation" sample is ≥0.85. If the identification result is "yes," the safety belt compliance detection branch is triggered, which determines whether the worker is wearing a safety belt based on the bounding box prediction result. If the identification result is "no," the execution of this algorithm branch is terminated. This algorithm first uses a Transformer architecture model combined with quantitative criteria, namely, the vertical coordinate of the target detection box occupies ≥0.7 and the cosine similarity is ≥0.85, to identify "rooftop operation" behavior, which is indistinguishable from traditional methods. The seatbelt detection method can accurately screen out the work scenarios that need to be detected, avoiding invalid seatbelt detection for non-roof-work personnel, significantly reducing the false alarm rate of the system. Moreover, the seatbelt compliance detection branch is only triggered when "roof-work" behavior is identified; if not identified, the execution of this algorithm branch is terminated, which can reduce unnecessary computing power consumption, improve the algorithm running efficiency of a single video stream, and adapt to the resource scheduling requirements of parallel execution of multiple algorithms. At the same time, the quantified identification judgment conditions make the identification results of "roof-work" behavior quantifiable and verifiable. Compared with traditional identification methods that rely on human experience or fuzzy features, the detection results are more stable and accurate, effectively identifying violations of not wearing seatbelts in roof-work scenarios. This solves the core problems of inaccurate scene identification and redundant alarms in the safety detection of high-altitude operations in industrial plants, improving the pertinence and effectiveness of safety supervision.
[0037] The model iteration optimization of the few-shot self-learning algorithm training engine adopts an incremental learning strategy. After receiving new sample data from the field each time, only the parameters of the prediction head layer and the last two layers of the Transformer decoder are updated, while the encoder layer parameters remain fixed. The learning rate for parameter updates is set to 1 / 10 of the initial training learning rate, and an early stopping strategy is adopted: when the validation set loss does not decrease for three consecutive iterations, the iteration is terminated. This design, which updates only the parameters of the prediction head layer and the last two layers of the Transformer decoder while keeping the encoder layer parameters fixed, maximizes the reuse of the model's learned global feature extraction capabilities and avoids the waste of computational power caused by full parameter updates. Practical verification shows that compared to full parameter updates, this approach is more efficient. The incremental learning strategy reduces computational cost for iterative training by updating parameters, and the minute-level training time is suitable for the rapid iteration needs of industrial sites. Setting the learning rate for parameter updates to 1 / 10 of the initial training learning rate allows for optimization of model adaptability using new samples while avoiding excessively high learning rates that could cause model parameter oscillations and deviations from the already adapted basic characteristics of the industrial scenario, ensuring the stability of the model after iteration. Finally, the early stopping strategy terminates iteration when the validation set loss does not decrease for three consecutive iteration cycles, effectively preventing overfitting during iterations with few samples. This ensures the model has good generalization ability for newly added risk scenarios in industrial sites, such as newly added rooftop operation safety detection rules, while avoiding ineffective iterations that consume computational resources, further shortening the iteration cycle. This incremental learning strategy enables the model to quickly adapt to changes in industrial site scenarios while balancing computational efficiency and model stability, solving the rigidity problem of traditional algorithms that are developed once and remain unchanged. It allows the algorithm to continuously adapt to the dynamically changing safety supervision needs of industrial sites.
[0038] The sensors connected to the perception layer include temperature sensors, smoke sensors, and vibration sensors. The AI intelligent analysis center fuses and analyzes visual detection data with sensor data. The fusion rule is as follows: when visual detection detects "suspected fireworks" and the smoke sensor concentration value is ≥5ppm, or when visual detection detects "abnormal equipment vibration" and the vibration sensor amplitude is ≥0.5g, a level one alarm is triggered. When only a single dimension detects an anomaly, a level two alarm is triggered. This fusion rule combines the scene recognition capability of visual images with the precise numerical detection capability of sensors. By verifying through "visual + sensor" dual dimensions, the accuracy of alarms is improved. Taking fireworks detection as an example, visual recognition of "suspected fireworks" may be misjudged due to factors such as light and obstruction. By combining the numerical condition of smoke sensor concentration ≥5ppm, the false alarm rate of fireworks detection can be reduced by more than 80%. The hierarchical alarm rule triggers alarms for dual-dimensional anomalies. Issuing Level 1 alarms and triggering Level 2 alarms based on single-dimensional anomalies can accurately distinguish risk levels. Level 1 alarms can quickly trigger on-site emergency response mechanisms, while Level 2 alarms prompt staff to review and investigate, avoiding the waste of emergency resources or delayed response to high-risk events caused by indiscriminate alarms in existing technologies. Simultaneously, for abnormal equipment vibration scenarios, visual detection and recognition of "abnormal equipment vibration" behavior combined with the quantitative indicator of vibration sensor amplitude ≥0.5g can accurately capture precursors to equipment failure. Compared to detection methods that rely solely on sensor values, this can exclude vibration anomalies caused by non-equipment components, such as personnel collisions or falling objects, providing early warning of equipment failure risks. Furthermore, this fusion analysis mode adapts to the multi-dimensional safety supervision needs of "people, machines, and environment" in industrial plants, achieving an upgrade from single-indicator detection to multi-dimensional collaborative judgment, improving the accuracy and response efficiency of security systems in identifying risks in complex industrial scenarios.
[0039] The position encoding of the Transformer encoder in S2 adopts a two-dimensional position encoding formula customized for industrial scenarios:
[0040] Where (x, y) are the coordinates of a pixel on the feature map, H and W are the height and width of the feature map of the industrial plant monitoring screen, respectively, and i is the position encoding dimension index. The feature dimension of the Transformer model is set to 512. It deeply combines the two-dimensional coordinates (x, y) of the feature map pixels with the actual size features (H, W) of the industrial plant monitoring screen. Compared to traditional location encoding methods that rely on only a single dimension or general size, this method can accurately capture the spatial location features of targets in industrial scenes. For example, it captures the high vertical (y-axis) proportion of rooftop workers and the distribution of aerial work equipment in local areas (specific x / y coordinates), avoiding the loss of location information caused by general encoding ignoring screen size features. Furthermore, the formula calculates the location code by fusing sine / cosine functions of the x and y dimensions, simultaneously representing the spatial relationship of the target in both horizontal and vertical directions. This solves the problem that traditional one-dimensional location encoding cannot fully describe the spatial orientation of targets in industrial scenes. For example, it can accurately distinguish between the different location information of "the left fire lane is blocked" and "the right fire lane is blocked" in the image. In addition, the feature dimension... The value is fixed at 512, which matches the feature extraction dimension of the Transformer architecture, ensuring the efficiency of fusion between location encoding and image features. In complex industrial scenarios, such as multiple overlapping targets and large-scale area monitoring, it can improve the global contextual relevance of target features. Compared with general location encoding, it improves the recognition accuracy of Transformer models for targets in industrial scenarios, and greatly enhances the scene adaptability and detection accuracy of few-shot training algorithms in industrial areas.
[0041] The risk scenario description text received by the natural language interface of the few-shot self-learning algorithm training engine is converted into structured algorithm training parameters by the natural language processing module. These structured parameters include the detection target category, the detection area coordinate range, and the alarm trigger threshold. This transforms unstructured natural language descriptions, such as "warn the security guard at the construction site entrance for not wearing a safety helmet," into standardized parameters that can be recognized by machines. Users do not need professional algorithm training knowledge, solving the problem of complex manual parameter configuration by professionals required for traditional algorithm customization, and significantly improving the ease of use for ordinary maintenance personnel. The detection target category in the structured parameters can accurately pinpoint the core recognition object for algorithm training, and the detection area coordinate range can limit... The effective analysis area of the algorithm, the alarm trigger threshold, and the clear risk judgment criteria, combined with these three elements, enable model training to move beyond vague scene descriptions and instead be based on quantified parameters. Compared to the traditional method of training solely based on sample labeling, the algorithm's identification targeting is improved, effectively avoiding false detections of irrelevant areas and targets. Furthermore, standardized structured parameters can be directly integrated with a multi-algorithm fusion scheduling manager, achieving seamless integration of algorithm customization and resource scheduling. For example, in the scenario of "fire lane obstruction," the converted detection area coordinate range parameters can directly guide the scheduling manager to allocate computing power to the analysis task in that area, further improving the system's collaborative operation efficiency and adapting to the diverse and refined safety supervision needs of industrial plants.
[0042] To make the objectives, technical solutions, and advantages of this invention clearer, the following provides a detailed illustrative description of the AI vision-based smart security system and method for industrial plants, using a specific industrial plant application scenario. This embodiment takes a smart security renovation project of a large chemical industrial park as an example. The park covers three core areas: a hazardous materials storage area, a production and processing area, and a loading and unloading area. It needs to achieve intelligent detection of four major categories of safety risks: smoke and fire detection, personnel crossing boundaries, compliance of safety belts for rooftop operations, and abnormal equipment vibration. The following details the entire process of system deployment, method execution, and key parameter configuration.
[0043] I. System Deployment and Hardware Configuration (a) Deployment of the perception layer Equipment selection: The sensing layer uses network cameras (4MP resolution, 25fps) compatible with ONVIF, GB / T28181, and RTSP protocols. 12 cameras are deployed in the hazardous materials storage area, 20 in the production and processing area, and 8 in the loading and unloading area, for a total of 40 cameras. Simultaneously, temperature sensors (measurement range -20℃~150℃, accuracy ±0.5℃), smoke sensors (detection range 0~20ppm, response time ≤3s), and vibration sensors (range 0~2g, accuracy ±0.01g) are also connected. Smoke sensors and temperature sensors are deployed in pairs in the hazardous materials storage area (each... One set of vibration sensors is deployed on key equipment such as core reactors and transfer pumps in the production and processing area (30 units in total).
[0044] Interface adaptation and data transmission: Unified access to data from cameras and sensors is achieved through an industrial gateway. The gateway is equipped with a dual-protocol conversion module to convert ONVIF / GB / T28181 protocol data into the internal standard TCP / IP protocol. The data transmission bandwidth is configured to 4Mbps per camera, and sensor data is uploaded at a frequency of 1Hz. All sensing devices are connected to the internal local area network and are physically isolated from the external network.
[0045] (ii) Deployment of the computing layer The AI Intelligent Analysis Center's hardware configuration includes four industrial-grade AI servers (each equipped with two Intel Xeon 8375C processors, eight NVIDIA A100 graphics cards, 128GB of RAM, and 4TB of SSD storage). Two of these servers serve as primary computing nodes, responsible for algorithm training and parallel execution of multiple algorithms, while the other two act as backup nodes to ensure system redundancy. The servers are deployed in a dedicated cabinet in the park's security control room and are equipped with independent UPS power supplies, providing at least four hours of continuous operation during power outages.
[0046] Algorithm Engine and Scheduler Deployment: The AI server is pre-installed with a Docker-based multi-algorithm container environment, which includes a built-in few-shot self-learning algorithm training engine (developed based on the PyTorch framework, supporting Transformer architecture model training) and a multi-algorithm fusion scheduler (based on Kubernetes for computing power scheduling); the park's safety level rules are entered into the preset rule base: smoke and fire detection safety level 10, abnormal equipment vibration level 9, rooftop operation safety belt detection level 8, and personnel crossing the boundary level 7; the resource utilization baseline values are: smoke and fire detection 0.4, abnormal equipment vibration 0.35, rooftop operation safety belt detection 0.3, and personnel crossing the boundary 0.25; the real-time requirements are: smoke and fire detection 1.0, abnormal equipment vibration 0.9, rooftop operation safety belt detection 0.8, and personnel crossing the boundary 0.7.
[0047] (III) Platform Layer Deployment ASS Intelligent Control Center Setup: The ASS Intelligent Control Center is built using a B / S architecture, with two application servers configured for dual-machine hot standby (Intel Xeon 6348 processor, 64GB memory, 1TB storage). The front-end uses the Vue.js framework to develop a visual interface, supporting four core functions: device status monitoring, algorithm task management, alarm information display, and model training configuration. The back-end uses the Spring Boot framework to connect to the algorithm interface of the AI server in the computing layer and the device data interface in the perception layer.
[0048] Interactive interface configuration: The natural language interactive interface supports both voice input (recognition accuracy ≥95%) and text input. It can parse natural language commands such as "smoke detection in hazardous materials storage area" and "safety belt detection for rooftop operations in loading and unloading area" and automatically convert them into structured parameters.
[0049] (iv) Deployment of Data Security Protection Unit Physical isolation: All devices in the perception layer, computing layer, and platform layer are deployed on an independent security intranet, completely isolated from the campus office network and the Internet through a physical firewall. Only one physical access port is reserved for operation and maintenance, and access is only possible with dual authorization of two people and two keys.
[0050] Encryption and Access Control Configuration: The encryption module uses the AES-256 algorithm based on a proprietary protocol to encrypt the external interfaces of the AI intelligent analysis center (such as the communication interface with the ASS intelligent control center); the authentication module configures hierarchical usernames and keys (administrator level, operator level, viewer level) for different operation and maintenance roles, and the keys are forcibly updated every 90 days; the IP whitelist module only allows the fixed IP range (192.168.10.0 / 24) of the security control room to access the data interface, and all other IPs are denied access.
[0051] II. Security Method Implementation Process Step 1: Data Acquisition at the Perception Layer The cameras in the perception layer acquire video streams from various areas in real time. Each frame of video is preprocessed using a convolutional neural network backbone (using ResNet50 as the backbone network) to extract keyframes (one frame is extracted from every two frames). Sensors collect data at preset frequencies: the temperature sensor uploads a temperature value every 10 seconds, the smoke sensor uploads a concentration value every 5 seconds, and the vibration sensor uploads an amplitude value every 2 seconds. All video streams and sensor data are transmitted to the AI intelligent analysis center through an industrial gateway. During transmission, the data is encapsulated using a proprietary protocol to prevent data tampering.
[0052] Step 2: Training the model with few samples Taking the training of the algorithm for "safety belt detection on the roof of the loading and unloading area" as an example, the specific execution is as follows: Scenario Description and Sample Upload: Maintenance personnel input "Alarm triggered when rooftop workers in loading / unloading area are not wearing seat belts" through the natural language interaction interface of the ASS Intelligent Control Center. The system automatically converts this description into structured parameters: the detection target category is "rooftop workers + seat belt", the detection area coordinate range is (x:100~1920, y:200~1080) of the camera image in the loading / unloading area, and the alarm trigger threshold is "Alarm triggered if no seat belt bounding box is detected". Then, 30 positive samples (rooftop workers wearing seat belts) and 30 negative samples (rooftop workers not wearing seat belts / non-rooftop workers) are uploaded. The sample annotation module automatically annotates the bounding boxes of the samples (annotation accuracy ±5 pixels).
[0053] Localized Model Training: The few-shot self-learning algorithm training engine initiates localized training. Based on the Transformer architecture, the large visual model performs the following steps: S1 Feature Extraction: The input sample images (with a uniform resolution adjusted to 1024×1024) are processed by the ResNet50 backbone network to output a 64×64×2048 two-dimensional feature map set; S2Transformer encoding / decoding processing: Encoder processing: The feature map is reduced in dimensionality (from 2048 to 512) and flattened (converted to a 64×64×512 one-dimensional sequence). The position code is calculated using a customized two-dimensional position encoding formula for industrial scenarios (H=64, W=64, d_model=512, i takes values from 0 to 255). The encoding formula is as follows: After encoding by a 6-layer encoder, an image feature sequence containing global context information is obtained; Decoder processing: Input 100 object query vectors (the number is adapted to the target density in the industrial scene), and perform 6 layers of decoding through self-attention mechanism (8 heads) and cross-attention mechanism (8 heads), and output object instance prediction features that embed target location and category information; S3 Prediction Head and Forward Computation: The 512-dimensional embedded feature output from the decoder is input into the feedforward network (hidden layer dimension 2048). The prediction head adopts a two-branch structure of classification branch + bounding box regression branch. The classification branch outputs the probability distribution of three classes: "roof worker", "seat belt" and "background". The bounding box branch outputs the (x,y,ω,h) coordinates of the target bounding box. S4 bipartite graph matching and loss calculation: Construct the cost matrix: Set (Classification loss weights) (Bounding box loss weights) Using the cross-entropy loss function, The CIoU loss function is used to calculate the matching cost between each ground truth target and the predicted bounding box: Loss function calculation: settings (Without target matching loss weights), the total loss function is: Training iterations: The initial learning rate was set to 1e-4, the batch size to 8, and the number of training iterations to 50 rounds. The training process was executed on a local AI server and took about 18 minutes to complete the training of the "rooftop operation safety belt detection" algorithm model.
[0054] Step 3: Parallel execution of multiple algorithms The AI intelligent analysis center's multi-algorithm fusion scheduling manager allocates computing resources to single video streams in the loading and unloading area. Three types of algorithm tasks—"safety belt detection on the roof," "personnel crossing the boundary," and "smoke and fire detection"—need to be executed in parallel. The specific resource allocation is as follows: Overall weight calculation: α=0.7 (security level weight), β=0.2 (resource consumption weight), γ=0.1 (real-time performance weight). The algorithm task parameters and weight calculation results are shown in the table below:
[0055] Priority calculation: Total weight sum = 7.18 + 5.82 + 5.02 = 18.02, and the priority of each algorithm is as follows: Fireworks detection: 7.18 / 18.02 × 100% ≈ 39.84% Safety belt test results for rooftop work: 5.82 / 18.02 × 100% ≈ 32.30% Personnel crossing the boundary: 5.02 / 18.02×100%≈27.86% Computing resource allocation: The total computing resources for a single video stream are 32 TOPS, and the allocation result is as follows: Fireworks detection: 39.84% × 32 ≈ 12.75 TOPS Safety belt test results for rooftop work: 32.30% × 32 ≈ 10.34 TOPS Personnel crossing the boundary: 27.86% × 32 ≈ 8.91 TOPS Parallel Algorithm Execution: According to the above computing power allocation results, the scheduler starts three types of algorithms to be executed in parallel on a single video stream. The processing time of each frame of video is ≤40ms, which meets the real-time requirement of 25fps. When the "Roof Operation Seatbelt Detection" algorithm is executed, it first uses the Transformer model to identify targets in the image whose vertical coordinate ratio of the target detection box is ≥0.7 and whose cosine similarity with the feature vector of the "Roof Operation" sample is ≥0.85. After determining that it is a "Roof Operation" behavior, the seatbelt compliance detection branch is triggered. If no seatbelt is detected, the branch is terminated, and only the personnel crossing the boundary and smoke detection algorithms are retained.
[0056] Step 4: Tiered Alarms and Result Display Alarm rule configuration: Level 1 alarm trigger conditions: Visual detection of "suspected smoke" and smoke sensor concentration ≥ 5ppm; or visual detection of "abnormal equipment vibration" and vibration sensor amplitude ≥ 0.5g; Level 2 alarm trigger conditions: visual detection of "suspected smoke / fire" / "abnormal equipment vibration" / "unsafety belt on rooftop" / "personnel crossing the boundary"; or sensor detection of smoke concentration ≥ 5 ppm / vibration amplitude ≥ 0.5 g.
[0057] Alarm Display and Response: The ASS Intelligent Control Center displays the algorithm detection results in real time, including the event name (e.g., "Working on the roof without a seatbelt"), target bounding box coordinates (e.g., x:500, y:600, ω:100, h:200), and risk level (Level 1 / Level 2). When a Level 1 alarm is triggered, the system automatically activates the on-site audible and visual alarms (loudness ≥110dB), pushes an SMS to the park safety administrator (delivered within 5 seconds), and marks the alarm location on the electronic map. When a Level 2 alarm is triggered, a pop-up notification is only displayed in the ASS Intelligent Control Center, requiring operator review and processing. Example: When the camera in the hazardous materials storage area detects "suspected smoke" (visual recognition confidence level 98%), and the smoke sensor in the corresponding area uploads a concentration value of 6.2ppm, the system triggers a Level 1 alarm. The audible and visual alarms in the control room are activated, the administrator receives an SMS message "Smoke risk in the northeast corner of the hazardous materials storage area, Level 1 alarm," the electronic map locates the area, and maintenance personnel arrive on-site for verification within 3 minutes.
[0058] Step 5: Localized security protection and model iterative optimization Data security protection: All video streams, sensor data, and model training data are stored locally on the AI server within the security intranet and are not transmitted to external networks; access to the ASS intelligent control center requires triple verification of username + key + IP whitelist. The key uses a combination of "letters + numbers + special symbols" and is ≥16 characters long; the external interface of the AI intelligent analysis center is encrypted using the AES-256 algorithm, and data transmission packets are equipped with timestamps and checksums to prevent tampering.
[0059] Model Iterative Optimization: After running for one month, collect new samples from on-site feedback (such as 20 negative samples of "workers on the roof are not wearing seat belts because they are obstructed by equipment" and 20 positive samples of "wearing seat belts in compliance with regulations"), and initiate an incremental learning strategy: Parameter update scope: Only the parameters of the model prediction head layer and the last two layers of the Transformer decoder are updated; the encoder layer parameters are fixed. Learning rate settings: The initial training learning rate is 1e-4, and the iterative learning rate is adjusted to 1e-5; Early stopping strategy: If the validation set loss does not decrease for three consecutive iterations, the iteration is terminated. This iteration was executed for a total of 20 rounds, taking about 10 minutes. After the iteration, the model's accuracy in detecting seat belts in occluded scenarios increased from 85% to 94%.
[0060] The above are merely embodiments of the present invention. Commonly known structures and characteristics of the solutions are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the description of specific embodiments in the specification can be used to interpret the content of the claims.
Claims
1. An AI vision-based intelligent security system for industrial plants, characterized in that, It includes a perception layer, a computing layer, a platform layer, and a data security protection unit; The perception layer is configured with a data access interface, which supports ONVIF protocol, GB / T28181 and RTSP protocol, and is used to connect to cameras and sensors. The computation layer includes an AI intelligent analysis center, which has a built-in multi-algorithm container, a few-shot self-learning algorithm training engine, and a multi-algorithm fusion scheduling manager. The few-shot self-learning algorithm training engine includes a natural language interactive interface, a sample annotation module, and a localized training unit, which is used to perform model training operations. The multi-algorithm fusion scheduling manager has a built-in preset rule library for allocating computing resources to a single video stream, enabling multiple algorithm tasks to be executed in parallel on a single video stream. The platform layer is the ASS Intelligent Control Center, which is equipped with a management module for centralized control of system functions. The data security protection unit includes a physical isolation module, an encryption module, an authentication module, and an IP whitelist module. The physical isolation module is used to isolate all data processing units of the system from the external network. The encryption module uses the AES algorithm based on a private protocol to encrypt the external interface of the large model. The authentication module configures usernames and keys. The IP whitelist module controls the IP addresses used to obtain data interfaces.
2. The intelligent security system and method for industrial plants based on AI vision according to claim 1, characterized in that, The few-shot self-learning algorithm training engine performs model training operations based on a large visual model with the Transformer architecture, specifically including the following steps: S1. Feature Extraction: The input video frames are processed by the convolutional neural network backbone to output a two-dimensional feature map set. S2, Transformer encoder-decoder processing: The encoder performs multi-layer encoding on the feature map in sequence, including dimensionality reduction, flattening, and position encoding, to obtain an image feature sequence containing global context information; the decoder takes the object query vector and the image feature sequence as input, and outputs object instance prediction features with rich embedded information through multi-layer decoding of self-attention mechanism and cross-attention mechanism. S3. Prediction Head and Forward Computation: The embedded features output from the decoder are input into the feedforward network, and the prediction head completes the category prediction and bounding box prediction. S4. Bipartite Graph Matching and Loss Calculation: Construct the cost matrix, calculate the loss value after solving for the optimal matching, and complete the model training iteration.
3. The intelligent security system and method for industrial plants based on AI vision according to claim 1, characterized in that, According to the system of claim 2, the cost matrix construction and loss calculation for bipartite graph matching in step S4 adopt the following algorithm formula: Cost matrix formula Loss function formula in, The cost of matching the i-th real target with the j-th predicted bounding box; The classification loss weight coefficient has a value range of [1,5]. is the bounding box loss weight coefficient, with a value range of [1, 10]; Let cross-entropy be the classification loss function. For the actual target category label, To predict the probability distribution of categories; The CIoU bounding box loss function is... The coordinates of the actual target bounding box. To predict the bounding box coordinates; This represents the total loss value. This is the loss weight coefficient when there is no target matching, and its value ranges from [0.1, 1]. This indicates an empty category with no real target match.
4. The intelligent security system and method for industrial plants based on AI vision according to claim 1, characterized in that, The dynamic resource allocation algorithm of the multi-algorithm fusion scheduler adopts the following formula: Resource allocation weight formula Single-stream video algorithm execution priority formula Resource allocation formula in, Let be the comprehensive weight value of the k-th algorithm task; α is the security level weight coefficient, with a value range of [0.5, 0.8]. β represents the algorithm task safety level, ranging from 1 to 10; β is the resource consumption weight coefficient, ranging from [0.1, 0.3]. γ represents the resource utilization rate of the algorithm task, with a value ranging from 0 to 1; γ is the real-time weight coefficient, with a value ranging from [0.1, 0.2]. To meet the real-time requirements of the algorithm task, the value ranges from 0 to 1; is the execution priority of the k-th algorithm task; n is the total number of algorithm tasks executed in parallel on a single video stream; The amount of computing resources allocated to the k-th algorithm task; This represents the total computing resources required for a single video stream.
5. The intelligent security system and method for industrial plants based on AI vision according to claim 1, characterized in that, The AI intelligent analysis center's built-in dedicated algorithm model library includes a rooftop operation safety belt detection algorithm. This algorithm achieves deep integration of logical rules and visual recognition: it uses a Transformer architecture model to identify whether "rooftop operation" exists in the image. The identification criteria are that the vertical coordinate of the target detection box occupies ≥0.7 of the image height, and the cosine similarity between the target feature vector and the feature vector of the "rooftop operation" sample is ≥0.
85. If the identification result is "yes", the safety belt compliance detection branch is triggered, and the bounding box prediction result is used to determine whether the worker is wearing a safety belt. If the identification result is "no", the execution of this algorithm branch is terminated.
6. The intelligent security system and method for industrial plants based on AI vision according to claim 1, characterized in that, The model iteration optimization of the few-shot self-learning algorithm training engine adopts an incremental learning strategy. After receiving new sample data from the field each time, only the parameters of the prediction head layer and the last two layers of the Transformer decoder are updated, while the encoder layer parameters are kept fixed. The learning rate for parameter updates is set to 1 / 10 of the initial training learning rate, and an early stopping strategy is adopted: the iteration is terminated when the validation set loss does not decrease for three consecutive iteration cycles.
7. The intelligent security system and method for industrial plants based on AI vision according to claim 1, characterized in that, The sensors connected to the perception layer include temperature sensors, smoke sensors, and vibration sensors. The AI intelligent analysis center fuses and analyzes visual detection data with sensor data. The fusion rules are as follows: when the visual system detects "suspected smoke" and the smoke sensor concentration value is ≥5ppm, or when the visual system detects "abnormal equipment vibration" and the vibration sensor amplitude is ≥0.5g, a level one alarm is triggered; when an abnormality is detected only in a single dimension, a level two alarm is triggered.
8. The intelligent security system and method for industrial plants based on AI vision according to claim 2, characterized in that, The position encoding of the Transformer encoder in S2 adopts a two-dimensional position encoding formula customized for industrial scenarios: Where (x, y) are the coordinates of a pixel on the feature map, H and W are the height and width of the feature map of the industrial plant monitoring screen, respectively, and i is the position encoding dimension index. This represents the feature dimension of the Transformer model, with a value of 512.
9. The intelligent security system and method for industrial plants based on AI vision according to claim 1, characterized in that, The risk scenario description text received by the natural language interactive interface of the few-shot self-learning algorithm training engine is converted into structured algorithm training parameters by the natural language processing module. The structured parameters include the detection target category, the detection area coordinate range, and the alarm trigger threshold.
10. A smart security method for industrial plants based on AI vision, characterized in that, The system implementation according to any one of claims 1-10 includes the following steps: S1. The perception layer collects video streams and sensor data from the industrial plant area and transmits them to the AI intelligent analysis center of the computing layer through standard protocols. S2. Users input a description of the security risk scenario through the platform's natural language interaction interface, upload 30 positive and 30 negative sample images to complete the annotation, trigger the few-shot self-learning algorithm training engine, and complete minute-level model training on the local AI server. The S3 AI Intelligent Analysis Center uses a multi-algorithm fusion scheduling manager to allocate computing resources to a single video stream and execute algorithm tasks completed in multiple rounds of training in parallel. S4. The platform layer displays the algorithm detection results in real time, including event name, coordinate frame, and risk level. When an anomaly is detected, a graded alarm is triggered according to preset rules. S5: All data processing, model training, and algorithm execution are completed within the internal LAN. Data security is protected through AES encryption and IP whitelisting. The model is also regularly optimized through iterative small-sample testing to adapt to changes in the factory environment.
Citation Information
Patent Citations
Safety production supervision system and method based on artificial intelligence and big data analysis
CN120494507A