Factory traffic AI video monitoring method, system, equipment and medium

By improving the Mask-RCNN model and combining 5G edge computing and network slicing technology, the identification accuracy and real-time nature of the factory traffic monitoring system in complex scenarios is solved, accurate identification and real-time alarm for specific targets are achieved, and the intelligent level of factory traffic management is improved.

CN120495962APending Publication Date: 2025-08-15CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510669312.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing factory traffic monitoring system has poor adaptability and single functions in complex scenarios, and cannot accurately identify specific targets. It lacks real-time and stability, making it difficult to meet the needs of intelligent traffic management.

Method used

Adopting the improved Mask-RCNN model, an attention mechanism and a customized target classification system are introduced, combined with 5G edge computing and network slicing technology, the hardware resource allocation and deployment redundant architecture are optimized, and high-precision target recognition and real-time decision-making are achieved.

Benefits of technology

It improves the intelligence level of factory traffic management, ensures traffic safety and logistics efficiency, achieves accurate identification and real-time alarms of specific goals, and improves the stability and real-time nature of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495962A_ABST
    Figure CN120495962A_ABST
Patent Text Reader

Abstract

The invention provides a factory traffic AI video monitoring method and system, electronic equipment and a storage medium, and aims to solve the problems that existing video monitoring is poor in complex scene adaptability and single in function, and the method comprises the following steps: collecting factory traffic data, and carrying out data preprocessing and labeling to obtain a data set; a Mask-RCNN model is selected as a vehicle detection and segmentation network, an attention mechanism is introduced into the model, a target classification system and a recognition strategy are customized, a TensorFlow deep learning framework is utilized to train the model through a data set, and the training process is optimized and adjusted; performing target detection and segmentation on a to-be-detected factory traffic video by using the trained model, and automatically identifying a target; and data processing and analysis are carried out on an identification result, so that intelligent decision-making of factory traffic monitoring is realized. According to the invention, the intelligent level of factory traffic management is improved, and the factory traffic safety and logistics efficiency are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of target recognition technology, and in particular to an AI video monitoring method for factory traffic, an AI video monitoring system for factory traffic, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the continuous development of industrial production, factory traffic management faces many challenges, such as large vehicle and personnel flows, numerous traffic safety hazards, and limited logistics efficiency. Traditional video surveillance systems mainly rely on manual monitoring and simple image analysis, which is difficult to meet the needs of modern factories for intelligent and efficient traffic management. In recent years, with the development of artificial intelligence technology, especially the application of deep learning algorithms in image recognition and analysis, AI video surveillance technology has gradually become an effective means to solve factory traffic problems. However, existing factory traffic monitoring systems mostly rely on traditional video analysis or basic deep learning models (such as YOLO and Faster R-CNN), which have the following limitations:

[0003] (1) Poor adaptability to complex scenarios: The recognition accuracy of traditional algorithms drops significantly in complex factory scenarios such as occlusion, lighting changes, and dense traffic.

[0004] (2) Single function: Existing systems mostly focus on target detection and lack support for behavior analysis, rule matching, and multimodal data fusion.

[0005] Therefore, a new video surveillance solution is urgently needed. Summary of the Invention

[0006] To at least address the existing issues of poor adaptability to complex scenarios and single functionality, the present disclosure provides an AI video surveillance method for factory traffic, an AI video surveillance system for factory traffic, an electronic device, and a computer-readable storage medium. These methods improve the target recognition algorithm, enhance the system's real-time performance and stability, and utilize customized traffic behavior analysis capabilities to enhance the intelligent level of factory traffic management, ensuring factory traffic safety and logistics efficiency.

[0007] In a first aspect, the present disclosure provides a method for AI video monitoring of factory traffic, the method comprising:

[0008] By deploying collection equipment in key areas of the factory, we collect factory traffic data, pre-process and label the data, and generate a data set.

[0009] We selected the Mask-RCNN model as the network for vehicle detection and segmentation. We introduced an attention mechanism into the Mask-RCNN model, customized the target classification system and recognition strategy, trained the model on the dataset using the TensorFlow deep learning framework, and optimized and adjusted the training process.

[0010] Use the trained Mask-RCNN model to detect and segment objects in the factory traffic video to be inspected, and automatically identify the objects;

[0011] The identification results are processed and analyzed to realize intelligent decision-making for factory traffic monitoring.

[0012] Furthermore, the introduction of the attention mechanism into the Mask-RCNN model includes:

[0013] Add a CBAM (Convolutional Block Attention Module) after the convolutional layer of each ResNet (Residual Network) residual block in the Mask-RCNN model to adjust the channel and spatial attention of the output feature map;

[0014] The CBAM module is inserted between the feature fusion layers of FPN (Feature Pyramid Networks) to enhance the expression capability of multi-scale features.

[0015] Furthermore, the customized target classification system and identification strategy include:

[0016] Determine the target classification system and identification strategy based on the factory traffic rules and actual needs;

[0017] Add custom categories to the classification branch of Mask-RCNN based on the target classification system and recognition strategy, and perform hierarchical classification design;

[0018] For hazardous materials signs and helmet areas, local feature learning is strengthened in the feature extraction stage, and target detection and attribute recognition are jointly trained through multi-task learning.

[0019] Furthermore, the method further comprises:

[0020] The factory traffic video to be inspected, collected by the acquisition equipment, is transmitted to the back-end MEP (Mobile Edge Platform) server. During the transmission process, network slicing technology is used to ensure the network performance of video data transmission.

[0021] Build an AI computing resource pool on the MEP server to rationally allocate computing resources.

[0022] Furthermore, the method further comprises:

[0023] The real-time data of various types of equipment in the factory area are integrated through the GIS (Geographic Information System) map as the factory area traffic video data to be detected.

[0024] Furthermore, the method further comprises:

[0025] If an abnormal event is detected after data processing and analysis, a real-time alarm is triggered and the event details are automatically recorded.

[0026] Furthermore, the method further comprises:

[0027] Regularly poll the device video signal to detect interruption, blur or blockage faults;

[0028] If abnormal data is found, it will be marked and the fault type will be analyzed. A percentage chart will be generated and an alert will be pushed to the operation and maintenance personnel.

[0029] In a second aspect, the present disclosure provides an AI video surveillance system for factory traffic, the system comprising:

[0030] The collection and processing module is configured to collect factory traffic data through collection devices deployed in key areas of the factory, and perform data preprocessing and annotation to obtain a data set;

[0031] An optimization and training module, which is configured to select the Mask-RCNN model as the network for vehicle detection and segmentation, introduce an attention mechanism into the Mask-RCNN model, customize the target classification system and recognition strategy, use the TensorFlow deep learning framework to train the model on the dataset, and optimize and adjust the training process;

[0032] The detection module is configured to use the trained Mask-RCNN model to detect and segment objects in the factory traffic video to be inspected, and automatically identify the objects;

[0033] The analysis module is configured to process and analyze the identification results to realize intelligent decision-making for factory traffic monitoring.

[0034] In a third aspect, the present disclosure provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program. When the processor runs the computer program stored in the memory, the processor executes the factory traffic AI video monitoring method as described in any one of the first aspects.

[0035] In a fourth aspect, the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the factory traffic AI video monitoring method described in any one of the first aspects above is implemented.

[0036] Beneficial effects:

[0037] This disclosure provides an AI-powered video surveillance method, system, electronic device, and storage medium for factory traffic. By improving the Mask-RCNN model and introducing an attention mechanism, the method improves target recognition accuracy and segmentation in complex factory scenarios. Customized target classification systems and recognition strategies enable more accurate identification of specific types of vehicles and pedestrians. This enhances the system's real-time performance and stability, improves the intelligence of factory traffic management, and ensures factory traffic safety and logistics efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 A flowchart of an AI video monitoring method for factory traffic provided in the first embodiment of the present disclosure;

[0039] Figure 2 This is an architecture diagram of a factory traffic AI video surveillance system based on Mask-RCNN provided by an embodiment of the present disclosure;

[0040] Figure 3 A schematic diagram of integrating multiple types of devices through a GIS map provided by an embodiment of the present disclosure;

[0041] Figure 4 A schematic diagram of a video monitoring-real-time preview process provided by an embodiment of the present disclosure;

[0042] Figure 5 A schematic diagram of an emergency response process provided by an embodiment of the present disclosure;

[0043] Figure 6 A schematic diagram of a video quality diagnosis process provided by an embodiment of the present disclosure;

[0044] Figure 7 A schematic diagram of an information publishing process provided by an embodiment of the present disclosure;

[0045] Figure 8 A schematic diagram of a user rights management process provided by an embodiment of the present disclosure;

[0046] Figure 9 A schematic diagram of the Mask-RCNN model structure provided in an embodiment of the present disclosure;

[0047] Figure 10 A schematic diagram of an application of a system provided by an embodiment of the present disclosure in a practical environment;

[0048] Figure 11 This is an architecture diagram of an AI video surveillance system for factory traffic provided in Example 3 of the present disclosure;

[0049] Figure 12 This is an architectural diagram of an electronic device provided in Example 4 of the present disclosure. DETAILED DESCRIPTION

[0050] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments and drawings described herein are only used to explain the present disclosure, rather than to limit the present disclosure.

[0051] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence; and, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be arbitrarily combined with each other.

[0052] The terms used in the embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit the present disclosure. The singular forms "a," "an," "the," and "the" used in the embodiments of the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0053] In the subsequent description, suffixes such as "module," "component," or "unit" used to represent elements are used only to facilitate the description of the present disclosure and have no specific meaning. Therefore, "module," "component," or "unit" may be used interchangeably.

[0054] Existing factory traffic monitoring systems mostly rely on traditional video analysis or basic deep learning models, and still have the following problems:

[0055] (1) The model has insufficient generalization ability and is not adaptable to special scenarios. It cannot accurately identify specific targets in the factory area (such as dangerous goods transport vehicles and people not wearing safety helmets). In a complex factory environment, the types of vehicles and people are diverse, the background is complex and there are occlusions, resulting in the target recognition accuracy of existing technologies being difficult to meet actual needs.

[0056] (2) The cloud processing mode results in poor real-time performance and is difficult to meet the factory's immediate alarm needs.

[0057] (3) Lack of deep integration with 5G network slicing and edge computing, unable to guarantee the efficiency of high-concurrency video stream processing.

[0058] (4) Insufficient real-time performance and stability of the system: The existing system has a slow response speed when accessing high-concurrency video streams and processing large amounts of data, and is prone to stability problems during long-term operation, affecting the actual application effect.

[0059] (5) Lack of customized analysis of factory traffic rules: Existing technologies fail to fully integrate factory traffic rules and actual needs, and the analysis of traffic behavior is not accurate enough to effectively assist factory traffic management.

[0060] To address the above shortcomings, the present invention aims to provide an AI video surveillance method for factory traffic, based on an improved Mask-RCNN model, to achieve:

[0061] (1) By improving the Mask-RCNN model, the target recognition accuracy and segmentation effect in complex factory scenes are improved.

[0062] (2) Combine 5G edge computing and network slicing technology to achieve low-latency, highly reliable data processing and decision feedback.

[0063] (3) Build a multimodal analysis framework to support vehicle behavior rule matching, real-time warning of safety hazards, and localized data storage.

[0064] This paper improves the target recognition algorithm, enhances the real-time performance and stability of the system, customizes traffic behavior analysis and other functions, improves the intelligence level of factory traffic management, and ensures factory traffic safety and logistics efficiency.

[0065] The relevant technical terms and contents appearing in this disclosure are explained as follows:

[0066] Mask RCNN follows the idea of Faster RCNN and uses the feature extraction

[0067] The Mask RCNN architecture is based on the ResNet-FPN architecture, with an additional mask prediction branch for binary mask prediction. It not only detects objects in images but also provides high-quality segmentation results for each object. It can also be extended to other tasks such as keypoint detection. Mask RCNN has achieved state-of-the-art results in a range of challenging COCO tasks, such as object detection, instance segmentation, and human keypoint detection, with excellent performance metrics.

[0068] TensorFlow is widely used to build and train neural networks. It performs calculations through data flow graphs, where nodes represent mathematical operations and edges represent multidimensional data arrays (i.e., tensors). This design enables TensorFlow to perform efficient calculations on a variety of platforms. Training the Mask-RCNN model using the TensorFlow framework mainly involves: installing necessary dependencies and configuring the environment, converting labeled datasets (such as COCO format) into TFRecord format and enhancing it; adjusting the configuration (number of categories, optimization strategy) based on the pre-trained model, building the input pipeline and defining the loss function; starting training and monitoring metrics (such as loss, mAP (mean Average Precision), and optimizing performance through data enhancement, learning rate scheduling, and model structure adjustment; and finally exporting the trained model and deploying it to a server or edge device for efficient object detection and instance segmentation.

[0069] The following is a detailed description of the technical solutions of the present invention and how the technical solutions of the present invention solve the technical problems in the prior art with specific embodiments. It will be appreciated that, in the embodiments of the present application, the execution subject may perform some or all of the steps in the embodiments of the present application, and these steps or operations are merely examples. The embodiments of the present application may also perform other operations or variations of various operations. In addition, the various steps may be performed in different orders as presented in the embodiments of the present application, and it may not be necessary to perform all the operations in the embodiments of the present application. Furthermore, the following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in certain embodiments.

[0070] Figure 1 This is a flow chart of a factory traffic AI video monitoring method provided in the first embodiment of the present disclosure, such as Figure 1 As shown, the method includes:

[0071] Step S101: Collect factory traffic data through collection equipment deployed in key areas of the factory, and perform data preprocessing and annotation to obtain a data set;

[0072] Step S102: Selecting a Mask-RCNN model as the network for vehicle detection and segmentation, introducing an attention mechanism into the Mask-RCNN model, customizing the target classification system and recognition strategy, using the TensorFlow deep learning framework to train the model using the dataset, and optimizing and adjusting the training process;

[0073] Step S103: Use the trained Mask-RCNN model to perform target detection and segmentation on the factory traffic video to be detected, and automatically identify the target;

[0074] Step S104: Process and analyze the identification results to achieve intelligent decision-making for factory traffic monitoring.

[0075] The factory traffic AI video monitoring method of the disclosed embodiment is implemented based on the factory traffic AI video monitoring system, and the system includes a front-end video acquisition module, a back-end server, an AI analysis module, a user interface, and a data storage module. The front-end video acquisition module collects factory traffic video data through a high-definition camera and transmits it to the back-end server. The back-end server is equipped with a high-performance computing device and a GPU (Graphics Processing Unit) accelerator card to support the efficient operation of the AI algorithm. The AI analysis module performs target detection and segmentation on the video data based on the improved Mask-RCNN model, and performs behavior analysis in combination with the factory traffic rules. The user interface provides a friendly visual interface, which is convenient for managers to view the monitoring video and analysis results in real time. The data storage module is used to store video data and analysis results, and ensure the security and reliability of the data.

[0076] Figure 2 This is the system architecture diagram, showing the connection relationship and data flow of the front-end video acquisition module, back-end server, AI analysis module, user interface and data storage module. Architecture Description:

[0077] (1) Front-end user interface

[0078] Based on web or desktop applications, it provides interactive features such as GIS map display, real-time video preview, device management, and emergency command. Technology stack: Vue.js / React (web), Electron (desktop).

[0079] (2) Backend service layer

[0080] Handles business logic, including video streaming distribution, AI analysis, device status monitoring, and data storage and querying. Technology stack: Python Flask / Django, RTSP (Real Time Streaming Protocol) processing, and RESTful API (Application Programming Interface).

[0081] (3) Database and Storage

[0082] Storage device information, user permissions, plan configuration, event records, video metadata, etc. Technology stack: MySQL (relational data), MongoDB (unstructured data), MinIO (video storage).

[0083] (4) AI model processing layer

[0084] Traffic object detection (vehicles, pedestrians) and abnormal event recognition based on Mask-RCNN. Technology stack: TensorFlow / PyTorch, OpenCV (video frame processing), GPU acceleration.

[0085] (5) Hardware device interface layer

[0086] Integrates cameras (box cameras, dome cameras), weather stations, and communication stations, acquiring data via ONVIF / RTSP protocols. Supports real-time monitoring and alarming of device status.

[0087] The specific implementation process of the method includes:

[0088] 1. Large amounts of factory traffic data were collected through data collection equipment deployed in key areas (such as intersections, loading and unloading areas, and pedestrian walkways). This data was then preprocessed and labeled to prepare for subsequent model training. During the data preparation phase, a large amount of image and video data of factory traffic scenes was collected, including targets such as different types of vehicles (such as trucks, forklifts, and cars), pedestrians, and common obstacles. This data was carefully annotated with information such as target category, location, bounding box, and instance segmentation mask, providing accurate supervision information for model training.

[0089] 2. Model Training: The Mask-RCNN model was selected as the network for vehicle detection and segmentation. The model was trained using deep learning frameworks such as TensorFlow, and the training process was optimized and adjusted. During model training, the dataset was divided into training, validation, and test sets according to a certain ratio. The Mask-RCNN model was trained using the training set, and the model parameters were continuously adjusted through the backpropagation algorithm to enable the model to learn the characteristic representation of the target. During the training process, various optimization strategies were adopted, such as learning rate adjustment, momentum optimization, and weight decay, to improve the model's convergence speed and generalization ability. At the same time, to prevent model overfitting, data augmentation techniques such as random cropping, flipping, scaling, and color jittering were used to increase the diversity of the training data.

[0090] After multiple rounds of training, the model is evaluated using the validation set. Model hyperparameters are adjusted based on evaluation metrics (such as accuracy, recall, and F1 value) to select the optimal model. Finally, the final model is tested on the test set to obtain the target recognition accuracy metric.

[0091] 3. Vehicle detection and segmentation: Use the trained Mask-RCNN model to detect and segment vehicles in factory traffic videos, and automatically identify vehicle type, color, license plate, and other information.

[0092] 4. Data processing and decision-making: Vehicle identification results are processed and analyzed to enable intelligent decision-making for factory traffic monitoring. For example, changes in vehicle type and number can be used to determine whether factory areas need to be adjusted.

[0093] To improve target recognition accuracy and adaptability to complex scenarios, the present disclosure optimizes the Mask-RCNN model. Specific improvements include:

[0094] Introducing the attention mechanism: By introducing the attention mechanism, the model can pay more attention to the key feature areas of the target, improving the recognition ability in complex backgrounds and occlusion situations.

[0095] Customized target classification system: Based on factory traffic rules and actual needs, a customized target classification system and recognition strategy can more accurately identify specific types of vehicles and pedestrians, such as hazardous chemical transport vehicles and workers without helmets.

[0096] Data enhancement technology: Use data enhancement techniques such as random cropping, flipping, scaling, and color jittering to increase the diversity of training data and improve the generalization ability of the model.

[0097] This disclosed embodiment improves the Mask-RCNN model and introduces an attention mechanism to enhance target recognition accuracy and segmentation in complex factory scenarios. Customized target classification systems and recognition strategies enable more accurate identification of specific types of vehicles and pedestrians. This enhances the system's real-time performance and stability, improves the intelligence of factory traffic management, and ensures safe and efficient factory traffic.

[0098] Furthermore, the introduction of the attention mechanism into the Mask-RCNN model includes:

[0099] Add a CBAM module after the convolutional layer of each ResNet residual block in the Mask-RCNN model to adjust the channel and spatial attention of the output feature map;

[0100] The CBAM module is inserted between the feature fusion layers of FPN to enhance the expression ability of multi-scale features.

[0101] One of the core improvements to the Mask-RCNN model is the introduction of an attention mechanism, which enhances the model's focus on key features of the target and improves recognition accuracy in complex scenarios (such as occlusion, lighting changes, and dense traffic).

[0102] CBAM (Convolutional Block Attention Module) is a module that combines channel attention and spatial attention, which can adaptively adjust the importance weights of different channels and spatial positions in the feature map. Its structure is as follows:

[0103] Channel attention: Extract channel features through global average pooling and maximum pooling, generate channel weight vectors, and amplify the contribution of important channels.

[0104] Spatial attention: Based on the channel attention output, a spatial weight matrix is generated through convolution operation to highlight the key areas of the target.

[0105] The backbone network of Mask-RCNN is usually ResNet-FPN (Feature Pyramid Network), which is used to extract multi-scale features. The CBAM module is embedded in the residual block of ResNet. The specific design is as follows:

[0106] Insertion at the end of the residual block: Add a CBAM module after the convolutional layer of each ResNet residual block to adjust the channel and spatial attention of the output feature map.

[0107] Original residual block process: input → convolutional layer → batch normalization → activation function → output.

[0108] Improved process: input → convolutional layer → batch normalization → activation function → CBAM module → output.

[0109] FPN inter-layer embedding: The CBAM module is inserted between the feature fusion layers of FPN (such as P2-P5) to enhance the expression ability of multi-scale features.

[0110] After introducing the attention mechanism, the training and tuning strategies need to be adjusted, including:

[0111] Loss function adjustment:

[0112] Based on the original Mask-RCNN loss function (classification loss, bounding box regression loss, mask loss), an attention weight regularization term is added to prevent the attention module from overfitting.

[0113] Learning rate optimization:

[0114] Using a hierarchical learning rate strategy, different learning rates are set for the backbone network (ResNet) and the newly added CBAM module (such as 1e-4 for the backbone network and 1e-3 for the CBAM module) to accelerate the convergence of the attention module.

[0115] The attention mechanism can enhance the model's focus on the key features of the target and improve the recognition accuracy in complex scenarios (such as occlusion and background interference).

[0116] Furthermore, the customized target classification system and identification strategy include:

[0117] Determine the target classification system and identification strategy based on the factory traffic rules and actual needs;

[0118] Add custom categories to the classification branch of Mask-RCNN based on the target classification system and recognition strategy, and perform hierarchical classification design;

[0119] For hazardous materials signs and helmet areas, local feature learning is strengthened in the feature extraction stage, and target detection and attribute recognition are jointly trained through multi-task learning.

[0120] This system combines factory traffic regulations with actual needs to conduct customized traffic behavior analysis. For example, based on information such as vehicle and personnel trajectory, speed, and dwell time, it can determine whether there are any violations or safety hazards, and promptly alert managers to address any abnormalities. In this way, the system can more accurately assist factory traffic management, improving management efficiency and safety.

[0121] To accurately identify specific targets within the factory (such as hazardous chemical transport vehicles and personnel not wearing helmets), a classification system must be systematically customized from four levels: demand analysis, data design, model optimization, and training strategy. The following is the specific implementation process:

[0122] 1. Requirements analysis and rule mapping, including:

[0123] Identify plant management needs, such as:

[0124] Safety rules: Identify vehicles transporting hazardous materials (special passes are required), workers not wearing hard hats, illegally parked forklifts, etc.

[0125] Efficiency rules: monitor the dwell time of logistics vehicles, vehicle density in congested areas, etc.

[0126] Develop a classification labeling system

[0127] Basic categories: Use general categories (such as "vehicle" and "pedestrian").

[0128] Customized categories: Vehicles: dangerous goods transport vehicles, ordinary trucks, forklifts, engineering vehicles Pedestrians: workers wearing safety helmets, workers not wearing safety helmets, visitors (without work clothes).

[0129] Behavior tags: speeding, illegal parking, and not driving according to the route.

[0130] 2. Data collection and labeling, including:

[0131] Data collection must meet the following requirements:

[0132] Scene coverage: Collect image and video data for typical factory scenes (loading and unloading areas, warehouse passages, intersections).

[0133] Multimodal data: combining camera video, traffic statistics from the intermodal station, and RFID (Radio Frequency Identification) tags (hazardous goods vehicle identification).

[0134] Refined annotation:

[0135] Labeling tools: Use tools such as Label Studio and CVAT to label the following:

[0136] Target category: Distinguish dangerous goods transport vehicles from ordinary trucks (through body markings and dangerous goods signs).

[0137] Attribute tags: vehicle color, license plate number, worker helmet color.

[0138] Behavior labeling: Annotate speeding trajectories (track speed through consecutive frames).

[0139] Labeling rules: Dangerous goods transport vehicles: The vehicle body must be labeled with dangerous goods identification (such as UN number, flammable symbol). No helmet: The head area must be marked as not covered by a helmet.

[0140] 3. Model classification head optimization, including:

[0141] Classification header structure adjustment

[0142] New output node: Add custom categories (such as "hazardous goods transport vehicle" and "worker not wearing a helmet") to the classification branch of Mask-RCNN.

[0143] Hierarchical classification design: Using a tree-like classification structure, first distinguish between "vehicles / pedestrians" and then subdivide them into subcategories (such as vehicles → hazardous materials transport vehicles).

[0144] Feature enhancement strategy

[0145] Focus on key areas: For hazardous materials signs and safety helmet areas, strengthen local feature learning during the feature extraction stage.

[0146] Multi-task learning: Jointly train object detection (classification + localization) and attribute recognition (such as license plate, helmet color).

[0147] 4. Training strategy optimization, including:

[0148] Data enhancement targeted design, such as:

[0149] Simulate factory scenes: add obstructions (such as cargo blocking vehicles), lighting changes (factory lights at night), and blur effects (motion blur).

[0150] Small sample enhancement: For scarce samples such as dangerous goods transport vehicles, GAN (Generative Adversarial Networks) is used to generate synthetic data.

[0151] Loss function adjustment

[0152] Class-weighted loss: Assigns higher weights to classes with small samples (such as hazardous materials transport vehicles) to alleviate the class imbalance problem.

[0153] Attribute supervision loss: Added auxiliary task of helmet detection (binary cross entropy loss).

[0154] Transfer learning and fine-tuning

[0155] Pre-trained model: Mask-RCNN is pre-trained based on the COCO dataset, retaining general object detection capabilities.

[0156] Fine-tuning in stages:

[0157] The first stage: freeze the backbone network and only train the classification head to adapt to new categories.

[0158] The second stage: unfreeze some backbone layers and jointly optimize feature extraction and classification.

[0159] 5. Verification and scenario adaptation, including:

[0160] Offline testing:

[0161] Confusion Matrix Analysis: Check the false detection rates of hazardous materials transport vehicles and ordinary trucks, and optimize feature extraction accordingly.

[0162] Key Metrics:

[0163] Accuracy: The accuracy of identifying dangerous goods vehicles must be ≥95%.

[0164] Recall rate: The missed detection rate of workers not wearing safety helmets is ≤5%.

[0165] Online tuning,

[0166] Dynamic rule engine: Encodes factory traffic rules into logical conditions (such as "hazardous goods vehicle enters non-designated area → alarm").

[0167] Feedback loop: Collect false positive samples (such as misclassifying an ordinary truck as a hazardous goods vehicle) and iteratively update the training data.

[0168] Taking hazardous materials vehicle monitoring in a chemical plant as an example, the recognition logic is as follows: detect the vehicle → classify it as a "hazardous materials transport vehicle" → verify the vehicle's UN number (OCR (Optical Character Recognition)) → check access permissions (compared with a database). The alarm rules are: if the hazardous materials vehicle does not follow the designated route or stops for an extended period, an audible and visual alarm is triggered and the safety officer is notified.

[0169] Through demand-driven classification system design, refined data annotation, model structure optimization, and targeted training strategies, a customized target classification system can accurately identify specific targets within a factory. Combined with a rules engine and feedback mechanism, the system dynamically adapts to different factory scenarios, ultimately achieving both safety and efficiency improvements.

[0170] Furthermore, the method further comprises:

[0171] Transmit the factory traffic video to be inspected, collected by the acquisition equipment, to the back-end MEP server. During the transmission process, network slicing technology is used to ensure the network performance of video data transmission.

[0172] Build an AI computing resource pool on the MEP server to rationally allocate computing resources.

[0173] To improve the real-time performance and stability of the system, this paper adopts the following optimization methods:

[0174] Network slicing technology: Combined with 5G edge cloud computing, network slicing technology is used to ensure the network performance of video data transmission and ensure the real-time performance of the system when high-concurrency video streams are accessed. 5G network slicing divides the physical network into multiple logically independent virtual network slices, allocating dedicated slices (such as URLLC slices) for video streams to ensure low latency (≤20ms) and high reliability (99.999% SLA (Service Level Agreement). Edge computing collaboration is achieved by deploying MEC (Mobile edge computing) nodes on the base station side to process video data locally and reduce the round-trip transmission delay in the cloud. Combining 5G edge computing and network slicing technology, low-latency, highly reliable data processing and decision feedback are achieved.

[0175] Hardware resource optimization: Build an AI computing resource pool on the MEP server to rationally allocate computing resources and improve the system's processing efficiency and stability. The MEP server is equipped with NVIDIA A100 GPUs (supporting multi-instance GPU technology) and FPGA accelerator cards. GPU resources are managed through the Kubernetes cluster, and computing power is allocated on demand (for example, one GPU instance is allocated to each video stream). Task priority scheduling is implemented, for example, real-time video analysis tasks (such as hazardous materials vehicle identification) have higher priority than offline tasks (such as log analysis). Use a real-time operating system (RTOS) or Linux kernel preemptive scheduling (PREEMPT_RT). Dynamic load balancing: Based on the Prometheus monitoring system, GPU / CPU utilization is collected in real time and tasks on overloaded nodes are automatically migrated. This improves the system's processing efficiency and stability.

[0176] System redundancy design: Redundancy design is adopted to ensure that when hardware equipment or network failure occurs, the system can quickly switch to backup equipment to ensure uninterrupted operation of the system.

[0177] Furthermore, the method further comprises:

[0178] The real-time data of various types of equipment in the factory are integrated through the GIS map as the factory traffic video data to be detected.

[0179] like Figure 3 As shown, this disclosure integrates real-time data from multiple types of equipment (cameras, traffic control stations, and weather stations) by loading a GIS map. Users can filter device types to view location and monitoring information (such as traffic flow and weather parameters). Clicking a device pops up a details panel displaying status and historical records. Map zooming and multi-device linkage are supported, enabling global monitoring and rapid location of factory traffic, improving data interaction efficiency and visualization capabilities.

[0180] like Figure 4 As shown, based on the factory traffic AI video monitoring system, the method can also realize video monitoring-real-time preview,

[0181] After the user selects a camera, the system verifies the device's online status and pulls the RTSP video stream, decoding it and displaying it in multi-grid or full-screen mode. It supports image capture, local recording, audio control, and pan / tilt (pan / tilt) operation (rotation and zoom). Abnormal devices are marked gray and unavailable. Recordings are saved locally by default, and key operation records are synchronized to the database, ensuring flexible real-time monitoring and data traceability.

[0182] Furthermore, the method further comprises:

[0183] If an abnormal event is detected after data processing and analysis, a real-time alarm is triggered and the event details are automatically recorded.

[0184] like Figure 5 As shown, emergency response relies on AI to detect unusual events (such as accidents and congestion), triggering real-time alerts and automatically recording event details. The system then matches emergency plans, dispatches emergency resources (personnel and equipment), and initiates the command process. The response process records the progress of the incident, expert opinions, and instructions from superiors, ultimately generating an analysis report. This supports closed-loop event management and historical review, improving emergency response efficiency.

[0185] Furthermore, the method further comprises:

[0186] Regularly poll the device video signal to detect interruption, blur or blockage faults;

[0187] If abnormal data is found, it will be marked and the fault type will be analyzed. A percentage chart will be generated and an alert will be pushed to the operation and maintenance personnel.

[0188] like Figure 6 As shown in the figure, the system's video quality diagnosis module periodically polls device video signals to detect faults such as interruptions, blurring, or occlusions. Abnormal data is marked and analyzed for fault type, generating a percentage chart (bar chart or line graph) and sending alerts to operations and maintenance personnel. Fault details can be found and statistical reports can be exported, helping to quickly locate device issues and ensure stable operation of the monitoring system.

[0189] like Figure 7 As shown, the method also enables information release management, including: users edit information board content and submit it for review. Once approved, it is distributed to one or multiple information boards, with retry support in the event of failure. The published content records the status (success / failure) and provides preview and historical query functions. It supports common phrase templates and group messaging, and integrates IP broadcasting to achieve multimodal information synchronization, ensuring the timeliness and accuracy of factory notifications.

[0190] Furthermore, the system manages user permissions as follows Figure 8 As shown, administrators can add, import, or delete users, assign roles, and bind organization / device permissions. Permission control is granularized down to the function menu, supporting password changes and operation log auditing. Abnormal logins trigger alerts, and role permissions are dynamically adjusted to ensure system access security and compliance, meeting the permission management needs of multi-level organizations.

[0191] The disclosed embodiments systematically solve the problems of low recognition accuracy, insufficient real-time performance, and poor stability in complex scene monitoring in factory traffic monitoring by improving the Mask-RCNN model (introducing an attention mechanism and a customized classification system), integrating 5G edge computing and network slicing technology, optimizing hardware resource allocation, and deploying redundant architecture. The accuracy of target recognition is greatly improved, and end-to-end latency is reduced. The system has strong availability, supports concurrent processing of multiple video streams, and realizes precise supervision of hazardous materials vehicles and real-time alerts for violations in actual scenarios, reducing operation and maintenance costs, and providing an efficient, reliable, and intelligent full-stack solution for factory safety management. The intelligence level of factory traffic management is improved, ensuring factory traffic safety and logistics efficiency.

[0192] The second embodiment of the present disclosure also provides an AI monitoring method for factory traffic to solve the defects of poor adaptability and single function of existing video monitoring technology in complex scenes. The deep learning algorithm based on Mask-RCNN is used as the core target recognition technology. The structure diagram of the Mask-RCNN model is as follows: Figure 9 As shown in the figure, RoIAlign technology is used to precisely align the target areas in the image. Each aligned area is classified into a category (such as "vehicle" or "pedestrian"). The target position is annotated with a bounding box, and a segmentation mask may be superimposed for instance segmentation. On this basis, the attention mechanism and customized target classification system are introduced to improve the recognition accuracy of specific targets in the factory. In addition, the real-time performance, stability, and intelligence of the system are improved through functions such as edge computing and multimodal analysis, ultimately achieving the goal of optimizing the efficiency and safety of factory traffic management.

[0193] The method comprises:

[0194] 1. Data Preparation Phase

[0195] Objective: Collect and process high-quality training data to lay the foundation for model training.

[0196] Key steps:

[0197] Data collection:

[0198] Equipment deployment: High-definition cameras, dispatch stations, weather stations, and other equipment are deployed in key areas of the factory (such as intersections, loading and unloading areas, and pedestrian walkways) to ensure coverage of complex scenarios (such as occlusion, lighting changes, and dense traffic).

[0199] Data type: Collect multimodal data such as video streams, traffic flow statistics, meteorological parameters (temperature, humidity, wind speed), and equipment status data.

[0200] Data annotation:

[0201] Labeling tools: Use tools such as LabelImg and CVAT to perform bounding box annotation and instance segmentation mask annotation on targets (vehicles, pedestrians, hazardous materials transport vehicles, etc.) in video frames.

[0202] Labeling rules: Develop factory-specific classification labels (such as "personnel not wearing a hard hat" and "forklift speeding") to ensure labeling consistency and accuracy.

[0203] Data preprocessing:

[0204] Enhancement technology: Apply data enhancement techniques such as random cropping, flipping, and color jittering to increase data diversity.

[0205] Data partitioning: Divide the data into training set, validation set, and test set in a certain ratio (e.g., 7:2:1) to ensure the generalization ability of the model.

[0206] 2. Model development and optimization stage

[0207] Objective: Build an improved Mask-RCNN model to improve recognition accuracy in complex scenarios.

[0208] Key steps:

[0209] Model architecture improvements:

[0210] Attention mechanism embedding: CBAM (Convolutional Block Attention Module) is introduced into the ResNet-FPN backbone network of Mask-RCNN to enhance the model's attention to key features.

[0211] Customized classification head: Design fine-grained classifiers based on factory needs (e.g., distinguishing between hazardous materials transport vehicles and ordinary trucks).

[0212] Model training and tuning:

[0213] Training environment: Use the PyTorch / TensorFlow framework, configure multi-GPU parallel training, and accelerate computing.

[0214] Hyperparameter optimization: Adjust the learning rate, batch size, and loss function weights (such as the balance between classification loss and segmentation loss) through grid search or Bayesian optimization.

[0215] Prevent overfitting: Apply early stopping, dropout layers, and regularization techniques.

[0216] Model Validation:

[0217] Performance evaluation: Calculate indicators such as mAP (mean average precision) and IoU (intersection over union) on the test set to verify the improvement effect of the model in occluded and dense scenes.

[0218] Scenario adaptation: Perform transfer learning for different factory environments (such as chemical parks and logistics parks) and fine-tune model parameters.

[0219] 3. System Architecture Construction Phase

[0220] Objective: Build a factory monitoring system that supports low latency and high concurrency.

[0221] Key steps:

[0222] Front-end and back-end development:

[0223] Front-end interface: A visual interface developed based on Vue.js / React, integrating GIS maps, real-time video preview, alarm panel and other functions.

[0224] Backend services: Use the Flask / Django framework to build a RESTful API to implement video streaming distribution (RTSP / WebRTC), device status monitoring, and data storage (MySQL / MongoDB).

[0225] 5G edge computing integration:

[0226] Edge node deployment: Deploy MEC (Multi-Access Edge Computing) servers within the factory to run AI models locally and reduce cloud dependency.

[0227] Network slice configuration: Work with operators to allocate dedicated network slices (such as URLLC slices) for video streams to ensure bandwidth and low latency (≤50ms).

[0228] Functional module development:

[0229] Comprehensive monitoring module: Integrates data from multiple devices through GIS maps and supports multi-device linkage (such as automatically calling nearby cameras when the weather is abnormal).

[0230] Emergency response system: Develop a rule engine to realize automatic alarm, plan matching and resource scheduling for abnormal events (accidents, congestion).

[0231] Video quality diagnosis: Regularly detects camera interruptions, blurring, and other issues, generates operation and maintenance reports, and pushes alerts.

[0232] 4. System Integration and Testing Phase

[0233] Objective: To verify the system functional integrity and performance stability.

[0234] Key steps:

[0235] Module joint debugging:

[0236] Interface testing: Verify the correctness of front-end and back-end API communication, video streaming (RTSP / HLS), database reading and writing, and other functions.

[0237] Multimodal data fusion: Test the real-time synchronization and joint analysis capabilities of video data, exchange station data, and meteorological data.

[0238] Performance testing:

[0239] Stress testing: Simulates high-concurrency scenarios (such as accessing 200 video streams simultaneously) to detect system response time, resource utilization, and crash thresholds.

[0240] Stability test: Run continuously for 72 hours, recording the system failure rate and automatic switching (redundant design) efficiency.

[0241] User Acceptance Testing (UAT):

[0242] Scenario coverage: Test core functions such as vehicle identification, behavior analysis, and alarm triggering in a real factory environment.

[0243] Feedback optimization: Adjust interface interaction logic, alarm thresholds and other details based on user feedback.

[0244] 5. Deployment and Operations

[0245] Objective: To realize the practical application and continuous optimization of the system in the factory.

[0246] Key steps:

[0247] Hardware deployment:

[0248] Edge server installation: Deploy MEC servers in the factory computer room and configure a GPU computing resource pool.

[0249] Camera and sensor networking: Connect cameras via the ONVIF protocol to ensure that the device's online status can be monitored.

[0250] System online:

[0251] Grayscale release: Start with a trial run in a local area, and then gradually expand the coverage to avoid the risk of switching the entire factory at once.

[0252] User training: Provide system operation training to management personnel, focusing on functions such as alarm processing and emergency command.

[0253] Operation and maintenance and iteration:

[0254] Monitoring and maintenance: Use log analysis tools (such as ELK Stack) to monitor system health in real time and regularly update models and software versions.

[0255] Data closure: Collect new data from actual applications, continuously optimize models (such as incremental learning), and improve scenario adaptability.

[0256] The system of the disclosed technical solution has performed well in practical applications. Through the optimized Mask-RCNN model, the vehicle recognition accuracy rate reached 93%, and the pedestrian recognition accuracy rate reached 88%, which are significantly higher than the existing technical level. When the system is accessed by high-concurrency video streams and processes large amounts of data, the response time does not exceed 1 second, meeting the needs of real-time monitoring. During the long-term operation test, the system stability was good, and no system failures or abnormalities caused by environmental factors occurred. Through customized traffic behavior analysis, the system can effectively identify and warn of violations and safety hazards within the factory area, providing a strong guarantee for the company's safe production. Figure 10 This is a schematic diagram of the system's application in actual environments, demonstrating the system's real-time monitoring and behavior analysis capabilities for vehicles and personnel.

[0257] The third embodiment of the present disclosure also provides a factory traffic AI video monitoring system, such as Figure 11 As shown, the system includes:

[0258] The collection and processing module 11 is configured to collect factory traffic data through collection devices deployed in key areas of the factory, and perform data preprocessing and annotation to obtain a data set;

[0259] an optimization and training module 12, configured to select a Mask-RCNN model as a network for vehicle detection and segmentation, introduce an attention mechanism into the Mask-RCNN model, customize a target classification system and recognition strategy, train the model using a dataset using the TensorFlow deep learning framework, and optimize and adjust the training process;

[0260] A detection module 13 is configured to use the trained Mask-RCNN model to perform target detection and segmentation on the factory traffic video to be detected, and automatically identify the target;

[0261] The analysis module 14 is configured to process and analyze the identification results to achieve intelligent decision-making for factory traffic monitoring.

[0262] Furthermore, the optimization and training module 12 is specifically configured as follows:

[0263] Add a CBAM module after the convolutional layer of each ResNet residual block in the Mask-RCNN model to adjust the channel and spatial attention of the output feature map;

[0264] The CBAM module is inserted between the feature fusion layers of FPN to enhance the expression ability of multi-scale features.

[0265] Furthermore, the optimization and training module 12 is specifically configured as follows:

[0266] Determine the target classification system and identification strategy based on the factory traffic rules and actual needs;

[0267] Add custom categories to the classification branch of Mask-RCNN based on the target classification system and recognition strategy, and perform hierarchical classification design;

[0268] For hazardous materials signs and helmet areas, local feature learning is strengthened in the feature extraction stage, and target detection and attribute recognition are jointly trained through multi-task learning.

[0269] Furthermore, the system further includes a system optimization module 15;

[0270] The system optimization module 15 is configured to transmit the factory traffic video to be detected collected by the collection device to the back-end MEP server, and to ensure the network performance of the video data transmission during the transmission process through the network slicing technology; and

[0271] Build an AI computing resource pool on the MEP server to rationally allocate computing resources.

[0272] Furthermore, the system also includes a comprehensive monitoring module 16;

[0273] The comprehensive monitoring module 16 is configured to integrate the real-time data of various types of equipment in the factory area through a GIS map as the factory area traffic video data to be detected.

[0274] Furthermore, the system also includes an emergency response module 17;

[0275] The emergency handling module 17 is configured to trigger a real-time alarm and automatically record event details if an abnormal event is detected after data processing and analysis.

[0276] Furthermore, the system further includes a video quality diagnosis module 18;

[0277] The video quality diagnosis module 18 is configured to periodically poll the device video signal to detect interruption, blur or occlusion faults; and

[0278] If abnormal data is found, it will be marked and the fault type will be analyzed. A percentage chart will be generated and an alert will be pushed to the operation and maintenance personnel.

[0279] The factory traffic AI video monitoring system of the disclosed embodiment is used to implement the factory traffic AI video monitoring method in method embodiment 1 and embodiment 2, so the description is relatively simple. For details, please refer to the relevant description in the previous method embodiment, which will not be repeated here.

[0280] In addition, if Figure 12 As shown, the fourth embodiment of the present disclosure further provides an electronic device, including a memory 100 and a processor 200, wherein the memory 100 stores a computer program. When the processor 200 runs the computer program stored in the memory 100, the processor 200 executes the above-mentioned various possible methods.

[0281] The memory 100 is connected to the processor 200 . The memory 100 may be a flash memory, a read-only memory, or other memory. The processor 200 may be a central processing unit or a single-chip microcomputer.

[0282] In addition, an embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored, and the computer program is used by a processor to execute the above-mentioned various possible methods.

[0283] The computer-readable storage medium includes volatile or nonvolatile, removable or non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, computer program modules or other data). Computer-readable storage media include, but are not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable read only memory), flash memory or other memory technology, CD-ROM (Compact Disc Read-Only Memory), Digital Versatile Disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.

[0284] It is understood that the above embodiments are merely exemplary embodiments for illustrating the principles of the present disclosure, and the present disclosure is not limited thereto. Those skilled in the art may make various modifications and improvements without departing from the spirit and substance of the present disclosure, and such modifications and improvements are also considered to be within the scope of protection of the present disclosure.

Claims

1. A factory traffic AI video monitoring method, characterized in that: The method comprises: By deploying collection equipment in key areas of the factory, we collect factory traffic data, pre-process and label the data, and generate a data set. We selected the Mask-RCNN model as the network for vehicle detection and segmentation. We introduced an attention mechanism into the Mask-RCNN model, customized the target classification system and recognition strategy, trained the model on the dataset using the TensorFlow deep learning framework, and optimized and adjusted the training process. Use the trained Mask-RCNN model to detect and segment objects in the factory traffic video to be inspected, and automatically identify the objects; The identification results are processed and analyzed to realize intelligent decision-making for factory traffic monitoring.

2. The method according to claim 1, characterized in that The introduction of the attention mechanism into the Mask-RCNN model includes: A convolutional attention module (CBAM) is added after the convolutional layer of each residual network (ResNet) residual block in the Mask-RCNN model to perform channel and spatial attention adjustments on the output feature maps. The CBAM module is inserted between the feature fusion layers of the feature map pyramid network (FPN) to enhance the expression capability of multi-scale features.

3. The method according to claim 1, characterized in that The customized target classification system and identification strategy include: Determine the target classification system and identification strategy based on the factory traffic rules and actual needs; Add custom categories to the classification branch of Mask-RCNN based on the target classification system and recognition strategy, and perform hierarchical classification design; For hazardous materials signs and helmet areas, local feature learning is strengthened in the feature extraction stage, and target detection and attribute recognition are jointly trained through multi-task learning.

4. The method according to claim 1, wherein The method further comprises: The factory traffic video to be inspected, collected by the acquisition equipment, is transmitted to the back-end mobile edge computing platform MEP server. During the transmission process, network slicing technology is used to ensure the network performance of video data transmission. Build an AI computing resource pool on the MEP server to rationally allocate computing resources.

5. The method according to claim 1, characterized in that The method further comprises: The real-time data of various types of equipment in the factory are integrated through the Geographic Information System (GIS) map as the factory traffic video data to be detected.

6. The method according to claim 1, characterized in that The method further comprises: If an abnormal event is detected after data processing and analysis, a real-time alarm is triggered and the event details are automatically recorded.

7. The method according to claim 1, characterized in that The method further comprises: Regularly poll the device video signal to detect interruption, blur or blockage faults; If abnormal data is found, it will be marked and the fault type will be analyzed. A percentage chart will be generated and an alert will be pushed to the operation and maintenance personnel.

8. A factory traffic AI video monitoring system, characterized by: The system comprises: The collection and processing module is configured to collect factory traffic data through collection devices deployed in key areas of the factory, and perform data preprocessing and annotation to obtain a data set; An optimization and training module, which is configured to select the Mask-RCNN model as the network for vehicle detection and segmentation, introduce an attention mechanism into the Mask-RCNN model, customize the target classification system and recognition strategy, use the TensorFlow deep learning framework to train the model on the dataset, and optimize and adjust the training process; The detection module is configured to use the trained Mask-RCNN model to detect and segment objects in the factory traffic video to be inspected, and automatically identify the objects; The analysis module is configured to process and analyze the identification results to realize intelligent decision-making for factory traffic monitoring.

9. An electronic device, characterized in that: It includes a memory and a processor, wherein a computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the factory traffic AI video monitoring method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the factory traffic AI video monitoring method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Scene-based monitoring model customization method, equipment and medium

    CN121501255A