Intelligent monitoring system

Through the multi-unit collaborative work of the smart monitoring system, and using YOLOv5s and RFID technology, the problems of incomplete monitoring of the existing urban monitoring system in construction safety, crowd gathering, security and traffic management have been solved, and efficient and low-cost all-round monitoring management has been achieved.

CN120472388APending Publication Date: 2025-08-12哈尔滨应通科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510540986.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing urban monitoring system has problems of incomplete monitoring and high cost in terms of construction safety, crowd gathering, security, traffic management and road facility monitoring.

Method used

The intelligent monitoring system is adopted, including a safety helmet monitoring unit, a personnel gathering monitoring unit, an area intrusion monitoring unit, a vehicle illegal parking detection unit and a road facility monitoring unit, and the target recognition and detection are used by YOLOv5s, and the monitoring accuracy and efficiency are improved through RFID technology, time window filtering algorithm and weighted priority scheduling algorithm.

Benefits of technology

It realizes comprehensive monitoring of construction safety, crowd gathering, security and traffic management, reduces management costs, and promptly detects abnormal or dangerous signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472388A_ABST
    Figure CN120472388A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent monitoring system, and relates to the field of urban intelligent monitoring. A safety helmet monitoring unit is used for detecting the safety helmet wearing condition in a construction area; the personnel gathering monitoring unit is used for monitoring the crowd gathering condition and sending out gathering early warning when the personnel gathering condition occurs; the area intrusion monitoring unit is used for monitoring whether an area is intruded or not; the vehicle illegal parking detection unit is used for monitoring whether illegal parking vehicles exist in the parking forbidding area or not and sending out illegal parking signals. According to the invention, comprehensive monitoring in the monitoring area can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of urban intelligent monitoring, and in particular to an intelligent monitoring system. Background Art

[0002] The accelerated pace of urbanization has inevitably thrust cities to the center of the world stage, where they play a leading role. At the same time, cities also face challenges such as environmental pollution, traffic congestion, energy shortages, housing shortages, unemployment, and disease. In this new environment, addressing the numerous challenges posed by urban development and achieving sustainable development has become a crucial issue in urban planning and construction. Against this backdrop, "smart cities" have become a viable path to addressing urban issues and a key trend in future urban development. The accelerated development of smart cities will drive rapid growth in local economies and industries such as satellite navigation, the Internet of Things, intelligent transportation, smart grids, cloud computing, and software services, creating new opportunities for these sectors. Summary of the Invention

[0003] In view of the above-mentioned defects of the prior art, the present invention provides an intelligent monitoring system, comprising:

[0004] Safety helmet monitoring unit, used to detect safety helmet wearing in the construction area;

[0005] The crowd gathering monitoring unit is used to monitor crowd gatherings and issue a gathering warning when a crowd gathering occurs;

[0006] Regional intrusion monitoring unit, used to monitor whether an area has been invaded;

[0007] The vehicle illegal parking detection unit is used to monitor whether there are illegally parked vehicles in the prohibited parking area and issue an illegal parking signal.

[0008] Furthermore, the helmet monitoring unit identifies the category of the helmet being worn through YOLOv5s.

[0009] Furthermore, the crowd gathering monitoring unit detects the number of people in an area through YOLOv5s to obtain a crowd density value in the area, and outputs a warning result when the crowd density value is greater than a preset threshold.

[0010] Furthermore, the working process of the regional intrusion detection unit is: setting a closed target area, detecting the coordinates of a suspicious target, and issuing an intrusion warning message when the coordinates fall within the closed target area.

[0011] Furthermore, the vehicle illegal parking detection unit detects a vehicle outline area in an area, determines whether the vehicle outline area overlaps with the illegal parking area, and if so, issues a vehicle illegal parking warning message.

[0012] Furthermore, the monitoring system includes a road facility monitoring unit, which is used to monitor the covering status of road manhole covers and the overflowing status of garbage bins.

[0013] Furthermore, the smart monitoring system is installed on the urban lighting device, and the smart monitoring system includes a lighting monitoring unit for monitoring the operating status of the lighting device according to the real-time status parameters of the lighting device.

[0014] Compared with the prior art, the present invention has the following technical effects:

[0015] The present invention can realize comprehensive monitoring and management of construction safety, crowd gathering, security monitoring, traffic management and road facilities, can timely detect abnormal or dangerous signals, has more diverse and comprehensive functions, and reduces the monitoring and management costs in the area.

[0016] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a structural schematic diagram of a specific embodiment of the present invention;

[0018] Figure 2 This is a schematic diagram of the YOLOv5s network structure according to a specific embodiment of the present invention;

[0019] Figure 3 This is a schematic diagram of the structure of MobileNetV3 according to a specific embodiment of the present invention;

[0020] Figure 4 is a schematic diagram of the structure of an attention unit according to a specific embodiment of the present invention;

[0021] Figure 5 2 is a structural principle diagram of a dilute spatial pyramid pooling module according to a specific embodiment of the present invention. DETAILED DESCRIPTION

[0022] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.

[0023] like Figure 1As shown, in a specific embodiment, a smart monitoring system is provided, including a helmet detection unit, a crowd gathering monitoring unit, an area intrusion monitoring unit, a vehicle illegal parking detection unit, a road facility monitoring unit, and a lighting monitoring unit. The helmet monitoring unit is used to detect whether helmets are worn in the construction area; the crowd gathering monitoring unit is used to monitor crowd gatherings and issue a gathering warning when a crowd gathering occurs; the area intrusion monitoring unit is used to monitor whether an area has been invaded; and the vehicle illegal parking detection unit is used to monitor whether there are illegally parked vehicles in a prohibited parking area and issue a parking violation signal.

[0024] In one specific embodiment, RFID-assisted identification technology is integrated to enable target identification in obstructed scenarios. To further improve monitoring and management accuracy, a time window filtering algorithm is employed in one embodiment to prevent false alarms. For example, an alarm is triggered only after five consecutive frames of detection. To prevent interference between units, each detection unit in this embodiment maintains a synchronized clock error of less than 300ms.

[0025] The functions of each of the aforementioned units can be implemented simultaneously or scheduled as needed. In this embodiment, the hardware utilizes the NVIDIA Jetson AGX Xavier edge computing device, which boasts 32TOPS of computing power and 64GB of memory to store the AI models of each unit. This allows the device to run complex applications and enhance real-time processes, enabling this embodiment's edge computing device to achieve higher levels of computing density, energy efficiency, and AI reasoning capabilities at the edge. In terms of software architecture, multi-threaded processing is employed for each module, with the detection module based on YOLOv5s' multi-threaded concurrent execution. The threading module is used to handle reasoning tasks for multiple images or video streams.

[0026] In a specific embodiment, each unit is scheduled using the following scheduling scheme, which uses a weighted priority scheduling algorithm to calculate the scheduling priority. The scheduling priority calculation formula is:

[0027] Scheduling priority = BasePriority × Weight + DynamicAdjustment;

[0028] Among them, BasePriority is a static preset value with a value range of 0 to 100; Weight is a weight coefficient, which can be dynamically adjusted as needed; DynamicAdjustment is a correction value based on runtime feedback, which is set according to the operating needs and operating scenarios of the monitoring system.

[0029] The weight coefficient adjustment method can adopt the absolute weight method and the relative weight method. The absolute weight method is that the weight of each task is a fixed proportion. For example, when the main application is construction monitoring scenario, the weight value of the helmet detection unit is relatively large, and its corresponding priority will be relatively high. When it is necessary to monitor multiple working conditions at the same time, such as when it is necessary to monitor both the wearing of helmets and the condition of road manhole covers, the weight coefficients of the helmet detection unit and the road facility monitoring unit will be larger than those of other units. The relative weight method can adopt the relative weight method in the existing technology, but it is dynamically normalized according to the system load.

[0030] After calculating the scheduling priority of each unit, when the scheduling priority of a unit is high, when the high-priority task arrives, a full preemptive method or a partial preemptive method can be used. The full preemptive method is to immediately deprive the low-priority task resources and give priority to completing the tasks of the high-priority unit. The partial preemptive method is to allocate the remaining resources according to the weight ratio calculated according to the above embodiment to complete the tasks of each unit.

[0031] In order to improve the scheduling fairness of each unit, the virtual time T is calculated after each scheduling is completed. vruntime , the virtual time is:

[0032] T Vrunt i me (k) = T vrunt i me (k-1)+T exec ×(W NICE_0_LOAD / w);

[0033] Among them, T Vruntime (k) is the virtual time obtained by the k-th cumulative calculation; T Vruntime (k-1) is the virtual time obtained by the k-th cumulative calculation; T exec is the actual running time; W NICE_0_LOAD is the weight when nice is 0, its value is 1024, corresponding to the benchmark task; w is the weight coefficient.

[0034] In this example, a red-black tree is used to sort all tasks in ascending order by vruntime, with the leftmost node of the red-black tree being the next scheduling target. The scheduling strategy in this example adds an anti-starvation mechanism, which primarily involves two aspects: a priority ceiling, which sets a maximum wait time threshold (e.g., 500ms), and a weight decay mechanism, which reduces the weight of long-running tasks over time. In this example, the decay factor γ is set to 0.99 / second.

[0035] In a specific embodiment, in order to further improve the convenience of operation, cloud collaboration technology is adopted on the mobile phone and platform ends, Azure Cosmos DB is used to achieve real-time collaborative editing, and an Operational Ransformation (OT) + CRDT hybrid model algorithm is used to resolve conflicts, so that the operating status of each unit and all functions of the smart monitoring system can be displayed on the mobile terminal at the same time.

[0036] In a specific embodiment, in order to improve computational efficiency, the traditional OT algorithm is replaced with a neural optimal transport (Neural OT) algorithm. The traditional OT problem is transformed into a learnable deep network architecture. The cost matrix is dynamically generated through the neural network, breaking through the computational efficiency limitations of the traditional OT algorithm. The optimal transport (OT) problem formula is:

[0037] $$\min_{\pi\in\Pi(\mu,u)}\int_{\mathcal{X}\times\mathcal{Y}}c(x,y)d\pi(x,y)$$;

[0038] Among them, $\pi$ is the transport plan, which is a joint probability distribution; $\Pi(\mu,u)$ is the marginal distribution, which is the set of all joint probability distributions of $\mu$ and $u$; $\mu$: is the source distribution (the target distribution of the t-1 frame); $u$ is the target distribution (the observation distribution of the t frame); $c(x,y)$ is the cost function for transmitting from $x$ to $y$.

[0039] In this embodiment, the optimal transmission can be expressed by Monge. For continuous probability measures μ and ν, the Monge problem seeks a mapping T:X→Y to achieve transmission:

[0040] \[\inf_{T\#μ=ν}\int_Xc(x,T(x))dμ(x)\];

[0041] When the cost function c(x,y)=|xy| 2 / 2 (quadratic cost) and μ and ν have density, the optimal transmission map T is related to the gradient of a convex function u: The corresponding Monge-Ampère equation. In this case, the transformation condition T#μ=ν satisfied by the transmission mapping leads to the Monge-Ampère equation:

[0042]

[0043] Where: D 2u is the Hessian matrix of u; μ and ν are the probability density functions of the source domain and the target domain; u is a convex function (called the Brenier potential).

[0044] The cost function $c(x,y)$ is parameterized by a neural network, and a motion consistency constraint is added to the cost function:

[0045] $$c_{ij}=\underbrace{\lambda_1c_{app}(i,j)}_{\text{appearance cost}}+\underbrace{\lambda_2\|v_i-v_j\|^2}_{\text{motion consistency}}$$;

[0046] Among them, $c_{ij}$: the comprehensive matching cost between target $i$ and observation $j$; $c_{app}(i,j)$: appearance similarity measure (ReID feature cosine distance); $v_i$: the velocity vector of target $i$ predicted by the Kalman filter, embedding the non-differentiable OT process into the end-to-end training framework, $v_j$: the velocity estimate of observation $j$ (which can be calculated through the detection boxes of adjacent frames); $\lambda_1,\lambda_2$: hyperparameters, balancing the weights of appearance and motion constraints.

[0047] For the optimal transport problem of discrete point sets \(\{x_i\}\) and \(\{y_j\}\):

[0048] The discrete cost function \(c_{ij}\) corresponds to the continuous setting \(c(x,y)\). The decomposition \(\lambda_1c_{app}(x,y)+\lambda_2\|v_x-v_y\|^2\) can be regarded as a special case of the composite cost function Monge-Ampère expression. When the motion consistency term \(\|v_x-v_y\|^2\) is related to the position (such as \(v_x=x,v_y=y\)), And the appearance cost \(c_{app}(x,y)=0\), then the composite cost degenerates into the quadratic cost: \[c(x,y)=\lambda_2\|xy\|^2\]; at this time the corresponding standard Monge-Ampère equation is: \[\det(D^2u(x))=\frac{\mu(x)}{u(ablau(x))}\]; where \(\lambda_2\) is absorbed into the scale of the potential function \(u\).

[0049] For non-quadratic composite cost functions, it is necessary to introduce the generalized Monge-Ampère equation:

[0050] \[\det(D^2u(x)-A(x,ablau(x)))=\frac{\mu(x)}{u(ablau(x))}|\detD_{xy}^2c(x,ablau(x))|\];

[0051] Where \(A\) is a matrix related to the second-order derivative of the cost function, reflecting the non-quadratic nature of the composite cost.

[0052] Based on the standard evaluation dataset MOTChallenge, this embodiment reduces the number of IdentitySwitches by 35% compared to traditional matching methods based on IoU or appearance similarity.

[0053] In this embodiment, in order to achieve end-to-end training, the Sinkhorn approximation is used:

[0054] $\mathcal{L}_{OT}=\langle\pi^,c\rangle+\epsilonH(\pi^)$; where $\pi^*=\text{Sinkhorn}(c)$, $H$ is the entropy regularization term; since the entropy regularized OT solved by Sinkhorn is discrete, its similar expression in continuous form can be reflected by the heat kernel regularized Monge-Ampère equation:

[0055] \[\det(D^2u_\epsilon-\epsilon\log\rho_\epsilon)=\frac{\mu(x)}{u(ablau_\epsilon(x))}\]

[0056] where u_epsilon is the smoothed potential function (epsilon-approximation); rho_epsilon is the density correction term after thermal nuclear diffusion; this equation represents the effect of entropy regularization in the Monge-Ampère framework: the regularization parameter epsilon modifies the Hessian D2u to a more stable form, similar to the element-by-element exponential scaling of the matrix in the Sinkhorn algorithm.

[0057] The Sinkhorn approximation makes the originally non-differentiable OT distance \(\mathcal{L}_{OT}=\langle\pi^*,c\rangle\) differentiable because: the entropy regularization term \(H(\pi)\) makes the cost function strictly convex, and the solution \(\pi^*\) is unique and smoothly depends on the input \(c\). In the automatic differentiation framework, the gradient of the Sinkhorn iterative can be calculated by implicit differentiation or envelope theorem. In continuous form, this corresponds to:

[0058] \[\boxed{\lim_{\epsilon\to0}abla\mathcal{L}_{OT}^\epsilon=abla\left(\intu\,d\mu+\intv\,du\right)}\]; where \((u,v)\) is the Kantorovich potential function pair, satisfying \(u(x)+v(y)\leqc(x,y)\).

[0059] A hard hat is a protective headgear used during work, protecting the wearer's head from falling objects, small flying objects, and other factors. In recent years, workplace accidents caused by not wearing a hard hat or improperly wearing a hard hat have become increasingly frequent. The impact of these accidents is significant, causing significant pain not only to families but also to businesses. Ensuring that employees wear hard hats properly and protecting the interests of both employees and businesses has been a persistent goal for all parties. Therefore, research on hard hat wear monitoring algorithms is of great significance and has broad application value. Existing hard hat wear monitoring methods generally use a head detection and hard hat wear classification and recognition approach. This approach uses a face detector as a head detector, expanding the face detection frame by approximately two times to form the head detection frame area. When the wearer is facing away from the wearer, the face detector generally performs poorly and cannot identify whether the wearer is wearing a hard hat. To address this, the hard hat monitoring unit described in this embodiment uses YOLOv5s to identify the type of hard hat being worn.

[0060] The average precision of the helmet-wearing recognition method based on YOLOv5s target detection is mAP0.5 = 0.93, and mAP0.5:0.95 = 0.63, which basically meets the performance requirements of the business.

[0061] In a specific implementation, in order to deploy this function on the Android mobile platform, the YOLOv5s model was lightweighted. This allows real-time detection and recognition effects to be achieved on ordinary Android phones, with a CPU (4 threads) of about 30ms and a GPU of about 25ms, which can meet the performance requirements of the business.

[0062] In this embodiment, in the helmet wearing recognition method based on target detection, the helmet wearing category is directly treated as multiple target detection categories for training, and a one-stage method is used for direct end-to-end training, which has the advantages of simple tasks, fast speed, and simple deployment.

[0063] like Figure 2 As shown, the YOLOv5s network in this embodiment includes a Backbone unit, a CBAM attention unit, a neck unit, and a head unit. In order to reduce the number of parameters and FLOPs and to be more suitable for the unfamiliarity of mobile terminals, the Backbone unit in this embodiment is replaced by MobileNetV3. The structural principle diagram of MobileNetV3 is shown in FIG. Figure 3 As shown, the number of parameters is reduced by 30% and FLOPs is reduced by 25%.

[0064] The attention unit is as follows Figure 4 As shown in the figure, CBAM consists of two independent sub-modules, namely the Channel Attention Module (CAMQ) and the Spatial Attention Module (SAM). C*H*W , the one-dimensional channel attention module obtains a one-dimensional convolution M C ∈R C*1*1 , the two-dimensional spatial attention module obtains Ms∈R 1*H*W , so the entire attention process is:

[0065]

[0066] in, Represents element-wise multiplication.

[0067] In this embodiment, adding a CBAM attention unit between the backbone and the neck increases the mAP by 2-3%, which has the effect of suppressing background noise and enhancing the response of the target area.

[0068] In order to further improve the small target detection capability, in this embodiment, the SPPF module in the YOLOv5s network in the prior art is replaced by the ASPP (atrous spatial pyramid pooling) module, whose structure is as follows: Figure 5 As shown in the figure, multi-scale dilated convolution (rates = 6, 12, 18) is used instead of serial pooling to improve the small target detection capability through receptive field diversity and multi-scale feature fusion.

[0069] In recent years, with the continuous development of public transportation, the number of people who choose to travel by rail transit and other means has also increased, and stampedes are prone to occur in crowded environments. In addition, entertainment activities are becoming more and more abundant, and large-scale crowd gatherings are becoming more and more common, and incidents such as collisions and stampedes may occur; and during the epidemic, gatherings of people must be prohibited to prevent the spread of the epidemic. For the monitoring of the density of crowds in public activities, the current method of manual recognition of video surveillance is difficult to be on duty around the clock, so the method of using an intelligent image recognition system to detect crowd density has also become a solution to this problem. In a specific embodiment, the crowd gathering monitoring unit detects the number of people in an area through YOLOv5s to obtain a population density value in the area. When the population density value is greater than a preset threshold, an early warning result is output, so that intelligent reminders and guidance can be made in time when the crowd gathering degree is high.

[0070] The crowd gathering monitoring unit detects the number of people in an area from the collected pictures and videos. The area can be a place with a large flow of people, such as a road junction or a shopping mall. While introducing the algorithm principle, the Python implementation code, PyQt UI interface and training data set are given. For places such as shopping malls and road junctions where the flow of people needs to be controlled, the system uses YOLOv5 to detect the number of pedestrians and calculate the pedestrian density in the area.

[0071] The dataset used for training YOLOv5s in this embodiment is from CUHKOcclusionDataset, which is used to study activity analysis and crowded scenes. It also provides a labeled file. All labels have been converted to a txt format suitable for YOLO. Then the train.py program is executed for training. After the training is completed, the model training situation is observed by the curve of the loss function. The YOLOv5 training mainly includes three aspects of loss: rectangular box loss (box_loss), confidence loss (obj_loss) and classification loss (cls_loss). After the training is completed, a statistical graph of several training processes is generated in the logs directory. Generally, two indicators are encountered, namely recall and precision. Both indicators p and r are simply used to judge the quality of the model from one perspective. Both are values between 0 and 1, where closer to 1 indicates better performance of the model and closer to 0 indicates worse performance of the model. In order to comprehensively evaluate the performance of target detection, the mean average density map is generally used to further evaluate the quality of the model. By setting different confidence thresholds, we can obtain the p-value and r-value calculated by the model under different thresholds. Generally, the p-value and r-value are negatively correlated. Plotting them yields a corresponding curve, where the area of the curve is the AP. An AP value can be calculated for each target in the target detection model. Averaging all AP values yields the model's mAP value, and the optimal model is obtained after training. After obtaining the optimal model, the captured frame image is input into the optimal network for prediction, thereby obtaining the prediction result. With the prediction result, we can frame the pedestrians in the frame image, then use OpenCV drawing operations on the image, predefine the area of the current field of view, and then calculate the pedestrian density in the current image based on the predicted number of targets.

[0072] In this embodiment, the network structure of YOLOv5 mainly includes a Backbone module, a Neck module, and a Head module; the Backbone (backbone network) uses CSP (CrossStagePartial) Darknet53 as its backbone network. This structure helps to reduce the amount of calculation while maintaining the ability to extract features. The Neck (neck network) uses FPN (FeaturePyramidNetwork) and PAN (PathAggregationNetwork) structures, which help to fuse features at different scales, thereby improving the accuracy of detection. The Head (detection head) is used to generate the final detection results, including the category of the target and the position of the bounding box.

[0073] The YOLOv5 algorithm works as follows: YOLOv5 accepts images as input, typically resized to a uniform size (e.g., 640x640). Features are extracted from the input images through the Backbone network. Using the FPN and PAN architectures, features are fused across different layers of the network to capture information at different scales. The object detection results generated by the Head component include the category label and bounding box coordinates for each detected object.

[0074] The YOLOv5 model in this embodiment uses Mosaic data augmentation technology and adaptive anchor box calculation to improve the model's generalization ability and detection effect. The CIoU (Complete Intersection over Union) loss function optimizes the accuracy of bounding box regression.

[0075] The working process of the regional intrusion detection unit is: setting a closed target area, detecting the coordinates of a suspicious target, and issuing an intrusion warning message when the coordinates fall within the closed target area.

[0076] The method for setting the closed area includes: for the collected video, defining the monitoring area by setting boundary points, assigning a unique identifier to each area so that it can be distinguished in subsequent processing, drawing a closed area by a number of straight lines, and saving the image after drawing. If it is an online project, it can be saved to the database, and the data can be read out each time the camera is restarted, instead of being calibrated every time. The overall target area is an enclosed rectangle (if accuracy is required, instance segmentation can be added here). For example, when the calibrated area is a polygon, the polygon can be regular or irregular, but it must be closed. During the detection process, the calibrated area is directly cut off, and then only the calibrated area is detected.

[0077] Intrusion target determination steps: Use background subtraction, frame difference method or deep learning model (such as YOLO, SSD) to detect moving targets in the scene, track the target through Kalman filter, Hungarian algorithm or deep learning model (such as DeepSORT) to obtain the target's motion trajectory.

[0078] Steps to determine whether a target stays for a short time or a long time:

[0079] Residence time calculation: By recording the time when the target enters and leaves the area, the residence time of the target in the area is calculated.

[0080] Set judgment rules: Based on the preset threshold, determine whether the target is a short stay or a long stay. For example, if the stay time exceeds 10 seconds, it is a long stay, otherwise it is a short stay.

[0081] Target coordinate detection steps:

[0082] Coordinate acquisition: The target center coordinates or bounding box coordinates are obtained through the target tracking algorithm.

[0083] Coordinate Mapping: Maps the target's image coordinates to actual geographic coordinates (if necessary).

[0084] For example, a shopping mall's surveillance system can have a restricted VIP area. When someone enters this area, the system records their entry time and coordinates. If the person remains in the area for more than five minutes, the system sounds an alarm and notifies security personnel.

[0085] Through the above methods and examples, regional intrusion detection can be effectively implemented and adjusted and optimized according to actual needs.

[0086] In a specific embodiment, the vehicle illegal parking detection unit detects a vehicle outline area in an area, determines whether the vehicle outline area overlaps with the illegal parking area, and if so, issues a vehicle illegal parking warning message.

[0087] In this embodiment, the features of the entire image are extracted, and the extracted features are processed using multi-layer convolution operations and non-maximum suppression methods to detect the exact location of the target. Then, the accuracy of target detection is improved through inverse correction of the loss function. Finally, the overlap between the detection box and the illegal parking area box and the threshold are used to determine whether the vehicle is illegally parked. Experiments are conducted on the KITTI dataset and images on actual roads to verify that the algorithm can achieve accurate detection of targets.

[0088] The average vehicle average precision (mAP_0.5) based on YOLOv5s is 0.57192, and mAP_0.5:0.95 is 0.41403, which generally meets the performance requirements of the business. Furthermore, to enable deployment on Android mobile platforms, YOLOv5s has been lightweighted, achieving real-time detection and recognition on a typical Android phone in approximately 30ms on a CPU (4 threads) and 25ms on a GPU, essentially meeting the performance requirements of the business.

[0089] Furthermore, the monitoring system includes a road facility monitoring unit, which is used to monitor the covering status of road manhole covers and the overflowing status of garbage bins.

[0090] To monitor tilted, missing, or damaged manhole covers, we can use a YOLOv5s-based object detection method. If a manhole cover is detected to be tilted, missing, or damaged, or if there is trash on the surrounding roads or the trash bins are full, the monitoring results are notified to the relevant departments for processing.

[0091] The monitoring method in this embodiment includes manhole cover tilt detection, manhole cover loss detection and manhole cover damage detection.

[0092] The manhole cover tilt detection process includes:

[0093] Analyze through images or videos whether the edge of the manhole cover is parallel to the ground.

[0094] Identify the tilt angle of manhole covers using geometric features or deep learning models.

[0095] The manhole cover loss detection process includes:

[0096] Detect whether the manhole cover exists, that is, identify whether the wellhead is exposed.

[0097] It is possible to determine whether a manhole cover is covered by comparing it with the background or using a target detection model.

[0098] The manhole cover damage detection process includes:

[0099] Identify abnormal features such as cracks and gaps on the surface of manhole covers.

[0100] Use image segmentation or fine-grained detection methods to locate damaged areas.

[0101] In order to perform the above-mentioned manhole cover detection, YOLOv5s is used for detection in this embodiment. The specific detection process includes a data set collection step and a model training step.

[0102] The dataset collection step is used to collect image and video data containing manhole covers, including various conditions such as normal, tilted, missing, and damaged. If the data is insufficient, data augmentation techniques (such as rotation, scaling, and flipping) can be used to expand the dataset. Manhole covers are labeled using a labeling tool (such as LabelImg), with the labeling categories including "normal," "tilted," "missing," and "damaged." The labeling format is YOLO format (class_idx_centery_centerwidthheight).

[0103] The model training steps include:

[0104] 1) Training preparation:

[0105] Build a YOLOv5s pre-trained model for transfer learning; configure training parameters (such as learning rate, batch size, number of epochs, etc.), and specify the dataset path and category file.

[0106] 2) Start training:

[0107] Run the training script to fine-tune the model; monitor the loss function value and accuracy indicators (such as mAP) during training.

[0108] After training the YOLOv5s model using the above method, the real-time image or video stream is used as input, and the trained YOLOv5s model is used to detect the status of the manhole covers in the detection area. The model outputs the location, category, and confidence information of the manhole covers.

[0109] In order to further improve the detection accuracy of the output results, the "tilted" and "damaged" conditions of the manhole covers in the results output by the YOLOv5s model can be further detected in combination with edge detection, geometric analysis and other methods in existing technologies to improve the accuracy of the detection results and ultimately obtain the status information of the manhole covers.

[0110] When the detected manhole cover status is "tilted", "lost" or "broken", an alarm is immediately issued by sending a notification or recording a log.

[0111] YOLOv5s in this embodiment combines target detection and image segmentation to simultaneously detect the location of manhole covers and damaged areas. Using lightweight backbones such as MobileNet instead of CSPDarknet can improve the running speed. Other parts are the same as the YOLOv5s network structure used in the above embodiment.

[0112] The intelligent monitoring system, installed on urban lighting fixtures, includes a lighting monitoring unit that monitors the operating status of the lighting fixtures based on their real-time status parameters, including current, voltage, power, energy, frequency, and temperature. It provides 0-100% stepless dimming and remote on / off control of streetlights as needed. It also generates alarms for faults such as overvoltage, undervoltage, overcurrent, light source damage, and abnormal lighting. It also features precise positioning using the Beidou system.

[0113] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical principles disclosed herein are intended to be covered by the claims of the present invention.

Claims

1. An intelligent monitoring system, characterized in that: include: Safety helmet monitoring unit, used to detect safety helmet wearing in the construction area; The crowd gathering monitoring unit is used to monitor crowd gatherings and issue a gathering warning when a crowd gathering occurs; Regional intrusion monitoring unit, used to monitor whether an area has been invaded; The vehicle illegal parking detection unit is used to monitor whether there are illegally parked vehicles in the prohibited parking area and issue an illegal parking signal.

2. The intelligent monitoring system according to claim 1, characterized in that: The helmet monitoring unit uses YOLOv5s to identify the category of the helmet being worn.

3. The intelligent monitoring system according to claim 1, characterized in that: The crowd gathering monitoring unit detects the number of people in an area through YOLOv5s to obtain a crowd density value in the area. When the crowd density value is greater than a preset threshold, an early warning result is output.

4. The intelligent monitoring system according to claim 1, characterized in that: The working process of the regional intrusion detection unit is: setting a closed target area, detecting the coordinates of a suspicious target, and issuing an intrusion warning message when the coordinates fall within the closed target area.

5. The intelligent monitoring system according to claim 1, characterized in that: The vehicle illegal parking detection unit detects a vehicle outline area in an area and determines whether the vehicle outline area overlaps with the illegal parking area. If so, a vehicle illegal parking warning message is issued.

6. The intelligent monitoring system according to claim 1, characterized in that: The monitoring system includes a road facility monitoring unit, which is used to monitor the covering status of road manhole covers and the overflowing status of garbage bins.

7. The intelligent monitoring system according to claim 1, characterized in that: The smart monitoring system is installed on the urban lighting device, and the smart monitoring system includes a lighting monitoring unit for monitoring the operating status of the lighting device according to the real-time status parameters of the lighting device.