Dual-angle image processing monitoring system based on artificial intelligence

By adopting the feature fusion method of feature extraction network enhanced by attention mechanism and dynamic weight allocation in the dual-angle image processing monitoring system, combined with the generation adversarial network optimization fusion features, the shortcomings of the existing system in feature extraction and fusion are solved, and high-accuracy target recognition and flexible system applications are achieved.

CN120071253AInactive Publication Date: 2025-05-30HANGZHOU ZHONGXIANG CULTURE MEDIA CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510151222.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing dual-angle image processing monitoring system has shortcomings in feature extraction and fusion, resulting in low accuracy of target recognition, making the system difficult to utilize the spatial and semantic association of the target in complex scenarios, and relies on a large amount of labeled data, which is insufficient flexibility and practicality.

Method used

Using a dual-angle image processing monitoring system based on artificial intelligence, a feature extraction network enhanced by the attention mechanism and a feature fusion method based on dynamic weight allocation is obtained to obtain feature vectors that comprehensively reflect the monitoring scene, and a generative adversarial network is used to optimize the fusion feature.

Benefits of technology

It realizes accurate identification and detailed classification of targets in monitoring scenarios, and can keenly capture tiny anomalies and complex behavior patterns, significantly improving the accuracy of target recognition and system flexibility and practicality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071253A_ABST
    Figure CN120071253A_ABST
Patent Text Reader

Abstract

The invention discloses a dual-angle image processing monitoring system based on artificial intelligence. The system comprises the following steps of S1, dual-angle image acquisition, S2, image preprocessing, S3, feature extraction and fusion, S4, target identification and analysis, and S5, result output and feedback. According to the invention, through a feature fusion technology, multi-scale features and target relation information are comprehensively integrated, and targets in a monitoring scene can be identified extremely accurately and classified meticulously. By means of the method, whether tiny abnormal objects or complex behavior modes can be caught by the system, any detail possibly having potential safety hazards is not released, a highly reliable result is provided for monitoring work, and safe and stable operation of all fields is powerfully guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of monitoring devices, and in particular to a dual-angle image processing monitoring system based on artificial intelligence. Background Art

[0002] A dual-angle image processing monitoring system is an advanced monitoring solution that collects and processes images of a monitoring area from two different perspectives, providing more comprehensive and three-dimensional scene information. Combining artificial intelligence technology, the system can perform intelligent analysis on the collected images and has important application values in fields such as security and traffic management.

[0003] However, there are still many problems in the current dual-angle image processing monitoring system. In feature extraction, most adopt a single-scale method, which can only analyze images from a single scale, making it difficult to comprehensively obtain target features. For targets with large size differences, the features of small targets are easily lost, and the details of large targets are insufficiently extracted, resulting in low target recognition accuracy. In feature fusion, the relationships between targets are often ignored, and an unrelated feature fusion method is used, making it difficult for the system to utilize the spatial and semantic associations of targets in complex scenes, with a low recall rate for anomaly detection and limited multi-target detection capabilities. Moreover, the system relies on a large amount of labeled data for training. When new scenes and targets appear, re-labeling and training are required, which is costly and time-consuming, severely limiting the flexibility and practicality of the system.

[0004] Accordingly, this application proposes a dual-angle image processing monitoring system based on artificial intelligence. Summary of the Invention

[0005] The purpose of the present invention is to solve the drawbacks existing in the prior art, and a dual-angle image processing monitoring system based on artificial intelligence is proposed.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] A dual-angle image processing monitoring system based on artificial intelligence includes the following steps:

[0008] S1. Dual-angle image acquisition:

[0009] First, with the help of image acquisition devices set at two different angles, images of the monitoring area are simultaneously acquired at a predetermined time interval or trigger condition. The devices can adjust the focal length, aperture, and sensitivity to adapt to different scenarios;

[0010] S2. Image preprocessing:

[0011] Subsequently, the acquired dual-angle images are subjected to noise reduction, histogram equalization enhancement, and geometric correction processing to improve the image quality;

[0012] S3. Feature extraction and fusion:

[0013] Extract dual - angle image features and fuse them to obtain a feature vector that comprehensively reflects the monitoring scene, including a feature extraction network enhanced by an attention mechanism, a feature fusion method based on dynamic weight allocation, and optimization of the fused features using a generative adversarial network;

[0014] S4. Object recognition and analysis:

[0015] Input the fused feature vector into a pre - trained artificial intelligence object recognition model to identify the target object and analyze its behavior to determine whether there is an anomaly;

[0016] S5. Result output and feedback:

[0017] Finally, based on the results of object recognition and analysis, output comprehensive monitoring information, output monitoring information including the target category, location, and behavior status, alarm through sound, text message, email, etc. when an anomaly is detected, and store the monitoring data simultaneously.

[0018] Preferably, in the step S1, the adjustable parameters of the image acquisition device can be automatically adjusted according to the real - time light intensity and scene complexity to obtain the optimal image acquisition effect.

[0019] Preferably, in the step S2, the noise reduction process uses an adaptive filtering algorithm, which can automatically adjust the filtering parameters according to the local features of the image to more effectively remove noise.

[0020] Preferably, in the step S3, a deep convolutional neural network is used for feature extraction, which can extract higher - level and more abstract feature information of the image.

[0021] Preferably, in the step S3, when performing feature fusion, a combination of feature splicing and weighted fusion is used to give full play to the advantages of different features.

[0022] Preferably, in the step S3, dimensionality reduction processing is performed on the fused feature vector to reduce data redundancy and improve the efficiency of subsequent processing.

[0023] Preferably, in the step S4, the object recognition model uses a multi - modal fusion method, combining the visual features of the image and other information to improve the accuracy of object recognition.

[0024] Preferably, in the step S5, the alarm mechanism can perform hierarchical alarms according to the severity of the abnormal behavior to provide more targeted warnings for relevant personnel. In the step S5, the stored monitoring data uses encryption and compression technologies to ensure data security and save storage space.

[0025] The present invention has the following beneficial effects:

[0026] 1. Through the feature fusion technology, comprehensively integrating multi-scale features and target relationship information, it can extremely accurately identify and meticulously classify targets in the monitoring scenario. Whether it is a tiny abnormal object or a complex behavior pattern, the system can keenly capture them, not missing any detail that may pose a safety hazard, providing highly reliable results for the monitoring work and effectively ensuring the safe and stable operation of various fields.

[0027] 2. By deeply mining the internal relationships between targets in complex monitoring scenarios, it can make a rapid and flexible response to scene changes. Whether it is a crowded public place or a complex industrial environment, the system can effectively detect various abnormal behaviors and emergencies, ensuring no dead corners in the monitoring scope and providing timely and accurate information support for dealing with various complex situations.

[0028] 3. Through unsupervised feature fusion, it significantly reduces the dependence on a large amount of labeled data and can quickly adapt to new targets and scenarios in a short time. This not only greatly shortens the system deployment cycle but also reduces the human and material costs during the application process, making the system more convenient and efficient in actual applications and significantly improving the practicality and promotion value of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 FIG. is the overall flowchart of a dual-angle image processing monitoring system based on artificial intelligence proposed by the present invention;

[0030] Figure 2 FIG. is a partial code diagram for extracting and fusing different-scale features to capture image information in the first embodiment of the present invention;

[0031] Figure 3 FIG. is a partial code diagram for processing graph data and mining relationship features between nodes in the second embodiment of the present invention;

[0032] Figure 4 FIG. is a partial code diagram of an unsupervised feature fusion model based on adversarial learning in the third embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0034] A dual-angle image processing monitoring system based on artificial intelligence includes the following steps:

[0035] S1. Dual-angle image acquisition

[0036] With two image acquisition devices set at different angles, images of the monitoring area are acquired simultaneously at a predetermined time interval or trigger condition. The image acquisition devices have adjustable focal length, aperture, and sensitivity parameters, which can be automatically adjusted according to the real-time illumination intensity and scene complexity to obtain the optimal image acquisition effect and adapt to different scene requirements.

[0037] S2. Image preprocessing

[0038] Perform noise reduction, histogram equalization enhancement, and geometric correction on the acquired dual-angle images to improve image quality. Among them, the noise reduction process uses an adaptive filtering algorithm, which can automatically adjust the filtering parameters according to the local characteristics of the image and more effectively remove interference signals such as salt-and-pepper noise and Gaussian noise in the image.

[0039] S3. Feature extraction and fusion

[0040] Extract the features of the dual-angle images and perform fusion to obtain a feature vector that comprehensively reflects the monitoring scene, specifically including the following aspects:

[0041] Adopt a feature extraction network enhanced by an attention mechanism and perform feature extraction based on a deep convolutional neural network, which can extract higher-level and more abstract feature information of the image. This network introduces an attention mechanism on the basis of the traditional convolutional neural network, which can automatically focus on the key regions and key features in the image and suppress irrelevant information.

[0042] Based on the feature fusion method of dynamic weight allocation, when performing feature fusion, a combination of feature splicing and weighted fusion is adopted to give full play to the advantages of different features. Use a reinforcement learning algorithm to dynamically adjust the fusion weights of the feature vectors according to the importance of the dual-angle image features in different scenarios, so that the fused feature vector can more accurately reflect the monitoring scene information.

[0043] Use a generative adversarial network to optimize the fused features, and at the same time perform dimensionality reduction on the fused feature vectors to reduce data redundancy and improve the efficiency of subsequent processing. In the generative adversarial network, the generator takes the fused feature vector as the input and generates an optimized feature vector that approximates the real scene distribution, and the discriminator judges the difference between the generated feature vector and the real scene feature vector, and continuously optimizes the fused features through adversarial training.

[0044] S4. Object recognition and analysis

[0045] Input the fused feature vector into a pre-trained artificial intelligence object recognition model to identify the target object and analyze its behavior to determine whether there is an abnormality. The object recognition model adopts a multi-modal fusion method, combining the visual features of the image and other relevant information to improve the accuracy of object recognition.

[0046] S5. Result output and feedback

[0047] Based on the results of target identification and analysis, comprehensive monitoring information is output, including the category, location, behavior status, etc. of the target object. When an abnormality is detected, an alarm is issued through sound, text message, email, etc. The alarm mechanism can provide graded alarms according to the severity of the abnormal behavior, providing more targeted warnings for relevant personnel. At the same time, the stored monitoring data uses encryption and compression technology to ensure data security and save storage space.

[0048] Embodiment 1:

[0049] Step 1: Dual-angle image acquisition

[0050] High-definition cameras are selected and installed at diagonal positions of the shopping mall. Light sensors are deployed in various areas of the shopping mall to obtain light intensity data in real time. According to the empirical formula f = k 1 ×Iog(I)+b 1 To dynamically adjust the focal length, where k 1 =0.5, b 1 =5, I is the light intensity value (unit: lux). The aperture is adjusted according to the density of people. When D is greater than the set threshold (5 people per square meter), the aperture increases by two levels. The aperture value A adjustment formula is A=A 0 +2,A 0 The sensitivity ISO is adjusted in conjunction with the light intensity I and the aperture value A. Based on the exposure compensation formula, the exposure value E is kept within the appropriate range (such as E = 10). The camera is set to capture an image every 15 seconds to meet the real-time monitoring needs of personnel flow and product status.

[0051] Step 2: Image preprocessing

[0052] The bilateral filtering algorithm is used to remove image noise. The bilateral filtering formula is:

[0053] The spatial weight function Range Weight Function Standard deviation s =2,σ r =0.1. Adaptive histogram equalization (CLAHE) is used to enhance image details. CLAHE divides the image into 8×8 blocks. For each pixel in each block, its new gray value p′ is calculated by looking up the pre-calculated gray mapping table of the block, that is, p′=M(p). The geometric distortion of the image is corrected by perspective transformation. The perspective transformation formula is:

[0054]

[0055] By pre - marking multiple feature points in the mall scene, the perspective transformation matrix H is solved using the Direct Linear Transformation (DLT) algorithm.

[0056] Step 3: Feature extraction and fusion

[0057] Construct a feature extraction model based on the Feature Pyramid Network (FPN). First, use ResNet50 as the base network to downsample the dual - angle image to obtain feature maps with different resolutions, namely C2, C3, C4, and C5, and the corresponding downsampling ratios are 4, 8, 16, and 32. During the downsampling process, the convolution kernel size of the convolutional layer is 3×3, the stride is 2, and the padding is 1. Taking the C2 layer as an example, the feature map size H input and W input are the input image sizes. Then, through the top - down path, use 1x1 convolution to adjust the number of channels of the low - resolution feature map to be the same as that of the high - resolution feature Figure 1 map, then perform upsampling (nearest - neighbor interpolation method), and add them element - by - element to the high - resolution feature map to obtain the P2, P3, P4, and P5 feature maps. For example, for the feature map F i l (low - resolution) and F i h (high - resolution), the fused feature map F i = F i l + F i h . In this way, image features at different scales can be obtained, enhancing the recognition ability for target objects of different sizes.

[0058] Step 4: Object recognition and analysis

[0059] Input the fused feature vector into the object recognition model trained based on the SSD algorithm. The SSD model uses VGG16 as the backbone network and performs multi - scale detection on feature maps of different scales. For each position (x, y) on the feature map, a series of default boxes with different scales and aspect ratios are generated. The center coordinates (C x , C y ) of the default box are (x + 0.5, y + 0.5), and the scale s and aspect ratio r are preset according to different feature map levels. During the training process of the model, it is pre - trained using the COCO dataset, and then for the mall scene, 5000 images containing targets such as people and goods are collected for fine - tuning training. During training, the Stochastic Gradient Descent (SGD) optimizer is used, with an initial learning rate of 0.001, a momentum of 0.9, a weight decay of 0.0005, and the loss function is where x is the matching result, c is the predicted class confidence, l is the predicted bounding box, g is the ground truth bounding box, and L conf is the classification loss (softmax loss), and L lof is the regression loss (smoothL1 loss), α is the balancing coefficient (set to 1), and N is the number of default boxes matched. Through the behavior analysis module, the long short-term memory network (LSTM) is used to analyze the trajectory of the target object. The cell state update formula of the LSTM is:

[0060] i t = σ(W ii x t + b ii + W hi h t-1 + b hi )

[0061] f t = σ(W if x t + b if + W hf h t-1 + b hf )

[0062] σ t = σ(W io x t + b io + W ho h t-1 + b ho )

[0063] C t = tanh(W ic x t + b ic + W hc h t-1 + b hc )

[0064]

[0065] h t = o t tanh(C t )

[0066] where i t is the input gate, f t is the forget gate, o t is the output gate, C t is the cell state, and h tIn the hidden state, W is the weight matrix and b is the bias vector. By analyzing the hidden state sequence, it is determined whether there are abnormal behaviors such as a person lingering in a certain area for a long time (more than 5 minutes) or an item being abnormally moved (the moving distance exceeds the set threshold of 2 meters).

[0067] Step Five: Result Output and Feedback

[0068] When an abnormal behavior is detected, the system immediately issues an alarm through the internal broadcast system in the mall. The broadcast content includes the area and type of the abnormal behavior. At the same time, a text message notification is sent to the security personnel's mobile phones, and the text message content details the abnormal situation and location information. The monitoring images and analysis results are stored in real time on the Alibaba Cloud OSS cloud server in the storage formats of MP4 video format and JSON format analysis reports for convenient subsequent query and statistical analysis.

[0069] Embodiment Two:

[0070] Step One: Dual-Angle Image Acquisition:

[0071] High-definition cameras are installed at the entrance and exit of the parking lot respectively. The cameras are equipped with automatic focusing and aperture adjustment functions, and the parameters are adjusted in real time according to the vehicle entry and exit frequency F and light changes. When the light intensity I is lower than 50 lux, the aperture is automatically increased and the sensitivity is increased. The aperture adjustment formula is (A does not exceed the maximum aperture value), and the sensitivity adjustment formula is (ISO does not exceed the maximum sensitivity value). When a vehicle is detected to enter the shooting range, the focusing function is automatically triggered, and according to the focusing distance formula where f is the focal length, u is the object distance, and v is the image distance. By adjusting the focal length, a clear vehicle image is ensured to be captured. It is set that the camera captures an image every 10 seconds to meet the monitoring requirements for vehicle entry, exit, and parking status.

[0072] Step Two: Image Preprocessing

[0073] Median filtering is used to remove salt-and-pepper noise. The median filtering window size is set to 3×3, that is, the pixel value of each pixel point in the image is replaced with the median value of the pixel values in the neighborhood (3×3 neighborhood). Histogram equalization is used to enhance the image contrast. For a pixel point with a gray value of r in the image, its new gray value s is calculated by the formula where L is the total number of gray levels, M and N are the width and height of the image, and n i is the number of pixel points with a gray value of i. Affine transformation is used to correct the geometric distortion of the image. The affine transformation formula is By pre-marking multiple control points in the parking lot scene, the affine transformation matrix is solved using the least squares method.

[0074] Step Three: Feature Extraction and Fusion

[0075] First, use ResNet101 to extract features from the dual-angle images to obtain the feature vectors of the target objects (such as vehicles, pedestrians, etc.). Then, regard the target objects as nodes and construct a graph according to their spatial position relationships (such as Euclidean distance less than 5 meters and direction angle less than 30 degrees) and semantic relationships (such as the subordinate relationship between a vehicle and a driver). Use the graph convolutional network (GCN) to perform convolutional operations on the graph to update the feature representations of the nodes. The formula is:

[0076]

[0077] where A is the adjacency matrix, I is the identity matrix, is the degree matrix of, H (l) is the node feature matrix of the l-th layer, W (l) is the weight matrix of the l-th layer, and σ is the ReLU activation function. The GCN model is set with 3 layers, and the number of hidden units in each layer is 64, 32, and 16 respectively. Finally, concatenate and fuse the relational features obtained by the graph convolutional network with the traditional image features to improve the understanding and analysis ability of the target objects in the parking lot scene.

[0078] Step 4: Target recognition and analysis

[0079] Input the fused feature vectors into the target recognition model trained based on the YOLOv5 algorithm. During the training process of the YOLOv5 model, it is pre-trained using the COCO dataset, and then for the parking lot scene, 3000 images containing information such as vehicle types and license plate numbers are collected for fine-tuning training. During training, the Adam optimizer is used, with a learning rate of 0.0001, beta1 = 0.9, beta2 = 0.999, and the loss function is L = L cls + L obj + L box , where L cls is the classification loss (cross-entropy loss), L obj is the object confidence loss (binary cross-entropy loss), and L box is the bounding box regression loss (CIoU loss). The model can identify information such as vehicle types (such as sedans, SUVs, trucks, etc.) and license plate numbers, and judge behaviors such as the driving direction of the vehicle and whether it illegally parks (the parking time exceeds 15 minutes and it is not in the designated parking space) through the trajectory analysis module.

[0080] Step 5: Result output and feedback

[0081] When abnormal behavior is detected, the system automatically records information such as the license plate number, violation time, and location of the violating vehicle. The owner and the parking lot management staff are notified via SMS, and the SMS content includes the details of the violation and handling suggestions. At the same time, the monitoring data is stored in the MySQL database of the local server in a structured data format for convenient subsequent fee settlement and security management.

[0082] Embodiment Three:

[0083] Step 1: Dual-angle image acquisition

[0084] Install two industrial cameras at different positions in the factory workshop. The cameras have the functions of automatically adjusting the sensitivity and white balance. According to the operation rhythm of the production line and environmental changes, the cameras are set to collect images every 20 seconds. When the ambient light changes, the cameras detect the light intensity through the built-in light sensors and automatically adjust the sensitivity and white balance parameters to ensure clear images of the production equipment and workers' operations are collected.

[0085] Step 2: Image preprocessing

[0086] Use Gaussian filtering to remove Gaussian noise. Gaussian filtering uses a two-dimensional Gaussian function to convolve the image with a standard deviation σ = 1.5. The Retinex algorithm is used to enhance the brightness uniformity of the image. The Retinex algorithm calculates the illumination component and reflection component of the image and enhances the reflection component to improve the brightness uniformity of the image. Geometric correction is performed through polynomial transformation. By pre-marking multiple control points in the factory scene and using the least squares method to solve the polynomial coefficients, the geometric correction of the image is achieved.

[0087] Step 3: Feature extraction and fusion

[0088] Use the VGG16 network to extract features from the dual-angle images respectively, obtaining feature vectors F 1 and F 2 . Then construct a generative adversarial network (GAN). The generator G takes F 1 and F 2 as inputs and generates a fused feature vector. The discriminator D judges the difference between the generated fused feature vector and the feature vector in the real scene. The loss function of the generator is L G =-log(D(G(F 1 ,F 2 ))), and the loss function of the discriminator is

[0089] L D =-log(D(F real ))-log(1 - D(G(F 1 ,F 2 )))

[0090] where F real is the feature vector in the real scenario. During the training process, the Adam optimizer is adopted, and the learning rates of both the generator and the discriminator are 0.0001, beta1 = 0.5, and beta2 = 0.999. Through continuous adversarial training, the generated fused feature vector can better retain the key information of the two-angle images while removing redundancy and noise.

[0091] Step 4: Target recognition and analysis

[0092] The optimized fused feature vector is input into an anomaly detection model that combines LSTM and a convolutional neural network. The convolutional neural network uses the Inception module to extract the spatial features of the image, and LSTM is used to process the time series information to capture the time variation patterns of the equipment operation and the worker operation. During the training process of the model, 5000 groups of normal and abnormal data collected in the factory production scenario are used for training. The cross-entropy loss function and the Adam optimizer are adopted, and the learning rate is 0.001. The model can identify information such as the equipment operation status (such as normal operation, fault warning, fault occurrence), the worker operation actions (such as correct operation, violation operation), etc., and judge whether there are abnormal situations such as equipment faults and worker violation operations.

[0093] Step 5: Result output and feedback

[0094] When an abnormal situation is detected, the system immediately issues an alarm through the alarm in the workshop. The alarm sound lasts for 10 seconds, and the abnormal information and location are displayed on the display screen in the workshop. At the same time, the abnormal information and related monitoring data are stored in the Hadoop Distributed File System (HDFS) of the factory data center in the Parquet file format, which is convenient for subsequent fault troubleshooting and improvement.

[0095] It should be noted that in the comparative examples, in Comparative Example 1, traditional single-scale feature extraction is adopted, and the image features are obtained only from a single scale, making it difficult to comprehensively capture the features of targets of different sizes, resulting in an accuracy of only 85.06%. In Comparative Example 2, the traditional method without relational feature fusion is used, without considering the spatial and semantic relationships between target objects, and the accuracy is 88.21%. In Example 1, through the pyramid network based on multi-scale feature fusion, features are extracted and fused from different scales, and more comprehensive target information can be obtained, with an accuracy reaching 95.23%. In Example 2, the relational feature fusion combining the graph convolutional network is used to model the relationships between target objects, and the accuracy is 93.47%. In Example 3, the unsupervised feature fusion based on adversarial learning enables the features of the two angles to learn and fuse with each other in the confrontation, and the accuracy is 94.18%, as shown in Table 1 specifically.

[0096] Table 1: Comparison Table of Performance Metrics of Dual-Angle Image Processing Monitoring System Based on Artificial Intelligence

[0097]

[0098] It should be noted that in the comparative examples, Comparative Example 1 uses traditional single-scale feature extraction, and Comparative Example 2 uses a traditional method of non-relational feature fusion. In terms of the time-consuming of feature extraction, due to the simple calculation, Comparative Example 1 only takes 121.6 ms, and Comparative Example 2 takes 130.8 ms; Example 1 is based on multi-scale feature fusion, although the performance is improved but the calculation is complex, taking 152.3 ms, Example 2 combines graph convolutional network, the amount of calculation increases, taking 181.7 ms, Example 3 uses adversarial learning, the process is complex, taking 203.5 ms, and the time-consuming of the examples is generally longer than that of the comparative examples. In terms of the number of network layers, the VGG16 in Comparative Example 1 has only 16 layers, the ResNet101 in Comparative Example 2 has 101 layers, the ResNet50 in Example 1 plus FPN has a total of 53 layers, the ResNet101 in Example 2 plus GCN has a total of 104 layers, and the VGG16 in Example 3 plus GAN has a total of 19 layers. The number of network layers and performance are not simply positively correlated. For example, Example 2 has the most network layers, but the performance metrics are not the best. In terms of the loss function value, due to the poor effect of single-scale feature extraction, Comparative Example 1 is as high as 0.765, Comparative Example 2 has non-relational feature fusion, which is 0.632, and Examples 1 to 3 are 0.356, 0.421, and 0.389 respectively, as shown in Table 2 specifically.

[0099] Table 2: Table of Model Parameter Differences of Dual-Angle Monitoring System Based on Artificial Intelligence

[0100]

[0101] Specifically, in terms of target recognition accuracy, the first embodiment is based on a pyramid network with multi-scale feature fusion, reaching 95.23%; the second embodiment combines the relational feature fusion of the graph convolutional network, with an accuracy of 93.47%; the third embodiment is based on unsupervised feature fusion of adversarial learning, with an accuracy of 94.18%. However, the first comparative example using traditional single-scale feature extraction has an accuracy of only 85.06%, and the second comparative example without relational feature fusion has an accuracy of only 88.21%. This reflects the unique feature fusion method of the present invention, which can fully capture target information and accurately identify targets in complex scenarios. In terms of anomaly detection recall rate, the first embodiment is 92.15%, the second embodiment is 90.08%, and the third embodiment is 91.34%. The single-scale feature extraction of the first comparative example does not grasp enough information, only 80.02%, and the second comparative example lacks relational fusion and cannot analyze related behaviors, which is 83.17%. The embodiment captures abnormal behavior characteristics more comprehensively through multi-scale and relational feature fusion, and can detect anomalies more effectively. In terms of multi-target detection capability, the maximum number of detected targets in Example 1 is 50, in Example 2 is 45, and in Example 3 is 48. In comparison example 1, the maximum number of detected targets is only 30 due to the single-scale feature limitation, and in comparison example 2, the maximum number of detected targets is 35 due to the lack of relationship fusion. The fusion of multi-scale features and relationship features in the embodiment provides richer information, making it more advantageous in coping with multi-target detection in complex scenarios.

[0102] Specifically, the feature extraction time is 152.3ms, 181.7ms, and 203.5ms in Examples 1 to 3, respectively. Due to the use of complex feature fusion and network structure, the amount of calculation increases and the time consumption is long. Comparative Examples 1 and 2 are 121.6ms and 130.8ms, respectively. The traditional method is simple to calculate and takes less time. However, the increased time consumption of the embodiment has brought about a significant improvement in performance. In terms of the number of network layers, Example 1 has 53 layers, Example 2 has 104 layers, Example 3 has 19 layers, Comparative Example 1 has 16 layers, and Comparative Example 2 has 101 layers. The embodiments of the present invention focus on the rationality of the network structure, and do not simply rely on increasing the number of network layers to improve performance. For example, Example 2 has the most network layers, but the performance indicators are not optimal. The loss function values ​​of Examples 1 to 3 are 0.356, 0.421, and 0.389, respectively. Due to the better feature extraction and fusion methods, the training effect is better, and the difference between the model prediction value and the true value is smaller. For comparison example 1, the loss function value is as high as 0.765 due to the poor single-scale feature extraction effect, while for comparison example 2, the loss function value is 0.632 due to the unrelated feature fusion.

[0103] In general, the embodiments of the present invention far exceed the control ratio in key performance indicators such as target recognition, anomaly detection, and multi-target detection through innovative feature fusion methods and reasonable network construction. Although the time consumption of feature extraction has increased, the accuracy, stability and multi-target processing capabilities of the system have been effectively improved, and more accurate and reliable monitoring and analysis results can be provided in practical applications.

[0104] Furthermore, as Figure 2 shown, this code implements a pyramid network model for multi-scale feature fusion. Its core purpose is to more comprehensively capture information in images by extracting and fusing features at different scales, improving the accuracy of tasks such as object recognition. In terms of network structure, a PyramidFusionNet class is defined, which contains three convolutional layers (conv1, conv2, conv3) for feature extraction, one max pooling layer (pool) for downsampling, and a bilinear interpolation upsampling layer (upsample). In the forward method, features of different scales are extracted from the input image, and then the feature maps of different scales are upsampled to the same size and concatenated to output the fused feature map. In the training part, the mean squared error loss function (MSELoss) and Adam optimizer are used for 3 epochs of training, continuously adjusting the model parameters to reduce the error between the prediction result and the target.

[0105] Even further, as Figure 3 shown, the code constructs a relational feature fusion model combining a graph convolutional network for processing graph data, mining relational features between nodes, and is applicable to tasks that need to consider node relationships, such as social network analysis and object relationship modeling in object detection. In the network structure, the GCNRelationFusion class contains two graph convolutional layers (gcn1, gcn2) and a ReLU activation function layer (relu). In the forward method, input node features and edge indices are processed through two graph convolutional operations and activation functions to output node features fused with relational information. The training part also uses the mean squared error loss function and Adam optimizer for 3 epochs of training, updating the model parameters according to the error between the predicted node features and the target features.

[0106] Even further, as Figure 4 shown, the code implements an unsupervised feature fusion model based on adversarial learning, consisting of a generator and a discriminator. Through the adversarial training of the two, it learns the latent distribution of the data and achieves unsupervised feature fusion. In the network structure, the generator is a neural network containing two fully connected layers and a LeakyReLU activation function, used to map random noise into fake data; the discriminator is also a fully connected network containing a Sigmoid activation function, used to judge whether the input data is real data or generated fake data. The training part uses the binary cross-entropy loss function (BCELoss) for 3 epochs of training. In each epoch, the discriminator is trained first to accurately distinguish real data from fake data; then the generator is trained to make the generated fake data be able to deceive the discriminator.

[0107] In summary, the multi-scale feature fusion, graph convolutional network, and adversarial learning techniques implemented through code demonstrate powerful performance. The multi-scale feature fusion enables the system to simultaneously capture the overall contour and subtle features of the target when processing surveillance footage, greatly improving the accuracy of target recognition and anomaly detection. Whether it is a large suspicious target or a small abnormal item, it can be quickly detected. The graph convolutional network helps the system to mine the complex relationships between targets in the surveillance scene and convert them into valuable feature information, thereby enhancing the ability to understand and judge complex scenes and effectively identifying situations such as abnormal gatherings and illegal parking. Adversarial learning endows the system with the ability to autonomously learn the data distribution under unsupervised conditions, enabling it to quickly adapt to new surveillance environments and target types without a large amount of manual annotation and repeated training. These technologies cooperate with each other to provide an efficient, intelligent, and flexible solution for the surveillance system, significantly enhancing the practicality and reliability of the system.

[0108] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, should be covered by the protection scope of the present invention.

Claims

1. A dual-angle image processing monitoring system based on artificial intelligence, characterized in that: The following steps are involved: S1. Dual-angle image acquisition: First, with the help of two image acquisition devices set at different angles, images of the monitored area are collected simultaneously at predetermined time intervals or trigger conditions. The devices can adjust the focal length, aperture, and sensitivity to adapt to different scenes; S2. Image preprocessing: Then the collected dual-angle images are subjected to noise reduction, histogram equalization enhancement and geometric correction to improve image quality; S3, feature extraction and fusion: Extract and fuse dual-angle image features to obtain a feature vector that comprehensively reflects the monitoring scene, including a feature extraction network enhanced by an attention mechanism, a feature fusion method based on dynamic weight allocation, and a generative adversarial network to optimize fusion features; S4. Target identification and analysis: The fused feature vector is input into a pre-trained AI target recognition model to identify the target object and analyze its behavior to determine whether there is any anomaly; S5. Result output and feedback: Finally, based on the results of target identification and analysis, comprehensive monitoring information is output, including target category, location, and behavior status. When an abnormality is detected, an alarm is issued through sound, text message, email, etc., and the monitoring data is stored at the same time.

2. The dual-angle image processing monitoring system based on artificial intelligence according to claim 1 is characterized in that: In step S1, the adjustable parameters of the image acquisition device can be automatically adjusted according to the real-time light intensity and scene complexity to obtain the optimal image acquisition effect.

3. The dual-angle image processing monitoring system based on artificial intelligence according to claim 1 is characterized in that: In step S2, the noise reduction process uses an adaptive filtering algorithm, which can automatically adjust the filtering parameters according to the local features of the image to more effectively remove noise.

4. The dual-angle image processing monitoring system based on artificial intelligence according to claim 1 is characterized in that: In step S3, a deep convolutional neural network is used for feature extraction, which can extract more advanced and abstract feature information of the image.

5. The dual-angle image processing monitoring system based on artificial intelligence according to claim 1 is characterized in that: In step S3, when performing feature fusion, a combination of feature concatenation and weighted fusion is adopted to give full play to the advantages of different features.

6. The dual-angle image processing monitoring system based on artificial intelligence according to claim 1 is characterized in that: In step S3, the fused feature vector is subjected to dimensionality reduction processing to reduce data redundancy and improve the efficiency of subsequent processing.

7. The dual-angle image processing monitoring system based on artificial intelligence according to claim 1 is characterized in that: In step S4, the target recognition model adopts a multimodal fusion approach to combine the visual features of the image and other information to improve the accuracy of target recognition.

8. The dual-angle image processing monitoring system based on artificial intelligence according to claim 1 is characterized in that: In step S5, the alarm mechanism can perform graded alarms according to the severity of abnormal behavior to provide more targeted warnings to relevant personnel. In step S5, the stored monitoring data uses encryption and compression technology to ensure data security and save storage space.

Citation Information

Patent Citations

  • Multi-view face feature and audio feature fused emotion recognition method and system

    CN117312992A