A detection method and system based on the Vibe algorithm and artificial neural network

Through the combination of Vibe algorithm and generative adversarial network, the problem of low accuracy of computer vision detection under low light or low definition conditions is solved, efficient motion object detection is achieved, and hardware costs are reduced.

CN114943919BActive Publication Date: 2025-07-25CHINA UNICOM (GUANGDONG) IND INTERNET CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210599662.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-30
Publication Date
2025-07-25
Estimated Expiration
2042-05-30

AI Technical Summary

Technical Problem

The existing computer vision detection methods have low accuracy in motion objects under low light or low definition conditions, and the cost of improving detection accuracy is high.

Method used

The Vibe algorithm is used to extract monitoring image features, combine the generative adversarial network for super-resolution reconstruction, use the improved neural network model for target recognition, extract the moving target features through the Vibe algorithm, use the generative adversarial network to improve image resolution, and use the improved loss function to optimize the training process of the neural network model.

Benefits of technology

It significantly improves the accuracy of motion target detection under low definition conditions, avoids small target omissions, reduces the demand for hardware transformation, and improves the stability and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114943919B_ABST
    Figure CN114943919B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer vision, and more specifically, to a detection method and system based on the Vibe algorithm and artificial neural network. The present invention extracts moving targets in surveillance images through the Vibe background modeling algorithm, which can not only avoid the situation where small-volume moving targets are missed, but also greatly improve the accuracy of detecting moving targets. In addition, the present invention also reconstructs low-resolution images through a generative adversarial network, improving the resolution of moving target images. On the basis of the improved image resolution, the accuracy of detecting moving targets is significantly improved. Finally, the loss of ESRGAN is replaced with Wasserstein loss, making the training process of the generative adversarial network smoother and more stable. The improved Arcface loss significantly enhances the classification effect of the neural network classifier. The performance optimization of the generative adversarial network and the neural network further improves the accuracy of detecting moving targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and more particularly, to a detection method and system based on the Vibe algorithm and artificial neural network. Background Art

[0002] Computer vision is an imitation of biological vision using computers and related devices. Its main task is to process the captured pictures or videos to obtain three-dimensional information of the corresponding scenes, just as humans and many other biological species do every day.

[0003] Computer vision is a discipline on how to use cameras and computers to obtain the data and information of the objects we need to be photographed. Figuratively speaking, it is to install eyes (cameras) and brains (algorithms) on the computer so that the computer can perceive the environment. Computer vision is both an engineering field and a challenging and important research field in the scientific field. Computer vision is a comprehensive discipline that has attracted researchers from various disciplines to participate in its research. These include computer science and engineering, signal processing, physics, applied mathematics and statistics, neurophysiology, and cognitive science, etc.

[0004] The challenge of computer vision is to develop visual capabilities for computers and robots that are comparable to human levels. Machine vision requires image signals, texture and color modeling, geometric processing and reasoning, and object modeling. With the continuous development of computer technology, the functions of computer vision have become increasingly perfect and can perform different detection tasks. Among them, detecting moving targets is a key task in the field of computer vision. Most of the existing methods for detecting moving targets in computer vision are only based on neural networks, and the monitoring images are recognized through neural networks. This detection method based on neural networks can indeed achieve relatively good results when the conditions such as picture illumination, clarity, and the size of moving targets are relatively good. However, in real application scenarios, the illumination is often weak, the clarity is low, and the sizes of moving targets vary. Under these complex conditions, the detection method that purely relies on neural networks often has greatly reduced effects and cannot meet people's needs. Due to the limitations of real application scenario conditions, ordinary shooting equipment generally has difficulty obtaining high-clarity monitoring images, and the clarity of monitoring images is one of the main factors determining the detection accuracy of computer vision. Therefore, for moving targets, the detection accuracy of existing computer vision is not high. If the shooting equipment is modified or professional-level higher shooting equipment is used, the monitoring cost will be greatly increased. In actual scenarios with low light or no light, only low-clarity images can be collected. Affected by the low-clarity images, the detection accuracy of existing moving target detection has dropped severely, which has become one of the urgent problems to be solved in the field of computer vision. Therefore, there is an urgent need for a detection method and system based on the Vibe algorithm and artificial neural network that can improve the detection accuracy of moving targets in low-clarity images. Summary of the Invention

[0005] The present invention aims to overcome at least one defect of the above-mentioned prior art, and provides a detection method and system based on the Vibe algorithm and artificial neural network for solving the problem of low detection accuracy of moving targets in low-clarity images.

[0006] The technical solution adopted by the present invention is:

[0007] A detection method based on the Vibe algorithm and artificial neural network, comprising:

[0008] Obtaining a monitoring image of a moving target;

[0009] Performing feature extraction on the monitoring image based on the Vibe algorithm to obtain a first feature image; the feature is the feature of the moving target; the first feature image is an image of the moving target;

[0010] Performing super-resolution reconstruction on the first feature image based on a generative adversarial network to obtain a second feature image;

[0011] Perform moving target recognition on the second feature image based on a neural network model.

[0012] As a further solution of the present invention, feature extraction is performed on the monitoring image based on the Vibe algorithm to obtain a first feature image, including:

[0013] Select a single-frame image of the monitoring image for background modeling to obtain a background model;

[0014] Extract foreground pixel points from the monitoring image using the pixel points of the background model;

[0015] Perform dilation processing on the foreground pixel points;

[0016] Fuse the foreground pixel points after different dilation processes to obtain a connected domain;

[0017] Calculate the minimum bounding rectangle corresponding to the connected domain to obtain the moving target image.

[0018] As a further solution of the present invention, the generative adversarial network is established based on the ESRGAN algorithm and Wasserstein loss.

[0019] As a further solution of the present invention, the loss function of the generative adversarial network is:

[0020]

[0021] Where, is the overall loss function, is the mean square error weight coefficient, is the perceptual loss weight coefficient, is the adversarial loss weight coefficient;

[0022] The ;

[0023] Where, is the mean square error loss between the generated high-resolution image and the label high-resolution image pixels, E is the expectation, is the pixel value at the coordinate (x, y) of the label high-resolution image, is the high-resolution image generator of the adversarial neural network, is the pixel value at the coordinate (x, y) of the generated high-resolution image, is the low-resolution image, x is the abscissa of the image pixel, and y is the ordinate of the image pixel;

[0024] The ;

[0025] Where, is the perceptual loss, For the adversarial neural network feature extractor, is the labeled high-resolution image;

[0026] The ;

[0027] Wherein, is the difference between the generated high-resolution image and the discrimination confidence of the labeled high-resolution image, is the discriminator.

[0028] As a further aspect of the present invention, the loss function of the neural network model is:

[0029]

[0030] Wherein, N is the number of training samples, e is the base of the natural logarithm, s is the product of the modulus of the decision surface parameter and the modulus of the sample feature vector, is the angle between the sample feature vector and the labeled decision surface, yi is the i-th category, k is the total number of categories, m and m2 are constants, m > 0, m2 < 0.

[0031] The present invention also provides a detection system based on the Vibe algorithm and the artificial neural network, including:

[0032] A target monitoring module for obtaining a monitoring image of a moving target;

[0033] A target extraction module for extracting features from the monitoring image based on the Vibe algorithm to obtain a first feature image; the feature is the feature of the moving target; the first feature image is the moving target image;

[0034] An image reconstruction module for performing super-resolution reconstruction on the first feature image based on the generative adversarial network to obtain a second feature image;

[0035] A target recognition module for recognizing the moving target from the second feature image based on the neural network model.

[0036] As a further aspect of the present invention, the target extraction module includes:

[0037] A modeling unit for selecting a single-frame image of the monitoring image for background modeling to obtain a background model;

[0038] A foreground unit for extracting foreground pixel points from the monitoring image using the pixel points of the background model;

[0039] An expansion unit for performing expansion processing on the foreground pixel points;

[0040] A connectivity unit, which is used to fuse the foreground pixel points after dilation processing to obtain a connected domain;

[0041] A rectangle unit, which is used to calculate the minimum bounding rectangle corresponding to the connected domain to obtain the moving target image.

[0042] As a further solution of the present invention, the generative adversarial network is established based on the ESRGAN algorithm and Wasserstein loss.

[0043] As a further solution of the present invention, the loss function of the generative adversarial network is

[0044] ;

[0045] Among them, is the overall loss function, is the mean square error weight coefficient, is the perceptual loss weight coefficient, is the adversarial loss weight coefficient;

[0046] The ;

[0047] Among them, is the mean square error loss between the generated high-resolution image and the label high-resolution image pixels. E is the expectation, is the pixel value at the coordinate (x, y) of the label high-resolution image, is the high-resolution image generator of the adversarial neural network, is the pixel value at the coordinate (x, y) of the generated high-resolution image, is the low-resolution image, x is the abscissa of the image pixel, and y is the ordinate of the image pixel;

[0048] The ;

[0049] Among them, is the perceptual loss, is the feature extractor of the adversarial neural network, is the label high-resolution image;

[0050] The ;

[0051] Among them, is the difference between the discriminative confidence of the generated high-resolution image and the label high-resolution image, is the discriminator.

[0052] As a further solution of the present invention, the loss function of the neural network model is:

[0053] ;

[0054] Where N is the number of training samples, e is the base of the natural logarithm, s is the product of the modulus of the decision surface parameter and the modulus of the sample feature vector, is the angle between the sample feature vector and the label decision surface, yi is the i-th category, k is the total number of categories, m and m2 are constants, m > 0, m2 < 0.

[0055] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention extracts moving targets in surveillance images through the Vibe algorithm, which can not only avoid the omission of small-sized moving targets but also greatly improve the accuracy of detecting moving targets. In addition, the present invention also reconstructs low-resolution images through a generative adversarial network, improving the resolution of moving target images. On the basis of the improved image resolution, the accuracy of detecting moving targets is significantly improved. Finally, replacing the loss of ESRGAN with Wasserstein loss makes the training process of the generative adversarial network smoother and more stable. Improving the Arcface loss significantly enhances the classification effect of the classifier of the neural network model. The performance optimization of the generative adversarial network and the neural network model further improves the accuracy of detecting moving targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is the flowchart of the method of the present invention;

[0057] Figure 2 is the schematic diagram of the system of the present invention;

[0058] Figure 3 is the training architecture of the generative adversarial network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] The drawings of the present invention are only for illustrative purposes and should not be construed as limiting the present invention. For better illustrating the following embodiments, some components in the drawings will be omitted, enlarged or reduced, which do not represent the dimensions of the actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0060] Embodiment

[0061] The detection method for moving targets in the present invention combines the Vibe algorithm, generative adversarial network, and artificial neural network technologies, avoiding the over-dependence of a single artificial neural network on the input environment. Among them, the Vibe algorithm is a moving target detection algorithm that not only has no specific requirements for the video stream type, color space, and scene content, but also introduces a random selection mechanism into background modeling, describing the random volatility of the actual scene by randomly selecting samples to estimate the background model. On the one hand, the Vibe algorithm adjusts the temporal subsampling factor so that very few sample values can cover all background samples, taking into account both accuracy and computational load. On the other hand, it matches interference information such as noise with the background model before propagation, thus suppressing the propagation of interference information and having a strong noise suppression ability. The generative adversarial network (GAN) mainly consists of two parts, one is the generator G (Generator), and the other is the discriminator D (discriminator). The generator G is used to continuously generate samples and make the continuously generated samples closer and closer to the real samples. The discriminator D is a binary classification model used to continuously distinguish the authenticity of the samples generated by the generator G and make the ability to distinguish the authenticity of the samples higher and higher. Through training, the generator hopes that the generated samples can deceive the discriminator and achieve the purpose of passing off the false as the real. And the discriminator is also improving its ability to distinguish the true from the false through training. As the number of model training iterations increases, the generator and the discriminator learn in the process of mutual game and will eventually reach an equilibrium point. That is, a generator with very good effects that can generate data very close to the real samples is obtained. The artificial neural network (Artificial Neural Network, ANN) is an algorithm model that simulates the behavioral characteristics of animal neural networks to perform distributed parallel information processing. It abstracts the brain neuron network from the perspective of information processing, establishes a certain simple model, and forms different networks according to different connection methods. The artificial neural network is composed of a large number of nodes (or neurons) connected to each other. Each node represents a specific output function, called the activation function (activation function). The connection between each two nodes represents a weighted value for the signal passing through this connection, called the weight, which is equivalent to the memory of the artificial neural network. The output of the network varies depending on the connection method of the network, the weight value, and the activation function. And the network itself usually approximates a certain algorithm or function in nature, or may be an expression of a logical strategy. The artificial neural network has the abilities of self-learning, associative memory, and high-speed searching for optimal solutions, and can well handle problems where the information source is incomplete and contains false images, and give reasonable recognition and judgment.

[0062] As shown in the appendix Figure 1 In the figure, this embodiment provides a detection method based on the Vibe algorithm and artificial neural network. The specific steps include:

[0063] S10. Obtain the monitoring image of the moving target;

[0064] As a preferred embodiment of the present invention, pests are selected as the moving target, such as mice and cockroaches, etc. The night is selected as the time period for collecting the monitoring image, so as to simulate the real environment with poor shooting conditions. The monitoring image is collected in a complex environment with weak or no light, and a low-clarity image can be obtained to better train and verify the effect of the moving target detection model.

[0065] S20. Extract features from the monitoring image based on the Vibe algorithm to obtain a first feature image; the features are the features of the moving target; the first feature image is the moving target image;

[0066] As a further solution of the present invention, extracting features from the monitoring image based on the Vibe algorithm to obtain a first feature image includes:

[0067] Select a single-frame image of the monitoring image for background modeling to obtain a background model;

[0068] Use the pixel points of the background model to extract foreground pixel points from the monitoring image;

[0069] Perform dilation processing on the foreground pixel points;

[0070] Fuse the foreground pixel points after different dilation processes to obtain a connected domain;

[0071] Calculate the minimum bounding rectangle corresponding to the connected domain to obtain the moving target image.

[0072] Specifically, in this embodiment, the Vibe algorithm is adopted to extract moving targets from surveillance images. First, a frame is selected from the surveillance images as the initial image, and the background model is initialized according to the initial image. There are no moving targets in the initial image, and each pixel point of the background model is a sample point. Then, the pixel points of the background model are compared with the pixel points of each frame of the surveillance image. If the difference between the value of the pixel point of the surveillance image and the value of the pixel point of the background model is greater than a preset threshold, it is determined that the pixel point of the surveillance image is a foreground pixel point. Then, dilation processing is performed on all foreground pixel points in the surveillance image to connect and fuse the separated foreground pixel points in the same frame, obtaining a connected domain composed of foreground pixel points. Finally, the minimum bounding rectangle of the connected domain is calculated, and the bounding rectangle is set as the target detection frame. The moving target image is the surveillance image with the target detection frame. In a preferred embodiment, due to various factors such as the external environment and scene changes, to enable the background model to adapt to environmental changes within a certain period of time, the initial model must be continuously updated. The essence of background update is to use the model matched by the current frame to correct the model established by the past frames. The present invention extracts moving targets from surveillance images through the Vibe algorithm, which can not only avoid the situation of small-sized moving targets being missed, but also greatly improve the accuracy of detecting moving targets.

[0073] S30. Based on the generative adversarial network, perform super-resolution reconstruction on the first feature image to obtain a second feature image;

[0074] In this embodiment, the selected moving target is a pest. Since pests may appear in the field of view of the camera at any distance during the surveillance image shooting, the pixels of pests in the surveillance image are often low. In view of this, in this embodiment, the generative adversarial network is used to perform super-resolution reconstruction on the moving target in the surveillance image, that is, the pest, so as to improve the resolution of the moving target. The generative adversarial network has been adversarially trained before performing super-resolution reconstruction on the surveillance image. As shown in the appendix Figure 3 The adversarial training includes:

[0075] Collect moving target images to establish a training sample set; the training sample set includes: low-resolution images and high-resolution images corresponding to the low-resolution image samples;

[0076] A detection model is established based on a generative adversarial network, and the detection model is trained using the training sample set. The detection model includes: an adversarial neural network generator and an adversarial neural network discriminator. The low-resolution image is input into the adversarial neural network generator, and the adversarial neural network generator generates a high-resolution image. The generated high-resolution image and the real high-resolution image are respectively input into the adversarial neural network discriminator and the VGG16 feature extraction network. The adversarial neural network discriminator updates the Wasserstein loss through the generated high-resolution image and the real high-resolution image, and the VGG16 feature extraction network updates the MSE of the corresponding pixels of the feature map through the generated high-resolution image and the real high-resolution image, so as to improve the generation performance and discrimination performance of the detection model. In the present invention, the low-resolution image is reconstructed through a generative adversarial network, and the resolution of the moving target image is improved. On the basis of the improved image resolution, the accuracy of detecting the moving target is significantly improved.

[0077] S40. Perform moving target recognition on the second feature image based on the neural network model.

[0078] The neural network model serves as a classifier for the second feature image after super-resolution reconstruction, and is used to identify the second feature image, that is, whether there is a moving target in the foreground of the monitoring image, and classify the second feature image according to the recognition result. The neural network model has been trained before classifying the second feature image.

[0079] In the above step S30, the generative adversarial network is established based on the ESRGAN algorithm and the Wasserstein loss.

[0080] The ESRGAN (Enhanced Super-Resolution Generative Adversarial Networks) is obtained by improving the Super-Resolution Generative Adversarial Network (SRGAN). Compared with SRGAN, ESRGAN has been improved in three aspects: (1) improving the network structure, adversarial loss, and perceptual loss; (2) introducing the Residual-in-Residu Dense Block (RRDB); (3) using the VGG features before activation to improve the perceptual loss. Compared with SRGAN, the generated images of ESRGAN have more realistic and natural textures, further improving the visual quality.

[0081] GAN has problems such as difficult training, the loss of the generator and discriminator being unable to indicate the training process, and the lack of diversity in the generated samples. In view of this, professionals have proposed Wasserstein GAN. Wasserstein GAN has successfully solved the following problems: (1) Completely solved the problem of unstable GAN training, and there is no longer a need to carefully balance the training degrees of the generator and discriminator. (2) Basically solved the collapse mode problem, ensuring the diversity of the generated samples. (3) Finally, there is a numerical value like cross-entropy and accuracy during the training process to indicate the training progress. The smaller this value is, the better the GAN is trained, and the higher the quality of the images generated by the generator. (4) There is no need for a carefully designed network architecture, and the simplest multi-layer fully connected network can achieve it.

[0082] Since during the training process, the Wasserstein loss can make the model more stable and controllable compared to the loss of ESRGAN, is not prone to the collapse mode, and is more suitable for tasks such as pest detection, in this embodiment, the loss of ESRGAN is replaced with the Wasserstein loss.

[0083] The loss of the ESRGAN is:

[0084] ;

[0085] Among them, ;

[0086] ;

[0087] ;

[0088] The process of replacing the loss of ESRGAN with the Wasserstein loss is: Keep and unchanged, and replace with the following Wasserstein loss:

[0089]

[0090] After replacement, the loss function of the generative adversarial network is:

[0091]

[0092] Among them, is the overall loss function, is the mean square error weight coefficient, is the perceptual loss weight coefficient, is the anti-loss weight coefficient;

[0093] The ;

[0094] Wherein, is the mean square error loss between the generated high-resolution image and the pixels of the label high-resolution image. E is the expectation, is the pixel value at the coordinate (x, y) of the label high-resolution image, is the high-resolution image generator of the adversarial neural network, is the pixel value at the coordinate (x, y) of the generated high-resolution image, is the low-resolution image, x is the abscissa of the image pixel, and y is the ordinate of the image pixel;

[0095] The ;

[0096] Wherein, is the perceptual loss, is the feature extractor of the adversarial neural network, is the label high-resolution image;

[0097] Indicates: label high-resolution image and the low-resolution image The high-resolution image generated by the generator After that, the mean square error loss between all corresponding coordinate pixel values of the two feature maps and extracted by the adversarial neural network feature extractor φ( ).

[0098] The ;

[0099] Wherein, is the difference between the discriminant confidence of the generated high-resolution image and the label high-resolution image, is the discriminator.

[0100] The above formula represents the discriminator For the high-resolution image formed by the low-resolution image The discriminant confidence The expected value of And the discriminant confidence of the discriminator for the label high-resolution image The expectation of The difference between The expectation of between.

[0101] In the above step S40, the neural network model uses Arcfaceloss as the loss for binary classification.

[0102] The Arcfaceloss is as follows:

[0103]

[0104] In the original Arcfaceloss, only the angle between the decision boundary corresponding to the sample and its label class is added with a constant m (m >= 0). In the embodiment, in order to improve the training efficiency and strengthen the training effect, a constant m2 (m2 <= 0) is also added to the angle between the decision boundary of the sample and other classes. The modified loss function is:

[0105]

[0106] where N is the number of training samples, e is the base of the natural logarithm, s is the product of the modulus of the decision surface parameter and the modulus of the sample feature vector, is the angle between the sample feature vector and the label decision surface, yi is the i-th category, generally the label category, k is the total number of categories, and m and m2 are constants, m > 0, m2 < 0.

[0107] As shown in the appendix Figure 2 This embodiment also provides a detection system based on the Vibe algorithm and an artificial neural network, including:

[0108] A target monitoring module for obtaining a monitoring image of a moving target;

[0109] A target extraction module for extracting features from the monitoring image based on the Vibe algorithm to obtain a first feature image; the features are the features of the moving target; the first feature image is an image of the moving target;

[0110] As a further solution of the present invention, the target extraction module includes:

[0111] A modeling unit for selecting a single-frame image of the monitoring image for background modeling to obtain a background model;

[0112] A foreground unit for extracting foreground pixel points from the monitoring image using the pixel points of the background model;

[0113] An expansion unit for performing expansion processing on the foreground pixel points;

[0114] A connection unit for fusing the expanded foreground pixel points to obtain a connected domain;

[0115] A rectangle unit for calculating the minimum circumscribed rectangle corresponding to the connected domain to obtain the moving target image.

[0116] An image reconstruction module, configured to perform super-resolution reconstruction on the first feature image based on a generative adversarial network to obtain a second feature image;

[0117] A target recognition module, configured to perform moving target recognition on the second feature image based on a neural network model.

[0118] In the above image reconstruction module, the generative adversarial network is established based on the ESRGAN algorithm and Wasserstein loss, and the loss function of the generative adversarial network is:

[0119]

[0120] Wherein, is the overall loss function, is the mean square error weight coefficient, is the perceptual loss weight coefficient, is the adversarial loss weight coefficient;

[0121] The ;

[0122] Wherein, is the mean square error loss between the generated high-resolution image and the pixels of the label high-resolution image, E is the expectation, is the pixel value at the coordinate (x, y) of the label high-resolution image, is the high-resolution image generator of the adversarial neural network, is the pixel value at the coordinate (x, y) of the generated high-resolution image, is the low-resolution image, x is the abscissa of the image pixel, and y is the ordinate of the image pixel;

[0123] The ;

[0124] Wherein, is the perceptual loss, is the feature extractor of the adversarial neural network, is the label high-resolution image;

[0125] The ;

[0126] Wherein, is the difference between the discriminative confidence of the generated high-resolution image and the label high-resolution image, is the discriminator.

[0127] In the above target recognition module, the loss function of the neural network model is:

[0128]

[0129] where N is the number of training samples, e is the base of the natural logarithm, s is the product of the modulus of the decision surface parameter and the modulus of the sample feature vector, is the angle between the sample feature vector and the label decision surface, yi is the i-th category, k is the total number of categories, and m and m2 are constants, m > 0, m2 < 0.

[0130] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific implementation manners of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the claims of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A detection method based on the Vibe algorithm and artificial neural network, characterized in that Including: Obtain a monitoring image of a moving target; Extract features from the monitoring image based on the Vibe algorithm to obtain a first feature image; The features are the features of the moving target; The first feature image is an image of the moving target; Perform super-resolution reconstruction on the first feature image based on a generative adversarial network to obtain a second feature image; the generative adversarial network is established based on the ESRGAN algorithm and the Wasserstein loss; Perform moving target recognition on the second feature image based on a neural network model; Before performing super-resolution reconstruction on the monitoring image, the generative adversarial network has undergone adversarial training, and the adversarial training includes: Collect moving target images and establish a training sample set; the training sample set includes: low-resolution images and high-resolution images corresponding to the low-resolution image samples; Establish a detection model based on a generative adversarial network and train the detection model using the training sample set; the detection model includes: an adversarial neural network generator and an adversarial neural network discriminator: input the low-resolution image into the adversarial neural network generator, and the adversarial neural network generator generates a high-resolution image; input the generated high-resolution image and the real high-resolution image into the adversarial neural network discriminator and the VGG16 feature extraction network respectively. The adversarial neural network discriminator updates the Wasserstein loss through the generated high-resolution image and the real high-resolution image, and the VGG16 feature extraction network updates the MSE of the corresponding pixels of the feature map through the generated high-resolution image and the real high-resolution image; The loss function of the generative adversarial network is: Among them, is the overall loss function, is the mean square error weight coefficient, is the perceptual loss weight coefficient, is the adversarial loss weight coefficient, is the mean square error loss between the generated high-resolution map and the pixels of the label high-resolution image, is the perceptual loss, is the difference between the discriminant confidence of the generated high-resolution map and the label high-resolution map; Among them, is the mean square error loss between the generated high-resolution image and the pixels of the label high-resolution image. E is the expectation, is the pixel value at the coordinate (x, y) of the label high-resolution image, is the high-resolution image generator of the adversarial neural network, is the pixel value at the coordinate (x, y) of the generated high-resolution image, is the low-resolution image, x is the abscissa of the image pixel, and y is the ordinate of the image pixel; ; Among them, is the perceptual loss, is the adversarial neural network feature extractor, is the labeled high-resolution image, and E is the expectation, is the labeled high-resolution image is the feature map extracted by the adversarial neural network feature extractor at the pixel value of the (x, y) coordinate, is the high-resolution image generated from the low-resolution image, is the high-resolution image extracted and generated by the adversarial neural network feature extractor at the pixel value of the (x, y) coordinate of the feature map, is the low-resolution image; ; in, is the difference between the generated high-resolution image and the label high-resolution image discrimination confidence, is the discriminator, For expectations, For low-resolution images Generated high-resolution image The confidence level of the judgment, For expectations, Label high-resolution images The confidence level of the judgment.

2. The detection method based on the Vibe algorithm and artificial neural network according to claim 1, wherein Extracting features from the monitoring image based on the Vibe algorithm to obtain a first feature image includes: Select a single-frame image of the monitoring image for background modeling to obtain a background model; Extract foreground pixel points from the monitoring image using the pixel points of the background model; Perform dilation processing on the foreground pixel points; Fuse the different dilated foreground pixel points to obtain a connected domain; Calculate the minimum bounding rectangle corresponding to the connected domain to obtain the moving target image.

3. A detection method based on the Vibe algorithm and artificial neural network according to claim 1, characterized in that The loss function of the neural network model is: Where N is the number of training samples, e is the base of the natural logarithm, s is the product of the modulus of the decision surface parameter and the modulus of the sample feature vector, is the angle between the sample feature vector and the label decision surface, yi is the i-th category, k is the total number of categories, m and m2 are constants, m > 0, m2 < 0, is the cosine value of the angle between the feature vector of the i-th sample and its label category decision surface, k is the number of decision surfaces, j is the serial number of the decision surface, represents the angle between the feature vector of the i-th sample and the j-th non-its-category decision surface, is the cosine value of the angle between the feature vector of the i-th sample and the j-th non-its-label-category decision surface.

4. A detection system based on the Vibe algorithm and artificial neural network, characterized in that, For implementing the detection method based on the Vibe algorithm and the artificial neural network according to any one of claims 1-3, including: A target monitoring module for obtaining a monitoring image of a moving target; A target extraction module for extracting features from the monitoring image based on the Vibe algorithm to obtain a first feature image; the features are the features of the moving target; the first feature image is an image of the moving target; An image reconstruction module for performing super-resolution reconstruction on the first feature image based on a generative adversarial network to obtain a second feature image; the generative adversarial network is established based on the ESRGAN algorithm and the Wasserstein loss; A target recognition module for performing moving target recognition on the second feature image based on a neural network model.

5. According to a detection system based on the Vibe algorithm and the artificial neural network as described in claim 4, wherein The target extraction module includes: A modeling unit, configured to select a single-frame image of the monitoring image for background modeling to obtain a background model; A foreground unit, configured to extract foreground pixel points from the monitoring image by using the pixel points of the background model; An expansion unit, configured to perform an expansion process on the foreground pixel points; A connection unit, configured to fuse the foreground pixel points after the expansion process to obtain a connected domain; A rectangle unit, configured to calculate a minimum bounding rectangle corresponding to the connected domain to obtain the moving target image.

6. The detection system based on the Vibe algorithm and artificial neural network according to claim 4, characterized in that, The loss function of the generative adversarial network is: Among them, is the overall loss function, is the mean square error weight coefficient, is the perceptual loss weight coefficient, is the adversarial loss weight coefficient, is the mean square error loss between the generated high-resolution map and the pixels of the label high-resolution image, is the perceptual loss, is the difference between the discriminant confidence of the generated high-resolution map and the label high-resolution map; ; Among them, is the mean square error loss between the generated high-resolution image and the pixels of the label high-resolution image. E is the expectation. is the pixel value at the coordinate (x, y) of the label high-resolution image. is the high-resolution image generator of the adversarial neural network. is the pixel value at the coordinate (x, y) of the generated high-resolution image. is the low-resolution image, x is the abscissa of the image pixel, and y is the ordinate of the image pixel. ; Among them, is the perceptual loss, is the adversarial neural network feature extractor, is the labeled high-resolution image, and E is the expectation, is the labeled high-resolution image The feature map extracted by the adversarial neural network feature extractor at the pixel value at the (x, y) coordinate, is the high-resolution image generated from the low-resolution image is the adversarial neural network feature extractor The high-resolution image extracted and generated at the pixel value at the (x, y) coordinate of the feature map, is the low-resolution image; ; in, is the difference between the generated high-resolution image and the label high-resolution image discrimination confidence, is the discriminator, For expectations, For low-resolution images Generated high-resolution image The confidence level of the judgment, For expectations, Label high-resolution images The confidence level of the judgment.

7. The detection system based on the Vibe algorithm and artificial neural network according to claim 4, wherein The loss function of the neural network model is: Where N is the number of training samples, e is the base of the natural logarithm, s is the product of the modulus of the decision surface parameter and the modulus of the sample feature vector, is the angle between the sample feature vector and the label decision surface, yi is the i-th category, k is the total number of categories, m and m2 are constants, m > 0, m2 < 0, is the cosine value of the angle between the feature vector of the i-th sample and its label category decision surface, k is the number of decision surfaces, j is the serial number of the decision surface, is the angle representing the feature vector of the i-th sample and the j-th non-its-category decision surface, is the cosine value of the angle between the feature vector of the i-th sample and the j-th non-its-label-category decision surface.

Citation Information

Patent Citations

  • Multi-moving target tracking method based on improved Vibe model and BP neural network

    CN108198207A