Wild animal vehicle-mounted mobile monitoring method based on improved Yolo-v7 model

Through the improved Yolo-v7 model, combined with the super-resolution module and the small target improvement module, the problem of fuzzy targets and low recognition accuracy in on-board mobile monitoring of wild animals is solved, and high accuracy and high efficiency target detection is achieved.

CN120047968AInactive Publication Date: 2025-05-27MINISTRY OF ECOLOGY & ENVIRONMENT CENT FOR SATELLITE APPL ON ECOLOGY ENVIRONMENT
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510028273.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-01-06
Filing Date
2025-01-08
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In wildlife on-board mobile monitoring, animal targets shot from a distance are prone to blur or distortion, making it difficult to guarantee the accuracy of identification and counting.

Method used

The improved Yolo-v7 model is adopted, combining the super resolution module and the small target boost module. The super-resolution module realizes super-resolution reconstruction of images by generating adversarial networks to learn color texture features of animal features and background environments. The small-objective improvement module improves the model's ability to extract small-objectives and anti-interference by introducing an attention mechanism.

Benefits of technology

It improves the accuracy and recall of target detection, can accurately identify and count wild animal targets in complex wild environments, maintain a high inference speed, and show a good balance in parameter quantity and inference speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047968A_ABST
    Figure CN120047968A_ABST
Patent Text Reader

Abstract

The invention discloses a wildlife vehicle-mounted mobile monitoring method based on an improved Yo-v7 model, and relates to the field of target detection. A wild animal picture is obtained and serves as a target detection image to be input into the improved Yo-v7 model, and the improved Yo-v7 model comprises a super-resolution module and a small target lifting module; the super-resolution module learns common animal features in a mobile monitoring image and color texture feature transition of a background environment under different resolutions for a target detection image by using a process of adversarial training of a discriminator and a generator in a generative adversarial network, and realizes super-resolution reconstruction of a mobile ecological monitoring image; the small target lifting module introduces an attention mechanism into a classification weight parameter branch front end of Yo-v7; and the improved Yo-v7 model outputs the position and category information of the detected target and the confidence score, and stores the super-resolution reconstruction result of the corresponding batch of the super-resolution module. Effective technical support is provided for wild animal field monitoring in a natural state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of object detection, and more specifically, to a vehicle-mounted mobile monitoring method for wild animals based on an improved Yolo-v7 model. Background Technique

[0002] Due to the impacts of factors such as climate change, human activity interference, and destruction of biological habitats, global biodiversity is being severely threatened. The monitoring of wild animals can provide necessary basic data for biodiversity conservation, help conservation managers formulate scientific conservation management plans, and is a crucial part of biodiversity conservation. Wild animal monitoring methods are mainly divided into two categories: contact-based and non-contact-based. Contact-based mainly relies on manual monitoring and global positioning system (GPS) collar technology, which requires high professional skills, consumes a large amount of manpower, material resources, and financial resources, and has limited data collection, making it difficult to meet the current needs of global biodiversity conservation. Non-contact-based uses emerging wild animal monitoring technologies such as "3S" technology, digital imaging technology (such as automatic camera technology or trap camera), acoustic recording, and DNA barcoding to carry out wild animal local resource surveys, wild animal population number monitoring, wild animal activity habit research, biodiversity monitoring, and impact assessment of environmental changes on biodiversity, greatly improving the efficiency of wild animal monitoring. Among them, digital imaging technology has the advantages of little interference to animals (Schipper, 2007), strong species capture ability, and convenient archival retrieval of image data, and has become a common means for investigating species diversity, estimating animal population density, recording animal behavior patterns, and studying habitat selection, etc.

[0003] Digital imaging technology obtains wild animal image data, such as photos and videos, through a camera system, and extracts relevant information in the images through model algorithms and data analysis to analyze important information such as the species, population number, behavior, and habitat utilization of wild animals. Traditional trap cameras install digital cameras at hidden fixed positions and use active or passive methods to take static photos or dynamic images of animals. This method arranges the camera at a fixed position, with a fixed shooting angle, a short effective shooting distance, and poor shooting conditions, posing challenges to subsequent data processing and analysis. In recent years, some scholars have begun to install digital imaging devices such as infrared cameras and high-definition cameras on vehicles to obtain animal image data through vehicle fixed-point shooting and mobile cruising shooting. Compared with trap cameras, the mobile cruising shooting method of vehicles can achieve real-time shooting and tracking of wild animal individuals and population targets in long-distance and large-scale scenarios. At the same time, the vehicle-mounted mobile monitoring environment is complex and changeable, the shooting distance is far, the obtained animal targets are small, and the target images are prone to a certain degree of blur or distortion. In this case, how to quickly achieve accurate identification of species and accurate counting of population numbers in vehicle-mounted animal monitoring has become an urgent problem to be solved. Summary of the Invention

[0004] In view of this, the present invention provides a vehicle-mounted mobile monitoring method for wild animals based on an improved Yolo-v7 model to solve the problems in the background technology.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A vehicle-mounted mobile monitoring method for wild animals based on an improved Yolo-v7 model, comprising the following steps:

[0007] Obtain wild animal pictures and input them as target detection images into the improved Yolo-v7 model, where the improved Yolo-v7 model includes a super-resolution module and a small target enhancement module;

[0008] The super-resolution module uses the process of adversarial training between the discriminator and the generator in the generative adversarial network for the target detection images, learns the color texture feature transitions of common animal features and background environments in mobile monitoring images at different resolutions, and realizes the super-resolution reconstruction of mobile ecological monitoring images;

[0009] The small target enhancement module introduces an attention mechanism to the front end of the classification weight parameter branch of Yolo-v7 to improve the model's ability to extract small targets and anti-interference;

[0010] The improved Yolo-v7 model outputs the position and category information of the detected targets, as well as the confidence score, and simultaneously saves the super-resolution reconstruction results corresponding to the batches of the super-resolution module.

[0011] Optionally, the training iteration of super-resolution reconstruction includes training the discriminator and training the generator, specifically as follows:

[0012] Training the discriminator: Randomly extract a high-resolution reference sample x from the training set; Obtain a vector z extracted from the low-resolution input image, and use the generator network to synthesize a pseudo-high-resolution sample x*; Use the discriminator network to score x and x*, and calculate the feature similarity between the two; Calculate the image error and backpropagate the total error to update the trainable parameters of the discriminator, seeking to minimize the feature error;

[0013] Training the generator: Obtain a vector z extracted from the low-resolution input image, and use the generator network to synthesize a pseudo-high-resolution sample x*; Use the discriminator network to score x*; Calculate the feature error and backpropagate to update the trainable parameters of the generator, seeking to maximize the discriminator error.

[0014] Optionally, the super-resolution module is connected to the input branch, including an image retrieval and analysis module and a super-resolution reconstruction module. The image retrieval and analysis module is responsible for analyzing the feature information of the input mobile ecological monitoring image, then searching for a super-clear image with similar texture and color information from the high-resolution reconstruction reference image library, and extracting features to input into the super-resolution reconstruction module to provide feature reference for super-resolution reconstruction. The super-resolution reconstruction module is responsible for fusing the obtained high-resolution features with the low-resolution features in the neck network module to achieve super-resolution reconstruction.

[0015] Optionally, the main structure of the attention mechanism in the small target enhancement module includes a compression unit and an excitation unit connected in sequence. The output of the excitation unit and the original input are both connected to a scale unit for aligning the sizes of the input and output.

[0016] Optionally, the compression module is global average pooling, which is used to obtain the average value of all pixel information within a channel, realizing feature compression in the spatial dimension to obtain average features.

[0017] Optionally, the excitation module realizes channel mixing and correlation calculation by stacking two fully connected layers, and finally restricts the weight range through the Sigmoid activation function.

[0018] Optionally, the attention mechanism is introduced to the front end of the classification weight parameter branch of Yolo-v7, specifically the front ends of the neck network and the head network.

[0019] Optionally, the overall structure of the improved Yolo-v7 is that the target detection image is first preprocessed by the input module and then connected to the feature extraction network module and the super-resolution module respectively. The feature extraction network module inputs the extracted features into the neck network module and simultaneously copies them to the small target enhancement module. The neck network module aggregates the super-resolution reconstruction results and the position small target enhancement information to achieve accurate positioning of the target. Finally, the detection head module fuses the category small target enhancement information to achieve final positioning and classification for result output.

[0020] Optionally, the neck network module adopts the SPP-PAN Neck network, which is used to further extract the features of the image and simultaneously receive the outputs of the super-resolution reconstruction branch and the spatial domain enhancement; SPP-PAN includes a spatial pyramid pooling layer and a path aggregation layer.

[0021] As can be seen from the above technical solutions, compared with the prior art, the present invention provides a vehicle-mounted mobile monitoring method for wild animals based on an improved Yolo-v7 model. The super-resolution reconstruction module can effectively recover high-resolution detailed information from low-resolution images, improving the clarity of the images and the distinguishability of targets. The small target enhancement module can, through the attention mechanism, strengthen the model's feature extraction and localization capabilities for small-sized targets, effectively solving the problems of misdetection and missed detection of a large number of dense small targets during long-distance monitoring.

[0022] The Yolo-v7 model combining the two improvement methods not only improves the accuracy and recall rate of target detection in the wild animal monitoring task, but also, while maintaining a high inference speed, achieves accurate identification and statistics of wild animal targets in complex wild environments. When compared with mainstream detection models, the improved Yolo-v7 model adopted in this study has achieved excellent performance exceeding other models in terms of recall rate, accuracy, and mean average precision (mAP50), and has also shown a good balance in terms of the number of parameters and inference speed. This demonstrates its application potential in actual mobile ecological monitoring tasks. The research results of the present invention provide an effective technical support for the field monitoring of wild animals in their natural state, and have important practical significance for biodiversity conservation and wild animal research. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0024] Figure 1 It is a schematic diagram of the overall structure of the improved Yolo-v7 model of the present invention;

[0025] Figure 2 It is a schematic diagram of the super-resolution reconstruction process of the present invention;

[0026] Figure 3 It is a schematic diagram of the structure of the attention module of the present invention;

[0027] Figure 4 It is a graph showing the changes in precision, recall, and mean average precision (mAP50) during the iterative training process of 4 models;

[0028] Figure 5 It is a schematic diagram of the structure of the original Yolo-v7 model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0030] Embodiment 1

[0031] This embodiment discloses a vehicle-mounted mobile monitoring method for wild animals based on an improved Yolo-v7 model, including the following steps:

[0032] Obtain wild animal pictures and input them as target detection images into the improved Yolo-v7 model. The improved Yolo-v7 model includes a super-resolution module and a small target enhancement module;

[0033] The super-resolution module uses the process of adversarial training between the discriminator and the generator in the generative adversarial network for the target detection image, learns the color texture feature transition of common animal features and background environments in mobile monitoring images at different resolutions, and realizes the super-resolution reconstruction of mobile ecological monitoring images;

[0034] The small target enhancement module introduces an attention mechanism to the front end of the classification weight parameter branch of Yolo-v7 to improve the model's ability to extract small targets and resist interference;

[0035] The improved Yolo-v7 model outputs the position and category information of the detected targets, as well as the confidence score, and simultaneously saves the super-resolution reconstruction results corresponding to the batches of the super-resolution module.

[0036] The original Yolo-v7 network structure diagram is as Figure 5 shown. The entire network consists of an input end (InPut), a feature extraction network (Backbone), a neck network (Neck), and a multi-scale detection head (Head). When the algorithm runs, first, the input end preprocesses the RGB image sent into the network. The preprocessing operations include: data augmentation, image alignment, and reshaping the size;

[0037] The feature extraction network is mainly composed of consecutive stacks of multiple convolutional batch normalization activation (CBS) modules, efficient layer aggregation network (ELAN) modules, and super pooling downsampling (Max Pooling: MP) modules in the above figure, and downsamples the image by 32 times in total. A single convolutional batch normalization activation (CBS) module consists of a convolutional layer (Conv), a batch normalization layer Batchnormalization (BN), and an activation function SiLu. The ELAN module is composed of multiple stacked CBS modules. The entire structure contains rich gradient flow information, can effectively use network parameters, and accelerate network inference. The MP module consists of a CBS and a max pooling layer, and is mainly used for downsampling operations, which can effectively reduce feature loss;

[0038] The neck network mainly consists of modules such as Spatial Pyramid Pooling (SPPCSPC), Efficient Layer Aggregation Networks (ELAN), and CBS to form a top-down feature pyramid (Feature pyramid network: FPN) and a bottom-up path aggregation network (Path Aggregation Network: PAN). The efficient layer concatenation network (ELANW) adds two concatenation operations compared to the efficient layer aggregation network (ELAN). SPPCSPC is mainly used to increase the receptive field. FPN mainly upsamples features containing high-dimensional and strong semantic information to enhance semantic expressions at multiple scales. PAN mainly transmits shallow localization information to the deep layer to increase the localization ability at multiple scales, thereby achieving efficient feature fusion.

[0039] The multi-scale detection head is mainly of the REP structure. REP adopts a reparameterization design, trading off increased training cost for rapid inference. The fused multi-scale feature maps (80×80×128, 40×40×256, 20×20×512 respectively) are used to predict the anchor box positions and object classifications through the REP module, and finally the CIoU loss function is used for backpropagation to update the weight parameters, thereby enabling the model to detect targets of different scales.

[0040] Compared with general captured images, the content in the mobile ecological monitoring images shows characteristics such as low contrast and blurred details. In response to this characteristic, the present invention proposes a super-resolution reconstruction module (Super resolution), whose basic idea is to generate high-resolution images from low-resolution images in order to better display details and edges. Missing details are filled by using the information within the image. The super-resolution reconstruction process is as follows Figure 2As shown, it is essentially a small generative adversarial network model (GAN) that has been cropped. Structurally, it mainly consists of two parts: a generator and a discriminator. The entire reconstruction process is a process of adversarial training between the discriminator and the generator in the generative adversarial network, learning the transition of color texture features of common animal features and background environments in moving monitoring images at different resolutions, so as to achieve super-resolution reconstruction of moving ecological monitoring images.

[0041] For each training iteration in the super-resolution reconstruction process, the following operations are performed:

[0042] (1) Train the discriminator

[0043] a. Randomly extract a high-resolution reference sample x from the training set.

[0044] b. Obtain a vector z extracted from the low-resolution input image, and use the generator network to synthesize a pseudo-high-resolution sample x*.

[0045] c. Score x and x* using the discriminator network and calculate the feature similarity between the two.

[0046] d. Calculate the image error and backpropagate the total error to update the trainable parameters of the discriminator, seeking to minimize the feature error.

[0047] (2) Train the generator

[0048] a. Obtain a vector z extracted from the low-resolution input image, and use the generator network to synthesize a pseudo-high-resolution sample x*.

[0049] b. Score x* using the discriminator network.

[0050] c. Calculate the feature error and backpropagate to update the trainable parameters of the generator, seeking to maximize the discriminator error.

[0051] In the actual application process, each sub-module in Superresolution contains a depthwise separable convolutional layer, a batch normalization layer (BatchNormalization), and a non-linear activation layer (Leaky ReLU). Among them, the convolutional layer for feature extraction uses a convolutional kernel with a stride of 1, and the convolutional layer for feature dimensionality reduction uses a convolutional kernel with a stride of 2. Then, the fully connected layer is used to project the extracted features into a 512-dimensional vector, and the 512-dimensional vector is projected into a fixed value. Before passing through the Sigmoid function, it is normalized to avoid falling into the saturation region of the Sigmoid function. Then, the result after passing through the Sigmoid function is output as the discrimination score of the input image. After the improvement, through a few simple rounds of training, the detailed information of the image can be effectively improved, which not only reduces the training time but also makes the captured animal images clearer and more realistic, providing a more accurate data basis for subsequent animal species classification and quantity statistics.

[0052] Small targets refer to target objects with small sizes and large quantities that need to be detected and located in an image. Various scholars usually adopt methods such as multi-scale representation, introducing context information, and adopting region candidate strategies to improve the accuracy of small target detection. However, the above methods mainly target two-stage detection methods, which inevitably require increasing the number of layers of the model. As a result, more hardware resources need to be purchased to cope with the increase in the number of model parameters. Moreover, the improvement effect is limited on complex and low-resolution field ecological monitoring images, and it will also lead to a decrease in the inference running speed, making it impossible to process a large number of monitoring videos and photos in a timely manner.

[0053] In deep learning, the attention mechanism is used to help the model acquire the ability to focus on important information. It is often a specific structure that can enhance the network's ability to distinguish the importance of features and exclude unnecessary interference information. The attention mechanism in the visual field is mainly studied in classification models. Usually, the attention module is first added to the backbone network of the classification model, and then it is judged whether the attention module is effective according to the classification accuracy.

[0054] Aiming at the problem of dense small target detection in the mobile ecological monitoring video scenario, the small object optim module proposed in the present invention is essentially an improved method based on the fusion attention mechanism. The attention mechanism is introduced to the front end of the classification weight parameter branch of Yolo-v7, that is, the Neck module and the Head module mentioned above, to improve the model's ability to extract small targets and resist interference.

[0055] The attention module structure designed in the present invention is as Figure 3As shown in the figure, the main body consists of a compression unit and an excitation unit. Among them, the compression module is a simple global average pooling, which is used to obtain the average value of all pixel information within a channel, achieving feature compression in the spatial dimension. It is equivalent to obtaining the average feature with a global-sized receptive field. After obtaining the overall features of all channels, the excitation module is responsible for calculating the weight values corresponding to each channel based on the correlation of these features, and explicitly learning and modeling the correlation between feature channels through a learnable structure. The excitation module uses a fully connected layer with learnable parameters to calculate the correlation coefficients between different channels. At the same time, in order to introduce non-linearity, an activation function is added after the overall structure. The excitation module realizes channel mixing and correlation calculation by stacking two fully connected layers, and finally limits the weight range through the Sigmoid activation function. The module aligns the sizes of the input and Scale output, enabling it to be flexibly inserted into the backbone network feature fusion and detection head. The entire small target improvement process only requires 2 layers of SENet, which act on the spatial domain information and channel domain information respectively, that is, responsible for improving the position and category information of small targets. Such a structural design not only reduces redundancy but also can be flexibly set for different-sized feature maps. It can be plug-and-play in the Yolo-v7 model. The training process only focuses on small targets, enabling it to only focus on the target attention learning under a certain scale threshold, further reducing the training difficulty. Although general multi-layer convolutional neural networks also have structures for extracting main features, reducing interference, and implicitly focusing on the learning function of key regions, this ability needs to be manifested through a large amount of specific data training, with high training difficulty and high hardware costs. The attention mechanism can use a special module to complete this task. Relatively speaking, the training is simpler, and the model also converges better for dense small targets.

[0056] Embodiment 2

[0057] To improve the deficiencies of the current wildlife monitoring in natural scenes mainly relying on traditional fixed camera means, this study constructs a mobile monitoring dataset based on a photo collection of four animals, namely sheep, birds, antelopes, and deer, taken by a vehicle-mounted camera monitoring system.

[0058] Aiming at the problems of low picture resolution and low wildlife monitoring ability in long-distance scenes for mobile monitoring, an improved Yolo-v7 model combining a super-resolution reconstruction module and a small target improvement module is proposed.

[0059] The complete structure of the improved Yolo-v7 is as Figure 1As shown in the figure, during the entire detection process, the input image first undergoes InPut preprocessing and is then connected to the Backbone module and the super-resolution module respectively. The Backbone module is responsible for extracting features and inputting them into the Neck module, and at the same time copying them to the small object enhancement module. The Neck module aggregates the super-resolution reconstruction results and the small object enhancement information of the first step to achieve accurate positioning of the target. Finally, the detection head Head fuses the small object enhancement information of the second step to achieve final positioning and classification and outputs the results.

[0060] The specific processing flow of the network is as follows:

[0061] InPut: The image to be detected is used as the input, and the size of the input image can be arbitrary.

[0062] Super-resolution module: Connects to the InPut branch and includes the Image retrieval analysis and Super resolution modules. The Image retrieval analysis is responsible for analyzing the feature information of the mobile ecological monitoring image input in the InPut, and then searching for a high-definition image with similar texture and color information from the high-resolution reconstruction reference image library, extracting features and inputting them into the Super resolution module to provide feature references for super-resolution reconstruction; The Super resolution module is responsible for fusing the obtained high-resolution features with the low-resolution features in the Neck module to achieve super-resolution reconstruction.

[0063] Backbone network: A pre-trained convolutional neural network is used as the Backbone network to extract the features of the image.

[0064] Small object enhancement module (Small object optim): Two groups of SENT modules are responsible for strengthening the small object feature information initially extracted in the Backbone and converting it into position and category information, which are respectively sent to the Neck and Head modules to improve the accuracy of the overall Yolo-v7 model for small objects.

[0065] Neck network: Based on the Backbone network, Yolo-v7 uses a Neck network called SPP-PAN to further extract the features of the image, and at the same time receives the outputs of the super-resolution reconstruction branch and the spatial domain enhancement. The SPP-PAN network includes a spatial pyramid pooling (SPP) layer and a path aggregation (PAN) layer, which can effectively improve the accuracy of object detection.

[0066] Head Network: On the output enhanced in the Neck network domain channel domain, Yolo-v7 uses the Head network to predict the positions and classes of objects in the image. The Head network contains some convolutional layers and fully connected layers, as well as some special operations such as Focal Loss and CloU Loss used in Yolo-v7, which can improve the accuracy of object detection.

[0067] NMS: In the prediction results output by the Head network, the non-maximum suppression (NMS) algorithm is used to remove overlapping bounding boxes.

[0068] Output Results: Finally, Yolo-v7 will output the position and class information of the detected objects, as well as their confidence scores, and save the super-resolution reconstruction results corresponding to the batches of the super-resolution module.

[0069] To further verify the effectiveness of the improved Yolo-v7 model in the present invention, the following experiments were conducted:

[0070] This experiment was carried out on a server equipped with 128GB of memory, an Intel Core i9-13900k CPU, and 2 NVIDIA RTX 3090 graphics cards. Using the Python programming language, Pytorch was used as the deep model framework, and CDUA and cuDNN adapted to the graphics card driver were configured to call GPU acceleration for training. The specific experimental environment is shown in Table 1.

[0071] Table 1 Experimental Environment and Hyperparameter Registration

[0072]

[0073] To better evaluate the specific role played by the super-resolution reconstruction module and the small object enhancement module after being added to the Yolo-v7 model in the wild animal object detection task, this study designed four comparative experimental models, which are: Yolo-v7 (raw) without any modification, Yolo-v7 (sr) with only the super-resolution reconstruction module added, Yolo-v7 (sm) with only the small object enhancement module added, and the improved Yolo-v7 (sr_sm) with both the super-resolution reconstruction module and the small object enhancement module added. The maximum number of iterations during model training was set to 300 times, and the four models were trained respectively using the dataset of the present invention. The specific hyperparameter configuration for model training is shown in Table 1.

[0074] Figure 4Shows the changes in precision, recall, and mean average precision at 50 (mAP50) during the iterative training process of four models. It can be seen from the figure that compared with the results of the original Yolo-v7 (raw) model, both Yolo-v7 (sr) and Yolo-v7 (raw) have shown obvious improvements, and the introduction of the super-resolution reconstruction method has a more obvious improvement on the object recognition effect of the Yolo-v7 model. However, there are several relatively large fluctuations in precision. The reason is analyzed that there is no matching reference for super-resolution image reconstruction corresponding to the low-resolution input image. At the same time, from the comprehensive three precision evaluation indicators, it can be seen that Yolo-v7 (sr_sm) with both super-resolution reconstruction module and small object improvement module has achieved the most balanced results.

[0075] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined in the present invention can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown in the present invention, but will conform to the widest scope consistent with the principles and novel features disclosed in the present invention.

Claims

1. A method for monitoring wildlife movement in a vehicle based on an improved Yolo-v7 model, characterized in that: The following steps are involved: Obtain wildlife pictures as target detection image input to improve the Yolo-v7 model, wherein the improved Yolo-v7 model includes a super-resolution module and a small target enhancement module; The super-resolution module uses the process of adversarial training between the discriminator and the generator in the generative adversarial network to learn the color and texture feature transition of common animal features and background environments in mobile monitoring images at different resolutions, thereby realizing super-resolution reconstruction of mobile ecological monitoring images; The small target enhancement module introduces the attention mechanism into the front end of the classification weight parameter branch of Yolo-v7 to improve the model's ability to extract and resist interference with small targets; The improved Yolo-v7 model outputs the location and category information of the detected targets, as well as the confidence score, while saving the super-resolution reconstruction results of the corresponding batch of the super-resolution module.

2. According to claim 1, a method for monitoring the movement of wild animals in a vehicle based on an improved Yolo-v7 model is characterized in that: The training iterations of super-resolution reconstruction include training the discriminator and training the generator, as follows: Train the discriminator: randomly extract a high-resolution reference sample x from the training set; obtain a vector z extracted from the low-resolution input image, and use the generator network to synthesize a pseudo high-resolution sample x*; use the discriminator network to score the reference sample x and the pseudo high-resolution sample x*, and calculate the feature similarity between the two; calculate the image error and back-propagate the total error to update the trainable parameters of the discriminator, seeking to minimize the feature error; Train the generator: obtain a vector z extracted from a low-resolution input image, use the generator network to synthesize a pseudo high-resolution sample x*; use the discriminator network to score the pseudo high-resolution sample x*; The feature error is calculated and back-propagated to update the trainable parameters of the generator, seeking to maximize the discriminator error.

3. According to claim 1, a method for monitoring the movement of wild animals in a vehicle based on an improved Yolo-v7 model is characterized in that: The super-resolution module is connected to the input branch and includes an image retrieval analysis module and a super-resolution reconstruction module. The image retrieval analysis module is responsible for analyzing the feature information of the input mobile ecological monitoring image, and then searching for super-clear images with similar texture and color information from the high-resolution reconstruction reference image library, extracting features and inputting them into the super-resolution reconstruction module to provide feature references for super-resolution reconstruction; The super-resolution reconstruction module fuses the obtained high-resolution features with the low-resolution features in the neck network module to achieve super-resolution reconstruction.

4. According to claim 1, a method for monitoring wildlife movement in a vehicle based on an improved Yolo-v7 model is characterized in that: The main structure of the attention mechanism in the small target improvement module includes a compression unit and an excitation unit connected in sequence, and the output and the original input of the excitation unit are connected to the scale unit to align the sizes of the input and output.

5. According to claim 4, a method for monitoring the movement of wild animals in vehicles based on an improved Yolo-v7 model is characterized in that: The compression module is a global average pooling, which is used to obtain the average value of all pixel information in a channel, realize feature compression in the spatial dimension, and obtain average features.

6. According to claim 4, a method for monitoring the movement of wild animals in vehicles based on an improved Yolo-v7 model is characterized in that: The excitation module realizes channel mixing and correlation calculation by superimposing two fully connected layers, and finally limits the weight range through the Sigmoid activation function.

7. The method for monitoring wildlife movement in a vehicle based on an improved Yolo-v7 model according to claim 1, characterized in that: The attention mechanism is introduced into the front end of the classification weight parameter branch of Yolo-v7, specifically the front end of the neck network and the head network.

8. The method for monitoring wildlife movement in a vehicle based on an improved Yolo-v7 model according to claim 1 is characterized in that: The overall structure of the improved Yolo-v7 is that the target detection image is first pre-processed by the input module and then connected to the feature extraction network module and the super-resolution module respectively. The feature extraction network module inputs the extracted features into the neck network module and copies them to the small target enhancement module at the same time. The neck network module summarizes the super-resolution reconstruction results and the position small target enhancement information to achieve accurate positioning of the target. Finally, the detection head module fuses the category small target enhancement information to achieve the final positioning and classification for result output.

9. The method for monitoring wildlife movement in a vehicle based on the improved Yolo-v7 model according to claim 8 is characterized in that: The neck network module adopts the Neck network of SPP-PAN to further extract the features of the image, while receiving the output of the super-resolution reconstruction branch and the spatial domain enhancement; SPP-PAN includes a spatial pyramid pooling layer and a path aggregation layer.

Citation Information

Patent Citations

  • Single-frame infrared weak and small target detection method based on improved YOLOv7

    CN117611911A

  • Urban road traffic sign detection and identification method based on improved LEP-YOLO v7

    CN117636296A

  • PCB surface defect detection method based on improved YOLOv7

    CN118071715A

  • Generation method, system and apparatus capable of visual resolution enhancement, and storage medium

    WO2022242029A1