A wild animal species identification method suitable for low-light environments

By combining wavelet transform and fast Fourier transform with residual networks, disjoint weights in the Fourier domain and dynamic weights in the spatial domain are generated, which solves the problem of accuracy in identifying wild animal species in low-light environments and improves the recognition rate.

CN120954057BActive Publication Date: 2026-01-06HANGZHOU MEARI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511481299.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2026-01-06
Estimated Expiration
2045-10-16

AI Technical Summary

Technical Problem

In low-light environments, the accuracy of wildlife species identification is insufficient due to the loss of detailed information.

Method used

By combining wavelet transform and fast Fourier transform with a residual network, high- and low-frequency images are acquired losslessly. Spatial domain dynamic weights are generated using disjoint weights in the Fourier domain and an attention mechanism, and image recognition is performed by combining information such as geographic location and time.

Benefits of technology

It effectively preserves high-frequency detail information in low-light environments, improving the accuracy and recognition rate of wildlife species identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954057B_ABST
    Figure CN120954057B_ABST
Patent Text Reader

Abstract

The application relates to the field of image recognition classification, and discloses a wild animal species identification method suitable for a low-illumination environment, which comprises the following steps: target image acquisition, acquiring a target image in which wild animals exist in real time; image processing, performing image processing on the acquired target image through wavelet transform; image recognition, using deep learning to learn and acquire weights of the processed image, performing Fourier transform on the weights, acquiring Fourier domain information, combining image meta-information, performing vectorization on the meta-information through a nonlinear embedded multilayer full-connection neural network, and combining the information after vectorization to perform image recognition. Through the improved method of superimposing frequency domain information on space domain information, combining time information and geographical position information of animal appearance, and under the premise of increasing a small amount of parameters, the outdoor wild animal recognition capability is effectively improved, and especially the wild animal fine classification capability under low illumination and low resolution is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition and classification, and in particular to a method for identifying wild animal species in low-light environments. Background Technology

[0002] Currently, most outdoor cameras only record video, requiring manual filtering of footage for wildlife tracking, which is time-consuming and labor-intensive. Combining IoT, deep learning, and computer vision technologies to detect and identify outdoor wildlife not only provides automated monitoring and real-time analysis capabilities but also effectively reduces human capital investment, significantly improves work efficiency, and contributes to wildlife conservation and sustainable ecological development.

[0003] Common detection and classification techniques use fixed convolution kernels in the spatial domain for downsampling, feature extraction, and other operations. While these techniques can effectively capture local information, they often lose high-frequency components (edges, details, etc.) and are not satisfactory in representing global information.

[0004] Wild animals in the wild are mostly active at night, and the images captured by the equipment are of poor quality with significant loss of detail. Combined with the loss of detail in the algorithm, this seriously affects the accuracy of wildlife detection and identification.

[0005] As disclosed in prior art CN202011637663.6, an image super-resolution processing method includes: acquiring a training image set, which includes multiple image groups, each image group including a first image and a second image corresponding to each other, wherein the resolution of the first image is lower than that of the second image; building a first model or a second model based on a Fourier domain feature channel attention mechanism and a convolutional neural network; training the first model or the second model using the training image set; and performing super-resolution processing on the image to be processed using the trained first model or the second model. Summary of the Invention

[0006] This invention addresses the problem in existing technologies where the loss of detailed information affects the accuracy of wildlife species identification in low-light environments, and provides a method for wildlife species identification suitable for low-light environments.

[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0008] A method for identifying wild animal species in low-light environments, the method comprising:

[0009] Acquisition of target images, including real-time acquisition of target images containing wild animals;

[0010] Image processing involves performing wavelet transform on the acquired target image.

[0011] Image recognition uses deep learning to learn weights from the processed image, performs Fourier transform on the weights to obtain Fourier domain information, combines the image's metadata, and vectorizes the image's metadata through a multi-layer fully connected neural network with nonlinear embedding. Finally, it combines the vectorized information to perform image recognition.

[0012] Preferably, the method also includes improvements to the residual network, using deep learning to generate an improved image from the processed image through the improved residual network.

[0013] Preferably, for image processing, image processing of the acquired target image using wavelet transform includes:

[0014] The acquisition of wavelet domain images involves obtaining images of different frequencies through wavelet transform, and using a lossless method to obtain high and low frequency images, i.e., wavelet domain images.

[0015] The generation of Fourier domain weights involves using a fast Fourier transform on the learned weights to generate mutually disjoint Fourier domain weights.

[0016] The generation of spatial domain information utilizes the disjoint weights of the Fourier domain obtained in the Fourier domain, and obtains the corresponding spatial domain weights through inverse Fourier transform. Combined with the attention mechanism, dynamic weights of the spatial domain are generated. Spatial domain information includes dynamic weights of the spatial domain.

[0017] Image information fusion involves fusing the acquired wavelet domain image, Fourier domain weights, and spatial domain information.

[0018] Preferably, the wavelet domain image is obtained by dividing the image into four low- and high-frequency images using Haar wavelet transform: LL image, LH image, HL image, and HH image. LL represents the horizontal low-frequency and vertical low-frequency images; LH represents the horizontal low-frequency and vertical high-frequency images; HL represents the horizontal high-frequency and vertical low-frequency images; and HH represents the horizontal high-frequency and vertical high-frequency images. The resolution of each image is half that of the input image.

[0019] Preferably, the generation of Fourier domain weights involves using a Fast Fourier Transform (FFT) on the learned weights to generate mutually disjoint Fourier domain weights. Obtaining mutually disjoint Fourier domain weights includes:

[0020] The input image is first subjected to a Fourier transform, and then the Fourier parameters are convolved. The total number of parameters that need to be learned is... Where k is the convolution kernel, C in For the input image parameters, C outGiven the output image parameters; divide these parameters into N disjoint groups; reorganize the parameters of the Fast Fourier Transform into... p represents the recombined parameters; and each parameter is associated with a Fourier coordinate (u, v); where u is the Fourier abscissa and v is the Fourier ordinate.

[0021] Calculate (u, v) using the L2 norm. The values ​​are sorted from highest to lowest, dividing the frequency domain parameters into n disjoint groups, i.e. ,in, This is the set of parameters corresponding to different groups after grouping.

[0022] As a preferred method, the generation of spatial domain information utilizes the non-overlapping weights of the Fourier domain learned in the Fourier domain, obtains the corresponding spatial domain weights through inverse Fourier transform, and generates dynamic spatial domain weights by combining an attention mechanism. Spatial domain information includes dynamic spatial domain weights.

[0023] Using the Inverse Discrete Fourier Transform (IDFT) to convert each set of parameters Mapping to the spatial domain using formulas ,

[0024] Where k is the convolution kernel, Let be the corresponding parameter of the i-th group in the frequency-disjoint group and whose Fourier domain coordinates are (u, v). This represents the transformation result of the spatial domain point (p, q), where p is the abscissa of the spatial domain and q is the ordinate of the spatial domain; if If there is a response, then ,otherwise, ;

[0025] Weights for each group , By cutting and recombining, Transformed into a nucleus of size The number of image patches is , get n parameter, Weights are determined by Obtained in the frequency domain.

[0026] As a preferred approach, the image's longitude, latitude, altitude, and metadata are combined, and the metadata is vectorized using a multi-layer fully connected neural network with nonlinear embedding.

[0027] Longitude, latitude, and altitude information are mapped in the following way: [latitude, longitude, altitude] -> [x, y, z], where x is the latitude information, y is the longitude information, and z is the altitude information.

[0028] The acquired time is mapped using sin and cos trigonometric functions.

[0029] Where month represents months and hour represents hours;

[0030] Normalize the temperature and humidity information.

[0031] ,

[0032] Where temperature is the temperature information and humidity is the humidity information;

[0033] Obtain the longitude, latitude, altitude, and metadata of the location, and use nonlinear embedding to map the longitude, latitude, altitude, and metadata into vectors.

[0034] Preferably, improvements to the residual network using wavelet transform and fast Fourier transform include:

[0035] The acquisition of wavelet domain images involves obtaining images of different frequencies through wavelet transform, and using a lossless method to obtain high and low frequency images, i.e., wavelet domain images.

[0036] The generation of Fourier domain weights involves using the Fast Fourier Transform to learn the learned weights and generating mutually disjoint Fourier domain weights.

[0037] The generation of spatial domain information utilizes the non-overlapping weights of the Fourier domain learned in the Fourier domain. The corresponding spatial domain weights are obtained through inverse Fourier transform. Combined with the attention mechanism, dynamic weights of the spatial domain are generated. Spatial domain information includes dynamic weights of the spatial domain.

[0038] Image information fusion involves fusing the acquired wavelet domain image, Fourier domain weights, and spatial domain information.

[0039] This invention, by adopting the above technical solutions, has significant technical effects:

[0040] This invention uses Fast Fourier Transform to divide the parameters corresponding to different frequencies into disjoint groups and generate Fourier disjoint weights; through an attention mechanism, it uses frequency domain parameters to generate dynamic weights, thereby realizing the function of dynamic convolution in the spatial domain.

[0041] In low-light environments, this invention utilizes wavelet transform without stride to replace the commonly used downsampling method, preserving high-frequency detail information as much as possible. It also combines FDW to replace ordinary convolution, improving the classification network and enhancing classification accuracy.

[0042] This invention references the method used by humans to identify subcategories of species by combining other information; it utilizes information such as the geographical location, time, temperature, and humidity of the hardware device, and uses a nonlinearly embedded multilayer fully connected neural network to vectorize the metadata, combining it with an improved network for comprehensive identification, thereby improving the identification accuracy.

[0043] This invention improves the ability to identify outdoor wildlife by superimposing frequency domain information with spatial domain information, and by combining information such as the time and geographical location of animal appearance, with the addition of a small number of parameters, especially the ability to classify wildlife in detail under low light and low resolution conditions. Attached Figure Description

[0044] Figure 1 This is a flowchart of the present invention.

[0045] Figure 2 This is a flowchart of embodiment 3 of the present invention.

[0046] Figure 3 This is a flowchart of the dynamic FDW convolution module based on the attention mechanism of the present invention.

[0047] Figure 4 This is a schematic diagram of the improved residual network of the present invention.

[0048] Figure 5 This is a schematic diagram of the high-frequency and low-frequency information of the present invention.

[0049] Figure 6 This is a schematic diagram of the Fourier non-intersecting weights of the present invention.

[0050] Among them, FDW (Fourier Disjoint Weight), NS (No Stride), HWD (Haar Wavelet DownSample), and IDFT (Inverse Discrete Fourier Transform) are used. Detailed Implementation

[0051] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. Example 1

[0052] A method for identifying wild animal species in low-light environments, the method comprising:

[0053] Acquisition of target images, including real-time acquisition of target images containing wild animals;

[0054] Image processing involves performing wavelet transform on the acquired target image.

[0055] Image recognition uses deep learning to learn weights from the processed image, performs Fourier transform on the weights to obtain Fourier domain information, combines the image's metadata, and vectorizes the image's metadata through a multi-layer fully connected neural network with nonlinear embedding. Finally, it combines the vectorized information to perform image recognition.

[0056] An improvement is made using ResNet18 as the base residual network. To reduce the loss of detail information caused by downsampling, images of different frequencies are acquired in the wavelet domain. High- and low-frequency images are obtained using a no-offset stride (NS) method. A Fast Fourier Transform (FDW) is performed on the images, and the parameters of disjoint groups in the Fourier domain are learned. An inverse Fourier transform is then used to reorganize the frequency-domain learned parameters in the spatial domain. Combined with an attention mechanism, dynamic convolution weights in the spatial domain are generated, enhancing the ability to capture global information. Wavelet domain, Fourier domain, and spatial domain information are fused to obtain a feature map NSFDW that retains high-frequency detail information.

[0057] Figure 1 In the input image, the number of channels is C, the length is H, and the width is W. It is then transformed into a non-intersecting image with C1 channels, the length is H / 2, and the width is W / 2 through the non-intersecting weights of the stepless Fourier domain. Figure 4 In this process, by improving the residual network, which requires four improved residual modules, the channel is converted to C2, the length is H / 32, and the width is W / 32. Then, average pooling is performed, with the channel being C2, the length being 1, and the width being 1. Meta-information is then acquired, which includes event occurrence time, altitude, longitude, latitude, temperature, and humidity. Feature extraction is performed on the meta-information, and the extracted features are combined with the features after average pooling for feature concatenation.

[0058] For image processing, image processing of the acquired target image using wavelet transform includes:

[0059] Wavelet domain image acquisition: Wavelet transform is used to obtain images of different frequencies, and high and low frequency images are obtained in a lossless manner.

[0060] The generation of Fourier domain weights involves using a fast Fourier transform on the learned weights to generate mutually disjoint Fourier domain weights.

[0061] The generation of spatial domain information utilizes the non-overlapping weights of the Fourier domain learned in the Fourier domain. The corresponding spatial domain weights are obtained through inverse Fourier transform. Combined with the attention mechanism, dynamic weights of the spatial domain are generated. Spatial domain information includes dynamic weights of the spatial domain.

[0062] Image information fusion involves fusing the acquired wavelet domain image, Fourier domain weights, and spatial domain information.

[0063] The NS downsampling module is a sampling method without stride displacement. In traditional residual networks, image downsampling is generally performed using pooling or convolution with stride=2. This method loses some detailed information of the original image. In outdoor wildlife detection and recognition, especially at night when wildlife is frequently active, the loss of high-frequency detailed information will seriously reduce the accuracy of traditional classification models and affect the recognition performance in other sub-domains such as species identification.

[0064] The wavelet domain image is obtained by using Haar wavelet transform to divide the image into four low- and high-frequency images: LL image, LH image, HL image, and HH image. LL represents the horizontal low-frequency and vertical low-frequency images; LH represents the horizontal low-frequency and vertical high-frequency images; HL represents the horizontal high-frequency and vertical low-frequency images; and HH represents the horizontal high-frequency and vertical high-frequency images. The resolution of each image is half that of the input image.

[0065] The generation of Fourier domain weights involves using a Fast Fourier Transform (FFT) on the learned weights to generate disjoint Fourier domain weights. Obtaining disjoint Fourier domain weights includes:

[0066] The input image is first subjected to a Fourier transform, and then the Fourier parameters are convolved. The total number of parameters that need to be learned is... Where k is the convolution kernel, C in For the input image parameters, C out Given the output image parameters; divide these parameters into N disjoint groups; reorganize the parameters of the Fast Fourier Transform into... p represents the recombined parameters; and each parameter is associated with a Fourier coordinate (u, v); where u is the Fourier abscissa and v is the Fourier ordinate.

[0067] Calculate (u, v) using the L2 norm. The values ​​are sorted from highest to lowest, dividing the frequency domain parameters into n disjoint groups, i.e. ,in, This is the set of parameters corresponding to different groups after grouping.

[0068] The generation of spatial domain information utilizes the disjoint weights learned in the Fourier domain, and obtains the corresponding spatial domain weights through the inverse Fourier transform. Combined with an attention mechanism, dynamic spatial domain weights are generated. Spatial domain information includes these dynamic weights: each set of parameters P is processed using the inverse discrete Fourier transform (IDFT). i Mapping to the spatial domain using formulas ,

[0069] Where k is the convolution kernel, Let be the corresponding parameter of the i-th group in the frequency-disjoint group and whose Fourier domain coordinates are (u, v). This represents the transformation result of the spatial domain point (p, q), where p is the abscissa of the spatial domain and q is the ordinate of the spatial domain; if If there is a response, then ,otherwise, ;

[0070] Weights for each group , By cutting and recombining, Transformed into a nucleus of size The number of image patches is , get n parameter, Weights are determined by Obtained in the frequency domain.

[0071] Example 2

[0072] A method for identifying wild animal species in low-light environments. Figure 2 The methods include:

[0073] Acquisition of target images, including real-time acquisition of target images containing wild animals;

[0074] Image processing involves performing wavelet transform on the acquired target image; improvement of residual networks involves using deep learning to learn from the processed image and generating an improved image through the improvement of residual networks.

[0075] Image recognition uses deep learning to learn weights from the processed image, performs Fourier transform on the weights to obtain Fourier domain information, combines the image's metadata, and vectorizes the image's metadata through a multi-layer fully connected neural network with nonlinear embedding. Finally, it combines the vectorized information to perform image recognition.

[0076] Without adding too many parameters, the residual blocks of the standard ResNet18 network are optimized by combining the advantages of FDW; the standard convolutional layers are replaced with FDW modules, and NSFDW replaces downsampling and standard convolution. With only a few additional parameters, the recognition rate of cameras for wild animals at night can be greatly improved.

[0077] The residual network is improved by wavelet transform and fast Fourier transform; wavelet domain image acquisition: images of different frequencies are obtained by wavelet transform, and high and low frequency images, i.e., wavelet domain images, are obtained in a lossless manner.

[0078] The generation of Fourier domain weights involves using a fast Fourier transform on the learned weights to generate mutually disjoint Fourier domain weights.

[0079] The generation of spatial domain information utilizes the disjoint weights of the Fourier domain obtained in the Fourier domain, and obtains the corresponding spatial domain weights through inverse Fourier transform. Combined with the attention mechanism, dynamic weights of the spatial domain are generated. Spatial domain information includes dynamic weights of the spatial domain.

[0080] Image information fusion involves fusing the acquired wavelet domain image, Fourier domain weights, and spatial domain information.

[0081] The acquired target image is processed using wavelet transform and fast Fourier transform.

[0082] Wavelet domain image acquisition: Wavelet transform is used to obtain images of different frequencies, and high and low frequency images are obtained in a lossless manner.

[0083] The generation of Fourier domain weights involves using a fast Fourier transform on the learned weights to generate mutually disjoint Fourier domain weights.

[0084] The generation of spatial domain information utilizes the non-overlapping weights of the Fourier domain learned in the Fourier domain. The corresponding spatial domain weights are obtained through inverse Fourier transform. Combined with the attention mechanism, dynamic weights of the spatial domain are generated. Spatial domain information includes dynamic weights of the spatial domain.

[0085] Image information fusion involves fusing the acquired wavelet domain image, Fourier domain weights, and spatial domain information.

[0086] The wavelet domain image is obtained by using Haar wavelet transform to divide the image into four low- and high-frequency images: LL image, LH image, HL image, and HH image. LL represents the horizontal low-frequency and vertical low-frequency images; LH represents the horizontal low-frequency and vertical high-frequency images; HL represents the horizontal high-frequency and vertical low-frequency images; and HH represents the horizontal high-frequency and vertical high-frequency images. The resolution of each image is half that of the input image.

[0087] The generation of Fourier domain weights involves using a Fast Fourier Transform (FFT) on the learned weights to generate disjoint Fourier domain weights. Obtaining these disjoint weights includes: inputting the image, performing a Fourier Transform, and then convolving the Fourier parameters using a convolution kernel of size k. The total number of parameters to be learned is then determined. These parameters are divided into N disjoint groups; in this embodiment, k=3; the parameters of the Fast Fourier Transform are reorganized into Each parameter is associated with a Fourier coordinate (u, v); where u is the Fourier abscissa and v is the Fourier ordinate; (u, v) is calculated using the L2 norm. The values ​​are sorted from highest to lowest, dividing the frequency domain parameters into n disjoint groups, i.e. ,in, This is the set of parameters corresponding to different groups after grouping. Figure 5 As shown, here n=2 for ease of illustration, P0 represents low frequency and P1 represents high frequency. In application, n can be chosen to have a larger value to obtain more diverse information without increasing the parameter weights.

[0088] The generation of spatial domain information utilizes the non-overlapping weights of the Fourier domain learned in the Fourier domain. The corresponding spatial domain weights are obtained through inverse Fourier transform. Combined with the attention mechanism, dynamic weights of the spatial domain are generated. Spatial domain information includes dynamic weights of the spatial domain.

[0089] ,

[0090] Where k is the convolution kernel, Let be the corresponding parameter of the i-th group in the frequency-disjoint group and whose Fourier domain coordinates are (u, v). This is the transformation result of the spatial domain point (p, q), where p is the abscissa of the spatial domain and q is the ordinate of the spatial domain.

[0091] if If there is a response, then ,otherwise, There is a response at the blue location, and the parameter is the one corresponding to the blue location. All other color positions are 0.

[0092] Weights for each group By cutting and recombining, Transformed into a nucleus of size The coefficients of the dynamic convolution kernel are π1……π n The number of image patches is , get n parameter, Weights are determined by It acquires data in the frequency domain. Compared to the frequency domain singularity problem of traditional convolution, it ensures the diversity of parameters and effectively improves the ability to capture high-frequency information. In Figure 3, an attention mechanism is used to process the n groups of data acquired during training. The weights use dynamic convolution. Input: Feature map or image data; Avgpool pooling reduces dimensionality and extracts global features; Fully Connected Layer (FC) performs a linear transformation; ReLU (Rectified Linear Unit) activation function introduces non-linearity; Attention Mechanism enhances important features; Another FC further processes the features. Softmax: Normalization function that transforms the output into a probability distribution. Discrete Fourier Transform (DFT) converts the signal from the time domain to the frequency domain. The normalized output and the frequency domain of the DFT are weighted and fused using the FDW frequency domain information processing module; Output: Processed features.

[0093] Example 3

[0094] A method for identifying wild animal species in low-light environments, the method comprising:

[0095] Acquisition of target images, including real-time acquisition of target images containing wild animals;

[0096] Image processing involves performing wavelet transform on the acquired target image; improvement of residual networks involves using deep learning to learn from the processed image and generating an improved image through the improvement of residual networks.

[0097] Image recognition uses deep learning to learn weights from the processed image, performs Fourier transform on the weights to obtain Fourier domain information, combines the image's metadata, and vectorizes the image's metadata through a multi-layer fully connected neural network with nonlinear embedding. Finally, it combines the vectorized information to perform image recognition.

[0098] The residual network is improved by wavelet transform and fast Fourier transform; wavelet domain image acquisition: images of different frequencies are obtained by wavelet transform, and high and low frequency images, i.e., wavelet domain images, are obtained in a lossless manner.

[0099] The generation of Fourier domain weights involves using a fast Fourier transform on the learned weights to generate mutually disjoint Fourier domain weights.

[0100] The generation of spatial domain information utilizes the disjoint weights of the Fourier domain obtained in the Fourier domain, and obtains the corresponding spatial domain weights through inverse Fourier transform. Combined with the attention mechanism, dynamic weights of the spatial domain are generated. Spatial domain information includes dynamic weights of the spatial domain.

[0101] Image information fusion involves fusing the acquired wavelet domain image, Fourier domain weights, and spatial domain information.

[0102] Figure 4 In this model, the residual network, through skip connections, allows information to be directly transmitted across multiple layers, effectively mitigating the vanishing or exploding gradient problem in deep networks and ensuring stable gradient propagation to deeper network layers. Without increasing too many parameters, the residual blocks of the standard ResNet18 network are optimized by combining the advantages of Fast Fourier Transform (FDW). Replacing standard convolutional layers with FDW modules and using NSFWD instead of downsampling and standard convolution significantly improves the camera's nighttime wildlife recognition rate with only a small increase in parameters. The acquired target images are processed using wavelet transform and Fast Fourier Transform.

[0103] Wavelet domain image acquisition: Wavelet transform is used to obtain images of different frequencies, and high and low frequency images are obtained in a lossless manner; this method is used to replace the downsampling process in conventional networks.

[0104] The generation of Fourier domain weights involves using a fast Fourier transform on the learned weights to generate mutually disjoint Fourier domain weights.

[0105] The spatial domain information is generated by using the non-overlapping weights of the Fourier domain learned in the Fourier domain, obtaining the corresponding spatial domain weights through inverse Fourier transform, and combining the attention mechanism to generate dynamic spatial domain weights. Spatial domain information includes dynamic spatial domain weights; this method replaces conventional convolution.

[0106] Image information fusion involves fusing the acquired wavelet domain image, Fourier domain weights, and spatial domain information.

[0107] The wavelet domain image is obtained by using Haar wavelet transform to divide the image into four low- and high-frequency images: LL image, LH image, HL image, and HH image. LL represents the horizontal low-frequency and vertical low-frequency images; LH represents the horizontal low-frequency and vertical high-frequency images; HL represents the horizontal high-frequency and vertical low-frequency images; and HH represents the horizontal high-frequency and vertical high-frequency images. The resolution of each image is half that of the input image.

[0108] The generation of Fourier domain weights involves using a Fast Fourier Transform (FFT) on the learned weights to generate disjoint Fourier domain weights. Obtaining disjoint Fourier domain weights includes:

[0109] The input image is first subjected to a Fourier transform, and then the Fourier parameters are convolved with a kernel size of k. The total number of parameters to be learned is then determined. Divide these parameters into N disjoint groups; reorganize the parameters of the Fast Fourier Transform into... Each parameter is associated with a Fourier coordinate (u, v); where u is the Fourier abscissa and v is the Fourier ordinate; (u, v) is calculated using the L2 norm. The values ​​are sorted from highest to lowest, dividing the frequency domain parameters into n disjoint groups, i.e. ,in, This is the set of parameters corresponding to different groups after grouping.

[0110] Figure 5 As shown, for ease of illustration, n=2, P0 represents low frequency, and P1 represents high frequency. In application, n can be chosen to have a larger value to obtain more diverse information without increasing parameter weights. Spatial domain information is generated by using the non-overlapping weights learned in the Fourier domain, obtaining the corresponding spatial domain weights through inverse Fourier transform, and combining this with an attention mechanism to generate dynamic spatial domain weights. Spatial domain information includes these dynamic weights.

[0111] ,

[0112] Where k is the convolution kernel, Let be the corresponding parameter of the i-th group in the frequency-disjoint group and whose Fourier domain coordinates are (u, v). This is the transformation result of the spatial domain point (p, q), where p is the abscissa of the spatial domain and q is the ordinate of the spatial domain.

[0113] if If there is a response, then ,otherwise, ; Figure 5 In the middle, there is a response at the blue position, and the parameter is the one corresponding to the blue position. All other color positions are 0.

[0114] Weights for each group By cutting and recombining, Transformed into a nucleus of size The number of image patches is , get n parameter, Weights are determined by It acquires data in the frequency domain. Compared to the frequency domain singularity problem of traditional convolution, it ensures the diversity of parameters and effectively improves the ability to capture high-frequency information. In Figure 3, an attention mechanism is used to process the n groups of data acquired during training. The weights are applied using dynamic convolution.

[0115] By combining the longitude, latitude, altitude, and metadata of the image, the metadata is vectorized through a multi-layer fully connected neural network with nonlinear embedding.

[0116] The longitude, latitude, and altitude information are mapped in the following way: [latitude, longitude, altitude] -> [x, y, z], where x is the latitude information, y is the longitude information, and z is the altitude information.

[0117] The acquired time is mapped using sin and cos trigonometric functions.

[0118] ,

[0119] Where month represents months and hour represents hours;

[0120] Normalize the temperature and humidity information.

[0121] ,

[0122] Where temperature is the temperature information and humidity is the humidity information;

[0123] Obtain the longitude, latitude, altitude, and metadata of the location, and use nonlinear embedding to map the longitude, latitude, altitude, and metadata into vectors.

Claims

1. A wild animal species identification method suitable for low-light environments, the method comprising: target image acquisition, acquiring a target image in real time in which a wild animal exists; image processing, performing image processing on the acquired target image through wavelet transform; wavelet domain image acquisition, acquiring images of different frequencies through wavelet transform, and acquiring high and low frequency images, i.e., wavelet domain images, in a lossless manner; image recognition, using deep learning to learn the processed image to obtain weights, performing Fourier transform on the weights to obtain Fourier domain information, and using a nonlinear embedded multilayer fully connected neural network to vectorize the meta information of the image, wherein the meta information includes event occurrence time, altitude information, longitude information, latitude information, temperature information, and humidity, and the image is recognized by combining the vectorized information, specifically including: Fourier domain weight generation, using fast Fourier transform on the learned weights to generate Fourier domain mutually disjoint weights, and obtaining Fourier domain mutually disjoint weights; spatial domain information generation, using the Fourier domain mutually disjoint weights learned in the Fourier domain to obtain corresponding spatial domain weights through inverse Fourier transform, generating spatial domain dynamic weights in combination with an attention mechanism, and the spatial domain information includes the spatial domain dynamic weights; image information fusion, fusing the wavelet domain images, Fourier domain information, and spatial domain information to obtain a feature map that retains high frequency detail information.

2. The method for identifying wild animal species in low-light environments according to claim 1, wherein, Wavelet domain image acquisition uses Haar wavelet transform to divide the image into four low and high frequency images, namely LL image, LH image, HL image, and HH image; wherein LL is a horizontal low frequency and vertical low frequency image; LH is a horizontal low frequency and vertical high frequency image; HL is a horizontal high frequency and vertical low frequency image; and HH is a horizontal high frequency and vertical high frequency image; and the resolution of each image is one-half of the input image.

3. The method for identifying wild animal species in low-light environments according to claim 1, wherein, The Fourier domain weight is generated, a fast Fourier transform is used for learning the obtained weight, Fourier domain non-intersecting weights are generated, and the Fourier domain non-intersecting weights are obtained, including inputting image information, performing Fourier transform first, and performing convolution on Fourier parameters, so that the total parameter amount to be learned is wherein k is a convolution kernel, C in is an input image parameter, C out is an output image parameter; the parameter amount is divided into N non-intersecting groups; parameters of the fast Fourier transform are recombined as p is a recombined parameter; and each parameter is associated with a Fourier coordinate (u, v); wherein u is a Fourier horizontal coordinate, and v is a Fourier vertical coordinate; (u, v) is calculated by using L2 norm The parameters in the frequency domain are divided into n groups which are mutually disjoint by sorting the values from high to low, that is , wherein is the parameter set corresponding to different groups after grouping.

4. The method for identifying wild animal species in low-light environments according to claim 1, wherein, The generation of the spatial domain information utilizes the Fourier domain disjoint weights learned in the Fourier domain, and the corresponding spatial domain weights are obtained through inverse Fourier transform. The spatial domain dynamic weights are generated by combining the attention mechanism. The spatial domain information includes the spatial domain dynamic weights: each group of parameters is mapped into the spatial domain by the formula, , where k is a convolution kernel, is the corresponding parameter for the i-th group in the frequency disjoint group and the Fourier domain coordinate is (u, v), is the transform result for the spatial domain point (p, q), p is the spatial domain horizontal coordinate, q is the spatial domain vertical coordinate; if has a response, then , otherwise, ; Weights for each group , By cutting and recombining, Transformed into a nucleus of size The number of image patches is , get n parameter, Weights are determined by Obtained in the frequency domain.

5. The method for identifying wild animal species in low-light environments according to claim 1, wherein, Vectorizing the meta information through a nonlinear embedded multilayer fully connected neural network; Mapping the latitude, longitude, and altitude, the mapping method being: [latitude, longitude, altitude] -> [x, y, z], wherein x is the latitude information, y is the longitude information, and z is the altitude information; Mapping the acquired time using sin and cos trigonometric functions, wherein month is the month and hour is the hour; normalizing the temperature information and the humidity information, where temperature is the temperature information and humidity is the humidity information.

Citation Information

Patent Citations

  • Image super-resolution processing method

    CN112614056A

  • Image recognition method and device for wavelet domain CNN learning

    CN113762290A

  • Road damage detection method and system

    CN117541881A