A deep learning-based method for tracking giant pandas in the wild

By combining the Lab color space and dynamic tracking and localization model with a deep learning method based on the contrasting values ​​of green, red, blue, and yellow, the problems of lighting changes and background interference in the field tracking of giant pandas were solved, achieving real-time and accurate target detection and tracking.

CN121074155BActive Publication Date: 2026-03-24SICHUAN FORESTRY RES INST (SICHUAN FORESTRY IND RES & DESIGN INST)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Traditional RGB color-based target detection algorithms are easily affected by changes in lighting, vegetation occlusion, and background similarity when tracking giant pandas in the wild, leading to missed detections or false detections. Furthermore, deep learning models are difficult to achieve both real-time performance and lightweight design in wild scenarios.

Method used

The Lab color space is used to separate brightness and chromaticity. By constructing a dynamic tracking and localization model, and combining the contrasting values ​​of green and red and blue and yellow, a deep learning model composed of convolutional layers, pooling layers, and attention layers is used to track giant pandas and dynamically adjust the brightness and color features.

Benefits of technology

It improves the distinction between giant panda targets and backgrounds, enhances sensitivity to neutral color regions, enables real-time and accurate field tracking, and supports an integrated air-ground wildlife tracking system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121074155B_ABST
    Figure CN121074155B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on deep learning's giant panda field tracking method, belong to image processing technical field, including the following steps: S1, using the camera installed in fixed position to collect real-time field environment image, and extract the Lab triplet of pixel in real-time field environment image;S2, using contrast matrix to update the L channel value of pixel Lab triplet;S3, according to the a channel value and b channel value of pixel in real-time field environment image, calculate green red color relative standing value and blue yellow color relative standing value;S4, based on the green red color relative standing value and blue yellow color relative standing value of real-time field environment image and the L channel value of pixel after updating, constructs dynamic tracking positioning model, determines the tracking result of real-time field environment image using dynamic tracking positioning model.The application monitors the activity of giant panda in real time, provides data support for protection strategy, and is conducive to realizing the construction of air-ground integrated wildlife tracking system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of image processing, and particularly relates to a giant panda wild tracking method based on deep learning. BACKGROUND

[0002] In the field of wildlife protection, as a flagship species, the monitoring of the wild behavior of giant pandas and the protection of their habitats are of great significance to the maintenance of biodiversity. In the wild scene, the light conditions are variable (such as overcast, foggy and direct light), the vegetation is obstructive (such as dense bamboo forest areas), and the giant panda is similar to the background (black and white fur and rock, shadow confusion), which leads to the fact that traditional target detection algorithms based on RGB color (such as YOLO, Faster R-CNN) are prone to miss detection or false detection. Existing methods mostly rely on the RGB color space, but this space is sensitive to light changes and does not explicitly model the human visual perception mechanism. The Lab color space, as a uniform color space, has the characteristics that the L channel (brightness) and the a / b channel (color opposite value) are decoupled, which is more in line with the human eye's perception of color and brightness, but it has not been fully applied to the field of wild animal tracking.

[0003] At the same time, deep learning models perform well in target tracking, but the wild scene requires the model to have the dual characteristics of real-time and lightness (embedded device deployment). Traditional methods either sacrifice accuracy for speed or rely on complex network structures, resulting in excessive consumption of computing resources. SUMMARY

[0004] The application is proposed to solve the above problems, and provides a giant panda wild tracking method based on deep learning.

[0005] The technical scheme of the application is: a giant panda wild tracking method based on deep learning, comprising the following steps:

[0006] S1, using a camera installed at a fixed position to collect real-time wild environment images, and extracting Lab triplets of pixel points in the real-time wild environment images;

[0007] S2, updating the L channel value of the Lab triplet of the pixel points using a comparison matrix;

[0008] S3, calculating the green-red relative opposite value and the blue-yellow relative opposite value according to the a channel value and the b channel value of the pixel points in the real-time wild environment images;

[0009] S4, constructing a dynamic tracking positioning model based on the green-red relative opposite value and the blue-yellow relative opposite value of the real-time wild environment images and the updated L channel value of the pixel points, and determining the tracking result of the real-time wild environment images using the dynamic tracking positioning model.

[0010] Further, S2 comprises the following sub-steps:

[0011] S21, extracting the L channel value corresponding to the Lab triplet of the pixel point in the real-time field environment image;

[0012] S22, constructing the forward contrast matrix and the turning contrast matrix of the pixel point according to the L channel value of the pixel point;

[0013] S23, taking the maximum value between the maximum eigenvalue of the forward contrast matrix and the maximum eigenvalue of the turning contrast matrix as the brightness update coefficient;

[0014] S24, updating the L channel value of the pixel point according to the brightness update coefficient of the pixel point.

[0015] The beneficial effects of the above further scheme are: in the present application, the Lab space separates the brightness (L channel) from the chrominance (a / b channel), so that the brightness adjustment will not affect the color information. This is crucial for the accurate segmentation of the black and white fur of the giant panda, avoiding the contour blur caused by color distortion. By filling the four-neighborhood pixels (up, down, left and right), the brightness changes in the horizontal and vertical directions are captured, reflecting the local edge information (such as the intersection of the panda outline and the background). By filling the diagonal neighborhood pixels, the diagonal brightness changes are captured, supplementing the structural information in the orthogonal direction (such as the diagonal stripes of the fur texture).

[0016] The rest of the positions are filled with 1, highlighting the contrast of the neighborhood pixels and suppressing the interference of irrelevant areas.

[0017] The maximum eigenvalue of the matrix corresponds to the main direction of the data, reflecting the most significant mode of the neighborhood brightness change.

[0018] Further, in S22, the forward contrast matrix and the turning contrast matrix of the pixel point are both set as three-row and three-column square matrices;

[0019] When constructing the forward contrast matrix of the pixel point, the L channel values of the four-neighborhood pixels of the pixel point are sequentially filled into the corresponding positions of the three-row and three-column square matrix, and the rest of the positions are filled with 1, generating the forward contrast matrix;

[0020] When constructing the turning contrast matrix of the pixel point, the L channel values of the D-neighborhood pixels of the pixel point are sequentially filled into the corresponding positions of the three-row and three-column square matrix, and the rest of the positions are filled with 1, generating the turning contrast matrix.

[0021] The expression of the forward contrast matrix of the pixel point in the first row and the first column is: The expression of the forward contrast matrix of the pixel point in the first row and the first column is:

[0022]

[0023] In the formula, indicates the​​ row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point;

[0024] the row and column pixel point, the row and column pixel point, the turning contrast matrix of the row and column pixel point is expressed as:

[0025]

[0026] wherein, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point; the L channel value of the row and column pixel point, the L channel value of the row and column pixel point, the L channel value of the row and column pixel point;

[0027] Further, in S24, the calculation formula of the updated L channel value of the pixel point is:

[0028] ;

[0029] wherein, the L channel value of the pixel point, a random value between 0 and 1, a luminance update coefficient, a sign function, the average of the L channel values of the 4-neighbor and D-neighbor pixel points of the pixel point in the real-time field environment image.

[0030] The beneficial effect of the further scheme is that in the application, the horizontal, vertical and diagonal information is integrated to avoid single direction deviation.

[0031] Further, S3 comprises the following sub-steps:

[0032] S31, extracting the a channel value and the b channel value corresponding to the Lab triplet of the pixel point in the real-time field environment image;

[0033] S32, calculating the green-red relative standing value according to the a channel value of the pixel point;

[0034] S33, calculating the blue-yellow relative standing value according to the b channel value of the pixel point.

[0035] The beneficial effect of the further scheme is that in the application, the global color contrast is quantified to provide key color features for the dynamic tracking model and capture the contrast intensity of green-red and blue-yellow colors in the image.

[0036] Further, in S32, the green-red relative standing value is calculated according to the following formula:

[0037] ;

[0038] In the formula, a i represents the a channel value of the i th pixel point, represents taking the maximum value, represents taking the minimum value, represents a constant; In S33, the calculation formula of the blue-yellow relative standing value is as follows:

[0039] ;

[0040] In the formula, b i represents the b channel value of the i th pixel point.

[0041] Since the theoretical range of the a / b channel is [−128, 127], the relative standing value is finally normalized to the range of [0, 2], so that and can be directly used as attention weights.

[0042]

[0043] ​​​​Further, in S4, the dynamic tracking positioning model comprises a convolutional layer, a pooling layer, a multiplier M1, a multiplier M2, a first attention layer, a second attention layer, and an output layer.

[0044] The input end of the convolutional layer is configured to input a real-time field environment image; the first output end of the convolutional layer is connected with the input end of the pooling layer; the second output end of the convolutional layer is connected with the input end of the second attention layer; the first output end and the second output end of the pooling layer are connected with the input end of the multiplier M1 and the input end of the multiplier M2 respectively; the output end of the multiplier M1 and the output end of the multiplier M2 are connected with the input end of the first attention layer; the output end of the first attention layer is connected with the input end of the second attention layer; the output end of the second attention layer is connected with the input end of the output layer; and the output end of the output layer is configured to output a tracking result of the real-time field environment image.

[0045] The convolutional layer is the first layer of the model, responsible for extracting primary visual features from the input real-time field environment image. These features include edges, textures, corner points, and simple shape patterns, providing a foundation for subsequent high-level semantic understanding. The convolutional layer contains multiple learnable convolutional kernels, each sliding over the input image to capture local features through a local receptive field. Each convolutional kernel generates a feature map, where each element is the weighted sum of the convolutional kernel at the corresponding position of the input image, reflecting the specific visual pattern of the image at that position. The convolutional layer has two output ends: one connected to the pooling layer to pass the extracted features to the pooling layer for downsampling; the other directly connected to the second attention layer to provide the original feature input for the attention mechanism. The pooling layer is used to reduce the spatial dimension of the feature map, thereby reducing the computational load and enhancing the model's robustness to minor transformations of the input image such as translation, rotation, and scaling. Typically, max-pooling or average-pooling operations are used. The pooling layer has two output ends: one connected to the multiplier M1 to pass the pooled features to the multiplier M1 for processing; the other connected to the multiplier M2 to pass the pooled features to the multiplier M2 for processing.

[0046] The multipliers are used to element-wise multiply the outputs of the pooling layer with the green-red color opponent value and the blue-yellow color opponent value. This operation allows the model to adjust the importance of the pooled features based on the color opponent values, thereby enhancing the focus on specific color regions. The green-red color opponent value and the blue-yellow color opponent value reflect the contrast information of green-red and blue-yellow colors in the image. Through the multipliers, these color contrast information is integrated into the pooled features, allowing the model to pay more attention to regions with high color contrast.

[0047] The first attention layer combines the outputs of the multipliers M1 and M2 and the probabilities that the pixel points a channel and b channel are 0 to generate weighted features. By dynamically adjusting the influence of color opposite values on the pooling features, the first attention layer can enhance the sensitivity to the neutral color area of the panda. This layer uses the Softmax function to calculate the weights, which reflect the distribution of neutral color pixels (pixels with a channel and b channel close to 0) in the image. The output of the first attention layer is connected to the second attention layer, and the fused feature map is passed to the second attention layer for further processing.

[0048] The second attention layer further processes the output of the first attention layer and combines the updated L channel value (brightness information) to generate the final feature representation, strengthening the detection of the panda's edges and contours. By combining the brightness information, the second attention layer can modulate the color-space features and enhance the response to high-contrast areas such as the edges of the panda.

[0049] The output of the second attention layer is connected to the output layer, which passes the final feature representation to the output layer for decoding. The output layer is responsible for converting the feature representation of the second attention layer into the final tracking result. Depending on the specific task, the output layer can be a fully connected layer, a convolutional layer, or other types of decoder.

[0050] Further, the expression of the first attention layer is:

[0051] ;

[0052] In the formula, represents the processing result of the first attention layer, represents the green-red opposite value, represents the blue-yellow opposite value, represents the probability of a pixel point with a channel value of 0 in the real-time outdoor environment image, represents the probability of a pixel point with a channel value of 0 in the real-time outdoor environment image, represents the output result of the pooling layer, represents the exponential.

[0053] The beneficial effects of the above further scheme are: the opposite value is weighted and fused with the pooling features, enhancing the response to color-sensitive areas and improving the sensitivity to the neutral color area of the panda.

[0054] Further, the expression of the second attention layer is:

[0055] ;

[0056] In the formula, represents the processing result of the second attention layer, represents the The updated pixel point, The bias term is represented, The processing result of the first attention layer is represented, The number of pixel points of the real-time field environment image is represented.

[0057] The beneficial effects of the above further scheme are: for each pixel, the output result of the first attention layer is multiplied by the brightness, multi-modal fusion is realized, the responses of all pixels are accumulated and biased, the feature response is adjusted by the updated L channel value, and the panda edge (high contrast area) is strengthened.

[0058] The beneficial effects of the present application are:

[0059] (1) The present application separates the brightness (L channel) from the color (a / b channel) by extracting the Lab triplet of the pixel points in the real-time field environment image, avoids the interference of light changes on the color features, and designs a contrast matrix to dynamically update the L channel, strengthens the panda edge (such as the brightness mutation area of the fur and the background), and improves the discrimination degree of the target and the background.

[0060] (2) The present application defines the green-red relative value and the blue-yellow relative value, quantifies the global color contrast, inputs the two relative values into the dynamic tracking model, enhances the attention to the neutral color area of the panda through dynamic weight distribution, and suppresses the background interference.

[0061] (3) The model of the present application fuses the features extracted by each layer to determine whether there is a giant panda in the field environment, monitors the giant panda activity in real time, provides data support for the protection strategy, and is conducive to realizing the construction of the air-ground integrated wild animal tracking system. BRIEF DESCRIPTION OF DRAWINGS

[0062] Figure 1 The flowchart of the giant panda field tracking method based on deep learning;

[0063] Figure 2 The structure diagram of the dynamic tracking positioning model. DETAILED DESCRIPTION

[0064] The embodiments of the present application will be further described below with reference to the accompanying drawings.

[0065] As Figure 1 shown, the present application provides a giant panda field tracking method based on deep learning, which comprises the following steps:

[0066] S1, using a camera installed at a fixed position to collect a real-time field environment image, and extracting a Lab triplet of pixel points in the real-time field environment image;

[0067] S2, updating the L channel value of the Lab triplet of the pixel point by using the contrast matrix;

[0068] S3, calculating the green-red color relative value and the blue-yellow color relative value according to the a channel value and the b channel value of the pixel point in the real-time field environment image;

[0069] S4, constructing a dynamic tracking positioning model based on the green-red color relative value and the blue-yellow color relative value of the real-time field environment image and the updated L channel value of the pixel point, and determining the tracking result of the real-time field environment image by using the dynamic tracking positioning model.

[0070] In the embodiment of the present application, S2 comprises the following sub-steps:

[0071] S21, extracting the L channel value corresponding to the Lab triplet of the pixel point in the real-time field environment image;

[0072] S22, constructing the forward contrast matrix and the turning contrast matrix of the pixel point according to the L channel value of the pixel point;

[0073] S23, taking the maximum value between the maximum eigenvalue of the forward contrast matrix and the maximum eigenvalue of the turning contrast matrix as the brightness update coefficient;

[0074] S24, updating the L channel value of the pixel point according to the brightness update coefficient of the pixel point.

[0075] In the present application, the Lab space separates the brightness (L channel) from the chrominance (a / b channel), so that the brightness adjustment will not affect the color information. This is crucial for the accurate segmentation of the black and white fur of the giant panda, avoiding the contour blur caused by color distortion. By filling the four-neighborhood pixels (up, down, left and right), the brightness changes in the horizontal and vertical directions are captured, reflecting the local edge information (such as the intersection of the panda outline and the background). By filling the diagonal neighborhood pixels, the diagonal brightness changes are captured, supplementing the orthogonal direction structure information (such as the diagonal stripes of the fur texture).

[0076] The remaining positions are filled with 1, highlighting the contrast of the neighborhood pixels and suppressing the interference of irrelevant areas.

[0077] The maximum eigenvalue of the matrix corresponds to the main direction of the data, reflecting the most significant mode of the neighborhood brightness change.

[0078] In the embodiment of the present application, in S22, the forward contrast matrix and the turning contrast matrix of the pixel point are both set as three-row and three-column square matrices;

[0079] When constructing the forward contrast matrix of the pixel point, the L channel values of the four-neighborhood pixels of the pixel point are sequentially filled into the corresponding positions of the three-row and three-column square matrix, and the remaining positions are filled with 1, generating the forward contrast matrix;

[0080] When constructing the steering contrast matrix of the pixel point, the L channel value of the D neighborhood pixel point of the pixel point is sequentially filled into the corresponding position of the three-row three-column square matrix, and the remaining positions are filled with 1, to generate the steering contrast matrix.

[0081] The expression of the forward contrast matrix of the pixel point in the i-th row and the j-th column is: The expression of the forward contrast matrix of the pixel point in the i-th row and the j-th column is:

[0082]

[0083] In the formula, L(i, j) represents the L channel value of the pixel point in the i-th row and the j-th column, In the formula, L(i, j) represents the L channel value of the pixel point in the i-th row and the j-th column, In the formula, L(i, j) represents the L channel value of the pixel point in the i-th row and the j-th column, In the formula, L(i, j) represents the L channel value of the pixel point in the i-th row and the j-th column, In the formula, L(i, j) represents the L channel value of the pixel point in the i-th row and the j-th column, In the formula, L(i, j) represents the L channel value of the pixel point in the i-th row and the j-th column, In the formula, L(i, j) represents the L channel value of the pixel point in the i-th row and the j-th column, In the formula, L(i, j) represents the L channel value of the pixel point in the i-th row and the j-th column;

[0084] The expression of the steering contrast matrix of the pixel point in the i-th row and the j-th column is:

[0085]

[0086] In the formula, L(i, j) represents the L channel value of the pixel point in the i-th row and the j-th column, In the formula, L(i, j) represents the L channel value of the pixel point in the i-th row and the j-th column, In the formula, L(i, j) represents the L channel value of the pixel point in the i-th row and the j-th column, In the formula, L(i, j) represents the L channel value of the pixel point in the i-th row and the j-th column, In the formula, L(i, j) represents the L channel value of the pixel point in the i-th row and the j-th column, In the formula, L(i, j) represents the L channel value of the pixel point in the i-th row and the j-th column; In the formula, L(i, j) represents the L channel value of the pixel point in the i-th row and the j-th column,

[0087] In the embodiment of the present application, the calculation formula of the updated L channel value L(i, j) of the pixel point in S24 is: ​​​​​​​​​​​​​​​​​​​​​​​

[0088] ;

[0089] In the formula, This represents the L channel value of a pixel. Represents a random value between 0 and 1. Indicates the brightness update factor. Represents a symbolic function. This represents the average L-channel value of the 4-neighbor and D-neighbor pixels of a pixel in a real-time field environment image.

[0090] In this invention, horizontal, vertical, and diagonal information are integrated to avoid deviations in a single direction. The update coefficient is determined by neighborhood feature values, with drastic adjustments in high-contrast areas (such as black-and-white boundaries) and gentle adjustments in low-contrast areas (such as bamboo forest backgrounds). The sign function, combined with mean constraints, enables intelligent increases and decreases in brightness, avoiding overexposure or underexposure caused by global mean.

[0091] In this embodiment of the invention, S3 includes the following sub-steps:

[0092] S31. Extract the a-channel and b-channel values ​​corresponding to the Lab triplet of pixels in real-time field environment images;

[0093] S32. Calculate the green-red contrast value based on the a channel value of the pixel;

[0094] S33. Calculate the contrast value of blue and yellow based on the b channel value of the pixel.

[0095] In this invention, by quantifying global color contrast, key color features are provided for the dynamic tracking model, capturing the contrast intensity of green-red and blue-yellow colors in the image.

[0096] In this embodiment of the invention, in S32, the green and red opposing values... The calculation formula is:

[0097] ;

[0098] In the formula, Indicates the first The a-channel value of each pixel, This indicates taking the maximum value. This indicates taking the minimum value. Represents a constant;

[0099] In S33, the contrasting values ​​of blue and yellow The calculation formula is:

[0100] ;

[0101] In the formula, represents the b channel value of the i-th pixel point.

[0102] Since the a / b channel theoretical range is [-128, 127], the relative opposite values are finally normalized to the range of [0, 2], so that and can be directly used as attention weights.

[0103] In the embodiment of the present application, in S4, as shown in the figure, Figure 2 the dynamic tracking positioning model includes a convolutional layer, a pooling layer, a multiplier M1, a multiplier M2, a first attention layer, a second attention layer, and an output layer.

[0104] The input end of the convolutional layer is used to input the real-time field environment image; the first output end of the convolutional layer is connected with the input end of the pooling layer; the second output end of the convolutional layer is connected with the input end of the second attention layer; the first output end and the second output end of the pooling layer are connected with the input end of the multiplier M1 and the input end of the multiplier M2 respectively; the output end of the multiplier M1 and the output end of the multiplier M2 are both connected with the input end of the first attention layer; the output end of the first attention layer is connected with the input end of the second attention layer; the output end of the second attention layer is connected with the input end of the output layer; and the output end of the output layer is used to output the tracking result of the real-time field environment image.

[0105] The convolutional layer is the first layer of the model, responsible for extracting primary visual features from the input real-time field environment image. These features include edges, textures, corner points, and simple shape patterns, providing a foundation for subsequent high-level semantic understanding. The convolutional layer contains multiple learnable convolutional kernels, each of which slides over the input image, capturing local features through a local receptive field. Each convolutional kernel generates a feature map, where each element is the weighted sum of the convolutional kernel at the corresponding position of the input image, reflecting the specific visual pattern of the image at that position. The convolutional layer has two output ends: one output end is connected to the pooling layer, passing the extracted features to the pooling layer for downsampling; the other output end is directly connected to the second attention layer, providing the original feature input for the attention mechanism. The pooling layer is used to reduce the spatial dimension of the feature map, thereby reducing the computational load and enhancing the model's robustness to minor transformations of the input image, such as translation, rotation, and scaling. Typically, maximum pooling or average pooling operations are used. The pooling layer has two output ends: one output end is connected to the multiplier M1, passing the pooled features to the multiplier M1 for processing; the other output end is connected to the multiplier M2, passing the pooled features to the multiplier M2 for processing.

[0106] ​The multipliers are used to element-wise multiply the output of the pooling layer with the green-red color-opponent value and the blue-yellow color-opponent value. This operation enables the model to adjust the importance of the pooled features according to the color-opponent values, thereby enhancing the focus on specific color regions. The green-red color-opponent value and the blue-yellow color-opponent value reflect the contrast information of green-red and blue-yellow colors in the image. Through the multipliers, these color contrast information is integrated into the pooled features, enabling the model to pay more attention to regions with high color contrast.

[0107] The first attention layer combines the outputs of the multipliers M1 and M2, as well as the probabilities of the pixel points a channel and b channel being 0, to generate weighted features. By dynamically adjusting the influence of color-opponent values on pooled features, the first attention layer can enhance the sensitivity to the neutral color regions of the panda. This layer uses the Softmax function to calculate weights, which reflect the distribution of neutral color pixels (pixels with a channel and b channel close to 0) in the image. The output of the first attention layer is connected to the second attention layer, passing the fused feature map to the second attention layer for further processing.

[0108] The second attention layer further processes the output of the first attention layer, combining the updated L channel value (brightness information) to generate the final feature representation, strengthening the detection of the panda's edges and contours. By combining brightness information, the second attention layer can modulate the color-space features, enhancing the response to high-contrast regions such as the edges of the panda.

[0109] The output of the second attention layer is connected to the output layer, passing the final feature representation to the output layer for decoding. The output layer is responsible for converting the feature representation of the second attention layer into the final tracking result. Depending on the specific task, the output layer can be a fully connected layer, a convolutional layer, or other types of decoder.

[0110] In an embodiment of the present application, the expression of the first attention layer is:

[0111] ;

[0112] In the formula, represents the processing result of the first attention layer, represents the green-red color-opponent value, represents the blue-yellow color-opponent value, represents the probability of the pixel point a channel being 0 in the real-time outdoor environment image, represents the probability of the pixel point a channel being 0 in the real-time outdoor environment image, represents the output result of the pooling layer, represents the exponential.

[0113] The opposite value and the pooling feature are weighted and fused to enhance the response to the color sensitive region and improve the sensitivity to the neutral color region of the panda.

[0114] In the embodiment of the present application, the expression of the second attention layer is:

[0115] ;

[0116] In the formula, represents the processing result of the second attention layer, represents the updated pixel point of the first attention layer, represents the bias term, represents the processing result of the first attention layer, represents the number of pixel points of the real-time field environment image.

[0117] For each pixel, the output result of the first attention layer is multiplied by the brightness to realize multi-modal fusion, the responses of all pixels are accumulated and biased, the feature response is adjusted by the updated L channel value, and the edge of the panda (high contrast region) is strengthened.

[0118] Those skilled in the art will appreciate that the embodiments described herein are intended to help the reader understand the principles of the present application and should be understood as not limiting the scope of protection of the present application to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations according to the technical inspiration disclosed in the present application without departing from the essence of the present application, and these modifications and combinations are still within the scope of protection of the present application.

Claims

1. A deep learning-based method for tracking giant pandas in the wild, characterized in that, Includes the following steps: S1. Use a camera installed at a fixed location to collect real-time field environment images and extract Lab triples of pixels in the real-time field environment images; S2. Update the L channel value of the Lab triplet of the pixel using the contrast matrix; S3. Calculate the green-red contrast value and the blue-yellow contrast value based on the a-channel and b-channel values ​​of the pixels in the real-time field environment image; S4. Based on the contrast values ​​of green and red and blue and yellow in real-time field environment images and the L channel values ​​after pixel updates, construct a dynamic tracking and positioning model, and use the dynamic tracking and positioning model to determine the tracking results of real-time field environment images. In S4, the dynamic tracking and localization model includes a convolutional layer, a pooling layer, a multiplier M1, a multiplier M2, a first attention layer, a second attention layer, and an output layer. The input of the convolutional layer is used to input real-time field environment images; the first output of the convolutional layer is connected to the input of the pooling layer; the second output of the convolutional layer is connected to the input of the second attention layer; the first and second outputs of the pooling layer are respectively connected to the inputs of multiplier M1 and multiplier M2; the outputs of multiplier M1 and multiplier M2 are both connected to the input of the first attention layer; the output of the first attention layer is connected to the input of the second attention layer; the output of the second attention layer is connected to the input of the output layer; the output of the output layer is used to output the tracking results of the real-time field environment images. The expression for the first attention layer is: ; In the formula, This represents the processing result of the first attention layer. This represents the opposite values ​​of green and red. This represents the contrasting values ​​of blue and yellow. This represents the probability of a pixel with a channel value of 0 appearing in a real-time outdoor environment image. This represents the probability of a pixel with a channel value of 0 appearing in a real-time outdoor environment image. This represents the output of the pooling layer. Indicates an exponent; The expression for the second attention layer is: ; In the formula, This indicates the processing result of the second attention layer. Indicates the first After updating by one pixel, Indicates the bias term. This represents the processing result of the first attention layer. This indicates the number of pixels in a real-time field environment image.

2. The deep learning-based giant panda tracking method in the wild according to claim 1, characterized in that, S2 includes the following sub-steps: S21. Extract the L channel value corresponding to the Lab triplet of the pixel in the real-time field environment image; S22. Based on the L channel values ​​of the pixels, construct the forward contrast matrix and the turning contrast matrix of the pixels; S23. Use the maximum value between the largest eigenvalue of the forward contrast matrix and the largest eigenvalue of the turning contrast matrix as the brightness update coefficient; S24. Update the L channel value of the pixel according to the pixel brightness update coefficient.

3. The deep learning-based giant panda tracking method according to claim 2, characterized in that, In step S22, both the forward contrast matrix and the turning contrast matrix of the pixel are set as a three-row, three-column square matrix. When constructing the positive contrast matrix of a pixel, the L channel values ​​of the four neighboring pixels of a pixel are sequentially filled into the corresponding positions of the three-row, three-column square matrix, and the remaining positions are filled with 1 to generate the positive contrast matrix. When constructing the pixel orientation contrast matrix, the L channel values ​​of the D neighboring pixels of a pixel are sequentially filled into the corresponding positions of the three-row, three-column matrix, and the remaining positions are filled with 1 to generate the orientation contrast matrix.

4. The deep learning-based giant panda tracking method in the wild according to claim 2, characterized in that, In step S24, the updated L channel value of the pixel The calculation formula is: ; In the formula, This represents the L channel value of a pixel. Represents a random value between 0 and 1. Indicates the brightness update factor. Represents a symbolic function. This represents the average L-channel value of the 4-neighbor and D-neighbor pixels of a pixel in a real-time field environment image.

5. The deep learning-based giant panda tracking method in the wild according to claim 1, characterized in that, S3 includes the following sub-steps: S31. Extract the a-channel and b-channel values ​​corresponding to the Lab triplet of pixels in real-time field environment images; S32. Calculate the green-red contrast value based on the a channel value of the pixel; S33. Calculate the contrast value of blue and yellow based on the b channel value of the pixel.

6. The deep learning-based giant panda tracking method in the wild according to claim 5, characterized in that, In S32, the green and red opposing values The calculation formula is: ; In the formula, Indicates the first The a-channel value of each pixel, This indicates taking the maximum value. This indicates taking the minimum value. Represents a constant; In S33, the contrasting values ​​of blue and yellow are... The calculation formula is: ; In the formula, Indicates the first The b-channel value of each pixel.

Citation Information

Patent Citations

  • High-precision multi-target intelligent identifying, positioning and tracking method and system based on unmanned aerial vehicle

    CN113538585A

  • Image color cast detection method based on deep learning and color space transformation

    CN117132544A