Small target object recognition method and device based on dark environment and terminal equipment

Through multi-scale feature enhancement and feature fusion decoding calculation of Yolov8 module, the recognition accuracy and efficiency of small target objects in dim environments are improved, and the problem of small data volume and insignificant features in dim environments is solved.

CN120339589APending Publication Date: 2025-07-18E SURFING VISION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510480088.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art has problems such as scarce data sets, insignificant features, low recognition accuracy and high false alarm rate in small target objects recognition in dim environments. Especially in small animal recognition tasks at night, deep learning models are difficult to effectively train and recognize.

Method used

The multi-scale feature enhancement module is used to enhance the feature information of the image, combined with Yolov8's Neck module for feature fusion and weighting, and decoding and calculation is performed through Yolov8's Head decoding module to identify the position coordinates and confidence of small target objects.

Benefits of technology

It improves the recognition accuracy and efficiency of small target objects in dim environments, and solves the problems of low recognition accuracy and high false alarm rate caused by small data volume and inconspicuous features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339589A_ABST
    Figure CN120339589A_ABST
Patent Text Reader

Abstract

The invention relates to a small target object recognition method and device based on a dark environment and terminal equipment. The method comprises the steps of obtaining a template image and a target image; respectively carrying out image feature information enhancement processing on the template image and the target image by adopting a multi-scale feature enhancement module to obtain a template enhanced image and a target enhanced image; respectively carrying out feature extraction on the template enhanced image and the target enhanced image to obtain a template feature map and a target feature map; performing similarity calculation according to the template feature map and the target feature map to obtain a similarity value set; performing feature fusion processing on the target feature map by adopting a Yolov8 Neck module to obtain a multi-channel feature map; according to all the similarity value sets and all the multi-channel feature maps, weighting processing is correspondingly carried out to obtain final feature maps of different scales; and a Yolov8 Head decoding module is adopted to decode, calculate and identify all the final feature maps to obtain the position coordinates and confidence of the small target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of computer vision and artificial intelligence, and particularly to a method, device, and terminal device for identifying small target objects in a dim environment. Background Art

[0002] The identification of small animals such as mice and cockroaches at night as target objects is an important task for health and safety monitoring in areas such as kitchens and catering back kitchens. However, due to the small size of mice, their frequent nocturnal activities, and poor lighting conditions, there are many challenges in existing methods for identifying and detecting small animals such as mice. First, the dataset of small animals at night is scarce, and the amount of data samples that can be collected is very small, which cannot support the training of a deep learning model to a relatively good accuracy. In existing technologies, the most common method for identifying small animals at night is to directly adopt some object detection methods based on deep learning, such as the YOLO series. However, most of these methods rely on a large amount of labeled data, so they perform mediocrely in the detection task of small animals at night with a small number of samples, and the accuracy of the finally trained algorithm is generally average. Second, the imaging area of small animals at night is small, the night light is dim, and there is a lot of image noise. Therefore, it is more difficult for the deep learning model to extract the feature information of small animals, and the ability to classify based on the existing feature information is poor. Ultimately, it affects the recognition accuracy of the deep learning model for small animals. In actual projects, it is often found that the recognition algorithm misidentifies other objects in the monitoring area as target objects and then continuously triggers alarms, causing great trouble to users.

[0003] Therefore, the problems existing in the recognition of small animals as target objects by the existing deep learning model are as follows: how to improve the problem that the features of small animals in night images are not obvious and difficult to extract; and how to train a model with relatively high accuracy through a small amount of data due to the difficulty in collecting positive samples and the small dataset. Summary of the Invention

[0004] This application provides a method, device, and terminal device for identifying small target objects in a dim environment, which is used to solve the technical problems of low recognition accuracy and high false alarm rate of target objects caused by the difficulty in collecting pictures of target objects at night, small data volume, and unclear features of target objects at night in the existing task of identifying small animals as target objects.

[0005] To achieve the above object, this application provides the following technical solutions:

[0006] On the one hand, it provides a method for identifying small target objects in a dim environment, including the following steps:

[0007] Obtain a template image and a target image of the small target object in a dim environment;

[0008] The multi-scale feature enhancement module is used to perform image feature information enhancement processing on the template image and the target image respectively, obtaining a template enhanced image corresponding to the template image and a target enhanced image corresponding to the target image;

[0009] Feature extraction is performed on the template enhanced image and the target enhanced image respectively, obtaining template feature maps corresponding to different scales of the template enhanced image and target feature maps corresponding to different scales of the target enhanced image;

[0010] Similarity calculation is performed based on the template feature maps and the target feature maps of different scales, obtaining a set of similarity values of different scales; and the Neck module of Yolov8 in the object detection algorithm based on deep learning is used to perform feature fusion processing on the target feature maps of different scales, obtaining multi-channel feature maps corresponding to different scales;

[0011] Weighted processing is performed corresponding to all the sets of similarity values and all the multi-channel feature maps, obtaining final feature maps of different scales;

[0012] The Head decoding module of Yolov8 in the object detection algorithm based on deep learning is used to perform decoding calculation on all the final feature maps, identifying the position coordinates and confidence levels of small target objects.

[0013] Preferably, the multi-scale feature enhancement module includes a first branch sub-module and a second branch sub-module connected in parallel, and the first branch sub-module includes a first branch, a second branch, a third branch, a fourth branch, and a fifth branch connected in parallel;

[0014] The second branch and the third branch are used to increase the receptive field of the image and extract image feature information of different scales and shapes;

[0015] The fourth branch is used to enhance the edge high-frequency information and texture high-frequency information of the image;

[0016] The fifth branch is used to process the image with a dilation rate of 5;

[0017] Among them, the second branch sub-module is used to process the input image through a 1×1 convolutional layer, obtaining a first processed image, and the first branch sub-module is used to process the input image through the first branch, the second branch, the third branch, the fourth branch, and the fifth branch, obtaining a second processed image; the multi-scale feature enhancement module also processes the second processed image through a 1×1 convolutional layer and then merges it with the first processed image, obtaining an enhanced image with enhanced image feature information.

[0018] Preferably, the first branch includes a 3×3 ordinary convolutional layer; the second branch includes a 1×3 deformable convolutional layer and a 3×3 dilated convolutional layer with a dilation rate of 2; the third branch includes a 3×1 deformable convolutional layer and a 3×3 dilated convolutional layer with a dilation rate of 2; the fourth branch includes a 1×1 ordinary convolutional layer and a max pooling layer; the fifth branch includes a dilated convolutional layer with a dilation rate of 2.

[0019] Preferably, the small target object recognition method based on a dim environment includes: using a cross-stage local network to respectively perform feature extraction on the template enhanced image and the target enhanced image, obtaining template feature maps with different scales corresponding to the template enhanced image and target feature maps with different scales corresponding to the target enhanced image.

[0020] Preferably, calculating similarity according to the template feature maps and the target feature maps with different scales, obtaining a set of similarity values with different scales, including:

[0021] Obtaining the feature values at each position from the template feature map and the target feature map of the same scale, obtaining corresponding R template feature values and R target feature values;

[0022] Calculating a position coefficient according to the template feature value and the target feature value at the same position using an activation function;

[0023] Calculating a position distance according to the template feature value, the target feature value and the position coefficient at the same position; forming the set of similarity values of the same scale with all the position distances of the same scale.

[0024] Preferably, performing weighted processing on all the sets of similarity values and all the multi-channel feature maps correspondingly, obtaining final feature maps with different scales, including: performing multiplication weighted processing on the position distances of the set of similarity values of the same scale and the channel data at the same position in the multi-channel feature map of the corresponding scale, obtaining the final feature map of this scale.

[0025] Preferably, using the Head decoding module of Yolov8 in the object detection algorithm based on deep learning to perform decoding calculation on all the final feature maps, identifying the position coordinates and confidence of the small target object, including:

[0026] Using the Head decoding module of Yolov8 to process all the final feature maps, obtaining the prediction data output by each prediction box;

[0027] Calculating according to the prediction data of each prediction box, obtaining the coordinate data, confidence and class probability corresponding to each prediction box;

[0028] Select the prediction box with the largest numerical value from the class probabilities of all the prediction boxes as the recognition prediction box, and use the coordinate data of the recognition prediction box as the position coordinates of the recognized small target object, and use the confidence of the recognition prediction box as the confidence of the recognized small target object.

[0029] In another aspect, a small target object recognition device based on a dim environment is provided, including an image acquisition unit, an image enhancement unit, a feature extraction unit, a calculation fusion unit, a weighting processing unit, and a calculation recognition unit;

[0030] The image acquisition unit is used to acquire a template image and a target image of a small target object in a dim environment;

[0031] The image enhancement unit is used to perform image feature information enhancement processing on the template image and the target image respectively by using a multi-scale feature enhancement module to obtain a template enhanced image corresponding to the template image and a target enhanced image corresponding to the target image;

[0032] The feature extraction unit is used to perform feature extraction on the template enhanced image and the target enhanced image respectively to obtain template feature maps with different scales corresponding to the template enhanced image and target feature maps with different scales corresponding to the target enhanced image;

[0033] The calculation fusion unit is used to calculate the similarity according to the template feature maps and the target feature maps with different scales to obtain a set of similarity values with different scales; and use the Neck module of Yolov8 in the object detection algorithm based on deep learning to perform feature fusion processing on the target feature maps with different scales to obtain multi-channel feature maps corresponding to different scales;

[0034] The weighting processing unit is used to perform weighting processing according to all the sets of similarity values and all the multi-channel feature maps to obtain final feature maps with different scales;

[0035] The calculation recognition unit is used to perform decoding calculation on all the final feature maps by using the Head decoding module of Yolov8 in the object detection algorithm based on deep learning to recognize the position coordinates and confidence of the small target object.

[0036] Preferably, the multi-scale feature enhancement module includes a first branch sub-module and a second branch sub-module connected in parallel, and the first branch sub-module includes a first branch, a second branch, a third branch, a fourth branch, and a fifth branch connected in parallel;

[0037] The second branch and the third branch are used to increase the receptive field of the image and extract image feature information with different scales and shapes;

[0038] The fourth branch is used to enhance the edge high-frequency information and texture high-frequency information of the image;

[0039] The fifth branch is used to process the image with a dilation rate of 5;

[0040] Among them, the second branch sub-module is used to process the input image through a 1×1 convolutional layer to obtain a first processed image, and the first branch sub-module is used to process the input image through the first branch, the second branch, the third branch, the fourth branch and the fifth branch to obtain a second processed image; the multi-scale feature enhancement module also processes the second processed image through a 1×1 convolutional layer and then merges it with the first processed image to obtain an enhanced image with enhanced image feature information.

[0041] On the other hand, a terminal device is provided, including a processor and a memory;

[0042] The memory is used to store program code and transmit the program code to the processor;

[0043] The processor is used to execute the above-mentioned small target object recognition method based on a dim environment according to the instructions in the program code.

[0044] The small target object recognition method, device and terminal device based on a dim environment, the small target object recognition method based on a dim environment includes obtaining a template image and a target image of a small target object in a dim environment; using a multi-scale feature enhancement module to perform image feature information enhancement processing on the template image and the target image respectively to obtain a template enhanced image corresponding to the template image and a target enhanced image corresponding to the target image; performing feature extraction on the template enhanced image and the target enhanced image respectively to obtain template feature maps of different scales corresponding to the template enhanced image and target feature maps of different scales corresponding to the target enhanced image; calculating similarity according to the template feature maps and target feature maps of different scales to obtain a set of similarity values of different scales; and using the Neck module of Yolov8 in the object detection algorithm based on deep learning to perform feature fusion processing on the target feature maps of different scales to obtain multi-channel feature maps corresponding to different scales; performing weighted processing according to all the sets of similarity values and all the multi-channel feature maps corresponding to them to obtain final feature maps of different scales; using the Head decoding module of Yolov8 in the object detection algorithm based on deep learning to perform decoding calculation on all the final feature maps to identify the position coordinates and confidence of the small target object.

[0045] As can be seen from the above technical solutions, the present application has the following advantages: The small target object recognition method based on a dim environment performs image enhancement, feature extraction, similarity calculation, feature fusion, weighted processing on the template image and the target image, and then performs feature decoding and recognition to obtain the position coordinates and confidence of the small target object, improving the recognition accuracy and efficiency, and solving the technical problems of low recognition accuracy and high false alarm rate of target objects in the existing task of recognizing small animals and other target objects at night, due to the difficulty in collecting pictures of target objects at night, small data volume, and unclear features of target objects at night.

[0046] The small target object recognition device based on a dim environment realizes image enhancement, feature extraction, similarity calculation, feature fusion, weighted processing on the template image and the target image through an image acquisition unit, an image enhancement unit, a feature extraction unit, a calculation and fusion unit, a weighted processing unit and a calculation and recognition unit, and then performs feature decoding to recognize the position coordinates and confidence of the small target object, improving the recognition accuracy and efficiency. Description of the Drawings

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0048] Figure 1 It is a flowchart of the steps of the small target object recognition method based on a dim environment described in the embodiments of the present application;

[0049] Figure 2 It is a flowchart of the small target object recognition method based on a dim environment described in the embodiments of the present application;

[0050] Figure 3 It is a schematic structural framework diagram of the multi-scale feature enhancement module in the small target object recognition method based on a dim environment described in the embodiments of the present application;

[0051] Figure 4 It is a schematic architecture diagram of the cross-stage local network in the small target object recognition method based on a dim environment described in the embodiments of the present application;

[0052] Figure 5 It is a schematic framework diagram of the small target object recognition device based on a dim environment described in the embodiments of the present application;

[0053] Figure 6 It is a schematic diagram of the terminal device described in the embodiments of the present application. Detailed Embodiments

[0054] In order to make the inventive purpose, features, and advantages of the present application more obvious and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the embodiments described below are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without making creative efforts belong to the scope of protection of the present application.

[0055] In the description of the embodiments of the present application, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, the meaning of "a plurality" is two or more, unless otherwise clearly and specifically defined.

[0056] In the embodiments of the present application, unless otherwise clearly specified and limited, terms such as "installation", "connection", "connection", "fixation", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to specific circumstances.

[0057] Explanation of the patent terms of the present application:

[0058] Receptive field: When a receptor is stimulated and excited, nerve impulses (various sensory information) are transmitted to the higher center through the centripetal neurons in the sensory organ. The stimulus area to which a neuron responds (dominates) is called the receptive field (receptive field) of the neuron. Also translated as receptive field. The terminal sensory neurons, relay nucleus neurons, and neurons in the sensory area of the cerebral cortex all have their own receptive fields.

[0059] Ordinary convolution: The ordinary convolution in the neural network CNN is a three-dimensional convolution. The number of channels of each convolution kernel is the same as the number of channels of the input data, and the number of convolution kernels is the number of channels of the output data. The specific calculation process is to perform the operation of multiplying and adding the corresponding positions of the convolution kernel and the input data of the corresponding receptive field in each channel to obtain an output value. N different convolution kernels result in N output values. Connecting these values along the depth direction gives the output of a point; sliding the convolution kernel along the space gives the final convolution result.

[0060] Deformable convolution means that the convolution kernel is no longer rectangular and can be of any shape. Traditional convolution kernels are generally rectangular or square, but the existing technology puts forward a rather counterintuitive view that the shape of the convolution kernel can be variable. A deformable convolution kernel allows it to only look at the image regions of interest, resulting in better recognized features.

[0061] Atrous convolution, also known as dilated convolution, is mainly used in the field of object segmentation. A standard 3x3 convolution kernel can only see an area of 3x3 in the corresponding region. However, to enable the convolution kernel to see a larger range, dilated convolution makes it possible. The atrous convolution introduces a new parameter called the dilation rate, which defines the spacing between values when the convolution kernel processes data. In other words, compared with the original standard convolution, dilated convolution has an additional hyperparameter of the dilation rate, which refers to the number of intervals before the four points of the kernel. The dilation rate of ordinary convolution is 1.

[0062] The max pooling layer is a dimensionality reduction operation mainly used to reduce the amount of data and reduce overfitting. It takes the maximum value of each small rectangular region in the feature map output by the previous layer as the output. Specifically, the max pooling layer divides the input matrix into blocks and selects the maximum value from each block for output. These output matrices are stacked together to form a new feature map.

[0063] The embodiments of the present application provide a method, device, and terminal device for identifying small target objects based on a dim environment, which solve the technical problems of low accuracy and high false alarm rate in the task of identifying target objects such as small animals at night, due to the difficulty in collecting pictures of target objects at night, small data volume, and unclear features of target objects at night.

[0064] Embodiment 1:

[0065] Figure 1 is the flowchart of the steps of the method for identifying small target objects based on a dim environment described in the embodiments of the present application, Figure 2 is the framework flowchart of the method for identifying small target objects based on a dim environment described in the embodiments of the present application.

[0066] As Figure 1 and Figure 2 shown, the embodiments of the present application provide a method for identifying small target objects based on a dim environment, including the following steps:

[0067] S1. Obtain a template image and a target image of the small target object in a dim environment.

[0068] It should be noted that step S1 is to obtain the template image and the target image of the small target object in a dim environment, providing data for subsequent steps. In this embodiment, the small target object can be animals such as mice and cockroaches. The dim environment can be at night or an environment with insufficient light and being dim. This method for identifying small target objects based on a dim environment will be described taking a mouse as an example.

[0069] S2. Use a multi-scale feature enhancement module to perform image feature information enhancement processing on the template image and the target image respectively, obtaining a template enhanced image corresponding to the template image and a target enhanced image corresponding to the target image.

[0070] It should be noted that step S2 is to perform image feature information enhancement on the template image and the target image respectively through the multi-scale feature enhancement module MFE. In this embodiment, in actual scenario applications, the imaging area of a mouse at night is small, and since night pictures are grayscale pictures, the difference between the features of the mouse and the background features is small, and a large amount of background feature information increases the noise, making it more difficult to extract effective feature information from the image. Secondly, the installation heights and distances of the imaging devices (such as cameras) for obtaining images in different scenarios are different, so the sizes of the mouse features are different, resulting in poor robustness of the trained model. To improve the effective extraction of the feature information of the mouse at night, this method for identifying small target objects based on a dim environment uses the multi-scale feature enhancement module MFE to perform image feature information enhancement on the template image and the target image respectively, providing data for more comprehensively capturing local and global feature information in subsequent steps.

[0071] S3. Perform feature extraction on the template enhanced image and the target enhanced image respectively, obtaining template feature maps of different scales corresponding to the template enhanced image and target feature maps of different scales corresponding to the target enhanced image.

[0072] It should be noted that in step S3, feature extraction is performed on the template enhanced image and the target enhanced image obtained in step S2 respectively, obtaining template feature maps and target feature maps of different scales, providing data for weight sharing in subsequent steps.

[0073] S4. Calculate the similarity based on the template feature maps and target feature maps of different scales, obtaining a set of similarity values of different scales; and use the Neck module of Yolov8 in the object detection algorithm based on deep learning to perform feature fusion processing on the target feature maps of different scales, obtaining multi-channel feature maps corresponding to different scales.

[0074] It should be noted that in step S4, first, the similarity sets of different scales are obtained by calculating the feature maps based on the template feature map and the target feature obtained in step S3, providing correction data for the subsequent weighted processing; second, according to the Neck module of Yolov8 in the object detection algorithm based on deep learning, the target feature maps of different scales are subjected to feature fusion processing to obtain multi-channel feature maps corresponding to different scales. In this embodiment, the small target object recognition method based on a dim environment calculates the similarity set through the template feature map and the target feature map that provide more obvious contour feature picture feature information, so as to correct the feature extraction of the backbone network during the subsequent training process, and solve the problem of insufficient model training caused by a small number of samples. In this embodiment, the Neck module of YOLOv8 is mainly responsible for fusing and enhancing the target feature maps from different scales to obtain multi-channel feature maps corresponding to different scales. The multi-channel feature maps include information at different levels, including low-level features and high-level features. Low-level features usually contain detailed information such as edges and textures, which are helpful for locating small targets. High-level features contain rich semantic information, which is beneficial for judging the target category and overall shape. The Nec module fuses these feature maps from different scales together through upsampling, concatenation (Concat), and further convolutions (such as the C2f module and the PAN / FPN structure) to generate a more unified and expressive feature representation. The features of the multi-channel feature maps obtained by the small target object recognition method based on a dim environment retain both the low-level localization information and fuse the high-level semantic information, providing better input for the subsequent detection head (Head), thereby improving the accuracy and robustness of the detection.

[0075] S5. Perform weighted processing on all similarity value sets and all corresponding multi-channel feature maps to obtain final feature maps of different scales.

[0076] It should be noted that in step S5, a weighted multiplication operation is performed on the similarity value set obtained in step S4 and the feature information after the Neck module of Yolov8 processes the multi-channel feature maps to obtain the final feature maps of different levels of information, providing data for the subsequent steps.

[0077] S6. Use the Head decoding module of Yolov8 in the object detection algorithm based on deep learning to perform decoding calculations on all the final feature maps to identify the position coordinates and confidence levels of small target objects.

[0078] It should be noted that in step S6, the final feature maps obtained in step S5 are subjected to decoding calculations using the Head decoding module of Yolov8 to identify the position coordinates and confidence levels of small target objects.

[0079] In the embodiment of the present application, the small target object recognition method based on a dim environment combines a multi-scale feature enhancement module (MFE) with feature extraction and Yolov8 to implement an improved end-to-end nighttime mouse detection framework, aiming to solve the problem of insufficient accuracy in model training during nighttime mouse recognition.

[0080] A small target object recognition method based on a dim environment provided by the present application includes: obtaining a template image and a target image of a small target object in a dim environment; using a multi-scale feature enhancement module to perform image feature information enhancement processing on the template image and the target image respectively, to obtain a template enhanced image corresponding to the template image and a target enhanced image corresponding to the target image; performing feature extraction on the template enhanced image and the target enhanced image respectively, to obtain template feature maps of different scales corresponding to the template enhanced image and target feature maps of different scales corresponding to the target enhanced image; calculating similarity based on the template feature maps and target feature maps of different scales to obtain a set of similarity values of different scales; and using the Neck module of Yolov8 in a deep learning-based object detection algorithm to perform feature fusion processing on the target feature maps of different scales to obtain multi-channel feature maps corresponding to different scales; performing weighted processing on all the sets of similarity values and all the multi-channel feature maps correspondingly to obtain final feature maps of different scales; using the Head decoding module of Yolov8 in a deep learning-based object detection algorithm to perform decoding calculation on all the final feature maps to identify the position coordinates and confidence of the small target object. By performing image enhancement, feature extraction, similarity calculation, feature fusion, weighted processing, and then feature decoding and recognition on the template image and the target image, the small target object recognition method based on a dim environment obtains the position coordinates and confidence of the small target object, improves the recognition accuracy and efficiency, and solves the technical problems of low recognition accuracy and high false alarm rate of target objects in the existing task of recognizing small animal target objects at night, due to the difficulty in collecting pictures of target objects at night, small data volume, and unclear features of target objects at night.

[0081] Figure 3 It is a schematic structural framework diagram of the multi-scale feature enhancement module in the small target object recognition method based on a dim environment described in the embodiment of the present application.

[0082] As Figure 3 shown, in an embodiment of the present application, the multi-scale feature enhancement module includes a first branch sub-module and a second branch sub-module connected in parallel. The first branch sub-module includes a first branch, a second branch, a third branch, a fourth branch, and a fifth branch connected in parallel;

[0083] The second branch and the third branch are used to increase the receptive field of the image and extract image feature information of different scales and shapes;

[0084] The fourth branch is used to enhance the high-frequency edge information and high-frequency texture information of the image;

[0085] The fifth branch is used to process the image with a dilation rate of 5;

[0086] Among them, the second-branch sub-module is used to process the input image through a 1×1 convolutional layer to obtain a first processed image, and the first-branch sub-module is used to process the input image through the first branch, the second branch, the third branch, the fourth branch, and the fifth branch to obtain a second processed image; the multi-scale feature enhancement module also processes the second processed image through a 1×1 convolutional layer and then merges it with the first processed image to obtain an enhanced image with enhanced image feature information.

[0087] It should be noted that, as Figure 3 shown, the first branch includes a 3×3 ordinary convolutional layer Conv; the second branch includes a 1×3 variable convolutional layer Variable Conv and a 3×3 dilated convolutional layer Dilated Conv with a dilation rate of 2; the third branch includes a 3×1 variable convolutional layer and a 3×3 dilated convolutional layer with a dilation rate of 2; the fourth branch includes a 1×1 ordinary convolutional layer and a max pooling layer; the fifth branch includes a dilated convolutional layer with a dilation rate of 2. In this embodiment, the third branch is composed of a 3x1 variable convolutional layer and a 3x3 dilated convolutional layer with a dilation rate of 2. The functions of the second branch and the third branch are both to increase the receptive field and extract feature information of different scales and shapes. The fourth branch is composed of a 1x1 ordinary convolutional layer and a max pooling layer, which is used to enhance the high-frequency information such as the edges and textures of the image. The fifth branch is composed of a dilated convolutional layer with a dilation rate of 5, which is used to reduce the model parameters and the computational amount. After the outputs of multiple branches are connected, the first-branch sub-module is merged with the input feature map through a 1x1 ordinary convolutional layer. Through the multi-scale feature enhancement module MFE, the input image can be enhanced with image feature information, which is convenient for the subsequent feature extraction network to more easily extract the information of features related to the task, and solves the problem that the features of small target objects (such as mice) in night images are not obvious.

[0088] Figure 4 This is a schematic diagram of the architecture of the cross-stage local network in the small target object recognition method based on a dim environment described in the embodiment of the present application. In Figure 4Among them, "conv" is the convolutional layer, "concat" is the concatenation, "Dense Layer" is the fully connected layer, and "Partial Dense Block" is an important part of CSPNet, aiming to solve the problem of low computational efficiency in the inference process of traditional deep neural networks. The Partial Dense Block divides the feature map into two parts. One part is processed by the traditional convolutional layer, and the other part is directly passed to the subsequent layers through cross-stage connections, thus reducing the computational amount and memory overhead. The "Partial Transition Layer" is an important part of the DenseNet network structure, mainly used to reduce the size and number of channels of the feature map. The Partial Transition Layer usually consists of a 1x1 convolution and a 2x2 average pooling. Its purpose is to reduce the depth of the feature map while maintaining the spatial resolution of the feature map, thereby reducing the computational amount and memory usage. The "copy" in CSPNet refers to a network design method, mainly used to reduce the computational amount and improve the running speed while maintaining the accuracy of the model.

[0089] In one embodiment of the present application, the small target object recognition method based on a dim environment includes: using a cross-stage local network CSPDarknet to respectively extract features from the template enhanced image and the target enhanced image, obtaining template feature maps corresponding to different scales of the template enhanced image and target feature maps corresponding to different scales of the target enhanced image.

[0090] It should be noted that when the cross-stage local network CSPNet (Cross Stage Partial Network) extracts features from an image, the input image is split into two parts: one part is processed through a series of convolutional or residual blocks to extract deep features; the other part directly bypasses these processes of a series of convolutional or residual blocks and is retained as bypass information. Finally, these two parts of features are integrated together through concatenation (or fusion) and a subsequent 1×1 convolutional layer to obtain the feature map of the extracted features. By extracting features from the image through the cross-stage local network to obtain the corresponding feature map, it can not only avoid the redundant calculation caused by the repeated transmission of gradients in the network, but also enhance the reusability of features, improve the expression ability of feature extraction, and the inference efficiency. For example: Figure 4As shown in the figure, initial feature extraction (such as the Focus layer): First, the input image passes through a Focus module. The Focus module "compresses" the spatial information onto the channels by sampling the image at intervals (for example, taking a value every other pixel), thus generating a feature map with more channels, which provides preliminary low-level features for subsequent extraction; Stagewise downsampling and feature fusion: The image enters multiple stages (such as dark2, dark3, dark4, dark5). At the beginning of each stage, a convolutional layer is used for downsampling to reduce the spatial size while increasing the number of channels; then it enters the CSPLayer module. Inside the CSPLayer module, the feature map is divided into two parts: the main branch and the bypass branch. The main branch further extracts features through a series of residual modules (such as the Bottleneck module) to capture deep semantic information. The bypass branch refers to directly retaining the original features without complex processing. Finally, the features obtained from the main branch and the bypass branch are integrated through concatenation or addition, and then passed through a 1×1 convolutional layer to output a feature map that combines shallow details and deep semantic information, providing data for subsequent data sharing.

[0091] In an embodiment of the present application, since the commonly used similarity calculation methods are Euclidean distance and cosine similarity at present, cosine similarity can ensure that the directions of sample features and the feature center are basically the same, but there is no limit to the similarity of the magnitudes between features, which may lead to classification errors. Euclidean distance can ensure that the geometric distance between two features is very close, but when the Euclidean distance between two features is non-zero, the angle between them is not unique. It can be seen that these two commonly used measurement methods have their respective limitations. To address such limitations, the small target object recognition method based on a dim environment calculates similarity according to template feature maps and target feature maps of different scales, and obtains a set of similarity values of different scales, including:

[0092] Obtain the feature values at each position from the template feature map and the target feature map of the same scale, and obtain the corresponding R template feature values and R target feature values;

[0093] Calculate the position coefficient according to the template feature value and the target feature value at the same position using an activation function;

[0094] Calculate the position distance according to the template feature value, the target feature value, and the position coefficient at the same position; form the similarity value set sim = {d i} of the same scale from all the position distances of the same scale.

[0095] It should be noted that the position distance is calculated using a distance formula according to the template feature value, the target feature value, and the position coefficient at the same position. In this embodiment, the activation function is:

[0096]

[0097] The distance formula is:

[0098] In the formula, cos() is the cosine function, sigmoid() is the activation function, and d i is the position distance of the i-th position, is the position coefficient of the ith position, is the target feature value at the i-th position, is the template feature value at the i-th position. In this embodiment, use The Euclidean distance is weighted. When the angle between the center of the template feature map and the target feature map is small, but the Euclidean distance is large, sim is affected by the parameter The impact is relatively large, which can balance the limitations of a single Euclidean distance, reduce some classification errors, and improve the prediction accuracy of similarity calculation. i}: sim is a matrix. The feature maps output by the cross-stage local network CSPDarknet are feature maps of three different scales, which are 80*80*256 (256 is the number of channels), 40*40*512 (512 is the number of channels), and 20*20*512 (512 is the number of channels). When calculating the similarity value, the target feature map and the template feature map of different scales are calculated for similarity respectively, that is, the features of the corresponding positions of the two 80*80*256 feature maps are calculated for similarity, and the obtained similarity value is a matrix of 80*80; the processing methods of 40*40*512 and 20*20*512 feature maps are similar. Therefore, the sim values of the similarity value sets are finally obtained, which are three matrices of 80*80, 40*40, and 20*20.

[0099] In one embodiment of the present application, weighted processing is performed on all similarity value sets and all multi-channel feature maps to obtain final feature maps of different scales, including: weighted multiplication of position distances of similarity value sets of the same scale with channel data at the same position in the multi-channel feature map of the corresponding scale to obtain the final feature map of the scale.

[0100] It should be noted that the features processed by the Neck module of Yolov8 are actually multi-channel feature maps (for example, 64*64*80, where 64*64 is the size of the feature map and 80 is the number of channels in the depth). The specification of the similarity values is the same as that of the feature map. The similarity values at the same positions are multiplied by the values of the feature map on the Neck module of Yolov8 to output the final feature map of the corresponding scale. In this embodiment, during the process of multiplying and weighting the similarity with the feature map output by the Neck module, the feature maps output by the Neck module are also multi-channel feature maps of three scales: 80*80*256 (256 is the number of channels), 40*40*512 (512 is the number of channels), and 20*20*512 (512 is the number of channels). The set of similarity values of 80*80 is fused with the multi-channel feature map of 80*80*256 (the 256 channels at the same position are all multiplied by the distance at the same position), so the finally output is still the final feature map of the scale of 80*80*256.

[0101] In the embodiment of the present application, a self-matching branch module is composed of a multi-scale feature enhancement module, a cross-stage partial network CSPDarknek, and similarity calculation Similarity. During the process of processing the template image and the target image by the small target object recognition method based on a dim environment, by calculating the similarity between the features of the template feature map and the features of the target feature map, the features of small target objects (such as mice) are made clearer during the subsequent training process. Through the backpropagation of the similarity values, the parameters of the backbone network of Yolov8 are automatically corrected, the correct extraction of the feature information of small target objects (such as mice) by the network is improved, the loss function value is reduced, and thus the accuracy of the recognition target is improved, and the problem of insufficient training accuracy caused by insufficient number of positive samples is improved.

[0102] In an embodiment of the present application, the Head decoding module of Yolov8 in the object detection algorithm based on deep learning is used to perform decoding calculations on all final feature maps, and the position coordinates and confidence levels of small target objects identified include:

[0103] The Head decoding module of Yolov8 is used to process all final feature maps to obtain the prediction data output by each prediction box;

[0104] According to the prediction data of each prediction box, the coordinate data, confidence level, and class probability corresponding to each prediction box are calculated;

[0105] The prediction box with the largest value is selected from the class probabilities of all prediction boxes as the recognition prediction box, and the coordinate data of the recognition prediction box is used as the position coordinates of the identified small target object, and the confidence level of the recognition prediction box is used as the confidence level of the identified small target object.

[0106] It should be noted that according to the final feature maps obtained at three scales of 80*80, 40*40, and 20*20, each grid on each final feature map predicts a fixed number of bounding boxes (or prediction boxes), such as 3 boxes per grid. The predicted data output for each bounding box includes four coordinate offsets, the object existence reporting rate conf, and multiple class probabilities. The four coordinate offsets are denoted as: t x , t y , t w and t h . The multiple class probabilities are denoted as c1, c2,..., c n . In this embodiment, the output of the Head decoding module of Yolov8 is also three feature maps of different scales, but the number of channels of the feature map is related to the number of target classes to be detected. This small target object recognition method based on a dim environment is used to detect small target objects (such as mice), so the target classes are two: mice and non-mice. Each grid (feature map value) predicts 4 target position parameters (x, y, w, h) and 2 class scores, and the number of channels is 6 (4 + 2). That is, the output of the head decoding module is three final feature maps of scales 80*80*2 (2 is the number of channels), 40*40*2 (2 is the number of channels), and 20*20*2 (2 is the number of channels). Among them, the final feature map of the 80×80 scale is used to detect small targets; the final feature map of the 40×40 scale is used to detect medium targets; the final feature map of the 20×20 scale is used to detect large targets. The total number of predictions after splicing these three final feature maps is (80×80 + 40×40 + 20×20) = 8400 prediction boxes, and the final output dimension is 1×6×8400.

[0107] In the embodiment of this application, according to the prediction data calculation of each prediction box, the coordinate data, confidence, and class probability corresponding to each prediction box are obtained, including:

[0108] Obtain the coordinate data (c x , c y ) of the upper left corner of the current grid where each prediction box is located; and determine the stride according to the scale of the final feature map where the prediction box is located and the scale of the image;

[0109] According to the coordinate offsets, coordinate data, and stride, use the coordinate data formula to calculate to obtain the coordinate data (b x , b y );

[0110] According to the object existence reporting rate, use the confidence formula to calculate to obtain the confidence;

[0111] According to the multiple class probabilities, use the class probability formula to calculate to obtain the class probability.

[0112] It should be noted that determining the stride according to the scale where the prediction box is located can be understood as follows: the stride of the feature map relative to the input image (for example, if the scale of the input image is 640x640 and the scale of the feature map is 80x80, then the stride = 640 / 80 = 8). The coordinate data formula is b x = (sigmoid(t x )+c x )×stridee x , b y = (sigmoid(t y )+c y )×stridee y ; b w =anchor w ×e tw , b h =anchor h ×e th ; In the formula, sigmmoid() restricts the coordinate offset to (0, 1) to ensure that the center point falls within the current grid, and anchor w and anchor h are the preset anchor box width and height respectively (YOLOv8 may adopt adaptive anchor boxes or anchor-free design. Here, the classic YOLO is taken as an example). e tw and e th are both used to ensure that the width and height are positive numbers and to scale the anchor box size; stride is the stride. The confidence formula is C = sigmoid(conf), where C is the confidence. The class probability formula is P = max(sigmoid(c1), sigmoid(c2),..., sigmoid(c n ), where P is the class probability. For example: The prediction data at the grid (5, 5) of a certain 80x80 final feature map are: t x = 0.3, t y = -0.2, t w = 0.1, t h = 0.4; conf = 2.5; c = 1.8. Decoding steps: Center point coordinates: b x = (sigmoid(0.3)+5)×8 ≈ (0.574+5)×8 = 44.6, b y = (sigmoid(-0.2)+5)×8 ≈ (0.450+5)×8 = 43.6, b w = anchor w ×e 0.1 = 20×1.105 = 22.1, b h = anchor h ×e 0.4= 30 × 1.492 = 44.8; C = sigmoid(2.5) ≈ 0.924. P = sigmoid(1.8) ≈ 0.858. The bounding box coordinates of the small target animal are identified as: (x = 44.6, y = 43.6, w = 22.1, h = 44.8), the confidence level is 0.924, and the probability of identifying the small target object is 0.858.

[0113] Embodiment 2:

[0114] Figure 5 It is a framework schematic diagram of the small target object recognition device based on a dim environment described in the embodiments of the present application.

[0115] As Figure 5 shown, the embodiments of the present application provide a small target object recognition device based on a dim environment, including an image acquisition unit 10, an image enhancement unit 20, a feature extraction unit 30, a calculation fusion unit 40, a weighting processing unit 50, and a calculation recognition unit 60;

[0116] The image acquisition unit 10 is used to acquire a template image and a target image of a small target object in a dim environment;

[0117] The image enhancement unit 20 is used to perform image feature information enhancement processing on the template image and the target image respectively by using a multi-scale feature enhancement module to obtain a template enhanced image corresponding to the template image and a target enhanced image corresponding to the target image;

[0118] The feature extraction unit 30 is used to perform feature extraction on the template enhanced image and the target enhanced image respectively to obtain template feature maps of different scales corresponding to the template enhanced image and target feature maps of different scales corresponding to the target enhanced image;

[0119] The calculation fusion unit 40 is used to perform similarity calculation according to the template feature maps and target feature maps of different scales to obtain a set of similarity values of different scales; and perform feature fusion processing on the target feature maps of different scales by using the Neck module of Yolov8 in the object detection algorithm based on deep learning to obtain multi-channel feature maps corresponding to different scales;

[0120] The weighting processing unit 50 is used to perform weighting processing according to all the sets of similarity values and all the multi-channel feature maps correspondingly to obtain final feature maps of different scales;

[0121] The calculation recognition unit 60 is used to perform decoding calculation on all the final feature maps by using the Head decoding module of Yolov8 in the object detection algorithm based on deep learning to identify the position coordinates and confidence level of the small target object;

[0122] Among them, the multi-scale feature enhancement module includes a first branch sub-module and a second branch sub-module connected in parallel. The first branch sub-module includes a first branch, a second branch, a third branch, a fourth branch, and a fifth branch connected in parallel;

[0123] The second branch and the third branch are used to increase the receptive field of the image and extract image feature information of different scales and shapes;

[0124] The fourth branch is used to enhance the edge high-frequency information and texture high-frequency information of the image;

[0125] The fifth branch is used to process the image with a dilation rate of 5;

[0126] Among them, the second branch sub-module is used to process the input image through a 1×1 convolutional layer to obtain a first processed image. The first branch sub-module is used to process the input image through the first branch, the second branch, the third branch, the fourth branch, and the fifth branch to obtain a second processed image. The multi-scale feature enhancement module also processes the second processed image through a 1×1 convolutional layer and then merges it with the first processed image to obtain an enhanced image with enhanced image feature information.

[0127] It should be noted that the content of the modules in the device of Embodiment 2 has been described in the content of the steps in Embodiment 1, and the content of the modules of the small target object recognition device based on a dim environment will not be repeated in this embodiment. In this embodiment, the small target object recognition device based on a dim environment realizes image enhancement, feature extraction, similarity calculation, feature fusion, weighted processing, and then feature decoding on the template image and the target image through an image acquisition unit, an image enhancement unit, a feature extraction unit, a calculation fusion unit, a weighted processing unit, and a calculation recognition unit, and identifies the position coordinates and confidence of the small target object, improving the recognition accuracy and efficiency.

[0128] Embodiment 5:

[0129] Figure 6 It is a schematic diagram of the terminal device described in the embodiment of the present application.

[0130] As Figure 6 shown, the embodiment of the present application provides a terminal device, including a processor and a memory;

[0131] The memory is used to store program code and transmit the program code to the processor;

[0132] The processor is used to execute the above-mentioned small target object recognition method based on a dim environment according to the instructions in the program code.

[0133] It should be noted that the processor is used to execute the steps in the above-mentioned embodiment of the small target object recognition method based on a dim environment according to the instructions in the program code. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the above-mentioned system / device embodiments.

[0134] Exemplarily, the computer program can be divided into one or more modules / units. One or more modules / units are stored in the memory and executed by the processor to complete the present application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the terminal device.

[0135] The terminal device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that this does not limit the terminal device, and it may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the terminal device may further include input / output devices, network access devices, a bus, etc.

[0136] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0137] The memory may be an internal storage unit of the terminal device, such as the hard disk or memory of the terminal device. The memory may also be an external storage device of the terminal device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device. Further, the memory may also include both the internal storage unit and the external storage device of the terminal device. The memory is used to store the computer program and other programs and data required by the terminal device. The memory may also be used to temporarily store the data that has been output or will be output.

[0138] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0139] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0140] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0141] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0142] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0143] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. A method for recognizing small target objects based on a dim environment, characterized in that, It includes the following steps: Obtain the template image and the target image of the small target object in a dim environment; Use a multi-scale feature enhancement module to perform image feature information enhancement processing on the template image and the target image respectively, to obtain a template enhanced image corresponding to the template image and a target enhanced image corresponding to the target image; Perform feature extraction on the template enhanced image and the target enhanced image respectively, to obtain template feature maps of different scales corresponding to the template enhanced image and target feature maps of different scales corresponding to the target enhanced image; Calculate the similarity according to the template feature maps and the target feature maps of different scales, to obtain a set of similarity values of different scales; and use the Neck module of Yolov8 in the object detection algorithm based on deep learning to perform feature fusion processing on the target feature maps of different scales, to obtain multi-channel feature maps corresponding to different scales; Perform weighted processing on all the sets of similarity values and all the multi-channel feature maps correspondingly, to obtain final feature maps of different scales; Use the Head decoding module of Yolov8 in the object detection algorithm based on deep learning to perform decoding calculation on all the final feature maps, and identify the position coordinates and confidence of the small target object.

2. The small target object recognition method based on a dim environment according to claim 1, wherein The multi-scale feature enhancement module includes a first branch sub-module and a second branch sub-module connected in parallel, and the first branch sub-module includes a first branch, a second branch, a third branch, a fourth branch and a fifth branch connected in parallel; The second branch and the third branch are used to increase the receptive field of the image and extract image feature information of different scales and shapes; The fourth branch is used to enhance the edge high-frequency information and texture high-frequency information of the image; The fifth branch is used to process the image with a dilation rate of 5; Wherein, the second branch sub-module is used to process the input image through a 1×1 convolutional layer to obtain a first processed image, and the first branch sub-module is used to process the input image through the first branch, the second branch, the third branch, the fourth branch and the fifth branch to obtain a second processed image; the multi-scale feature enhancement module also processes the second processed image through a 1×1 convolutional layer and then merges it with the first processed image to obtain an enhanced image with enhanced image feature information.

3. The small target object recognition method based on a dim environment according to claim 2, characterized in that, The first branch includes a 3×3 ordinary convolutional layer; the second branch includes a 1×3 deformable convolutional layer and a 3×3 atrous convolutional layer with a dilation rate of 2; the third branch includes a 3×1 deformable convolutional layer and a 3×3 atrous convolutional layer with a dilation rate of 2; the fourth branch includes a 1×1 ordinary convolutional layer and a max pooling layer; the fifth branch includes a 3×3 atrous convolutional layer with a dilation rate of 2.

4. The small target object recognition method based on a dim environment according to any one of claims 1-3, characterized in that It includes: Use a cross-stage local network to perform feature extraction on the template enhanced image and the target enhanced image respectively, to obtain template feature maps of different scales corresponding to the template enhanced image and target feature maps of different scales corresponding to the target enhanced image.

5. The small target object recognition method based on a dim environment according to any one of claims 1-3, characterized in that, Calculate the similarity based on the template feature maps and the target feature maps at different scales, and obtain the similarity value sets at different scales, including: Obtain the feature values at each position from the template feature map and the target feature map at the same scale, and obtain the corresponding R template feature values and R target feature values; Calculate the position coefficient according to the template feature value and the target feature value at the same position by using an activation function; Calculate the position distance according to the template feature value, the target feature value and the position coefficient at the same position; and form the similarity value set at the same scale from all the position distances at the same scale.

6. The small target object recognition method based on a dim environment according to any one of claims 1-3, characterized in that, Perform weighted processing on all the similarity value sets and all the multi-channel feature maps correspondingly, and obtain the final feature maps at different scales, including: perform multiplication weighted processing on the position distances in the similarity value set at the same scale and the channel data at the same position in the multi-channel feature map at the corresponding scale to obtain the final feature map at this scale.

7. The small target object recognition method based on a dim environment according to any one of claims 1-3, characterized in that Use the Head decoding module of Yolov8 in the object detection algorithm based on deep learning to perform decoding calculation on all the final feature maps, and identify the position coordinates and confidence of small target objects, including: Use the Head decoding module of Yolov8 to process all the final feature maps to obtain the prediction data output by each prediction box; Calculate according to the prediction data of each prediction box to obtain the coordinate data, confidence and class probability corresponding to each prediction box; Select the prediction box with the largest value from the class probabilities of all the prediction boxes as the recognition prediction box, and use the coordinate data of the recognition prediction box as the position coordinates of the recognized small target object, and use the confidence of the recognition prediction box as the confidence of the recognized small target object.

8. A small target object recognition device based on a dim environment, characterized in that, It includes an image acquisition unit, an image enhancement unit, a feature extraction unit, a calculation and fusion unit, a weighted processing unit and a calculation and recognition unit; The image acquisition unit is used to acquire the template image and the target image of the small target object in a dim environment; The image enhancement unit is used to perform image feature information enhancement processing on the template image and the target image respectively by using a multi-scale feature enhancement module to obtain a template enhanced image corresponding to the template image and a target enhanced image corresponding to the target image; The feature extraction unit is used to perform feature extraction on the template enhanced image and the target enhanced image respectively to obtain template feature maps at different scales corresponding to the template enhanced image and target feature maps at different scales corresponding to the target enhanced image; The calculation and fusion unit is used to calculate the similarity based on the template feature maps and the target feature maps at different scales to obtain the similarity value sets at different scales; and use the Neck module of Yolov8 in the object detection algorithm based on deep learning to perform feature fusion processing on the target feature maps at different scales to obtain multi-channel feature maps corresponding to different scales; The weighted processing unit is configured to perform weighted processing on all the similarity value sets and all the multi-channel feature maps correspondingly to obtain final feature maps of different scales; The calculation and recognition unit is configured to perform decoding calculations on all the final feature maps by using the Head decoding module of Yolov8 in the object detection algorithm based on deep learning, and recognize the position coordinates and confidence levels of small target objects.

9. The small target object recognition device based on a dim environment according to claim 8, characterized in that The multi-scale feature enhancement module includes a first branch sub-module and a second branch sub-module connected in parallel. The first branch sub-module includes a first branch, a second branch, a third branch, a fourth branch, and a fifth branch connected in parallel; The second branch and the third branch are used to increase the receptive field of the image and extract image feature information of different scales and shapes; The fourth branch is used to enhance the edge high-frequency information and texture high-frequency information of the image; The fifth branch is used to process the image with a dilation rate of 5; Among them, the second branch sub-module is configured to process the input image through a 1×1 convolutional layer to obtain a first processed image, and the first branch sub-module is configured to process the input image through the first branch, the second branch, the third branch, the fourth branch, and the fifth branch to obtain a second processed image; the multi-scale feature enhancement module also processes the second processed image through a 1×1 convolutional layer and then merges it with the first processed image to obtain an enhanced image with enhanced image feature information.

10. A terminal device, characterized in that, It includes a processor and a memory; The memory is configured to store program codes and transmit the program codes to the processor; The processor is configured to execute the small target object recognition method based on a dim environment according to any one of claims 1-7 according to the instructions in the program codes.