Target identification method, system and device based on improved YOLOv8 and medium

By improving the backbone and neck network structure of the YOLOv8 model, using the detail processing module and low-frequency enhancement filter to extract and enhance feature information, and perform multi-scale feature fusion, the problem of low target recognition accuracy of YOLOv8 in low brightness environments is solved, and the recognition accuracy is improved.

CN120495613APending Publication Date: 2025-08-15GUANGZHOU INST OF RAILWAY TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510345920.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The YOLOv8 model has low target recognition accuracy in low brightness environments, which is prone to missed detection and misjudgment, mainly due to incomplete feature extraction and noise interference.

Method used

Under low brightness conditions, the YOLOv8 model is improved, and the C3 module in the backbone network is replaced as a detail processing module and a low-frequency enhancement filter, and the downsampled C3 module in the neck network is a Laplace pyramid module. The feature information is extracted and enhanced through the detail processing module and a low-frequency enhancement filter, and the multi-scale feature fusion is used for multi-scale feature fusion.

Benefits of technology

The target recognition capability of the YOLOv8 model in low brightness environments is improved, feature extraction and recognition accuracy is enhanced, noise interference is reduced, and recognition accuracy is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495613A_ABST
    Figure CN120495613A_ABST
Patent Text Reader

Abstract

The invention relates to an improved YOLOv8-based target identification method, system and device and a medium, and the method comprises the steps: improving a YOLOv8 model, which comprises the steps: calling a detail processing module to replace a C3 module in a backbone network and additionally arranging a low-frequency enhancement filter, and calling a Laplacian pyramid module to replace a C3 module belonging to downsampling in a neck network; after the improved model is trained, image features are extracted through a detail processing module and a low-frequency enhancement filter in the model, feature information is enhanced, and feature images with different scale features are obtained; fusing feature information in the feature image through a neck network in the model to obtain an optimal feature image; and inputting the optimal feature image into the head network of the improved model to carry out target identification so as to obtain a target identification result. The technical problem of low target recognition accuracy caused by incomplete feature extraction of the YOLOv8 model in a low-brightness environment is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a target recognition method, system, device and medium based on improved YOLOv8. Background Art

[0002] In the field of artificial intelligence applications, the continuous optimization of target recognition algorithms is of great significance. Since the advent of the YOLO network, target detection algorithms have been widely studied. To date, many enhanced networks such as attention mechanisms, IOU overlap detection, and semantic segmentation networks have been developed from target recognition. Practice has proved that they do have certain recognition capabilities.

[0003] Since its introduction, the YOLO network has garnered widespread attention from deep learning researchers. Early versions of YOLO struggled with locating small and densely packed objects, resulting in inaccurate bounding box positioning and a shallower network structure, which limited its feature extraction capabilities. However, subsequent versions, such as YOLOv3 and YOLOv4, introduced multi-scale prediction to improve small object detection. YOLOv4 also employed the CIoU loss function to improve bounding box positioning accuracy and the CSPDarknet53. These deeper network structures enhanced feature extraction capabilities.

[0004] At present, the most advanced YOLOv8 target recognition network has achieved high accuracy and stability in general environments. However, the network has low target recognition accuracy in low-light environments and is prone to missed detections and misjudgments. This is mainly because the YOLOv8 target recognition network captures more noise in dark conditions, resulting in serious loss of feature extraction details. Summary of the Invention

[0005] (1) Technical issues to be resolved

[0006] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a target recognition method, system, device and medium based on an improved YOLOv8, which solves the technical problem of low image recognition target accuracy caused by incomplete feature extraction or noise interference of the YOLOv8 model in low-light environments.

[0007] (2) Technical solution

[0008] In order to achieve the above objectives, the main technical solutions adopted by the present invention include:

[0009] In a first aspect, an embodiment of the present invention provides a target recognition method based on an improved YOLOv8, comprising:

[0010] When the brightness of the input image is lower than the set brightness threshold, the preset YOLOv8 model is improved, including calling the preset detail processing module to replace the C3 module in the backbone network and adding a low-frequency enhancement filter in the backbone network, and calling the preset Laplacian pyramid module to replace the C3 module belonging to downsampling in the neck network;

[0011] After training the improved YOLOv8 model, the detail processing module is used to extract detail features from the image, and the low-frequency enhancement filter is used to perform multi-scale feature enhancement on the obtained initial feature image to obtain feature images with different scale features.

[0012] The Laplacian pyramid module is used to decompose the feature image into at least two sub-feature images with different resolutions. The convolutional layer of the neck network is used to deconvolve all sub-feature images to the same dimension and then perform tensor splicing to obtain the optimal feature image.

[0013] The optimal feature image is input into the head network of the improved YOLOv8 model for target recognition to obtain the target recognition result.

[0014] Optionally, before improving the preset YOLOv8 model, the following steps are also included:

[0015] Obtaining the image to be identified and the brightness information of the image through a preset shooting device;

[0016] Determine whether the brightness of the image is lower than the set brightness threshold;

[0017] If the brightness of the image is not lower than the set brightness threshold, the image is recognized by the preset YOLOv8 model;

[0018] If the brightness of the image is lower than the set brightness threshold, the preset YOLOv8 model is improved and the improved YOLOv8 model is used to perform target recognition on the image.

[0019] Optionally, after improving the preset YOLOv8 model, the following steps are also included:

[0020] Acquire image data below a brightness threshold;

[0021] Using a preset automatic marking tool, marking target objects in the image data to obtain marked image data;

[0022] The labeled image data is divided into a training set and a validation set in a set ratio, and the improved YOLOv8 model is trained using the training set and the validation set.

[0023] Optionally, extracting detail features from the image using a detail processing module, and performing multi-scale feature enhancement on the obtained initial feature image using a low-frequency enhancement filter to obtain feature images with different scale features includes:

[0024] Utilize the detail processing module to extract detail feature information including context information and edge information from the image to obtain the initial feature image;

[0025] The low-frequency enhancement filter is used to perform multi-scale feature enhancement on the initial feature image to obtain feature images with different scale features.

[0026] Optionally, extracting detail feature information including context information and edge information from the image using a detail processing module to obtain an initial feature image includes:

[0027] The long-distance connection channel or specific network structure configured by the context branch in the detail processing module is used to capture the long-range dependencies between different regions in the image to obtain the context information of the image;

[0028] The Sobel operator in the horizontal and vertical directions configured by the edge branch in the detail processing module is used to calculate the image gradient to obtain the edge information of the image.

[0029] Optionally, performing multi-scale feature enhancement on the initial feature image using a low-frequency enhancement filter to obtain feature images with different scale features includes:

[0030] The initial feature image is convolved through a convolution layer of low-frequency enhancement filters to generate an intermediate feature image with 32 channels;

[0031] The intermediate feature image is downsampled using the average pooling layer of the low-frequency enhancement filter to obtain the feature tensor after dimensionality reduction;

[0032] The feature tensor is divided into four sub-feature tensors along the channel dimension, and each sub-feature tensor is adaptively pooled using pooling windows of different sizes in the average pooling layer to generate the corresponding multi-scale feature tensor;

[0033] The multi-scale feature tensors are concatenated in the channel dimension, and the spatial dimension of the concatenated feature map is restored to the original size through deconvolution or upsampling;

[0034] The convolution operation is performed on the restored feature image to obtain feature images with different scale features.

[0035] Optionally, the Laplacian pyramid module is used to perform Gaussian blurring and downsampling on the input feature image according to a set sampling rule to obtain a target number of sub-feature images of different resolutions;

[0036] The downsampling rules include:

[0037] Each downsampled feature image is subjected to Gaussian blurring;

[0038] The pixels of the feature image processed by Gaussian blur are sampled at intervals so that the resolution of the output feature image is half of that of the input feature image;

[0039] The number of downsampling is consistent with the target number of sub-feature images.

[0040] In a second aspect, an embodiment of the present invention provides a target recognition system based on an improved YOLOv8, comprising:

[0041] The YOLOv8 model improvement module is used to improve the preset YOLOv8 model when the brightness of the input image is lower than the set brightness threshold. This includes calling the preset detail processing module to replace the C3 module in the backbone network, adding a low-frequency enhancement filter to the backbone network, and calling the preset Laplacian pyramid module to replace the C3 module in the neck network that belongs to downsampling;

[0042] A feature image extraction module is used to extract detail features from the image using the detail processing module after training the improved YOLOv8 model, and perform multi-scale feature enhancement on the obtained initial feature image using a low-frequency enhancement filter to obtain feature images with different scale features;

[0043] The feature fusion module is used to decompose the feature image into at least two sub-feature images with different resolutions using the Laplacian pyramid module, and deconvolute all sub-feature images to the same dimension using the convolution layer of the neck network, and then perform tensor splicing to obtain the optimal feature image;

[0044] The target recognition module is used to input the optimal feature image into the head network of the improved YOLOv8 model for target recognition and obtain the target recognition result.

[0045] In a second aspect, an embodiment of the present invention provides an electronic device, including:

[0046] processor;

[0047] A memory storing the target recognition method steps based on the improved YOLOv8 for controlling the processor.

[0048] In a second aspect, an embodiment of the present invention provides a computer-readable medium having computer-executable instructions stored thereon. When the executable instructions are executed by a processor, the above-mentioned target recognition method steps based on the improved YOLOv8 are implemented.

[0049] (3) Beneficial effects

[0050] The beneficial effects of the present invention are as follows: by replacing the C3 module in the backbone network of the YOLOv8 model with a DPM module, adding a low-frequency enhancement filter to the backbone network, and replacing the down-sampling C3 module in the neck network of the YOLOv8 model with a Laplacian pyramid module, the improved backbone network of the present invention can extract more detailed feature information and perform low-frequency enhancement on the extracted feature information. At the same time, the Laplacian pyramid module in the improved neck network can finely decompose the feature image to enhance and fuse the feature information at multiple scales, thereby improving the image target recognition capability of the YOLOv8 model. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 A flowchart of a target recognition method based on an improved YOLOv8 is provided in one embodiment of the present invention;

[0052] Figure 2 A framework diagram of an improved YOLOv8 model provided by one embodiment of the present invention;

[0053] Figure 3 A framework diagram of a detail processing module provided in one embodiment of the present invention;

[0054] Figure 4 A framework diagram of a low-frequency enhancement filter provided by one embodiment of the present invention;

[0055] Figure 5 A schematic diagram of an object recognition system based on improved YOLOv8 provided by one embodiment of the present invention;

[0056] Figure 6 A schematic diagram of the structure of a computer system of an electronic device provided by one embodiment of the present invention.

[0057] [Description of Reference Numerals]

[0058] 100: Target recognition system; 101: YOLOv8 model improvement module; 102: Feature image extraction module; 103: Feature fusion module; 104: Target recognition module;

[0059] 200: Computer system; 201: CPU; 202: ROM; 203: RAM; 204: Bus; 205: I / O interface; 206: Input part; 207: Output part; 208: Storage part; 209: Communication part; 210: Drive; 211: Removable medium. DETAILED DESCRIPTION

[0060] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation methods in conjunction with the accompanying drawings.

[0061] refer to Figure 1-4 As shown, an embodiment of the present invention proposes an improved YOLOv8-based target recognition method, which includes: when the brightness of an input image is lower than a set brightness threshold, improving a preset YOLOv8 model, including calling a preset detail processing module to replace the C3 module in the backbone network and adding a low-frequency enhancement filter in the backbone network, and calling a preset Laplacian pyramid module to replace the C3 module belonging to downsampling in the neck network; after training the improved YOLOv8 model, using the detail processing module to extract detail features in the image, and using the low-frequency enhancement filter to perform multi-scale feature enhancement on the obtained initial feature image to obtain feature images with different scale features; using the Laplacian pyramid module to decompose the feature image into at least two sub-feature images of different resolutions, and using the convolution layer of the neck network to deconvolve all the sub-feature images to the same dimension and then perform tensor splicing to obtain an optimal feature image; inputting the optimal feature image into the head network of the improved YOLOv8 model for target recognition to obtain a target recognition result.

[0062] This embodiment uses a DPM module to replace the C3 module in the backbone network of the YOLOv8 model, adds a low-frequency enhancement filter to the backbone network, and uses a Laplacian pyramid module to replace the downsampling C3 module in the neck network of the YOLOv8 model. Compared with the existing technology, the improved backbone network of this embodiment can extract more detailed feature information and perform low-frequency enhancement on the extracted feature information. At the same time, the improved Laplacian pyramid module in the neck network performs fine decomposition of the feature image, allowing for multi-scale enhancement and fusion of feature information, thereby improving the YOLOv8 model's ability to recognize image objects.

[0063] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0064] Specifically, refer to Figure 1 As shown, this embodiment proposes a target recognition method based on improved YOLOv8, which includes:

[0065] S100, when the brightness of the input image is lower than the set brightness threshold, the preset YOLOv8 model is improved, including calling the preset detail processing module to replace the C3 module in the backbone network and adding a low-frequency enhancement filter in the backbone network, and calling the preset Laplacian pyramid module to replace the C3 module belonging to the downsampling in the neck network. The framework of the improved YOLOv8 model is as follows Figure 2 shown.

[0066] In this embodiment, Figure 3 The detail processing module (DPM module) shown replaces the C3 module (CSP Bottleneck with 3 convolutions) of the backbone network in the original YOLOv8 model. The detail processing module extracts features from the input image instead of the original C3 module. The detail processing module enhances the YOLOv8 model's ability to capture fine-grained features through multi-scale feature fusion and extraction and efficient feature information flow, thereby improving the accuracy and robustness of target recognition.

[0067] In this embodiment, Figure 4 The low-frequency enhancement filter (LEF module) shown in the figure is added to the backbone network of the original YOLOv8 model. The low-frequency enhancement filter captures and enhances the low-frequency information in the image by filtering out high-frequency noise. It uses adaptive average pooling operations of different sizes to intercept the semantic information of low-frequency components to dynamically adapt to low-frequency information of different scales and semantics to ensure that the key details in the image are captured to the greatest extent possible, significantly improving the target recognition ability of the YOLOv8 model in low-light environments.

[0068] In this embodiment, the C3 module belonging to downsampling in the backbone network of the original YOLOv8 model is replaced by a Laplacian pyramid module. The Laplacian pyramid module can decompose the input feature image into sub-feature images of multiple resolutions. These sub-feature images contain feature information of different scales and can be further used for feature fusion. At the same time, they can also enhance the detail information of the feature image, thereby further enhancing the target recognition capability of the YOLOv8 model in low-light environments.

[0069] Furthermore, before step S100, the following steps are also included:

[0070] A100: Obtain an image to be recognized and brightness information of the image through a preset shooting device.

[0071] For example, high-definition cameras or other cameras with 360p or higher image quality installed on roads play a vital role in maintaining traffic order, ensuring public safety, and improving urban management efficiency. When capturing images and videos, these cameras simultaneously collect brightness information of the captured images or videos.

[0072] A200: Determine whether the brightness of the image is lower than a set brightness threshold.

[0073] In this embodiment, the brightness threshold is set differently depending on the application scenario. For example, the brightness threshold for outdoor images is set to 150 lx, and environments below 150 lx are considered low brightness environments; the brightness threshold for indoor images is set to 100 lx, and environments below 100 lx are considered low brightness environments.

[0074] A300a. If the brightness of the image is not lower than the set brightness threshold, the image is subjected to target recognition using the preset YOLOv8 model.

[0075] A300b. If the brightness of the image is lower than the set brightness threshold, the preset YOLOv8 model is improved, and the improved YOLOv8 model is used to perform target recognition on the image.

[0076] In this embodiment, the basic YOLOv8 model is used to identify targets in images taken under normal conditions, and the improved YOLOv8 model is used to enhance the target recognition capability of the YOLOv8 model, thereby improving the accuracy of target device recognition.

[0077] Furthermore, after step S100, the following steps are further included:

[0078] F100, obtain image data below a brightness threshold.

[0079] In this embodiment, the amount of image data used to train the improved YOLOv8 model ranges from several hundred to tens of thousands, depending on the complexity of the application environment. For example, when training for object recognition on suburban roads, a few hundred images may be sufficient; however, when training for object recognition on urban roads, thousands or even tens of thousands of images may be required to ensure effective training of the improved YOLOv8 model.

[0080] F200. Mark the target object in the image data using a preset automatic marking tool to obtain marked image data.

[0081] In this embodiment, the collected image data is imported into the T-REX website (an open source image target annotation website) to mark the target objects in the image (such as people, cars, etc.). Specifically, the target label is first defined, and then the annotation tool provided by the T-REX website is used to mark the image area containing the target with a rectangular box. The target category information and the normalized bounding box coordinates are then packaged and exported into an annotation file. Finally, the image file and the annotation file are organized into labeled image data in a set directory structure.

[0082] F300, divide the labeled image data into a training set and a validation set in a set ratio, and use the training set and the validation set to train the improved YOLOv8 model.

[0083] S200. After training the improved YOLOv8 model, extract detail features from the image using a detail processing module, and perform multi-scale feature enhancement on the obtained initial feature image using a low-frequency enhancement filter to obtain feature images with different scale features.

[0084] In this embodiment, the backbone network in the improved YOLOv8 model uses a detail processing module to extract fine-grained features and uses a low-frequency enhancement filter to perform low-frequency enhancement on the initial feature image output by the detail processing module. The detail processing module can obtain more feature information for texture feature extraction performance in low-brightness environments and dependency reasoning on global components. After enhancing the low-frequency information in the feature information, the feature information is easier to identify, thereby improving the target recognition ability of the model.

[0085] In this embodiment, step S200 may include sub-steps S210 to S220:

[0086] S210 , using a detail processing module to extract detail feature information including context information and edge information from the image to obtain an initial feature image.

[0087] Furthermore, step S210 may include sub-steps S211 to S212:

[0088] S211. Utilize the long-distance connection channel or specific network structure configured by the context branch in the detail processing module to capture the long-range dependency between different regions in the image to obtain the context information of the image.

[0089] In this embodiment, the long-distance connection channel or specific network structure configured by the context branch captures the remote dependency between different regions in the image to obtain the context information of the image, and performs global enhancement on the different category regions in the image. This includes: context information integration, combining the global context information in the image with local features to help the YOLOv8 model better understand the role of each feature in the overall image; feature representation enhancement processing, by increasing the number of channels of the feature map, the YOLOv8 model can learn more complex feature representations, thereby enhancing the description ability of components; spatial relationship establishment, establishing the spatial layout relationship (including position, size, direction, etc.) between cross-regions in the image through global information enhancement; category information perception enhancement, in the classification task, by enhancing the category-sensitive features in the image to increase the distinctiveness of different category features; context consistency constraint, used to constrain the target prediction results in the image to be semantically consistent with the global context information, thereby reducing misclassification and inconsistent predictions.

[0090] In one specific embodiment, residual blocks are used to process features before and after acquiring long-range dependencies. This includes: enhancing feature representations through residual block processing, enabling the network structure to better capture and utilize long-range dependencies; transmitting low-frequency information, such as context and global structure, directly through skip connections, ensuring that low-frequency information can be passed directly from input to output without passing through all convolutional and nonlinear activation layers, thereby reducing information loss; and adjusting feature dimensions, by changing the number of channels in the feature map, enabling the network structure to more efficiently process and transmit information. Residual learning using residual blocks allows for the transmission of rich low-frequency information through skip connections, which helps alleviate the vanishing gradient problem and allows the network structure to learn at a deeper level. The specific steps for using residual blocks are as follows: first, the input feature image is first passed through one or more convolutional layers to modify the feature image's dimensions; then, the feature image is directly passed to the output of the convolutional layer through an identity mapping (i.e., without any modification); finally, the output of the convolutional layer is added to the output of the identity mapping to obtain the final output feature image of the residual block.

[0091] For example, when two residual blocks are configured, the first residual block changes the channels of the feature image from 3 to 32 to increase the richness of the feature representation, allowing the network structure to learn more complex patterns; the second residual block changes the channels of the feature image from 32 to 3, which helps to reduce the dimension of the feature image, focus on the most relevant information, and use the features for subsequent processing or output. The expression for global enhancement of low-brightness images is:

[0092]

[0093] In formula (1), CB(x) represents the feature image enhanced by the context branch, x represents the input feature image, and γ represents a linear rectification function with leakage, which allows small gradients even in negative inputs. Represents the probability distribution after the input feature image is normalized, F1 represents the first convolution layer that operates on the input feature image, F2 represents the second convolution layer that operates on the input feature image, and ω represents the normalized exponential function, which is used to convert the input vector into a probability distribution.

[0094] S212 , calculating the image gradient using the Sobel operator in the horizontal and vertical directions configured by the edge branch in the detail processing module to obtain edge information of the image.

[0095] In this embodiment, the edge branch uses the Sobel operator in two different directions (horizontally and vertically) to calculate the gradient of the input image to obtain the edge information of the image to enhance the texture of the feature component. The Sobel operator is a discrete operator that uses Gaussian filtering and differential derivation at the same time, and can find edge information by calculating gradient approximation. The Sobel operator is used in the horizontal and vertical directions to re-extract edge information through convolution filters, and the residual is used to enhance the information flow. The mathematical expression of this process is:

[0096] CE(x)=F3(Sobelh(x)+Sobelw(x))+x (2)

[0097] In formula (1), CB(x) represents the feature image enhanced by the edge branch, F3 represents the third convolution layer that operates on the input feature image, and Sobel h (x) represents the gradient result obtained by using the horizontal Sobel operator on the feature image x. w (x) represents the gradient result obtained by applying the Sobel operator in the vertical direction to the feature image x.

[0098] S220 , performing multi-scale feature enhancement on the initial feature image using a low-frequency enhancement filter to obtain feature images with features at different scales.

[0099] Furthermore, step S220 may include sub-steps S221 to S225:

[0100] S221. Perform a convolution operation on the initial feature image through a convolution layer of a low-frequency enhancement filter to generate an intermediate feature image with 32 channels.

[0101] The initial feature image is convolved through the convolution layer of the low-frequency enhancement filter to capture most of the semantic information in the image, especially the low-frequency information, and output an intermediate feature image f with 32 channels, whose spatial dimension is h×w×32

[0102] S222. Downsample the intermediate feature image using the average pooling layer of the low-frequency enhancement filter to obtain a feature tensor after dimensionality reduction.

[0103] S223. Split the feature tensor into four sub-feature tensors along the channel dimension, and adaptively pool each sub-feature tensor using pooling windows of different sizes in the average pooling layer to generate a corresponding multi-scale feature tensor.

[0104] The feature tensor is split into four sub-feature tensors along the channel dimension, and each sub-feature tensor is adaptively pooled using pooling windows of different sizes to generate the corresponding multi-scale feature tensors f1, f2, f3, and f4. f1 is pooled using a 2x2 pooling window; f2 is pooled using a 3x3 pooling window; f3 is pooled using a 4x4 pooling window; and f4 is pooled using a 5x5 pooling window.

[0105] S224. Concatenate the multi-scale feature tensors in the channel dimension, and restore the spatial dimension of the concatenated feature map to the original size through deconvolution or upsampling.

[0106] The multi-scale feature tensors f1 to f4 are concatenated in the channel dimension, and the spatial dimension of the concatenated feature map is restored to the original size Rh×w through deconvolution or interpolation upsampling operations.

[0107] S225 . Perform a convolution operation on the restored feature image to obtain feature images with different scale features.

[0108] A 1×1 convolution operation is performed on the restored feature image to compress its channel number to 3, and the final low-frequency enhanced feature map f∈Rh×w×3 is output, where the low-frequency enhanced feature map retains the low-frequency components of the input image and suppresses high-frequency noise.

[0109] S300, using the Laplacian pyramid module to decompose the feature image into at least two sub-feature images with different resolutions, and using the convolution layer of the neck network to deconvolve all the sub-feature images to the same dimension and then perform tensor splicing to obtain the optimal feature image.

[0110] In this embodiment, the Laplacian pyramid module is used to perform Gaussian blurring and downsampling on the input feature image according to a set sampling rule to obtain a target number of sub-feature images of different resolutions.

[0111] Gaussian blur processing performs Gaussian blur processing on the image of the current layer and reduces high-frequency information. The kernel of Gaussian blur usually uses a 5×5 Gaussian filter. The calculation formula of the Gaussian kernel is:

[0112]

[0113] In formula (3), (x, y) represents the pixel position in the kernel, and σ represents the standard deviation, which is used to control the degree of blur.

[0114] The downsampling rules are as follows: each feature image that performs downsampling is Gaussian blurred; the pixels of the feature image that has been Gaussian blurred are sampled at intervals so that the resolution of the output feature image is half of that of the input feature image; the number of downsampling times is consistent with the target number of sub-feature images.

[0115] In a specific embodiment, when the feature image input to the Laplacian pyramid module is f∈Rh×w×3, four sub-feature images of different resolutions G1∈R·h / 2×w / 2×3, G2∈R·h / 4×w / 4×3, G3∈R·h / 8×w / 8×3, and G4∈R·h / 16×w / 16×3 are generated by downsampling.

[0116] S400: Input the optimal feature image into the head network of the improved YOLOv8 model to perform target recognition and obtain a target recognition result.

[0117] In this embodiment, the head network first adjusts the number of channels through 3×3 convolution, and then outputs the target category confidence through the classification branch, and the regression branch outputs the bounding box coordinates (center point, width and height) and existence probability. Next, the optimal feature image is matched through the preset Anchor template to generate an initial prediction box, and the output dimension of each detection head is H×W×3. Then, the normalized coordinates are converted to actual pixel positions through the decoder, and the confidence threshold is applied to screen the preliminary candidate boxes. Finally, non-maximum suppression (NMS) is performed, and redundant detection boxes are removed through the intersection-over-union (IoU) threshold, and the final target category, confidence and positioning information are output to achieve target recognition and detection.

[0118] also, Figure 5 A schematic diagram of the composition of a target recognition system based on improved YOLOv8 provided by the present invention is shown as follows: Figure 5 As shown, the present invention also provides a target recognition system 100 based on improved YOLOv8, which includes:

[0119] The YOLOv8 model improvement module 101 is used to improve the preset YOLOv8 model when the brightness of the input image is lower than the set brightness threshold, including calling the preset detail processing module to replace the C3 module in the backbone network and adding a low-frequency enhancement filter in the backbone network, and calling the preset Laplacian pyramid module to replace the C3 module belonging to downsampling in the neck network.

[0120] The feature image extraction module 102 is used to extract detail features from the image using the detail processing module after training the improved YOLOv8 model, and perform multi-scale feature enhancement on the obtained initial feature image using a low-frequency enhancement filter to obtain feature images with different scale features.

[0121] The feature fusion module 103 is used to decompose the feature image into at least two sub-feature images with different resolutions using the Laplacian pyramid module, and deconvolute all the sub-feature images to the same dimension using the convolution layer of the neck network, and then perform tensor splicing to obtain the optimal feature image.

[0122] The target recognition module 104 is used to input the optimal feature image into the head network of the improved YOLOv8 model to perform target recognition and obtain a target recognition result.

[0123] The functions of each module in the system are described in detail in the above method embodiment and will not be repeated here.

[0124] On the other hand, the present invention further provides an electronic device including a processor and a memory, wherein the memory stores steps for the processor to control the above-mentioned robot path planning method based on the improved fireworks algorithm.

[0125] Reference below Figure 6 , which shows a structural diagram of a computer system 200 suitable for implementing the robot path planning electronic device of an embodiment of the present application. Figure 6 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0126] like Figure 6 As shown, the computer system 200 includes a central processing unit (CPU) 201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 202 or a program loaded from a storage portion 208 into a random access memory (RAM) 203. Various programs and data required for the operation of the computer system 200 are also stored in the RAM 203. The CPU 201, the ROM 202, and the RAM 203 are connected to each other via a bus 204. An input / output (I / O) interface 205 is also connected to the bus 204.

[0127] The following components are connected to the I / O interface 205: an input section 206 including a keyboard, a mouse, and the like; an output section 207 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 208 including a hard disk; and a communication section 209 including a network interface card such as a LAN card or a modem. The communication section 209 performs communication processing via a network such as the Internet. A drive 210 is also connected to the I / O interface 205 as needed. A removable medium 211, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 210 as needed, so that computer programs read therefrom can be installed into the storage section 208 as needed.

[0128] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 209 and / or installed from the removable medium 211. When the computer program is executed by the central processing unit (CPU) 201, the above-mentioned functions defined in the system of the present application are performed.

[0129] It should be noted that the computer-readable medium shown in the embodiment of the present application may be a computer-readable signal medium or a computer-readable medium or any combination of the above two. The computer-readable medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiment of the present application, the computer-readable medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the embodiment of the present application, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.

[0130] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0131] The units involved in the embodiments described in this application may be implemented by software or hardware. The units described may also be provided in a processor, wherein the names of these units do not, in certain circumstances, constitute limitations on the units themselves.

[0132] In another aspect, an embodiment of the present invention further provides a computer-readable medium, which may be included in the electronic device described in the above embodiments, or may exist independently and not be incorporated into the electronic device. The computer-readable medium stores one or more computer-executable instructions. When the one or more computer-executable instructions are executed by a processor, the processor implements the following method steps:

[0133] S100. After training the improved YOLOv8 model, use the detail processing module to extract detail features in the image, and use the low-frequency enhancement filter to perform multi-scale feature enhancement on the obtained initial feature image to obtain feature images with different scale features.

[0134] S200, using the Laplacian pyramid module to decompose the feature image into at least two sub-feature images with different resolutions, and using the convolution layer of the neck network to deconvolve all the sub-feature images to the same dimension and then perform tensor splicing to obtain the optimal feature image.

[0135] S300: Input the optimal feature image into the head network of the improved YOLOv8 model for target recognition to obtain a target recognition result.

[0136] S400, merging the Gaussian mutation spark and the Levy mutation spark into the set of path solutions to be iterated, and after reaching the maximum number of iterations, selecting the optimal path solution as output based on the fireworks population fitness of each path solution in the iterated path solution set.

[0137] In summary, the target recognition method, system, device and medium based on the improved YOLOv8 proposed in this embodiment focus on the target recognition aspect in computer vision, and provide a solution for improving the accuracy of image recognition in low-light environments. Compared with the traditional target recognition network, which loses feature details in low-light environments, the improved YOLOV8 model has better detection performance in low-light environments, and compared with other large image detection models, the model has the characteristics of being lightweight, which has made certain contributions to the edge computing of mobile terminals. For example, in the field of autonomous driving, when facing intersection judgments, the improved YOLOV8 model can well identify obstacles at night and give early warnings, and the improved YOLOV8 model also makes it easier for traffic police to find vehicles that cause accidents at night.

[0138] Since the systems / devices described in the above embodiments of the present invention are systems / devices used to implement the methods of the above embodiments of the present invention, those skilled in the art will be able to understand the specific structures and variations of these systems / devices based on the methods described in the above embodiments of the present invention, and thus will not be described in detail here. All systems / devices used in the methods of the above embodiments of the present invention are within the scope of protection of the present invention.

[0139] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0140] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions.

[0141] It should be noted that, in the description of the present invention, the word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The present invention can be implemented by means of hardware comprising several distinct components and by means of a suitably programmed computer. The use of the words first, second, third, etc., is merely for convenience and does not imply any order. These words should be understood as part of the component name.

[0142] In addition, it should be noted that, in the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.

[0143] Although preferred embodiments of the present invention have been described, those skilled in the art will be able to make additional changes and modifications to these embodiments after obtaining the basic inventive concepts.

[0144] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the invention.

Claims

1. A target recognition method based on improved YOLOv8, characterized in that: include: When the brightness of the input image is lower than the set brightness threshold, the preset YOLOv8 model is improved, including calling the preset detail processing module to replace the C3 module in the backbone network and adding a low-frequency enhancement filter in the backbone network, and calling the preset Laplacian pyramid module to replace the C3 module belonging to downsampling in the neck network; After training the improved YOLOv8 model, the detail processing module is used to extract detail features from the image, and the low-frequency enhancement filter is used to perform multi-scale feature enhancement on the obtained initial feature image to obtain feature images with different scale features. The Laplacian pyramid module is used to decompose the feature image into at least two sub-feature images with different resolutions. The convolutional layer of the neck network is used to deconvolve all sub-feature images to the same dimension and then perform tensor splicing to obtain the optimal feature image. The optimal feature image is input into the head network of the improved YOLOv8 model for target recognition to obtain the target recognition result.

2. The method according to claim 1, wherein Before improving the default YOLOv8 model, it also includes: Obtaining the image to be identified and the brightness information of the image through a preset shooting device; Determine whether the brightness of the image is lower than the set brightness threshold; If the brightness of the image is not lower than the set brightness threshold, the image is recognized by the preset YOLOv8 model; If the brightness of the image is lower than the set brightness threshold, the preset YOLOv8 model is improved and the improved YOLOv8 model is used to perform target recognition on the image.

3. The method according to claim 1, wherein: After improving the preset YOLOv8 model, it also includes: Acquire image data below a brightness threshold; Using a preset automatic marking tool, marking target objects in the image data to obtain marked image data; The labeled image data is divided into a training set and a validation set in a set ratio, and the improved YOLOv8 model is trained using the training set and the validation set.

4. The method according to claim 1, wherein: The detail processing module is used to extract the detail features in the image, and the low-frequency enhancement filter is used to perform multi-scale feature enhancement on the obtained initial feature image to obtain feature images with different scale features, including: Utilize the detail processing module to extract detail feature information including context information and edge information from the image to obtain the initial feature image; The low-frequency enhancement filter is used to perform multi-scale feature enhancement on the initial feature image to obtain feature images with different scale features.

5. The method according to claim 4, wherein: The detail processing module is used to extract the detail feature information of the image containing context information and edge information, and the initial feature image is obtained including: The long-distance connection channel or specific network structure configured by the context branch in the detail processing module is used to capture the long-range dependencies between different regions in the image to obtain the context information of the image; The Sobel operator in the horizontal and vertical directions configured by the edge branch in the detail processing module is used to calculate the image gradient to obtain the edge information of the image.

6. The method according to claim 4, wherein: The low-frequency enhancement filter is used to perform multi-scale feature enhancement on the initial feature image to obtain feature images with different scale features, including: The initial feature image is convolved through a convolution layer of low-frequency enhancement filters to generate an intermediate feature image with 32 channels; The intermediate feature image is downsampled using the average pooling layer of the low-frequency enhancement filter to obtain the feature tensor after dimensionality reduction; The feature tensor is divided into four sub-feature tensors along the channel dimension, and each sub-feature tensor is adaptively pooled using pooling windows of different sizes in the average pooling layer to generate the corresponding multi-scale feature tensor; The multi-scale feature tensors are concatenated in the channel dimension, and the spatial dimension of the concatenated feature map is restored to the original size through deconvolution or upsampling; The convolution operation is performed on the restored feature image to obtain feature images with different scale features.

7. The method according to claim 1, wherein: The Laplacian pyramid module is used to perform Gaussian blurring and downsampling on the input feature image according to the set sampling rules to obtain the target number of sub-feature images of different resolutions; The downsampling rules include: Each downsampled feature image is subjected to Gaussian blurring; The pixels of the feature image processed by Gaussian blur are sampled at intervals so that the resolution of the output feature image is half of that of the input feature image; The number of downsampling is consistent with the target number of sub-feature images.

8. An object recognition system based on improved YOLOv8, characterized in that: include: The YOLOv8 model improvement module is used to improve the preset YOLOv8 model when the brightness of the input image is lower than the set brightness threshold. This includes calling the preset detail processing module to replace the C3 module in the backbone network, adding a low-frequency enhancement filter to the backbone network, and calling the preset Laplacian pyramid module to replace the C3 module in the neck network that belongs to downsampling; A feature image extraction module is used to extract detail features from the image using the detail processing module after training the improved YOLOv8 model, and perform multi-scale feature enhancement on the obtained initial feature image using a low-frequency enhancement filter to obtain feature images with different scale features; The feature fusion module is used to decompose the feature image into at least two sub-feature images with different resolutions using the Laplacian pyramid module, and deconvolute all sub-feature images to the same dimension using the convolution layer of the neck network, and then perform tensor splicing to obtain the optimal feature image; The target recognition module is used to input the optimal feature image into the head network of the improved YOLOv8 model for target recognition and obtain the target recognition result.

9. An electronic device, characterized in that: include: processor; A memory storing a target recognition method based on improved YOLOv8 for controlling the processor according to any one of claims 1 to 7.

10. A computer-readable medium having computer-executable instructions stored thereon, characterized in that: When the executable instructions are executed by the processor, the target recognition method steps based on the improved YOLOv8 are implemented as described in any one of claims 1 to 7.