Target identification method based on improved YOLOv8 model and application thereof

By adding C2f_SENet and Dyhead_DCNv3 modules to the YOLOv8 model, the identification accuracy and efficiency of the YOLOv8 model in complex lighting and occlusion are solved, and the recognition accuracy and efficiency of the electric vehicle is achieved.

CN120014407APending Publication Date: 2025-05-16WUXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510010746.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing YOLOv8 model reduces the accuracy and efficiency of electric vehicles when the light intensity is too high or too low and the electric vehicle is blocked.

Method used

Based on the original YOLOv8 model, the C2f_SENet module is added to enhance feature expression capabilities, and the Dyhead_DCNv3 module is added to improve multi-scale object detection capabilities to build an improved YOLOv8 model.

Benefits of technology

Through the improved YOLOv8 model, the accuracy and efficiency of target recognition are improved, especially in complex situations (such as light changes and occlusion) to identify electric vehicles more accurately and efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014407A_ABST
    Figure CN120014407A_ABST
Patent Text Reader

Abstract

The invention provides a target recognition method based on an improved YOLOv8 model, and relates to the technical field of image recognition. Firstly, a to-be-recognized image is acquired and preprocessed; then, an improved YOLOv8 model is constructed; on the basis of an original YOLOv8 model, a C2fSENet module used for enhancing the feature expression capability and a DyheadDCNv3 module used for improving the multi-scale target detection capability are added to the improved YOLOv8 model. And the improved YOLOv8 model is trained, and the trained improved YOLOv8 model is obtained. And finally, inputting the preprocessed image into the trained improved YOLOv8 model to obtain a target identification result. According to the method, the accuracy and efficiency of target recognition of the existing YOLOv8 model are improved, and the accuracy and efficiency of recognition of the electric vehicle in the elevator are improved when the method is applied to the complex conditions that the illumination intensity is too high or too low, and the electric vehicle is shielded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition, and more specifically relates to a target recognition method based on an improved YOLOv8 model and an application thereof. Background Art

[0002] As one of the fastest growing branches in the field of image processing, object recognition has realized recognition and processing tasks in a variety of states. With the rapid development of deep learning technology, the accuracy and efficiency of object detection have been significantly improved, enabling it to achieve efficient and accurate recognition in various complex scenarios.

[0003] The development of target recognition technology is based on existing algorithms and models. For example, convolutional neural network (CNN), as a deep learning model, extracts image features through multi-layer convolution and pooling operations, and uses these features for classification and target recognition. There are other algorithms and models such as support vector machine (SVM), deep belief network (DBN) and YOLO model, which also play an important role in target recognition.

[0004] In the specific application of target recognition, the recognition of electric vehicles entering the elevator is one of the important applications. Electric vehicles entering the elevator poses a major safety hazard because the battery of the electric vehicle may spontaneously ignite, causing fire and explosion. The fire and toxic gases spread rapidly in the confined space of the elevator, endangering people's lives. Therefore, it is urgent to accurately identify electric vehicles entering the elevator and take timely measures to prevent electric vehicles from entering the elevator.

[0005] At present, the detection of electric vehicles entering elevators has been proposed based on CNN, RCNN and other models, but these models are large in scale and slow in operation. The YOLO model has higher real-time performance due to its single-stage characteristics. However, the existing YOLOv8 model still faces defects such as large amount of calculation and parameters, low recognition accuracy, and difficulty in identifying small targets. If it is used as a means of target recognition for electric vehicles in elevators, the accuracy and efficiency of electric vehicle recognition will be reduced under complex conditions such as too high or too low light intensity and electric vehicles being blocked. Summary of the invention

[0006] In order to solve the problem that the accuracy and efficiency of electric vehicle recognition are reduced in the existing YOLOv8 model under complex conditions such as too high or too low light intensity and electric vehicles being blocked, the present invention proposes a target recognition method based on an improved YOLOv8 model to improve the accuracy and efficiency of target recognition.

[0007] In order to achieve the above technical effects, the technical solution of the present invention is as follows:

[0008] S1: Obtain the image to be recognized and perform preprocessing;

[0009] S2: construct an improved YOLOv8 model; based on the original YOLOv8 model, the improved YOLOv8 model adds a C2f_SENet module for enhancing feature expression capabilities and a Dyhead_DCNv3 module for improving multi-scale target detection capabilities;

[0010] S3: Train the improved YOLOv8 model to obtain a trained improved YOLOv8 model;

[0011] S4: Input the image preprocessed in step S1 into the trained improved YOLOv8 model to obtain the target recognition result.

[0012] Furthermore, the preprocessing operation is: removing duplicate images and invalid images that do not contain recognition targets from the images to be recognized, and performing exposure processing on the remaining images.

[0013] Furthermore, the improved YOLOv8 model consists of three parts: Backbone network, Neck network and Head network;

[0014] The Backbone network includes: a first Conv module, a second Conv module, a first C2f_SENet module, a third Conv module, a second C2f_SENet module, a fourth Conv module, a third C2f_SENet module, a fifth Conv module, a fourth C2f_SENet module and a first SPPF module connected in sequence.

[0015] Further, the Neck network includes: a sixth Conv module, a first Upsample module, a first Concat module, a first C2f module, a seventh Conv module, a second Upsample module, a second Concat module, a second C2f module, an eighth Conv module, a third Concat module, a third C2f module, a ninth Conv module, a fourth Concat module and a fourth C2f module connected in sequence; and also includes a tenth Conv module and an eleventh Conv module; the output end of the third Conv module is connected to the input end of the tenth Conv module, the output end of the tenth Conv module is connected to the second Concat module, the output end of the fourth Conv module is connected to the input end of the eleventh Conv module, the output end of the eleventh Conv module is connected to the input end of the first Concat module, the output end of the first SPPF module is connected to the input end of the sixth Conv module, the output end of the sixth Conv module is connected to the input end of the fourth Concat module, and the output end of the seventh Conv module is connected to the input end of the third Concat module.

[0016] Furthermore, the Head network includes: a first Dyhead_DCNv3 module, a second Dyhead_DCNv3 module and a third Dyhead_DCNv3 module; any one of the first Dyhead_DCNv3 module, the second Dyhead_DCNv3 module and the third Dyhead_DCNv3 module includes: a perception scale attention layer, a spatial perception attention layer and a task perception attention layer connected in sequence.

[0017] Furthermore, any one of the first C2f_SENet module, the second C2f_SENet module, the third C2f_SENet module and the fourth C2f_SENet module includes: a first Residual layer, a first Global Avg Pool layer, a first Fully Connected layer, a first Non-li near layer, a second Fully Connected layer, a first Sigmoid layer and a second Residual layer connected in sequence; the output end of the first Residual layer is also connected to the input end of the second Residual layer.

[0018] Based on the above technical means, C2f_SENet introduces squeezing and excitation operations to enhance feature expression capabilities and increase inter-channel dependencies, enabling the model to capture key information more effectively. It optimizes target recognition accuracy and improves the generalization ability of the model in different scenarios, thereby achieving better performance in complex tasks and improving the accuracy of electric vehicle recognition in complex situations.

[0019] Further, the perceptual scale attention layer includes: a first Avg Pool operation, a first DCNv3 operation, a first ReLU operation, and a first Hard sigmoid operation connected in sequence;

[0020] The spatial perception attention layer includes: a first Index operation, a second DCNv3 operation, a first Sigmoid operation and a first Offset operation; the output end of the first Index operation is connected to the input end of the second DCNv3 operation, and the output end of the second DCNv3 operation is respectively connected to the input ends of the first Sigmoid operation and the first Offset operation;

[0021] The task-aware attention layer includes: a second Avg Pool operation, a first Full Connected operation, a first RELU operation, a second Full Connected operation and a first Normalize operation, which are connected in sequence.

[0022] According to the above technical means, the Dyhead_DCNv3 module enhances the model's adaptability to target shape and size. At the same time, the multi-branch design of the Dyhead_DCNv3 module enhances the feature extraction capability and improves the accuracy and efficiency of target recognition.

[0023] Furthermore, a cross-stage cascade fusion structure is adopted in the Neck network. On the basis of the original YOLOv8 model, a Conv module is added before each Concat module to enhance the fusion effect of features of different scales.

[0024] The present invention also provides an application of a target recognition method based on an improved YOLOv8 model, wherein the target recognition method is applied to recognition of electric vehicles in an elevator, and the recognition target is an electric vehicle.

[0025] The present invention also provides an electronic device, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the steps of the target recognition method based on the improved YOLOv8 model described in any one of claims 1 to 8 through the computer program.

[0026] Compared with the prior art, the beneficial effects of this method are:

[0027] The present invention provides a target recognition method based on an improved YOLOv8 model and its application. First, an image to be recognized is obtained and preprocessed. Then, an improved YOLOv8 model is constructed; a C2f_SENet module is added on the basis of the original YOLOv8 model to improve the feature expression ability of the YOLOv8 model and increase the dependency between channels; a Dyhead_DCNv3 module is added on the basis of the original YOLOv8 model to improve the multi-scale target detection capability. The present invention uses an improved YOLOv8 model as a whole, improves the accuracy and efficiency of target recognition, and can be applied to improve the accuracy and efficiency of electric vehicle recognition in elevators in complex situations such as too high or too low light intensity and electric vehicles being blocked. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 A flowchart showing a target recognition method based on an improved YOLOv8 model proposed in an embodiment of the present invention;

[0029] Figure 2 A structural diagram showing an improved YOLOv8 model proposed in an embodiment of the present invention;

[0030] Figure 3 A schematic diagram showing the structure of a C2f_SNet module proposed in an embodiment of the present invention;

[0031] Figure 4A schematic diagram showing the structure of the DyHead_DCNv3 module proposed in an embodiment of the present invention;

[0032] Figure 5 An example diagram showing the image preprocessing of an electric vehicle entering an elevator proposed in an embodiment of the present invention;

[0033] Figure 6 A PR curve diagram showing target recognition using the original YOLOv8 model proposed in an embodiment of the present invention;

[0034] Figure 7 A PR curve diagram showing target recognition by the improved YOLOv8 model proposed in an embodiment of the present invention;

[0035] Figure 8 A schematic diagram showing an electronic device proposed in an embodiment of the present invention. DETAILED DESCRIPTION

[0036] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;

[0037] In order to better illustrate the present embodiment, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the actual size;

[0038] It is understandable to those skilled in the art that descriptions of certain well-known contents in the drawings may be omitted.

[0039] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0040] The positional relationships described in the drawings are only for illustrative purposes and should not be construed as limiting the present patent;

[0041] Example 1

[0042] This embodiment proposes a target recognition method based on an improved YOLOv8 model. Figure 1 The target recognition method proposed in this embodiment generally includes the following steps:

[0043] S1: Obtain the image of the electric vehicle entering the elevator to be identified and perform preprocessing;

[0044] S2: Build an improved YOLOv8 model; based on the original YOLOv8 model, add the C2f_SENet module for enhancing feature expression capabilities and the Dyhead_DCNv3 module for improving multi-scale target detection capabilities;

[0045] S3: Train the improved YOLOv8 model to obtain a trained improved YOLOv8 model;

[0046] S4: Input the image preprocessed in step S1 into the trained improved YOLOv8 model to obtain the target recognition result.

[0047] The composition of the improved YOLOv8 model is specifically described below. In this embodiment, Figure 2 The improved YOLOv8 model structure shown in the figure consists of three parts: Backbone network, Neck network and Head network.

[0048] The backbone network includes: a first Conv module, a second Conv module, a first C2f_SENet module, a third Conv module, a second C2f_SENet module, a fourth Conv module, a third C2f_SENet module, a fifth Conv module, a fourth C2f_SENet module and a first SPPF module are connected in sequence. Figure 3 In the structural diagram of the C2f_SNet module shown in the figure, any one of the first C2f_SENet module, the second C2f_SENet module, the third C2f_SENet module and the fourth C2f_SENet module includes: a first Residual layer, a first GLobal Avg Pool layer, a first Fully Connected layer, a first Non-li near layer, a second Fully Connected layer, a first Sigmoid layer and a second Residual layer connected in sequence; the output end of the first Residual layer is also connected to the input end of the second Residual layer.

[0049] The sizes of the output feature maps of the first C2f_SENet module, the second C2f_SENet module, the third C2f_SENet module, the fourth C2f_SENet module and the first SPPF module are 160×160×128, 80×80×256, 40×40×512, 20×20×1024, and 20×20×1024, respectively. The first SPPF module, i.e., the receptive field expansion module, is used to expand the receptive field range of the feature map.

[0050] The parameter expression of the C2f_SENet module is:

[0051]

[0052] In the formula, parameters represents the parameter quantity, r represents the reduction ratio, s represents the number of stages, represents the dimension of the output channel of stage s, N s Indicates a repeated block number.

[0053] C2f_SENet introduces squeezing and excitation operations to enhance feature expression capabilities and inter-channel dependency modeling, enabling the model to capture key information more effectively. It optimizes the accuracy of target recognition and improves the generalization ability of the model in different scenarios, thereby achieving better performance in complex tasks.

[0054] The sizes of the output feature maps of the first C2f_SENet module, the second C2f_SENet module, the third C2f_SENet module, the fourth C2f_SENet module and the first SPPF module are 160×160×128, 80×80×256, 40×40×512, 20×20×1024, and 20×20×1024, respectively. The first SPPF module, receptive field expansion, is used to expand the receptive field range of the feature map.

[0055] The neck network includes: a sixth Conv module, a first Upsample module, a first Concat module, a first C2f module, a seventh Conv module, a second Upsample module, a second Concat module, a second C2f module, an eighth Conv module, a third Concat module, a third C2f module, a ninth Conv module, a fourth Concat module and a fourth C2f module connected in sequence; and also includes a tenth Conv module and an eleventh Conv module; the output end of the third Conv module is connected to the input end of the tenth Conv module, the output end of the tenth Conv module is connected to the second Concat module, the output end of the fourth Conv module is connected to the input end of the eleventh Conv module, the output end of the eleventh Conv module is connected to the input end of the first Concat module, the output end of the first SPPF module is connected to the input end of the sixth Conv module, the output end of the sixth Conv module is connected to the input end of the fourth Concat module, and the output end of the seventh Conv module is connected to the input end of the third Concat module.

[0056] In the neck network, a convolution layer with a convolution kernel size of 1×1 and an output channel number of 256 is added before each Concat module, and a total of four are added. A consistent number of 256 channels is used throughout the entire neck part. Through the correlation analysis of cross-layer features, the fusion effect of features of different scales is further enhanced, solving the problem of feature loss that may be caused by the original neck network during information transmission. Compared with the original structure, CCFM can more efficiently utilize features at different levels, improve the accuracy of identifying multi-scale targets, enhance the robustness of the model, and reduce missed detections and false detections in complex scenarios.

[0057] The first Upsample module upsamples the feature map output by the sixth Conv module to obtain an upsampled feature map. The first Concat module fuses the upsampled feature map obtained by the first Upsample module and the feature map output by the eleventh Conv module. The second Concat module fuses the upsampled feature map output by the second Upsample module and the feature map output by the tenth Conv module. The third Concat module fuses the feature maps output by the eighth Conv module and the seventh Conv module. The fourth Concat module fuses the feature maps output by the ninth Conv module and the sixth Conv module. The Upsample module aggregates deep semantic information on a larger receptive field, and can adaptively design and reconstruct the kernel according to the content information on the specific feature map to achieve upsampling of the feature map.

[0058] The sizes of the output feature maps of the sixth Conv module, the first Upsample module, the first Concat module, the first C2f module, the seventh Conv module, the second Upsample module, the second Concat module, the second C2f module, the eighth Conv module, the third Concat module, the third C2f module, the ninth Conv module, the fourth Concat module, the fourth C2f module, the tenth Conv module and the eleventh Conv module are respectively 20×20×256, 40×40×256, 40×40×256, 40×40×256, 40×40×256, 80×80×256, 80×80×256, 80×80×256, 80×80×256, 80×80×256, 80×80×256, 80×80×256, 80×80×256, 80×80×256, 80×80×256, 40×40×256, 80×80×256.

[0059] The head network includes: a first Dyhead_DCNv3 module, a second Dyhead_DCNv3 module and a third Dyhead_DCNv3 module. Figure 4 As shown in the structural schematic diagram of the DyHead_DCNv3 module, any one of the first Dyhead_DCNv3 module, the second Dyhead_DCNv3 module and the third Dyhead_DCNv3 module includes: a perception scale attention layer, a spatial perception attention layer and a task perception attention layer connected in sequence.

[0060] The sizes of the output feature maps of the first Dyhead_DCNv module, the second Dyhead_DCNv module, and the third Dyhead_DCNv module are 80×80×255, 40×40×255, and 20×20×255, respectively.

[0061] The perceptual-scale attention layer includes: a first Avg Pool operation, a first DCNv3 operation, a first ReLU operation, and a first Hard sigmoid operation, which are connected in sequence.

[0062] The spatial perception attention layer includes: a first Index operation, a second DCNv3 operation, a first Sigmoid operation and a first Offset operation; the output end of the first Index operation is connected to the input end of the second DCNv3 operation, and the output end of the second DCNv3 operation is respectively connected to the input ends of the first Sigmoid operation and the first Offset operation.

[0063] The task-aware attention layer includes: a second Avg Pool operation, a first Fully Connected operation, a first RELU operation, a second Fully Connected operation, and a first Normalize operation, which are connected in sequence.

[0064] The Dyhead_DCNv3 module uses deformable convolution DCNv3 to enhance adaptability to target shape and size, and can adapt to complex environments and changing target shapes. In addition, the multi-branch design of Dyhead_DCNv3 optimizes the independent processing capabilities of the feature layer, making the feature expression more refined and significantly improving the accuracy and efficiency of target recognition.

[0065] Example 2

[0066] This embodiment describes the process of obtaining an image to be recognized and performing preprocessing.

[0067] The image to be identified obtained in this embodiment is an image of an electric vehicle entering an elevator, and the target object to be identified is an electric vehicle. The selected image of the electric vehicle entering the elevator is preprocessed, and the duplicate images and invalid images that do not contain the identification target in the electric vehicle entering the elevator data set are first eliminated. Then, the remaining images are exposed to -15% and 15% to simulate different lighting conditions and elevator reflection problems that elevators in different scenes may face when opening and closing doors. At the same time, in order to simulate the situation where the electric vehicle may be partially blocked by other objects when entering the elevator and in the elevator, the image of the electric vehicle entering the elevator is cropped twice at 20% to obtain the processed electric vehicle entering the elevator picture, which constitutes the electric vehicle entering the elevator data set.

[0068] In order to increase the diversity of the electric vehicle entering the elevator dataset, data enhancement techniques are used to expand the dataset. For example, rotation, flipping, scaling, cropping, and noise addition are performed, and then the dataset is randomly divided into training set, validation set, and test set in a ratio of 7:2:1. Figure 5The example images of the preprocessing of the electric car entering the elevator are shown in the figure, which are the original electric car entering the elevator image, the image with 15% less exposure and black frame occlusion, and the image with 15% more exposure and black frame occlusion. The black frame occlusion in the image is to simulate the situation of being occluded in reality, and the brightness of -15% and +15% is the brightness difference when the elevator door is opened and closed. This increases the diversity of the images.

[0069] Example 3

[0070] In this embodiment, the improved YOLOv8 model is trained using the training set to obtain a trained improved YOLOv8 model.

[0071] In this embodiment, when training the improved YOLOv8 model, the basic platform and parameters are as follows:

[0072] The improved YOLOv8 model network was trained in the PyTroch environment, using the SGD optimizer, a learning rate of 0.001, a momentum parameter of 0.937, 300 training rounds, and a batch_size value of 32; the final training mAP can reach 0.973, achieving an improvement in the performance of the improved YOLOv8 model.

[0073] Use the trained improved YOLOv8 model to test on the test set, compare the test results with the true labels, calculate the accuracy (Precision, P), recall (Recall, R), average accuracy (Average Precision, AP), mean average precision (mean Average Precision, mAP), and generate a PR curve.

[0074] The precision rate P refers to the proportion of all true positive samples that are correctly predicted as positive samples by the model. The calculation expression is:

[0075]

[0076] The recall rate R refers to the proportion of samples that are actually positive samples among the samples predicted by the model. The calculation expression is:

[0077]

[0078] AP refers to the average accuracy of each category, which is calculated by the area enclosed by the curve and the horizontal and vertical coordinates in the PR curve. The calculation expression is:

[0079]

[0080] mAP is obtained by averaging the AP of n categories, where n represents the number of categories; mAP represents the average accuracy of all categories, and the calculation expression is:

[0081]

[0082] Where TP represents the number of target frames correctly identified as electric vehicles, FP represents the number of target frames that mistakenly identify other classes as electric vehicles, FN represents the number of target frames that mistakenly identify electric vehicles as other classes, P represents the ratio of target frames correctly identified as electric vehicles to all target frames identified as electric vehicles, and R represents the ratio of target frames correctly identified as electric vehicles to all target frames of electric vehicles.

[0083] In this embodiment, the categories are: person, motorcycle, and bicycle.

[0084] In this embodiment, if Figure 6 The PR curve of target recognition using the original YOLOv8 model is shown in Figure 7 The PR curve of the improved YOLOv8 model for target recognition is shown in the figure. According to the PR curves of the original YOLOv8 model and the improved YOLOv8 model of the present invention, the mAP value of the improved YOLOv8 model is improved by 2.4% compared with the YOLOv8 model before improvement. And while improving the target recognition accuracy, the network parameters are reduced by 16.1% compared with the original YOLOv8, meeting the expected recognition requirements.

[0085] Example 4

[0086] This embodiment provides an application of a target recognition method based on an improved YOLOv8 model, wherein the target recognition method is applied to recognition of electric vehicles in an elevator, and the recognition target is an electric vehicle.

[0087] Example 5

[0088] like Figure 8 As shown, an embodiment of the present invention further proposes an electronic device, comprising a memory 101, a processor 102, and a computer program stored in the memory 101 and running on the processor 102, wherein when the processor 102 executes the computer program, the steps of the target recognition method based on the improved YOLOv8 model proposed in this embodiment are implemented.

[0089] Specifically, in the present embodiment, the processor 102 may include a central processing unit (CPU) or a specific integrated circuit, or be configured to implement one or more integrated circuits of the present embodiment, and the memory 101 may include a large-capacity memory for data or instructions. It may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 101 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 101 may be inside or outside the integrated gateway disaster recovery device.

[0090] The embodiments are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.

Claims

1. A target recognition method based on an improved YOLOv8 model, characterized in that: The following steps are involved: S1: Obtain the image to be recognized and perform preprocessing; S2: construct an improved YOLOv8 model; based on the original YOLOv8 model, the improved YOLOv8 model adds a C2f_SENet module for enhancing feature expression capabilities and a Dyhead_DCNv3 module for improving multi-scale target detection capabilities; S3: Train the improved YOLOv8 model to obtain a trained improved YOLOv8 model; S4: Input the image preprocessed in step S1 into the trained improved YOLOv8 model to obtain the target recognition result.

2. A target recognition method based on an improved YOLOv8 model according to claim 1, characterized in that: In step S1, the preprocessing operation is: removing duplicate images and invalid images that do not contain recognition targets from the images to be recognized, and performing exposure processing on the remaining images.

3. A target recognition method based on an improved YOLOv8 model according to claim 1, characterized in that: The improved YOLOv8 model consists of three parts: Backbone network, Neck network and Head network; The Backbone network includes: a first Conv module, a second Conv module, a first C2f_SENet module, a third Conv module, a second C2f_SENet module, a fourth Conv module, a third C2f_SENet module, a fifth Conv module, a fourth C2f_SENet module and a first SPPF module connected in sequence.

4. A target recognition method based on an improved YOLOv8 model according to claim 1, characterized in that: The Neck network includes: a sixth Conv module, a first Upsample module, a first Concat module, a first C2f module, a seventh Conv module, a second Upsample module, a second Concat module, a second C2f module, an eighth Conv module, a third Concat module, a third C2f module, a ninth Conv module, a fourth Concat module and a fourth C2f module connected in sequence; and also includes a tenth Conv module and an eleventh Conv module; the output end of the third Conv module is connected to the input end of the tenth Conv module, the output end of the tenth Conv module is connected to the second Concat module, the output end of the fourth Conv module is connected to the input end of the eleventh Conv module, the output end of the eleventh Conv module is connected to the input end of the first Concat module, the output end of the first SPPF module is connected to the input end of the sixth Conv module, the output end of the sixth Conv module is connected to the input end of the fourth Concat module, and the output end of the seventh Conv module is connected to the input end of the third Concat module.

5. The target recognition method based on the improved YOLOv8 model according to claim 3, characterized in that: The Head network includes: a first Dyhead_DCNv3 module, a second Dyhead_DCNv3 module and a third Dyhead_DCNv3 module; any one of the first Dyhead_DCNv3 module, the second Dyhead_DCNv3 module and the third Dyhead_DCNv3 module includes: a first perception scale attention layer, a first spatial perception attention layer and a first task perception attention layer connected in sequence.

6. A target recognition method based on an improved YOLOv8 model according to claim 3, characterized in that: Any one of the first C2f_SENet module, the second C2f_SENet module, the third C2f_SENet module and the fourth C2f_SENet module includes: a first Residual layer, a first GLobal Avg Pool layer, a first Fully Connected layer, a first Non-li near layer, a second Fully Connected layer, a first Sigmoid layer and a second Residual layer connected in sequence; the output end of the first Residual layer is also connected to the input end of the second Residual layer.

7. The target recognition method based on the improved YOLOv8 model according to claim 5, characterized in that: The perceptual scale attention layer includes: a first Avg Pool operation, a first DCNv3 operation, a first ReLU operation, and a first Hard sigmoid operation connected in sequence; The spatial perception attention layer includes: a first Index operation, a second DCNv3 operation, a first Sigmoid operation and a first Offset operation; the output end of the first Index operation is connected to the input end of the second DCNv3 operation, and the output end of the second DCNv3 operation is respectively connected to the input ends of the first Sigmoid operation and the first Offset operation; The task-aware attention layer includes: a second Avg Pool operation, a first Fully Connected operation, a first RELU operation, a second Fully Connected operation, and a first Normalize operation, which are connected in sequence.

8. The target recognition method based on the improved YOLOv8 model according to claim 4, characterized in that: The Neck network adopts a cross-stage cascade fusion structure. On the basis of the original YOLOv8 model, a Conv module is added before each Concat module to enhance the fusion effect of features of different scales.

9. An application of a target recognition method based on the improved YOLOv8 model according to any one of claims 1 to 8, characterized in that: The target recognition method is applied to the recognition of electric vehicles in an elevator, and the recognition target is an electric vehicle.

10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the steps of the target recognition method based on the improved YOLOv8 model described in any one of claims 1 to 8 through the computer program.