Pedestrian clothing identification method, device and equipment based on improved YOLOv8 model and storage medium

By improving the backbone network layer and feature fusion layer of the YOLOv8 model and combining it with the color and style classification branches of the detection head, the problem of insufficient accuracy of existing algorithms in pedestrian clothing recognition in complex scenes is solved, and accurate recognition of clothing types, colors and styles is achieved.

CN120656043AActive Publication Date: 2025-09-16HUNAN HUALIAN YUNCHUANG INFORMATION TECH CO LTD

Patent Information

Application Number
CN202511151724.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-09-16
Estimated Expiration
2045-08-18

AI Technical Summary

Technical Problem

The existing YOLOv3 and YOLOv5 algorithms are not sensitive enough to fine-grained clothing features in pedestrian clothing recognition, and YOLOv8 has limitations in processing multi-scale features, resulting in insufficient recognition capabilities in complex scenarios.

Method used

The improved YOLOv8 model uses a backbone network layer consisting of the CSPDarknet53 architecture, coordinate attention module, SCConv module, and Focus module. Combined with the PAN-FPN structure and a feature fusion layer with bidirectional cross-layer connections, it adds detection heads for color and style classification branches to achieve multi-scale feature extraction and accurate recognition.

Benefits of technology

It improves the accuracy of pedestrian clothing recognition in complex scenarios, can accurately identify clothing types, colors and styles, and enhances the adaptability and detection accuracy of targets of different scales.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656043A_ABST
    Figure CN120656043A_ABST
Patent Text Reader

Abstract

The invention discloses a pedestrian clothing identification method and device based on an improved YOLOv8 model, equipment and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining and preprocessing pedestrian image data, and generating high-quality input data through normalization; then, an improved YOLOv8 model is used for recognition, and the model comprises three core parts: a backbone network layer is responsible for feature extraction; the feature fusion layer is used for multi-scale feature fusion; the detection head is divided into a classification branch and a regression branch, and a color and style classification branch is additionally added, so that the type, color and style of the clothes are accurately recognized, and the accuracy of pedestrian clothes recognition in a complex scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a pedestrian clothing recognition method, device, equipment and storage medium based on an improved YOLOv8 model. Background Art

[0002] Currently, pedestrian clothing recognition primarily relies on deep learning-based object detection algorithms, such as the YOLO family of algorithms. The existing YOLOv3 and YOLOv5 algorithms have achieved good results in traffic sign recognition and pedestrian detection, but their backbone networks lack sensitivity to fine-grained clothing features (such as clothing patterns and colors). Furthermore, the existing YOLOv8 algorithm has limitations when processing multi-scale features, making it difficult to effectively extract clothing features of different scales simultaneously. This results in insufficient recognition of clothing in complex scenes. Therefore, there is an urgent need for a pedestrian clothing recognition method that can improve the accuracy of pedestrian clothing recognition in complex scenes. Summary of the Invention

[0003] The main purpose of this application is to provide a pedestrian clothing recognition method, device, equipment and storage medium based on an improved YOLOv8 model, aiming to solve the technical problem of how to improve the accuracy of pedestrian clothing recognition in complex scenarios.

[0004] To achieve the above objectives, this application proposes a pedestrian clothing recognition method based on an improved YOLOv8 model, including: Obtain pedestrian image data; Preprocessing the image data to obtain processed image data; The processed image data is input into a preset clothing recognition model for recognition to obtain a clothing recognition result. The preset clothing recognition model is an improved YOLOv8 model. The improved YOLOv8 model includes a backbone network layer, a feature fusion layer and a detection head. The backbone network layer is composed of a CSPDarknet53 architecture, a coordinate attention module, an SCConv module and a Focus module. The CSPDarknet53 architecture is composed of a first C2f module and a maximum pooling layer in series. The coordinate attention module is inserted into the first preset sublayer, the first sublayer and the second sublayer of the backbone network layer. Two preset sublayers and a third preset sublayer, the first C2f module includes two Bottleneck layers, each layer consists of a first convolution and a second convolution, and the Focus module consists of a third convolution; the feature fusion layer is composed of a PAN-FPN structure and a second C2f module, and the PAN-FPN structure and the second C2f module are both bidirectional cross-layer connections; the detection head is used for target detection and classification, and the detection head is divided into a classification branch and a regression branch, and the classification branch and the regression branch are both one-dimensional convolutions, and the classification branch also includes a color classification branch and a style classification branch.

[0005] In one embodiment, the step of inputting the processed image data into a preset clothing recognition model for recognition to obtain a clothing recognition result includes: Performing feature extraction on the processed image data through the backbone network layer to obtain clothing features; Performing multi-scale fusion on the clothing features through the feature fusion layer to obtain fusion features; The detection head classifies and regresses the fused features to obtain clothing recognition results.

[0006] In one embodiment, the step of extracting features from the processed image data through the backbone network layer to obtain clothing features includes: Slicing the processed image data through the Focus module and performing feature extraction using the third convolution to obtain an initial feature map; Input the initial feature map into the CSPDarknet53 architecture and perform feature extraction through the first Bottleneck layer of the first C2f module to obtain a first intermediate feature map, wherein the first Bottleneck layer performs feature transformation and feature extraction through the first convolution and the second convolution in sequence; Input the first intermediate feature map into the coordinate attention module and enhance it through the position-sensitive feature enhancement mechanism to obtain the second intermediate feature map; Input the second intermediate feature map into the CSPDarknet53 architecture and perform feature extraction through the second Bottleneck layer of the first C2f module to obtain a third intermediate feature map, wherein the second Bottleneck layer performs feature transformation and feature extraction through the first convolution and the second convolution in sequence; Passing the third intermediate feature map through a preset number of maximum pooling layers in sequence to perform multi-scale feature fusion to obtain a fourth intermediate feature map; The fourth intermediate feature map is input into the SCConv module for fast spatial pyramid pooling processing to obtain clothing features.

[0007] In one embodiment, the step of performing multi-scale fusion on the clothing features through the feature fusion layer to obtain fused features includes: Downsampling the clothing feature through a first downsampling path of a PAN-FPN structure to obtain a first downsampling feature; The first down-sampled features are sequentially subjected to a first convolution and a second convolution of a second C2f module to obtain a first cross-stage feature; Upsampling the first cross-stage feature using a nearest neighbor interpolation method through a first upsampling path of the PAN-FPN structure to obtain a first upsampled feature; Processing the first upsampled features by a second cross-stage part of a second C2f module to obtain a second cross-stage feature; Downsampling the second cross-stage feature through a second downsampling path of the PAN-FPN structure to obtain a second downsampled feature; The second down-sampled feature is concatenated with the clothing feature to obtain a fused feature.

[0008] In one embodiment, the step of classifying and regressing the fused features by the detection head to obtain clothing recognition results includes: Performing preliminary feature transformation on the fusion feature by the detection head to obtain an initial detection feature; Performing feature enhancement on the initial detection feature through a global attention mechanism module to obtain an enhanced detection feature; The initial detection features are input into the classification branch and the regression branch respectively to obtain the category probability distribution and target position information, wherein the category probability distribution is obtained based on the color category and the style category: fusing the category probability distribution and the target position information to obtain a clothing detection result; The clothing detection result is processed by a non-maximum suppression algorithm to obtain a clothing recognition result.

[0009] In one embodiment, before the step of inputting the processed image data into a preset clothing recognition model for recognition to obtain a clothing recognition result, the step includes: Acquire historical image data; Annotating the historical image data to generate a training set and a test set; Build an initial clothing recognition model; The clothing recognition model is trained using the training set to obtain a preset clothing recognition model.

[0010] In one embodiment, the step of training the clothing recognition model using the training set to obtain a preset clothing recognition model includes: Initializing weight and bias parameters of the initial clothing recognition model; Inputting the image data in the training set into the initial clothing recognition model for calculation to obtain predicted position and category probability; Calculating the error between the predicted position and the actual position according to a first loss function to obtain a distributed focus loss, wherein the actual position is calculated according to an external matrix; Calculating the cross entropy loss of the class probability and the true class label according to a second loss function to obtain the class probability loss, wherein the true class label is obtained by a preset coding table; Processing is performed according to the distribution focus loss and the category probability loss to obtain a total error value; Obtaining the gradients of the weight and bias parameters by back propagation algorithm calculation; The weight and bias parameters are iteratively updated according to the gradient through an optimization algorithm until a maximum number of iterations is reached or the total error value converges to a preset threshold, thereby obtaining a preset clothing recognition model.

[0011] In addition, to achieve the above objectives, the present application also proposes a pedestrian clothing recognition device based on an improved YOLOv8 model, wherein the pedestrian clothing recognition device based on the improved YOLOv8 model comprises: An acquisition module, used to acquire pedestrian image data; A processing module, configured to pre-process the image data to obtain processed image data; The result module is used to input the processed image data into a preset clothing recognition model for recognition to obtain clothing recognition results. The preset clothing recognition model is an improved YOLOv8 model. The improved YOLOv8 model includes a backbone network layer, a feature fusion layer and a detection head. The backbone network layer is composed of a CSPDarknet53 architecture, a coordinate attention module, an SCConv module and a Focus module. The CSPDarknet53 architecture adopts a first C2f module and a maximum pooling layer in series. The coordinate attention module is inserted into the first preset sub-layer of the backbone network layer. layer, a second preset sublayer and a third preset sublayer, the first C2f module includes two Bottleneck layers, each layer consists of a first convolution and a second convolution, and the Focus module consists of a third convolution; the feature fusion layer is composed of a PAN-FPN structure and a second C2f module, and the PAN-FPN structure and the second C2f module are both bidirectional cross-layer connections; the detection head is used for target detection and classification, and the detection head is divided into a classification branch and a regression branch, and the classification branch and the regression branch are both one-dimensional convolutions, and the classification branch also includes a color classification branch and a style classification branch.

[0012] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the pedestrian clothing recognition method based on the improved YOLOv8 model as described above are implemented.

[0013] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the pedestrian clothing recognition method based on the improved YOLOv8 model as described above.

[0014] This application acquires and preprocesses pedestrian image data, generating high-quality input data through normalization. It then uses an improved YOLOv8 model for recognition. This model consists of three core components: a backbone network layer responsible for feature extraction; a feature fusion layer for multi-scale feature fusion; and a detection head divided into classification and regression branches, with additional color and style classification branches. This allows for precise recognition of clothing type, color, and style, improving the accuracy of pedestrian clothing recognition in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0016] Figure 1 This is a flowchart of the first embodiment of the pedestrian clothing recognition method based on the improved YOLOv8 model of this application; Figure 2 This is a structural diagram of an improved YOLOv8 model according to the first embodiment of the pedestrian clothing recognition method based on the improved YOLOv8 model of this application; Figure 3 This is a flowchart of the second embodiment of the pedestrian clothing recognition method based on the improved YOLOv8 model of this application; Figure 4 This is a flowchart of the third embodiment of the pedestrian clothing recognition method based on the improved YOLOv8 model of this application; Figure 5 This is a schematic diagram of the module structure of the pedestrian clothing recognition device based on the improved YOLOv8 model in this application; Figure 6 Schematic diagram of the device structure of the hardware operating environment involved in the pedestrian clothing recognition method based on the improved YOLOv8 model in the embodiment of the present application.

[0017] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0018] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0019] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0020] In today's rapidly developing fields of public safety and technological surveillance, the demand for accurate identification of pedestrian clothing is growing. Accurately identifying a person's clothing plays a crucial role in ensuring public safety and optimizing the operational efficiency of intelligent transportation systems. While existing YOLOv3 and YOLOv5 algorithms have achieved good results in traffic sign recognition and pedestrian detection, their backbone networks lack sensitivity to fine-grained clothing features (such as clothing patterns and colors). Furthermore, the existing YOLOv8 algorithm has limitations in processing multi-scale features, making it difficult to effectively extract clothing features at different scales simultaneously. This results in insufficient clothing recognition capabilities in complex scenes.

[0021] Therefore, this application proposes a pedestrian clothing recognition method based on an improved YOLOv8 model to solve the above problems. The main solution of the embodiment of this application is: obtain pedestrian image data; preprocess the image data to obtain processed image data; input the processed image data into a preset clothing recognition model for recognition to obtain clothing recognition results. The preset clothing recognition model is an improved YOLOv8 model. The improved YOLOv8 model includes a backbone network layer, a feature fusion layer and a detection head. The backbone network layer is composed of a CSPDarknet53 architecture, a coordinate attention module, an SCConv module and a Focus module. The CSPDarknet53 architecture adopts a first C2f module and a maximum pooling layer connected in series. The coordinate attention module is inserted into the first preset sublayer, the second preset sublayer and the third preset sublayer of the backbone network layer. The first C2f module contains two Bottleneck layers, each layer consists of the first convolution and the second convolution, and the Focus module consists of the third convolution; the feature fusion layer is composed of the PAN-FPN structure and the second C2f module, and the PAN-FPN structure and the PAN-FPN structure are both bidirectional cross-layer connections; the detection head is used for target detection and classification. The detection head is divided into a classification branch and a regression branch. The classification branch and the regression branch are both one-dimensional convolutions. The classification branch also includes a color classification branch and a style classification branch.

[0022] Based on the above, the embodiment of the present application also provides a pedestrian clothing recognition method based on the improved YOLOv8 model, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the pedestrian clothing recognition method based on the improved YOLOv8 model of this application. In this embodiment, the pedestrian clothing recognition method based on the improved YOLOv8 model includes steps S10 to S30: Step S10: Obtain pedestrian image data.

[0023] It's important to note that cameras deployed at checkpoints, residential communities, and public areas collect raw pedestrian images from all angles and in all weather conditions. These devices cover a wide range of urban scenes, from bustling commercial streets to residential areas and major traffic arteries, ensuring diverse and representative data. When acquiring portrait image data, image processing techniques are employed to enhance data quality, taking into account interference from varying environmental conditions such as lighting variations, weather conditions (rain, fog, etc.), and background complexity. For example, the DehazeNet pre-trained model can effectively mitigate the impact of adverse weather conditions such as haze on image clarity, thereby improving the visibility of pedestrian features in the images.

[0024] Step S20: pre-process the image data to obtain processed image data.

[0025] It should be noted that the original pedestrian image data contains noise, uneven lighting, scale variations, and other issues. If these issues are not addressed, they will interfere with the model's learning process and reduce recognition accuracy. Therefore, before any deep learning model training, these image data must undergo a series of preprocessing operations.

[0026] Furthermore, step S20 also includes: performing preliminary noise reduction on the image data through median filtering to obtain first noise-reduced image data. Specifically, median filtering is an effective nonlinear noise reduction technique, particularly suitable for removing randomly occurring black and white point noise in images. Its basic principle is to smooth the image by replacing the value of each pixel with the median of all pixel values ​​in its surrounding neighborhood, thereby effectively suppressing noise while retaining edge information. In specific implementation, for each pixel position, a window of appropriate size (such as 3x3 or 5x5) is selected to cover it, and then the median of all pixel values ​​in the window is calculated, and this median value is used to replace the original value of the center pixel. This method can effectively protect the boundaries and detailed features of the image from being damaged, because median filtering does not blur the image edges like mean filtering. In addition, depending on the level of image noise, the size of the filter window can be flexibly adjusted to achieve the best noise reduction effect.

[0027] Next, the first denoised image data is subjected to secondary denoising through Gaussian filtering to obtain second denoised image data. Specifically, Gaussian filtering is a linear smoothing filtering technology based on Gaussian function, which can effectively remove Gaussian noise in the image while retaining the important features and details of the image as much as possible. For each pixel position, the average value of all pixel values ​​in its surrounding neighborhood weighted by Gaussian distribution is calculated, and this value is used to replace the original pixel value. In this way, not only can the random noise in the image be effectively reduced, but also the artifacts or unnecessary edge roughness that may be introduced by the first median filtering can be alleviated to a certain extent. It is worth noting that the reasonable adjustment of the parameters of the Gaussian filter is the key. Too strong filtering may cause the loss of image details, while too weak filtering cannot achieve the ideal noise reduction effect.

[0028] The second denoised image data is then normalized to produce normalized image data. Normalization involves scaling image pixel values ​​to a specific range (typically between 0 and 1), which helps accelerate neural network convergence and improve model training stability. Specifically, normalization involves linearly transforming the grayscale value or RGB channel value of each pixel to fit within a specified range. For most deep learning frameworks, scaling pixel values ​​to the range [0, 1] is a common practice. This can be achieved by simply dividing by 255 (assuming the original pixel value range is [0, 255]). In some cases, more complex normalization strategies may be employed, such as normalizing the input data to zero mean and unit variance based on the mean and standard deviation of the entire dataset. Normalized image data not only reduces differences between images due to factors such as lighting conditions and camera equipment, but also helps improve the convergence of the YOLOv8 model, especially in tasks involving color feature extraction.

[0029] Finally, a region-of-interest (ROI) method is used to crop key regions from the normalized image data to obtain processed image data. The region-of-interest (ROI) method crops the most valuable portion of a person—in this case, the pedestrian and their clothing information—to reduce interference from background noise and other irrelevant information. Specifically, determining the region of interest (ROI) involves identifying and locating the area within the image where the pedestrian resides. This can be achieved through methods such as a pre-trained object detection model or simple threshold segmentation. Once the pedestrian's approximate location is determined, the selection can be further refined to ensure that the cropped region encompasses both the complete pedestrian outline and sufficient contextual information for more accurate clothing attribute recognition. For example, appropriately expanding the cropping boundary can ensure that as much relevant detail as possible is captured, even when the pedestrian is in motion or partially occluded. For each resulting ROI, additional preprocessing steps can be applied, such as resizing to meet model input requirements, contrast enhancement, or brightness adjustment, to further optimize data quality. It is worth noting that reasonable cropping and scaling not only helps improve computational efficiency, but also avoids unnecessary computational overhead, which is especially important when processing large-scale video streams or real-time monitoring data.

[0030] The processed image data obtained through the above processing has reduced background noise and focuses on the key features of pedestrian clothing, which makes it easy to obtain accurate output values ​​for improving the YOLOv8 model.

[0031] In step S30 , the processed image data is input into a preset clothing recognition model for recognition to obtain a clothing recognition result.

[0032] It should be noted that if Figure 2As shown in the structure diagram of the improved YOLOv8 model, the improved YOLOv8 model includes a backbone network layer, a feature fusion layer, and a detection head. The backbone network layer consists of a CSPDarknet53 architecture, a coordinate attention module, an SCConv module, and a Focus module. The CSPDarknet53 architecture uses a first C2f module and a maximum pooling layer in series. The coordinate attention module is inserted into the first, second, and third preset sublayers of the backbone network layer. In this embodiment, the coordinate attention module is inserted after the third, sixth, and ninth layers of the backbone network layer, which helps enhance position-sensitive features and further improves the capture of detailed features. The first C2f module contains two Bottleneck layers, each consisting of a first and second convolution. Specifically, each layer consists of 1×1 and 3×3 convolutions, with the SiLU activation function. The Focus module consists of a third convolution, specifically a 6×6 convolution. The SCConv module improves the ability to extract color features, enabling the model to more accurately identify the color information of the portrait's clothing. Secondly, the feature fusion layer consists of a PAN-FPN structure and a second C2f module. Both modules utilize bidirectional cross-layer connections. The PAN-FPN structure incorporates bidirectional cross-layer connections, uses nearest neighbor interpolation for upsampling, and 3×3 convolutions (stride 2) for downsampling. The activation function is SiLU. This design improves computational efficiency while preserving multi-scale features. The use of the second C2f module further enhances the feature fusion effect, particularly for small-scale clothing detection. A lightweight SE attention module is added after the PAN-FPN upsampling path, and an RFE (receptive field enhancement) module is added to the downsampling path. These improvements work together to improve the model's adaptability to objects of different scales and its detection accuracy. Finally, the detection head performs object detection and classification. The detection head is divided into a classification branch and a regression branch. Both branches use one-dimensional convolutions. The classification branch also includes a color classification branch and a style classification branch. The detection head is divided into two branches, both of which use one-dimensional convolutions. The classification branch involves 1×1 convolutions and softmax operations to identify the pedestrian's clothing category. The regression branch, using 1×1 convolutions and a Dependent Fluent Loss, accurately locates the target. To further refine the recognition results, the classification branch also includes a color classification branch (implemented using a three-layer MLP) and a style classification branch (implemented using a two-way CNN). A GAM (global attention mechanism) module is inserted before the regression branch to optimize localization accuracy, ensuring that the model can accurately identify and locate the specific details of the pedestrian and their clothing.

[0033] Specifically, preprocessed image data is fed into the model to extract detailed information about pedestrian clothing. The improved YOLOv8 model, through its optimized backbone network layer, feature fusion layer, and detection head, can efficiently capture and analyze key features in images. First, the image data enters the backbone network layer. These components enhance the ability to extract clothing details and color features, thereby obtaining clothing features. Next, in the feature fusion layer, the SE attention and RFE modules are combined based on clothing features to further enhance the fusion effect of multi-scale features and obtain fused features. Finally, the fused features reach the detection head, which not only performs traditional classification and regression tasks, but also adds specialized color classification branches and style classification branches to ensure fine-grained recognition of clothing types, colors, and styles. The clothing recognition results output by this model cover comprehensive information from basic clothing categories to specific colors and styles.

[0034] This embodiment acquires and preprocesses pedestrian image data, generating high-quality input data through preprocessing. It then uses an improved YOLOv8 model for recognition. This model consists of three core components: a backbone network layer responsible for feature extraction; a feature fusion layer for multi-scale feature fusion; and a detection head divided into classification and regression branches, with additional color and style classification branches. This allows for precise recognition of clothing type, color, and style, improving the accuracy of pedestrian clothing recognition in complex scenarios.

[0035] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 3 The pedestrian clothing recognition method based on the improved YOLOv8 model further includes steps S201 to S204 before step S30: Step S201: Acquire historical image data.

[0036] It's important to note that historical image data can be obtained from a variety of sources, such as traffic surveillance cameras, community security systems, and public area surveillance. First, a large number of pedestrian images are collected using devices such as checkpoint cameras, community cameras, and public area cameras. These images not only cover a variety of pedestrians but are also captured at different times, weather conditions, and lighting conditions, providing rich training material for the model.

[0037] Step S202: annotate the historical image data to generate a training set and a test set.

[0038] It's important to note that professional annotation tools were used to meticulously annotate a large number of historical images collected from various surveillance devices. This process involved accurately delineating the clothing worn by pedestrians in each image and assigning the correct label to each sign. After the annotation process was complete, the dataset was then divided into training and test sets based on a specific ratio.

[0039] Specifically, this process is completed using professional software such as labelme, which allows users to intuitively mark the specific details of pedestrians and their clothing in each picture. During the labeling process, it is necessary to record the pedestrian's clothing information in detail, including but not limited to the color (such as red, black), type (such as short sleeves, long sleeves), brand (such as Nike, Adidas), and related attributes of shoes and pants. In order to ensure the quality and diversity of the dataset, images under various environmental conditions should be covered, such as different weather conditions, lighting conditions, etc. After the labeling is completed, the dataset is divided into training sets and test sets according to a certain ratio for subsequent model training and verification. In this implementation, the training set ratio is 80% and the test set ratio is 20%.

[0040] Step S203: constructing an initial clothing recognition model.

[0041] Step S204: Using the training set to train the initial clothing recognition model to obtain a preset clothing recognition model.

[0042] It's important to note that before training begins, the prepared training set must be imported into the training environment. This training set contains labeled, resized, denoised, and normalized image data. This process first involves inputting the labeled training set into an initial model based on a modified YOLOv8. Through a series of iterative optimizations, the model learns various characteristic representations of pedestrian clothing.

[0043] Furthermore, step S204 includes initializing the weights and bias parameters of the initial clothing recognition model. The image data from the training set is then input into the initial clothing recognition model to generate predicted bounding boxes and class probabilities. Specifically, the input image is processed through the model's feature extraction network (such as the MobileNetV3 inverse residual unit) based on the image data to generate a series of feature maps. These feature maps are then used to predict bounding boxes and their corresponding class probabilities. Specifically, each predicted bounding box consists of location information (x, y, w, h) and a confidence score, where (x, y) represents the location of the target center point, and (w, h) represents the width and height of the bounding box. The class probability is the predicted probability distribution for each possible style or color category. Next, a loss function is used to calculate the loss between the predicted bounding box and the annotated box, as well as the cross-entropy loss between the class probabilities and the true class labels, to obtain a total loss value. The annotated box is calculated using an external matrix, and the true class labels are obtained using a preset encoding table. Specifically, to evaluate the quality of the model's prediction results, it is necessary to define an appropriate loss function. There are two types of losses involved here: one is the loss between the predicted bounding box and the annotated box, and the other is the cross entropy loss between the predicted class probability and the true class label. The loss of the predicted bounding box is usually measured using mean squared error (MSE) or IoU loss. In this example, mean squared error is used for loss determination. The specific formula is: in is the position information of the real annotation box, To predict the location information of the bounding box. For classification tasks, cross entropy loss is used to quantify the difference between the class probability of a cluster and the true label. The specific formula is: in, is the probability distribution of true class labels (usually one-hot encoded), To predict the category probability, based on the above loss value, the total loss value is ,in and is the balancing coefficient. The backpropagation algorithm calculates the gradients of the weight and bias parameters. The backpropagation algorithm then calculates the gradients of each weight and bias parameter. These parameters are then updated based on the gradients using an optimization algorithm (such as SGD or Adam). The goal of the optimization algorithm is to minimize the total loss. Finally, the optimization algorithm iteratively updates the weight and bias parameters based on the gradients until the maximum number of iterations is reached or the total loss converges to a preset threshold, resulting in the preset clothing recognition model.

[0044] Furthermore, after step S204, the system further includes: validating the preset clothing recognition model using the test set to obtain a verification result, which includes recognition accuracy, missed detection rate, and false detection rate; performing a judgment based on the verification result and preset performance requirements to obtain a judgment result, where the preset performance requirements include an accuracy threshold, a missed detection threshold, and a false detection threshold; and if the judgment result indicates that the verification result does not meet the preset performance requirements, retraining the preset clothing recognition model using the training set until the verification result meets the preset performance requirements, thereby outputting the preset clothing recognition model. Specifically, after model training is completed, the preset clothing recognition model must be validated using the test set to assess whether its performance meets actual application requirements. This verification process primarily includes three core metrics: recognition accuracy, missed detection rate, and false detection rate. Recognition accuracy reflects the model's ability to correctly identify pedestrian clothing and is an important metric for measuring overall recognition performance. The missed detection rate assesses the proportion of target objects that the model fails to detect, which is particularly important for pedestrian clothing recognition in small objects or against complex backgrounds. The false detection rate reflects the frequency with which the model incorrectly identifies pedestrians in non-target areas, directly impacting the system's stability and practicality. During the validation process, images from the test set are fed into a pre-trained clothing recognition model. The model outputs corresponding detection results, including the location of each detected pedestrian and their predicted clothing category. These predictions are then compared with the manually annotated information in the test set to calculate the number of accurate recognitions, missed detections, and false detections. The three metrics mentioned above are then calculated based on these statistics. The validation results are then compared against pre-set performance requirements. These requirements typically include: recognition accuracy reaching a minimum threshold (e.g., 95%), missed detection rate remaining within a certain range (e.g., below 5%), and false detection rate not exceeding a certain upper limit (e.g., below 3%). If any of the validation results fall below the pre-set threshold, the model is deemed to be substandard. At this point, the model optimization mechanism is automatically triggered, and the model is retrained using the training set. During retraining, model performance can be improved by adjusting the learning rate, increasing the number of training rounds, introducing stronger data augmentation strategies, or optimizing the loss function weights. After training, validation is performed again using the test set, and this process is repeated until all performance metrics meet the pre-set requirements.

[0045] This embodiment generates a training set and a test set by acquiring and annotating historical image data, constructs an initial clothing recognition model, and uses the training set for model training to obtain a preset clothing recognition model, thereby improving the accuracy and robustness of pedestrian clothing recognition.

[0046] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above introduction and will not be described in detail later. Figure 4The pedestrian clothing recognition step S30 based on the improved YOLOv8 model further includes steps S301 to S303: Step S301: extract features from the processed image data through the backbone network layer to obtain clothing features.

[0047] It should be noted that the backbone network layer, as the core component of the model, is responsible for extracting high-level features from the input image that are helpful for subsequent classification and detection tasks. During this process, the processed image data first passes through a series of carefully designed modules, including the CSPDarknet53 architecture, the coordinate attention module, the SCConv module, and the Focus module.

[0048] Furthermore, step S301 includes: performing a slicing operation on the processed image data through the Focus module, and using the third convolution to perform feature extraction to obtain an initial feature map. Specifically, the processed image data is sliced ​​through the Focus module, and the input image is first split into four parts according to the spatial dimension and spliced ​​along the channel dimension to form a tensor with 2 times downsampling and 4 times the number of channels. Subsequently, a third convolution of 6×6, a step size of 2, and an output channel of 64 is used for feature extraction to obtain an initial feature map that retains the original details and has compressed dimensions. Next, the initial feature map is fed into the CSPDarknet53 architecture and extracted through the first Bottleneck layer of the first C2f module, yielding a first intermediate feature map. The first Bottleneck layer sequentially transforms and extracts features through first and second convolutions. Specifically, the initial feature map is fed into the CSPDarknet53 architecture, undergoes dimensionality reduction via a 1×1 first convolution, and then fed into the first Bottleneck layer of the C2f module. This layer sequentially performs a 1×1 first convolution to compress the channels, followed by a 3×3 second convolution to extract local texture. This layer then adds the feature map to the input via a residual connection, outputting a first intermediate feature map rich in fine-grained textures such as sleeves and collars. The first intermediate feature map is then fed into the coordinate attention module for enhancement via a position-sensitive feature enhancement mechanism, yielding a second intermediate feature map. Specifically, the first intermediate feature map is fed into the coordinate attention module, which performs global average pooling in both the horizontal and vertical directions to generate a position encoding vector. This vector is then multiplied channel-wise with the original feature map after a 1×1 convolution and sigmoid activation, achieving position-sensitive feature enhancement and resulting in a second intermediate feature map that highlights the spatial relationships of key regions. Next, the second intermediate feature map is fed into the CSPDarknet53 architecture and passed through the second Bottleneck layer of the first C2f module for feature extraction, yielding a third intermediate feature map. The second Bottleneck layer then undergoes feature transformation through the first and second convolutions for feature extraction. Specifically, the second intermediate feature map is fed back into the second Bottleneck layer of the C2f module of the CSPDarknet53 architecture, where the residual transformation of the 1×1 first convolution and the 3×3 second convolution is repeated to further refine detailed features such as the trouser seams and shoe uppers, outputting the third intermediate feature map. The third intermediate feature map is then passed through a preset number of max pooling layers for multi-scale feature fusion, yielding a fourth intermediate feature map. Specifically, the third intermediate feature map is passed through three cascaded 5×5 max pooling layers, each with a stride of 1, to gradually expand the receptive field and fuse multi-scale contextual information, yielding a fourth intermediate feature map that aggregates local and global information.Finally, the fourth intermediate feature map is input into the SCConv module for fast spatial pyramid pooling processing to obtain clothing features. Specifically, the fourth intermediate feature map is input into the SCConv module, and spatial convolution is first used to capture the correlation between channels, and then channel convolution is used to integrate spatial features. Finally, multi-scale fusion is completed through fast spatial pyramid pooling, and the output is clothing features with unified dimensions, rich details, and high sensitivity to color and texture, providing a reliable basis for subsequent classification and positioning.

[0049] Step S302: Perform multi-scale fusion on clothing features through a feature fusion layer to obtain fused features.

[0050] It should be noted that the feature fusion layer aims to integrate feature maps from different depths and resolutions to capture more comprehensive target information. This process usually includes operations such as upsampling, downsampling, and element-by-element addition or concatenation.

[0051] Furthermore, step S302 includes: downsampling the clothing features through the first downsampling path of the PAN-FPN structure to obtain first downsampled features. Specifically, through the first downsampling path of the PAN-FPN structure, a 3×3 convolution with a stride of 2 is performed on the clothing features to reduce the spatial size by half and simultaneously increase the channel dimension, thereby obtaining first downsampled features that retain edge and contour details. The first downsampled features are then sequentially subjected to the first convolution and second convolution of the second C2f module to obtain first cross-stage features. Specifically, the first downsampled features are fed into the second C2f module, where a 1×1 first convolution compresses the channels and reduces redundant parameters. A 3×3 second convolution then expands the receptive field while capturing local spatial relationships such as trouser folds and shoe texture. Finally, the compressed features are fused with the original features via a residual connection to output first cross-stage features rich in semantics and details. The first cross-stage features are then upsampled using the nearest neighbor interpolation method through the first upsampling path of the PAN-FPN structure to obtain the first upsampled features. Specifically, in the first upsampling path of the PAN-FPN structure, the first cross-stage features are upsampled by a factor of 2 using the nearest neighbor interpolation method, restoring the spatial resolution and introducing high-level semantics. This results in first upsampled features that are consistent with the scale of the original clothing features, thereby bridging the information gap between deep and shallow layers. The first upsampled features are then processed by the second cross-stage portion of the second C2f module to obtain second cross-stage features. Specifically, the second cross-stage portion of the C2f module performs a cross-stage partial connection on the first upsampled features: the features are first split using a 1×1 convolution and fed into two Bottleneck branches for deep feature extraction. These features are then concatenated with the main path features and fused using a 1×1 convolution to obtain second cross-stage features that balance color consistency and texture integrity. Finally, the second cross-stage features are downsampled through the second downsampling path of the PAN-FPN structure to obtain second downsampled features, which are then concatenated with the clothing features to obtain fused features. Specifically, a 3×3 convolution with a stride of 2 is performed on the second cross-stage features through the second downsampling path of the PAN-FPN structure, halving the feature size and increasing the channel dimension. This generates second downsampled features that carry contextual information while suppressing background noise. These second downsampled features are then concatenated with the clothing features in the channel dimension to form fused features that combine high-resolution details, medium-resolution textures, and low-resolution semantics. This provides the subsequent detection head with a multi-scale representation that combines both positioning accuracy and semantic discriminability, significantly improving the recognition accuracy of clothing categories, colors, and styles in complex scenes.

[0052] Step S303: The detection head classifies and regresses the fused features to obtain clothing recognition results.

[0053] It’s important to note that a detection head typically consists of two main components: a classification branch and a regression branch. The classification branch is responsible for determining the category of the object within each detection box (e.g., clothing type, color, etc.), while the regression branch focuses on adjusting the position and size of the detection box to more accurately enclose the target object.

[0054] Furthermore, step S303 includes: performing preliminary feature transformation on the fused features through the detection head to obtain initial detection features. Specifically, the detection head performs preliminary feature transformation on the fused features: first, using 1×1 convolution to unify the channel dimension to a preset value (e.g., 256), and then performing batch normalization and SiLU activation to obtain highly expressive initial detection features while maintaining gradient stability. Then, the initial detection features are enhanced through the global attention mechanism module to obtain enhanced detection features. Specifically, the initial detection features are sent to the global attention mechanism module, which first performs global average pooling along the channel dimension to capture global color distribution and style statistics. After 1×1 convolution compression, the module generates spatial-channel joint weights through Sigmoid. The weights are multiplied element-by-element with the original features to achieve global context enhancement, outputting enhanced detection features with prominent details and consistent semantics. Next, the enhanced detection features are fed into the classification branch and regression branch, respectively, to obtain a class probability distribution and target location information. The class probability distribution is based on color and style categories. Specifically, the enhanced detection features are fed into three decoupled branches in parallel: the classification branch outputs a color class probability distribution (red, blue, black, etc.) and a style class probability distribution (windbreaker, jeans, sneakers, etc.) through a 1×1 convolution followed by a softmax operation. The regression branch outputs the center point offset and width and height offsets through a 1×1 convolution followed by a Dependent Fluent Loss (DFL) loss, obtaining the target location information. Finally, the class probability distribution and target location information are fused to obtain the clothing detection result, which is then processed using a non-maximum suppression algorithm to obtain the clothing recognition result. Specifically, the color category probability distribution, style category probability distribution and target position information are fused: first, the three types of outputs are spliced ​​in the channel dimension, and then a weighted fusion layer is used to generate a comprehensive score according to the confidence threshold (0.5) to form a clothing detection result containing a "color-style-frame" triplet. The clothing detection result is then processed by the non-maximum suppression algorithm: the overlapping frames are filtered out with the IoU threshold (0.45), and the candidate frame with the highest confidence is retained. Finally, the structured clothing recognition result is output - including the attributes of the top "navy blue - windbreaker", the pants "light blue - jeans", the shoes "white - sneakers" and the corresponding pixel-level bounding boxes, achieving millisecond-level accurate analysis of single-frame images.

[0055] This embodiment uses a feature fusion layer to fuse clothing features at multiple scales, integrating feature maps at different depths to capture comprehensive target information and generate fused features. The detection head then performs classification and regression on these fused features, enabling accurate recognition of pedestrian clothing, including detailed information such as clothing type, color, and style. This improves the accuracy and granularity of clothing recognition and enhances the model's robustness and adaptability in complex scenarios.

[0056] Based on the first embodiment of the present application, the present application also provides a pedestrian clothing recognition device based on the improved YOLOv8 model, please refer to Figure 5 , the device comprises: The acquisition module 10 is used to acquire pedestrian image data.

[0057] The processing module 20 is used to pre-process the image data to obtain processed image data.

[0058] The result module 30 is used to input the processed image data into the preset clothing recognition model for recognition to obtain the clothing recognition result. The preset clothing recognition model is an improved YOLOv8 model. The improved YOLOv8 model includes a backbone network layer, a feature fusion layer and a detection head. The backbone network layer is composed of a CSPDarknet53 architecture, a coordinate attention module, an SCConv module and a Focus module. The CSPDarknet53 architecture adopts a first C2f module and a maximum pooling layer in series. The coordinate attention module is inserted into the first preset sublayer, the second preset sublayer and the third preset sublayer of the backbone network layer. The first C2f module includes two Bottleneck layers, each layer consists of a first convolution and a second convolution, and the Focus module consists of a third convolution; the feature fusion layer consists of a PAN-FPN structure and a second C2f module. The PAN-FPN structure and the second C2f module are both bidirectional cross-layer connections; the detection head is used to perform target detection and classification. The detection head is divided into a classification branch and a regression branch. The classification branch and the regression branch are both one-dimensional convolutions. The classification branch also includes a color classification branch and a style classification branch.

[0059] The pedestrian clothing recognition device based on the improved YOLOv8 model provided in this application, which employs the pedestrian clothing recognition method based on the improved YOLOv8 model in the above-mentioned embodiments, can solve the technical problem of how to improve the accuracy of pedestrian clothing recognition in complex scenarios. Compared with the prior art, the beneficial effects of the pedestrian clothing recognition device based on the improved YOLOv8 model provided in this application are the same as the beneficial effects of the pedestrian clothing recognition method based on the improved YOLOv8 model provided in the above-mentioned embodiments. Other technical features of the pedestrian clothing recognition device based on the improved YOLOv8 model are the same as those disclosed in the above-mentioned embodiments and are not further described here.

[0060] In one embodiment, the processing module 20 is further used to perform feature extraction on the processed image data through the backbone network layer to obtain clothing features; perform multi-scale fusion of the clothing features through the feature fusion layer to obtain fusion features; and perform classification and regression on the fusion features through the detection head to obtain clothing recognition results.

[0061] In one embodiment, the processing module 20 is further used to perform a slicing operation on the processed image data through the Focus module, and use the third convolution to perform feature extraction to obtain an initial feature map; the initial feature map is input into the CSPDarknet53 architecture and the first Bottleneck layer of the C2f module is used for feature extraction to obtain a first intermediate feature map, wherein the first Bottleneck layer is sequentially subjected to a first convolution and a second convolution for feature transformation to extract features; the first intermediate feature map is input into the coordinate attention module and enhanced through a position-sensitive feature enhancement mechanism to obtain a second intermediate feature map; the second intermediate feature map is input into the CSPDarknet53 architecture and the second Bottleneck layer of the C2f module is used for feature extraction to obtain a third intermediate feature map, wherein the second Bottleneck layer is sequentially subjected to a first convolution and a second convolution for feature transformation to extract features; the third intermediate feature map is sequentially subjected to a preset number of maximum pooling layers for multi-scale feature fusion to obtain a fourth intermediate feature map; the fourth intermediate feature map is input into the SCConv module for fast spatial pyramid pooling processing to obtain clothing features.

[0062] In one embodiment, the processing module 20 is further used to downsample the clothing feature through the first downsampling path of the PAN-FPN structure to obtain a first downsampled feature; the first downsampled feature is sequentially operated through the first convolution and the second convolution of the C2f module to obtain a first cross-stage feature; the first cross-stage feature is upsampled using the nearest neighbor interpolation method through the first upsampling path of the PAN-FPN structure to obtain a first upsampled feature; the first upsampled feature is processed through the second cross-stage part of the C2f module to obtain a second cross-stage feature; the second cross-stage feature is downsampled through the second downsampling path of the PAN-FPN structure to obtain a second downsampled feature; and the second downsampled feature is spliced ​​with the clothing feature to obtain a fusion feature.

[0063] In one embodiment, the processing module 20 is further used to perform preliminary feature transformation on the fused features through the detection head to obtain initial detection features; perform feature enhancement on the initial detection features through the global attention mechanism module to obtain enhanced detection features; input the enhanced detection features into the classification branch and the regression branch respectively to obtain category probability distribution and target position information, wherein the category probability distribution is obtained based on the color category and the style category; fuse the category probability distribution and the target position information to obtain a clothing detection result; and process the clothing detection result through a non-maximum suppression algorithm to obtain a clothing recognition result.

[0064] In one embodiment, the processing module 20 is further used to obtain historical image data; annotate the historical image data to generate a training set and a test set; construct an initial clothing recognition model; and train the clothing recognition model using the training set to obtain a preset clothing recognition model.

[0065] In one embodiment, the processing module 20 is also used to initialize the weights and bias parameters of the initial clothing recognition model; input the image data in the training set into the initial clothing recognition model for calculation to obtain the predicted position and category probability; calculate the error between the predicted position and the true position according to the first loss function to obtain the distribution focus loss, and the true position is calculated according to the external matrix; calculate the cross entropy loss of the category probability and the true category label according to the second loss function to obtain the category probability loss, and the true category label is obtained through a preset coding table; process according to the distribution focus loss and the category probability loss to obtain a total error value; calculate through the back propagation algorithm to obtain the gradient of the weight and bias parameters; iteratively update the weight and bias parameters through the optimization algorithm according to the gradient until the maximum number of iterations is reached or the total error value converges to a preset threshold, thereby obtaining a preset clothing recognition model.

[0066] The present application provides a pedestrian clothing recognition device based on an improved YOLOv8 model. The pedestrian clothing recognition device based on the improved YOLOv8 model includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the pedestrian clothing recognition method based on the improved YOLOv8 model in the above-mentioned embodiment one.

[0067] Reference below Figure 6, which shows a schematic structural diagram of a pedestrian clothing recognition device based on an improved YOLOv8 model suitable for implementing embodiments of the present application. The pedestrian clothing recognition device based on the improved YOLOv8 model in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The pedestrian clothing recognition device based on the improved YOLOv8 model shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0068] like Figure 6 As shown, a pedestrian clothing recognition device based on an improved YOLOv8 model may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. RAM 1004 also stores various programs and data required for the operation of the pedestrian clothing recognition device based on the improved YOLOv8 model. Processing device 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, hard disk, etc.; and communication devices 1009. Communication devices 1009 can allow the pedestrian clothing recognition device based on the improved YOLOv8 model to communicate wirelessly or wired with other devices to exchange data. Although various pedestrian clothing recognition devices based on the improved YOLOv8 model are shown in the figure, it should be understood that not all of the illustrated devices are required to be implemented or provided. More or fewer devices may be implemented or provided instead.

[0069] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0070] The pedestrian clothing recognition device based on the improved YOLOv8 model provided in this application, which employs the pedestrian clothing recognition method based on the improved YOLOv8 model in the above-mentioned embodiment, can solve the technical problem of how to improve the accuracy of pedestrian clothing recognition in complex scenarios. Compared with the prior art, the beneficial effects of the pedestrian clothing recognition device based on the improved YOLOv8 model provided in this application are the same as the beneficial effects of the pedestrian clothing recognition method based on the improved YOLOv8 model provided in the above-mentioned embodiment. The other technical features of the pedestrian clothing recognition device based on the improved YOLOv8 model are the same as those disclosed in the above-mentioned embodiment and are not further described here.

[0071] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0072] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0073] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, and the computer-readable program instructions are used to execute the pedestrian clothing recognition method based on the improved YOLOv8 model in the above embodiment.

[0074] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible storage medium that contains or stores a program that can be executed by instructions, used by a device, or used in conjunction with it. The program code contained on the computer-readable storage medium may be transmitted using any suitable storage medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0075] The above-mentioned computer-readable storage medium may be included in the pedestrian clothing recognition device based on the improved YOLOv8 model; or it may exist independently without being assembled into the pedestrian clothing recognition device based on the improved YOLOv8 model.

[0076] The computer-readable storage medium carries one or more programs. When executed by a pedestrian clothing recognition device based on the improved YOLOv8 model, the one or more programs enable the pedestrian clothing recognition device based on the improved YOLOv8 model to write computer program code for performing the operations of the present application in one or more programming languages, or a combination thereof. The programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0077] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based implementation that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0078] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0079] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned method for pedestrian clothing recognition based on the improved YOLOv8 model. This computer-readable storage medium addresses the technical problem of improving the accuracy of pedestrian clothing recognition in complex scenarios. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are similar to those of the method for pedestrian clothing recognition based on the improved YOLOv8 model provided in the aforementioned embodiments, and are not further elaborated here.

[0080] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned pedestrian clothing recognition method based on the improved YOLOv8 model.

[0081] The computer program product provided in this application can solve the technical problem of improving the accuracy of pedestrian clothing recognition in complex scenarios. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the pedestrian clothing recognition method based on the improved YOLOv8 model provided in the above embodiment, and will not be elaborated here.

[0082] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A pedestrian clothing recognition method based on an improved YOLOv8 model, characterized in that: include: Obtain pedestrian image data; Preprocessing the image data to obtain processed image data; The processed image data is input into a preset clothing recognition model for recognition to obtain a clothing recognition result. The preset clothing recognition model is an improved YOLOv8 model. The improved YOLOv8 model includes a backbone network layer, a feature fusion layer and a detection head. The backbone network layer is composed of a CSPDarknet53 architecture, a coordinate attention module, an SCConv module and a Focus module. The CSPDarknet53 architecture is composed of a first C2f module and a maximum pooling layer in series. The coordinate attention module is inserted into the first preset sublayer, the first sublayer and the second sublayer of the backbone network layer. Two preset sublayers and a third preset sublayer, the first C2f module includes two Bottleneck layers, each layer consists of a first convolution and a second convolution, and the Focus module consists of a third convolution; the feature fusion layer is composed of a PAN-FPN structure and a second C2f module, and the PAN-FPN structure and the second C2f module are both bidirectional cross-layer connections; the detection head is used for target detection and classification, and the detection head is divided into a classification branch and a regression branch, and the classification branch and the regression branch are both one-dimensional convolutions, and the classification branch also includes a color classification branch and a style classification branch.

2. The method according to claim 1, wherein The step of inputting the processed image data into a preset clothing recognition model for recognition to obtain a clothing recognition result includes: Performing feature extraction on the processed image data through the backbone network layer to obtain clothing features; Performing multi-scale fusion on the clothing features through the feature fusion layer to obtain fusion features; The detection head classifies and regresses the fused features to obtain clothing recognition results.

3. The method according to claim 2, wherein The step of extracting features from the processed image data through the backbone network layer to obtain clothing features includes: Slicing the processed image data through the Focus module and performing feature extraction using the third convolution to obtain an initial feature map; Input the initial feature map into the CSPDarknet53 architecture and perform feature extraction through the first Bottleneck layer of the first C2f module to obtain a first intermediate feature map, wherein the first Bottleneck layer performs feature transformation and feature extraction through the first convolution and the second convolution in sequence; Input the first intermediate feature map into the coordinate attention module and enhance it through the position-sensitive feature enhancement mechanism to obtain the second intermediate feature map; Input the second intermediate feature map into the CSPDarknet53 architecture and perform feature extraction through the second Bottleneck layer of the first C2f module to obtain a third intermediate feature map, wherein the second Bottleneck layer performs feature transformation and feature extraction through the first convolution and the second convolution in sequence; Passing the third intermediate feature map through a preset number of maximum pooling layers in sequence to perform multi-scale feature fusion to obtain a fourth intermediate feature map; The fourth intermediate feature map is input into the SCConv module for fast spatial pyramid pooling processing to obtain clothing features.

4. The method according to claim 2, wherein The step of performing multi-scale fusion on the clothing features through the feature fusion layer to obtain fusion features includes: Downsampling the clothing feature through a first downsampling path of a PAN-FPN structure to obtain a first downsampling feature; The first down-sampled features are sequentially subjected to a first convolution and a second convolution of a second C2f module to obtain a first cross-stage feature; Upsampling the first cross-stage feature using a nearest neighbor interpolation method through a first upsampling path of the PAN-FPN structure to obtain a first upsampled feature; Processing the first upsampled features by a second cross-stage part of a second C2f module to obtain a second cross-stage feature; Downsampling the second cross-stage feature through a second downsampling path of the PAN-FPN structure to obtain a second downsampled feature; The second down-sampled feature is concatenated with the clothing feature to obtain a fused feature.

5. The method according to claim 2, wherein The step of classifying and regressing the fused features by the detection head to obtain clothing recognition results includes: Performing preliminary feature transformation on the fusion feature by the detection head to obtain an initial detection feature; Performing feature enhancement on the initial detection feature through a global attention mechanism module to obtain an enhanced detection feature; The initial detection features are input into the classification branch and the regression branch respectively to obtain the category probability distribution and target position information, wherein the category probability distribution is obtained based on the color category and the style category: fusing the category probability distribution and the target position information to obtain a clothing detection result; The clothing detection result is processed by a non-maximum suppression algorithm to obtain a clothing recognition result.

6. The method according to claim 1, wherein Before the step of inputting the processed image data into a preset clothing recognition model for recognition to obtain a clothing recognition result, the method includes: Acquire historical image data; Annotating the historical image data to generate a training set and a test set; Build an initial clothing recognition model; The clothing recognition model is trained using the training set to obtain a preset clothing recognition model.

7. The method according to claim 6, wherein The step of using the training set to train the clothing recognition model to obtain a preset clothing recognition model includes: Initializing weight and bias parameters of the initial clothing recognition model; Inputting the image data in the training set into the initial clothing recognition model for calculation to obtain predicted position and category probability; Calculating the error between the predicted position and the actual position according to a first loss function to obtain a distributed focus loss, wherein the actual position is calculated according to an external matrix; Calculating the cross entropy loss of the class probability and the true class label according to a second loss function to obtain the class probability loss, wherein the true class label is obtained by a preset coding table; Processing is performed according to the distribution focus loss and the category probability loss to obtain a total error value; Obtaining the gradients of the weight and bias parameters by back propagation algorithm calculation; The weight and bias parameters are iteratively updated according to the gradient through an optimization algorithm until a maximum number of iterations is reached or the total error value converges to a preset threshold, thereby obtaining a preset clothing recognition model.

8. A pedestrian clothing recognition device based on an improved YOLOv8 model, characterized in that: The device comprises: An acquisition module, used to acquire pedestrian image data; A processing module, configured to pre-process the image data to obtain processed image data; The result module is used to input the processed image data into a preset clothing recognition model for recognition to obtain clothing recognition results. The preset clothing recognition model is an improved YOLOv8 model. The improved YOLOv8 model includes a backbone network layer, a feature fusion layer and a detection head. The backbone network layer is composed of a CSPDarknet53 architecture, a coordinate attention module, an SCConv module and a Focus module. The CSPDarknet53 architecture adopts a first C2f module and a maximum pooling layer in series. The coordinate attention module is inserted into the first preset sub-layer of the backbone network layer. layer, a second preset sublayer and a third preset sublayer, the first C2f module includes two Bottleneck layers, each layer consists of a first convolution and a second convolution, and the Focus module consists of a third convolution; the feature fusion layer is composed of a PAN-FPN structure and a second C2f module, and the PAN-FPN structure and the second C2f module are both bidirectional cross-layer connections; the detection head is used for target detection and classification, and the detection head is divided into a classification branch and a regression branch, and the classification branch and the regression branch are both one-dimensional convolutions, and the classification branch also includes a color classification branch and a style classification branch.

9. A pedestrian clothing recognition device based on an improved YOLOv8 model, characterized in that: The device includes: a memory, a processor, and a pedestrian clothing recognition program based on an improved YOLOv8 model stored in the memory and running on the processor, wherein the pedestrian clothing recognition program based on the improved YOLOv8 model is configured to implement the steps of the pedestrian clothing recognition method based on the improved YOLOv8 model as described in any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium stores a pedestrian clothing recognition program based on the improved YOLOv8 model. When the pedestrian clothing recognition program based on the improved YOLOv8 model is executed by the processor, the steps of the pedestrian clothing recognition method based on the improved YOLOv8 model as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Clothes identification method and device, equipment and medium

    CN119006874A

  • Real-time auricle identification and auricular point positioning method and system for multi-task feature sharing

    CN119523799A

  • Ship target detection method, device and equipment and storage medium

    CN119559374A

  • Remote sensing image small target detection method and device, medium and product

    CN119600436A

  • Image detection method based on improved YOLO algorithm

    CN120088208A

Cited By

  • Cigarette case specification identification method and device based on deep learning, and electronic equipment

    CN121747083A