Multi-scene target identification and detection method based on improved YOLOv5
By building multi-scene data sets and introducing adaptive reward function mechanisms, the YOLOv5 model is optimized, and the problem of degradation in multi-scene object detection performance in the existing technology is solved, and efficient and accurate multi-scene object detection is achieved.
Patent Information
- Application Number
- CN202510295744.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
The existing object detection methods lack dynamic perception and adaptive adjustment capabilities in multiple scenarios, resulting in a decline in detection performance in different environments and the inability to take into account the needs of multiple scenarios.
By constructing a multi-scene data set, a YOLOv5 multi-scene object recognition detection model is established, and an adaptive reward function mechanism is introduced, the YOLOv5 model is optimized, the model detection strategy is dynamically adjusted, and the object detection effect of different scenarios is achieved.
The adaptability of the YOLOv5 model in multiple scenarios and different stages is improved, efficient and accurate object detection is achieved, and the needs of multiple scenarios can be taken into account in real-time detection.
Smart Images

Figure CN120219964A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and specifically, it is a multi-scenario target recognition and detection method based on improved YOLOv5. Background Technique
[0002] With the rapid development of artificial intelligence technology, target detection algorithms have been widely applied in many fields, such as industrial defect detection, robot vision recognition, environmental monitoring, etc. The YOLOv5 algorithm has attracted much attention for its fast speed and high accuracy. The traditional YOLOv5 model performs well in a single scenario, but its generalization ability in multi-scenarios is limited. The lighting conditions, obstacle density, texture complexity, etc. in different scenarios vary greatly, resulting in a decline in the performance of the model when applied across scenarios. For example, in photovoltaic defect detection, the detection ability of small targets and distant targets needs to be improved; in service robot item recognition, real-time performance and accuracy need to be balanced; in water surface floating garbage tracking, the adaptability to complex scenarios is insufficient; in cotton apical bud detection, lightweight and anti-occlusion capabilities need to be enhanced. This makes the existing target detection methods lack the ability of dynamic perception and adaptive adjustment of scene features, difficult to maintain high-precision detection effects in different environments, unable to simultaneously meet the detection requirements of multiple scenarios in real time, and difficult to achieve efficient and accurate target detection.
[0003] Based on this, a multi-scenario target recognition and detection method based on improved YOLOv5 is now provided, which can eliminate the drawbacks of the existing technical solutions. Summary of the Invention
[0004] The purpose of the present invention is to provide a multi-scenario target recognition and detection method based on improved YOLOv5 to solve the problem that the existing target detection methods in the background technique lack the ability of dynamic perception and adaptive adjustment of scene features.
[0005] To achieve the above purpose, the present invention provides the following technical solutions:
[0006] A multi-scenario target recognition and detection method based on improved YOLOv5, and the specific use steps are as follows:
[0007] S1. Construct a multi-scenario data set and label the data set. The data set includes but is not limited to industrial defect detection categories, robot vision recognition categories, environmental monitoring categories, and is subdivided according to the task progress and environmental complexity of different scenarios;
[0008] S2. Establish a YOLOv5 multi-scenario target recognition and detection model, and use the above data set to train the YOLOv5 model to obtain a model capable of recognizing scene types;
[0009] S3. Introduce an adaptive reward function mechanism, integrate the adaptive reward function mechanism with the YOLOv5 object detection algorithm, optimize the YOLOv5 model, and automatically adjust the model detection strategy through a weight adjustment module according to the parameter characteristics of different scenarios;
[0010] S4. Set up a multi-scenario object recognition and detection device based on the improved YOLOv5. The detection device integrates functions such as picture detection, video detection, and real-time camera monitoring. Deploy the optimized YOLOv5 model inside the detection device to achieve the object detection effect for different scenarios.
[0011] Preferably, the construction of the multi-scenario dataset in step S1 includes the following steps:
[0012] S11. Collect the original picture materials of different scenarios, and perform data screening operations through Python to eliminate the pictures that do not meet the requirements;
[0013] S12. Perform data preprocessing operations on the pictures to extract the key scene parameter characteristics in the pictures. The scene parameter characteristics include lighting conditions, obstacle density, texture complexity, color distribution, and target size distribution;
[0014] S13. Mark the pictures through the LabelImg tool and establish corresponding labels. The labels include but are not limited to target categories, position coordinates, environmental complexity, and task progress;
[0015] S14. Input all the pictures according to the type of the dataset, review the marked pictures and the corresponding labels, and ensure that the picture materials are consistent with the dataset.
[0016] Preferably, step S2 is specifically as follows:
[0017] S21. Add a scene recognition branch on the basis of the YOLOv5 model. The input is the scene feature vector, and the output is the scene type;
[0018] S22. Divide the above dataset into a training set, a validation set, and a test set in a ratio of 70%:20%:10%, and uniformly adjust the sizes of the pictures in the dataset;
[0019] S23. Set training parameters. The parameters include but are not limited to the learning rate, the number of training rounds, and the batch size. Adopt a transfer learning strategy to initialize the parameters of the YOLOv5 model;
[0020] S24. For the scene recognition task, use the cross-entropy loss function to calculate the difference between the predicted scene category and the true scene category. For the object detection task, use the cross-entropy loss function and the mean squared error loss function to calculate the difference between the predicted bounding box and the true bounding box;
[0021] S25. Evaluate the YOLOv5 model using the validation set, adjust the training parameters according to the evaluation metrics, and prevent the model from overfitting;
[0022] S26. Evaluate multiple trained models using the test set to obtain the YOLOv5 multi-scene object recognition and detection model with the optimal comprehensive performance in object detection and scene recognition tasks.
[0023] Preferably, the specific steps of step S3 are as follows:
[0024] S31. Design an adaptive reward function using the scene parameter features and the environmental complexity;
[0025] S32. Incorporate the adaptive reward function into the calculation of the loss function, calculate the reward value according to the parameter features, and combine it with the loss function to obtain a new comprehensive loss function;
[0026] S33. Design a weight adjustment module, and obtain the parameter features and task progress of different scenes in real time through the weight adjustment module, and dynamically adjust the weight of the comprehensive loss function.
[0027] Preferably, the adaptive reward function in step S31 includes a target completion reward, a safety reward, and an agent behavior efficiency reward. The target completion reward is used to measure the degree of task completion of the agent, including the number of detected targets and the target distance. The safety reward is used to measure the safety performance of the agent during the task completion process. The agent behavior efficiency reward is used to measure the behavior efficiency of the agent, and the behavior efficiency includes detection speed, detection accuracy, and resource utilization rate.
[0028] Preferably, the relational expression of the adaptive reward function is as follows:
[0029] R = w1·R target + w2·R safety + w3·R efficiency
[0030] where R is the adaptive reward function, R target is the target completion reward, R safety is the safety reward, R efficiency is the agent behavior efficiency reward, and w1, w2, and w3 are dynamic weights.
[0031] Preferably, the adaptive reward function can dynamically adjust the weights according to the environmental complexity, task progress, and agent behavior:
[0032] Weight adjustment based on environmental complexity: Assume that the environmental complexity is C env , and the environmental complexity threshold is T env, if the target reward weight is w1 and the safety reward weight is w2, then:
[0033]
[0034] In a high-complexity environment, increase the safety reward weight, adjust w2 to w2″, decrease the target reward weight, adjust w1 to w1″; in a low-complexity environment, increase the target reward weight, adjust w1 to w1′, and decrease the safety reward weight, adjust w2 to w2′;
[0035] Weight adjustment based on task progress: Assume the distance between the agent and the task goal is d, and the target distance threshold is T d , the target reward weight is w1, and the safety reward weight is w2, then:
[0036] When d < T d At this time, w1 = w1″′, w2 = w2″′. When the distance between the agent and the task goal is small, that is, when the task is approaching completion, increase the target reward weight, adjust w1 to w1″′, and decrease the safety reward weight, adjust w2 to w2″′;
[0037] Weight adjustment based on agent behavior: Assume the success rate of the agent's behavior is SR, and the success rate threshold is T SR , the efficiency reward weight is w3, then:
[0038]
[0039] When the success rate of the agent's behavior is low, increase the efficiency reward weight, adjust w3 to w3′; when the success rate of the agent's behavior is high, decrease the efficiency reward weight, adjust w3 to w3″.
[0040] Preferably, the relational expression of the comprehensive loss function in step S32 is as follows:
[0041] L = w4·L cls + w5·L loc + w6·L conf + α·R
[0042] Where L is the comprehensive loss function, L cls is the classification loss, L loc is the localization loss, L conf is the confidence loss, w4, w5, w6 are the original loss function weights of the YOLOv5 model, which are fixed values, and α is the weight coefficient of the adaptive reward function.
[0043] Preferably, the weight adjustment module in the step S33 is used to obtain the parameter features in the pictures of different scenarios and the task progress in different task objectives in real time, calculate the weights of each loss term based on the above parameter features and task progress, and apply the dynamic weights to the comprehensive loss function.
[0044] Preferably, the multi-scenario target recognition and detection device based on the improved YOLOv5 includes:
[0045] A processor, which is used to implement the above-mentioned multi-scenario target recognition and detection method based on the improved YOLOv5 when executing a computer program;
[0046] A storage, which is used to store the computer program.
[0047] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0048] The present invention provides a multi-scenario target recognition and detection method based on the improved YOLOv5. Considering the requirements of different application scenarios, it is pre-trained on a data set containing various scenario data, and then fine-tuned for specific tasks in different scenarios, enabling the model to quickly adapt to new scenarios, reducing training time and data requirements. A weight adjustment module is set up, which can perceive the key parameter features of the current scenario in real time and dynamically adjust the weights of the loss function according to the environmental complexity, task progress and agent behavior, enabling the YOLOv5 model to better adapt to different scenarios and stages. The improved YOLOv5 model can take into account the requirements of multiple scenarios in real-time detection and achieve efficient and accurate target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a flowchart of the present invention.
[0050] Figure 2 is a flowchart of step S1 of the present invention.
[0051] Figure 3 is a flowchart of step S2 of the present invention.
[0052] Figure 4 is a flowchart of step S3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0054] In this embodiment, as Figures 1 - 4 shown, a multi-scenario target recognition and detection method based on the improved YOLOv5 is as follows:
[0055] S1. Build a multi-scenario dataset and label the dataset, which includes but is not limited to industrial defect detection, robot vision recognition, and environmental monitoring, and is subdivided according to the task progress and environmental complexity of different scenarios;
[0056] S2. Establish a YOLOv5 multi-scenario object recognition and detection model, and use the above dataset to train the YOLOv5 model to obtain a model capable of recognizing scenario types;
[0057] S3. Introduce an adaptive reward function mechanism, integrate the adaptive reward function mechanism with the YOLOv5 object detection algorithm, optimize the YOLOv5 model, and automatically adjust the model detection strategy through a weight adjustment module according to the parameter characteristics of different scenarios;
[0058] S4. Set up a multi-scenario object recognition and detection device based on the improved YOLOv5. The detection device integrates functions such as image detection, video detection, and camera real-time monitoring. Deploy the optimized YOLOv5 model inside the detection device to achieve the object detection effect for different scenarios;
[0059] Specifically, in step S1, first determine the data collection scope: for industrial defect detection, such as different types of factory workshops like electronic manufacturing factories and machining workshops, collect images of circuit board solder joint defects and chip surface scratches of electronic products, and images of gear wear and bearing cracks of mechanical parts. The collected images are targeted; for robot vision recognition, simulate the working scenarios of robots in different environments, such as warehouse goods handling and indoor navigation and obstacle avoidance, and collect images including but not limited to the appearance, shape, and placement position of goods, as well as obstacles and signs in the indoor environment; for environmental monitoring, collect images of different natural and urban environments, such as forests, rivers, mountains, vegetation coverage, traffic congestion, and crowd congestion.
[0060] Then use devices such as high-resolution industrial cameras, ordinary digital cameras, and surveillance cameras to collect images, and take pictures at different times (day and night) and weather conditions (sunny, rainy, cloudy) to extract key images and increase the data volume.
[0061] Finally, subdivide according to the task progress and environmental complexity. For example, for industrial defect detection, divide the task progress according to different stages of the production process (raw material inspection, semi-finished product processing, finished product assembly), and classify the environmental complexity according to factors such as light conditions and equipment layout density. For example, an area with dim light and dense equipment layout is a high-complexity environment, and an area with sufficient light and simple equipment layout is a low-complexity environment.
[0062] Such as Figures 1 - 2As shown, building a multi-scenario dataset in step S1 includes the following steps:
[0063] S11. Collect original image materials of different scenes, and perform data screening operations through Python to eliminate images that do not meet the requirements;
[0064] S12, performing data preprocessing operations on the image to extract key scene parameter features in the image, the scene parameter features including lighting conditions, obstacle density, texture complexity, color distribution, and target size distribution;
[0065] S13. Label the images using the LabelImg tool and create corresponding labels, including but not limited to target category, location coordinates, environment complexity, and task progress;
[0066] S14, input all images according to the type of data set, review the marked images and corresponding labels, and ensure that the image materials are consistent with the data set;
[0067] Specifically, Python is used to read the basic information of the image, such as size, color mode, etc., to determine whether the width and height of the image meet the set minimum values, and to filter the images according to the set conditions. For example, if the image is not clear or has insufficient resolution, the images that do not meet the conditions will be eliminated.
[0068] Perform data enhancement operations, such as random flipping, rotation, scaling, adding noise, etc., to expand the diversity of the data set and improve the generalization ability of the model;
[0069] Normalize the pixel values of the image to the range of [0, 1] or [-1, 1]. According to the requirements of the specific task, extract the key features of the image, classify the collected images, and mark the scene category to which each image belongs. If it is a target detection or image segmentation task, the location and category of the target also need to be marked.
[0070] like Figures 1 - 3 As shown, step S2 is specifically as follows:
[0071] S21. Add a scene recognition branch based on the YOLOv5 model. The input is the scene feature vector and the output is the scene type. The scene recognition branch can map the features to different scene categories, such as industrial scenes, robot scenes, environmental monitoring scenes, etc., and the subdivision types under each scene by adding a fully connected layer and a Softmax layer based on the existing feature extraction layer of the model.
[0072] S22. Divide the above dataset into a training set, a validation set, and a test set with a ratio of 70%:20%:10%. For example, if there are 1000 pieces of data in the original dataset, belonging to three types: industrial defect detection, robot vision recognition, and environmental monitoring, with a quantity ratio of 3:2:1, then in the training set (700 pieces), the validation set (200 pieces), and the test set (100 pieces), the quantity ratio of samples of types A, B, and C should also maintain the ratio of 3:2:1 to ensure a reasonable distribution of various samples, avoid unbalanced samples in the test set, make the model training and evaluation more representative, uniformly adjust the size of the pictures in the dataset, adjust all pictures to the same size, and uniformly convert pictures in different formats to common formats such as JPEG, PNG, etc., to ensure that the color mode of the pictures is consistent, usually converted to the RGB mode to meet the requirements of model input, and ensure that each set contains pictures of various scenarios;
[0073] S23. Set training parameters, including but not limited to the learning rate, the number of training epochs, and the batch size. Adopt a transfer learning strategy to initialize the parameters of the YOLOv5 model. The model adopts an adaptive learning rate strategy (including a learning rate decay strategy), and the Adam optimization algorithm can be selected to adaptively adjust the learning rate of each parameter. The learning rate decays exponentially, using the formula lr = lr0×γ t , where lr0 is the initial learning rate, which can be set to 0.001, γ is the decay factor, set to 0.99, t is the number of training epochs, usually set to 200 - 500 epochs according to the dataset scale and model convergence situation, and the batch size can be set according to the hardware memory;
[0074] S24. For the scene recognition task, use the cross - entropy loss function to calculate the difference between the predicted scene category and the true scene category. For the object detection task, use the cross - entropy loss function and the mean squared error loss function to calculate the difference between the predicted bounding box and the true bounding box;
[0075] S25. Use the validation set to evaluate the YOLOv5 model, and adjust the training parameters according to the evaluation metrics to prevent model overfitting. Assume that during the training process, every 10 epochs, use the validation set to evaluate the model performance, calculate the mAP of object detection and the accuracy of scene recognition. If the performance of the validation set has not improved for 5 consecutive epochs, then reduce the learning rate;
[0076] S26. Use the test set to evaluate multiple training models to obtain the YOLOv5 multi-scene target recognition detection model with the best comprehensive performance in target detection and scene recognition tasks. Use the test set to evaluate the saved multiple models, calculate the comprehensive score of each model in target detection and scene recognition tasks (such as the weighted sum of target detection mAP and scene recognition accuracy), and select the model with the highest comprehensive score as the final multi-scene target recognition detection model for the deployment of subsequent detection equipment.
[0077] like Figures 1 - 4 As shown, step S3 is specifically as follows:
[0078] S31. Using scene parameter features and environment complexity, design an adaptive reward function. According to the threshold of environment complexity, divide the environment into simple and complex levels. In simple environments, such as open and evenly illuminated indoor scenes, target detection is relatively easy. Complex environments have complex obstacles and illumination changes. The above extracted scene parameter features are used as input.
[0079] S32, incorporating the adaptive reward function into the calculation of the loss function, calculating the reward value according to the parameter characteristics, and combining it with the loss function to obtain a new comprehensive loss function;
[0080] S33. Design a weight adjustment module, through which parameter characteristics and task progress of different scenarios are obtained in real time, and the weight of the comprehensive loss function is dynamically adjusted;
[0081] Specifically, the weight adjustment module in step S33 is used to obtain parameter features in different scene pictures and task progress in different task objectives in real time, and calculate the weights of each loss item based on the above parameter features and task progress, and apply the dynamic weights to the comprehensive loss function. Through the weight adjustment module, the weights of each loss item in the comprehensive loss function can be dynamically adjusted according to the scene parameter features and task progress. Parameters such as lighting conditions, obstacle density, texture complexity, etc. of different scenes are obtained through sensors or cameras, and the weights of target completion rewards, safety rewards and efficiency rewards are dynamically adjusted, thereby optimizing the performance of the model in multi-scene target recognition and detection tasks. It can effectively improve the adaptability and detection accuracy of the model, is suitable for complex and changeable actual application scenarios, and optimizes the performance of the model in multi-scene target recognition and detection tasks.
[0082] Among them Figure 4As shown, the adaptive reward function in step S31 includes a target completion reward, a safety reward, and an agent behavior efficiency reward. The target completion reward is used to measure the degree of task completion of the agent, including the number of detected targets and the target distance. It is calculated based on the number of detected targets and the target distance. For example, the more targets are detected, the higher the reward value; the closer the target distance, the higher the reward value. The safety reward is used to measure the safety performance of the agent during task completion, and is calculated by detecting whether the agent collides or enters a dangerous area. The agent behavior efficiency reward is used to measure the behavior efficiency of the agent. The behavior efficiency includes detection speed, detection accuracy, and resource utilization rate, and is calculated based on the detection speed and resource utilization rate. For example, the faster the detection speed, the higher the reward value; the lower the resource utilization rate, the higher the reward value.
[0083] The relational expression of the adaptive reward function is as follows:
[0084] R = w1·R target + w2·R safety + w3·R efficiency
[0085] where R is the adaptive reward function, R target is the target completion reward, R safety is the safety reward, R efficiency is the agent behavior efficiency reward, and w1, w2, and w3 are dynamic weights, which are adjusted according to the scene characteristics and task requirements;
[0086] Specifically, in multi-scene target recognition and detection, in the face of different scenes such as industry, robot vision, and environmental monitoring, the adaptive reward function can enable the model to quickly adapt to scene changes. For example, in the industrial scene, it focuses on the agent behavior efficiency reward, and in the robot vision scene, it focuses on the target completion reward, enabling the model to effectively learn in each scene, improve the detection performance. The adaptive reward function can reduce ineffective exploration by dynamically adjusting the reward, accelerate the model to converge to the optimal strategy, shorten the training time, and improve the training efficiency.
[0087] Among them, as Figure 1 shown, the adaptive reward function can dynamically adjust the weights according to the environmental complexity, task progress, and agent behavior. For example, in a complex environment such as a city street with a large number of vehicles, pedestrians, and traffic signs, there are more interferences and uncertainties, while in a simple environment like an empty warehouse, there are fewer interference factors. The target recognition in different scenes has different emphases, and the parameter characteristics have different emphases. Through the adaptive reward function, computing resources can be reasonably allocated to improve the overall task completion efficiency;
[0088] Weight adjustment based on environmental complexity: Assume the environmental complexity is C env , and the environmental complexity threshold is T env, if the target reward weight is w1 and the safety reward weight is w2, then:
[0089]
[0090] In a high-complexity environment, increase the safety reward weight, adjust w2 to w2″, decrease the target reward weight, adjust w1 to w1″. In a low-complexity environment, increase the target reward weight, adjust w1 to w1′, and decrease the safety reward weight, adjust w2 to w2′;
[0091] The environmental complexity is a key factor for dynamically adjusting the weights. If the environmental complexity exceeds the threshold, w2 can be adjusted to 0.8 and w1 to 0.2. At this time, the weights pay more attention to safety risk management. In a simple environment, the weights pay more attention to the completion of the target task. At this time, w1 can be set to 0.8 and w2 to 0.2:
[0092]
[0093] Weight adjustment based on the task progress: Assume the distance between the agent and the task goal is d, and the target distance threshold is T d , if the target reward weight is w1 and the safety reward weight is w2, then:
[0094] When d < T d , w1 = w1″′, w2 = w2″′. When the distance between the agent and the task goal is small, i.e., the task is approaching completion, increase the target reward weight, adjust w1 to w1″′, and decrease the safety reward weight, adjust w2 to w2″′;
[0095] When the agent is approaching the task goal, the weights w1 and w2 can be adjusted to 0.9 and 0.1. At this time, more emphasis is placed on the target reward weight, and completing the task becomes the top priority;
[0096] Weight adjustment based on the agent's behavior: Assume the success rate of the agent's behavior is SR, and the success rate threshold is T SR , if the efficiency reward weight is w3, then:
[0097]
[0098] When the success rate of the agent's behavior is low, increase the efficiency reward weight, adjust w3 to w3′. When the success rate of the agent's behavior is high, decrease the efficiency reward weight, adjust w3 to w3″;
[0099] The success rate of the agent's behavior is an important indicator to measure its behavior. If the success rate of the agent's behavior is lower than the success rate threshold, it means that more attention is paid to efficiency in the reward, and the efficiency reward weight w3 can be adjusted to 0.8. If the success rate of the agent's behavior is higher than the success rate threshold, it means that other aspects (such as goals and safety) may be more concerned, and the efficiency reward weight w3 can be adjusted to 0.2. Then:
[0100]
[0101] According to the task requirements, set priorities for different weight parameters. For example: in the security risk scenario, the security reward weight w2 has the highest priority, and in the task-critical scenario, the target reward weight w1 has the highest priority, ensuring that the weights of critical tasks are adjusted first to adapt to complex scenarios.
[0102] The relational expression of the comprehensive loss function in step S32 is as follows:
[0103] L = w4·L cls + w5·L loc + w6·L conf + α·R
[0104] Where L is the comprehensive loss function, L cls is the classification loss, L loc is the localization loss, L conf is the confidence loss, w4, w5, w6 are the weights of the original loss function of the YOLOv5 model, which are fixed values, and α is the weight coefficient of the adaptive reward function;
[0105] Specifically, after introducing the adaptive reward function, the formula of the comprehensive loss function needs to combine the adaptive reward function with the original loss function of YOLOv5 (classification loss, localization loss, confidence loss). The adaptive reward function optimizes the performance of the model in different scenarios by dynamically adjusting the weights. The classification loss is used to measure the prediction error of the target class, the localization loss is used to measure the prediction error of the target bounding box, the confidence loss is used to measure the confidence error of the target existence, and the adaptive reward function is used to dynamically adjust the learning strategy of the model;
[0106] α is a hyperparameter used to balance the weights between the original loss function of YOLOv5 and the adaptive reward function. When the target recognition pays more attention to the accuracy of target detection, a smaller α (such as 0.1) can be set. At this time, the original loss function of YOLOv5 has a greater impact on the total loss, and the model will pay more attention to the accuracy of target detection. If the target recognition pays more attention to task completion, safety and efficiency, a larger α (such as 0.5, 1.0) can be set. At this time, the adaptive reward function has a greater impact on the total loss, and the model will pay more attention to task completion, safety and efficiency. α can be dynamically adjusted according to the performance of the model and the scenario requirements.
[0107] Among them, as Figure 1 shown, the multi-scenario object recognition and detection device based on the improved YOLOv5 includes:
[0108] A processor, when executing a computer program, implements the above-mentioned multi-scenario object recognition and detection method based on the improved YOLOv5;
[0109] A storage, used to store the computer program;
[0110] Specifically, the detection device should also be equipped with a high-resolution image acquisition device, a data transmission interface, a software system, and a monitoring module integrating functions of picture detection, video detection, and camera real-time monitoring. The image acquisition device includes industrial cameras, network cameras, etc., to ensure that clear image data can be obtained. The data transmission interface facilitates connecting to external devices and data transmission. The software system should have an operation panel interface to facilitate operators to set parameters, manage tasks, and view results. The monitoring module supports uploading pictures for object detection;
[0111] Deploy the optimized YOLOv5 model to the software system of the detection device, optimize the model using an adaptive reward function, improve the inference speed of the model, obtain the running status and detection results of the device in real time through the network, discover and solve device failures in a timely manner, and achieve real-time object detection to adapt to the changing scenario requirements and improve the detection performance.
[0112] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application should not easily think of changes or substitutions, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A multi-scene target recognition and detection method based on improved YOLOv5, characterized in that: The specific steps are as follows: S1. Construct a multi-scenario dataset and label the dataset. The dataset includes but is not limited to industrial defect detection, robot visual recognition, and environmental monitoring, and is segmented according to the task progress and environmental complexity of different scenarios. S2. Establish a YOLOv5 multi-scene target recognition and detection model, and use the above data set to train the YOLOv5 model to obtain a model that can recognize scene types; S3, introduce the adaptive reward function mechanism, integrate the adaptive reward function mechanism with the YOLOv5 target detection algorithm, optimize the YOLOv5 model, and automatically adjust the model detection strategy through the weight adjustment module according to the parameter characteristics of different scenarios; S4. Set up a multi-scenario target recognition detection device based on improved YOLOv5, which integrates image detection, video detection and camera real-time monitoring functions, and deploys the optimized YOLOv5 model inside the detection device to achieve target detection effects in different scenarios.
2. A multi-scenario target recognition detection method based on improved YOLOv5 according to claim 1, characterized in that: The construction of the multi-scene dataset in step S1 comprises the following steps: S11. Collect original image materials of different scenes, and perform data screening operations through Python to eliminate images that do not meet the requirements; S12, performing data preprocessing operations on the image to extract key scene parameter features in the image, wherein the scene parameter features include lighting conditions, obstacle density, texture complexity, color distribution, and target size distribution; S13. Label the image using the LabelImg tool to create corresponding labels, including but not limited to target category, location coordinates, environment complexity, and task progress; S14. Input all images according to the type of data set, review the marked images and corresponding labels, and ensure that the image materials are consistent with the data set.
3. A multi-scenario target recognition detection method based on improved YOLOv5 according to claim 2, characterized in that: The step S2 is specifically as follows: S21, based on the YOLOv5 model, add a scene recognition branch, the input is the scene feature vector, and the output is the scene type; S22, dividing the above dataset into a training set, a validation set, and a test set in a ratio of 70%:20%:10%, and uniformly resizing the images in the dataset; S23, setting training parameters, including but not limited to learning rate, number of training rounds, batch size, using transfer learning strategy, and initializing parameters of the YOLOv5 model; S24. For scene recognition tasks, the cross entropy loss function is used to calculate the difference between the predicted scene category and the real scene category. For object detection tasks, the cross entropy loss function and the mean square error loss function are used to calculate the difference between the predicted box and the real box. S25. Use the validation set to evaluate the YOLOv5 model and adjust the training parameters according to the evaluation indicators to prevent the model from overfitting. S26. Use the test set to evaluate multiple training models and obtain the YOLOv5 multi-scene object recognition detection model with the best comprehensive performance in object detection and scene recognition tasks.
4. A multi-scenario target recognition and detection method based on improved YOLOv5 according to claim 3, characterized in that: The step S3 is specifically as follows: S31. Design an adaptive reward function using scene parameter characteristics and environmental complexity; S32, incorporating the adaptive reward function into the calculation of the loss function, calculating the reward value according to the parameter characteristics, and combining it with the loss function to obtain a new comprehensive loss function; S33. Design a weight adjustment module, through which the parameter characteristics and task progress of different scenarios are obtained in real time, and the weight of the comprehensive loss function is dynamically adjusted.
5. A multi-scenario target recognition and detection method based on improved YOLOv5 according to claim 4, characterized in that: The adaptive reward function in step S31 includes a goal completion reward, a safety reward and an agent behavior efficiency reward. The goal completion reward is used to measure the degree of completion of the agent's task, including the number of detected targets and the target distance. The safety reward is used to measure the safety performance of the agent during the task completion process. The agent behavior efficiency reward is used to measure the agent's behavior efficiency, and the behavior efficiency includes detection speed, detection accuracy and resource utilization.
6. A multi-scenario target recognition and detection method based on improved YOLOv5 according to claim 5, characterized in that: The relational expression of the adaptive reward function is as follows: R=w1·R target +w2·R safety +w3·R efficiency Where R is the adaptive reward function, R target Reward for goal completion, R safety For safety bonus, R efficiency is the agent behavior efficiency reward, w1, w2, w3 are dynamic weights.
7. The multi-scenario target recognition and detection method based on improved YOLOv5 according to claim 6, characterized in that: The adaptive reward function can dynamically adjust the weights according to the complexity of the environment, the progress of the task, and the behavior of the agent: Weight adjustment based on environmental complexity: Assume that the environmental complexity is C env , the environmental complexity threshold is T env , the target reward weight is w1, the safety reward weight is w2, then: In a high-complexity environment, increase the weight of the safety reward, adjust w2 to w2″, reduce the weight of the target reward, adjust w1 to w1″; in a low-complexity environment, increase the weight of the target reward, adjust w1 to w1′, reduce the weight of the safety reward, and adjust w2 to w2′; Weight adjustment based on task progress: Assume that the distance between the agent and the task goal is d, and the goal distance threshold is T d , the target reward weight is w1, the safety reward weight is w2, then: When d<T d When w1=w1″′, w2=w2″′, when the distance between the agent and the task goal is small, that is, when the task is close to completion, increase the target reward weight, adjust w1 to w1″′, reduce the safety reward weight, and adjust w2 to w2″′; Weight adjustment based on agent behavior: Assume that the success rate of agent behavior is SR, and the success rate threshold is T SR , the efficiency reward weight is w3, then: When the success rate of the agent's behavior is low, the efficiency reward weight is increased and w3 is adjusted to w3′. When the success rate of the agent's behavior is high, the efficiency reward weight is reduced and w3 is adjusted to w3″.
8. A multi-scenario target recognition and detection method based on improved YOLOv5 according to claim 7, characterized in that: The relational expression of the comprehensive loss function in step S32 is as follows: L=w4·L cls +w5·L loc +w6·L conf +α·R Where L is the comprehensive loss function, L cls is the classification loss, L loc is the positioning loss, L conf is the confidence loss, w4, w5, w6 are the original loss function weights of the YOLOv5 model, which are fixed values, and α is the weight coefficient of the adaptive reward function.
9. A multi-scenario target recognition and detection method based on improved YOLOv5 according to claim 8, characterized in that: The weight adjustment module in step S33 is used to obtain parameter features in different scene images and task progress in different task objectives in real time, and calculate the weights of each loss item based on the above parameter features and task progress, and apply the dynamic weights to the comprehensive loss function.
10. A multi-scenario target recognition and detection method based on improved YOLOv5 according to claim 9, characterized in that: The multi-scenario target recognition detection device based on improved YOLOv5 includes: A processor, configured to implement the multi-scenario target recognition and detection method based on improved YOLOv5 as described in any one of claims 1 to 9 when executing a computer program; Memory used to store computer programs.
Citation Information
Cited By
Hydrological digital twinborn model construction and management and control method and device based on AI digital view
CN120744389A
Multi-scene-oriented robot collaborative path intelligent scheduling system and method
CN122044107A