Renewable resource intelligent sorting control method and system based on YOLO framework

Through the intelligent sorting control method for regeneration resources based on the YOLO framework, the target detection and feature extraction are used to use deep neural networks, and combined with the deep reinforcement learning algorithm to optimize sorting task scheduling, the problems of limited target recognition capabilities and multi-objective parallel sorting problems in the complex background in the existing technology are solved, and efficient and accurate sorting of regeneration resources is achieved.

CN119942424APending Publication Date: 2025-05-06HANGZHOU FULUN ECOLOGY TECH CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202411746247.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has limited ability to identify regenerative resource targets in complex contexts, and the sorting control method is difficult to cope with complex scenarios of multi-objective parallel sorting, resulting in insufficiency of sorting and inefficiency.

Method used

Using the intelligent sorting control method for regeneration resources based on the YOLO framework, images are collected through high-definition cameras, pre-trained YOLOv5 deep neural network is used for object detection and feature extraction, and combined with the deep reinforcement learning algorithm to optimize sorting task scheduling to generate sorting execution timing.

Benefits of technology

It improves the accuracy and efficiency of sorting of renewable resources, realizes the precise classification and delivery of renewable resources, and reduces the error rate of manual intervention and sorting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942424A_ABST
    Figure CN119942424A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent sorting control method and system for renewable resources based on a YOLO framework, and relates to the technical field of control, and the method comprises the steps: collecting a renewable resource image through a camera arranged above a conveyor belt, carrying out the size normalization processing, inputting an improved YOLOv5 deep neural network, carrying out the target detection, and obtaining the position and category information of the renewable resources. The YOLOv5 network adopts an improved CSPDarknet53 backbone network, and a double attention module is introduced into the YOLOv5 network. And according to the detection result, the predicted time for the targets to arrive at the sorting mechanism is calculated in combination with the speed parameters of the conveying belt, multi-target parallel sorting tasks are optimized and scheduled based on a deep reinforcement learning algorithm, a sorting execution time sequence is generated, and the pneumatic sorting mechanism is controlled to achieve accurate classified putting. According to the invention, the sorting efficiency and accuracy of renewable resources are improved, and the sorting cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to control technology, and in particular to a renewable resource intelligent sorting control method and system based on a YOLO framework. Background Art

[0002] The recycling of renewable resources is crucial to environmental protection and resource conservation. Traditionally, the sorting of renewable resources is mainly done manually, which is inefficient and costly, and is easily affected by human factors, resulting in the sorting accuracy and speed being unable to meet the growing demand for recycling of renewable resources. With the rapid development of automation technology, automated sorting equipment has gradually been applied to the field of renewable resource recycling.

[0003] Existing automated sorting technologies mainly include spectral sorting, magnetic sorting, eddy current sorting, and sorting based on machine vision. Among them, sorting technology based on machine vision has the advantages of multiple types of recognition, high sorting accuracy, and fast speed, and has become an important development direction for intelligent sorting of renewable resources. However, the existing renewable resource sorting technology based on machine vision still has some defects and shortcomings:

[0004] 1. Traditional machine vision algorithms have limited ability to identify renewable resource targets in complex backgrounds, especially in situations such as lighting changes and target occlusion, which can easily lead to misidentification or missed identification, resulting in reduced sorting accuracy.

[0005] 2. Most of the existing sorting control methods use simple rule control or sequential control, which is difficult to cope with the complex scenario of multi-target parallel sorting, resulting in low sorting efficiency.

[0006] 3. Lack of optimization of sorting execution sequence leads to problems such as frequent starting and stopping of pneumatic actuators and uncoordinated movements during the sorting process, affecting sorting efficiency and equipment life. Summary of the invention

[0007] The embodiments of the present invention provide a renewable resource intelligent sorting control method and system based on the YOLO framework, which can solve the problems in the prior art.

[0008] According to a first aspect of the embodiments of the present invention,

[0009] Provides a renewable resource intelligent sorting control method based on the YOLO framework, including:

[0010] The image of the renewable resources on the conveyor belt is collected in real time by a high-definition camera arranged above the renewable resources sorting conveyor belt, and the image of the renewable resources is normalized to obtain a standardized image, and the standardized image is input into a pre-trained YOLOv5 deep neural network. The YOLOv5 deep neural network adopts an improved CSPDarknet53 as the backbone network, and introduces a dual attention module of a channel attention mechanism and a spatial attention mechanism in the feature extraction network;

[0011] Using the YOLOv5 deep neural network to perform feature extraction and target detection on the input standardized image, obtain the position coordinate information and category probability information of the renewable resource target, determine the center point position, width and height of each renewable resource target according to the position coordinate information, classify the detected renewable resource targets into plastic, metal, glass, paper and other categories based on the category probability information, and calculate the confidence score of each target at the same time, retain the targets with confidence scores greater than the preset confidence threshold and generate target detection result data;

[0012] According to the position coordinate information and category information in the target detection result data, combined with the conveyor belt speed parameters, the estimated time for each renewable resource target to reach the sorting actuator is calculated, and the multi-target parallel sorting tasks are optimized and scheduled based on the deep reinforcement learning algorithm. The sorting execution sequence is generated, and the sorting control instructions are sent to the pneumatic sorting actuator through the programmable logic controller, wherein the sorting control instructions include the target category, execution time and pneumatic nozzle opening duration, so as to realize the accurate classification and delivery of renewable resources.

[0013] Using the YOLOv5 deep neural network to perform feature extraction and target detection on the input standardized image includes:

[0014] A dual attention mechanism is set after each stage network of the backbone network, and the dual attention mechanism includes a spatial attention module and a channel attention module, wherein the spatial attention module extracts the salient features of the target area and the global context features respectively through the maximum pooling branch and the average pooling branch, and the output features of the maximum pooling branch and the output features of the average pooling branch are input into a three-by-three convolutional layer and then fused to generate a spatial attention weight matrix, and the channel attention module calculates the importance weight of each channel using the global feature information, and the spatial attention weight matrix and the channel importance weight are adaptively fused through a learnable weight coefficient to obtain a first enhanced feature map;

[0015] Inputting the first enhanced feature map into a compressible excitation module in a feature pyramid network, the compressible excitation module comprising a global average pooling layer and a two-layer fully connected network, the global average pooling layer extracts channel dimension statistical features, the two-layer fully connected network learns the channel dependency relationship based on the channel dimension statistical features to obtain channel importance weights, and performing weighted fusion on different feature layers of the first enhanced feature map based on the channel importance weights to obtain a second enhanced feature map;

[0016] Inputting the second enhanced feature map into a dynamic anchor box optimization module in the detection head network, the dynamic anchor box optimization module includes an anchor box deformation network branch and an intersection-and-union prediction branch, wherein the anchor box deformation network branch predicts deformation parameters using the second enhanced feature map as input, deforms the initial anchor box based on the predicted deformation parameters to obtain a deformed anchor box, and the intersection-and-union prediction branch evaluates the matching degree between the deformed anchor box and the true target box and feeds back the matching degree information to the anchor box deformation network branch for iterative optimization to obtain an optimized detection box set;

[0017] The optimized detection frame set is post-processed, and the overlapping detection frames are screened using an improved soft non-maximum suppression algorithm, the improved soft non-maximum suppression algorithm extracts the depth feature vector in the overlapping detection frame, calculates the cosine distance of the depth feature vector to obtain the feature similarity, performs a weighted combination of the feature similarity and the geometric overlap of the detection frame to obtain a suppression score, and based on the suppression score, the overlapping detection frames in the optimized detection frame set are screened to obtain the final target detection result, wherein the weight coefficient of the suppression score is determined by optimizing the validation set.

[0018] Post-processing the optimized detection frame set and screening the overlapping detection frames using an improved soft non-maximum suppression algorithm includes:

[0019] Arrange the overlapping detection frame set in descending order according to the confidence, select the overlapping detection frames with confidence greater than a preset confidence threshold as the candidate frame set, and use the ROI Align operation to uniformly sample the detection frames in the candidate frame set as feature blocks;

[0020] Input the feature block into a two-layer fully connected network to obtain a deep feature vector, calculate the cosine distance of the deep feature vector corresponding to the overlapping detection frame in the candidate frame set to obtain feature similarity, and calculate the generalized intersection-over-union ratio of the overlapping detection frame to obtain geometric overlap;

[0021] A dynamic weight factor is set based on the confidence of the detection frame in the candidate frame set, and the feature similarity and the geometric overlap are weighted based on the dynamic weight factor to obtain a suppression score;

[0022] Softly suppressing the candidate frame set according to the suppression score, suppressing the corresponding detection frame when the suppression score is greater than a first suppression threshold, linearly attenuating the confidence of the corresponding detection frame when the suppression score is between the first suppression threshold and a second suppression threshold, and retaining the corresponding detection frame when the suppression score is less than the second suppression threshold, to obtain a screened detection frame set;

[0023] The position parameters of adjacent detection frames with similar confidence levels in the filtered detection frame set are weighted averaged and fused to obtain the final target detection result.

[0024] According to the position coordinate information and category information in the target detection result data, combined with the conveyor belt speed parameters, the estimated time for each renewable resource target to arrive at the sorting execution mechanism is calculated, and the multi-target parallel sorting tasks are optimized and scheduled based on the deep reinforcement learning algorithm. The sorting execution sequence is generated, including:

[0025] Obtain pixel coordinates and category information corresponding to the position coordinate information of the target in the target detection result data, convert the pixel coordinates into physical coordinates through camera calibration parameters, and calculate the estimated time for the target to reach the sorting actuator based on the speed data collected by the conveyor speed sensor and the distance from the detection area to the sorting actuator;

[0026] Construct a reinforcement learning environment model, taking the target category information, physical coordinates, estimated arrival time, working status information and location information of the sorting execution mechanism, and buffer occupancy information as the state space, and taking the target allocation decision, execution timing arrangement, and buffer usage strategy of the sorting execution mechanism as the action space;

[0027] The reinforcement learning environment model is trained using a dual Q learning network structure, wherein the dual Q learning network includes a strategy network and a target network, and a reward function is constructed based on sorting accuracy, throughput, and energy efficiency;

[0028] The priority score of the target is calculated according to the expected arrival time of the target, the unit time benefit and the complexity of the sorting action, the multi-target sorting tasks are sorted based on the priority score, and the time axis is discretized into time windows of fixed length; within each time window, the action decision based on the output of the policy network is combined with the kinematic constraints, safety spacing requirements and energy balance constraints of the sorting actuator to generate a local optimal scheduling plan, and the sorting execution timing is obtained by sliding the time window.

[0029] The reinforcement learning environment model is trained using a dual Q learning network structure, wherein the dual Q learning network includes a strategy network and a target network, and a reward function is constructed based on sorting accuracy, throughput, and energy efficiency, including:

[0030] Constructing a policy network and a target network, wherein the policy network and the target network have the same network architecture, the policy network is used for action selection and value function update, the target network is used for calculating the target Q value, the input layer dimensions of the policy network and the target network are consistent with the state vector dimensions, the policy network and the target network perform feature extraction through two hidden layers, and the output layer dimensions of the policy network and the target network correspond to the action space dimensions;

[0031] The parameters of the policy network are updated based on the temporal difference learning method. For the state and action, the target Q value is calculated. The target Q value is the sum of the product of the immediate reward and the discount factor and the smaller value of the two Q network prediction values. The discount factor is 0.99. The two Q network prediction values ​​are calculated based on the optimal action selected by the target network in the next state.

[0032] The parameters of the target network are updated by a soft update mechanism, the parameters of the target network are updated by an exponential sliding average of the policy network parameters, and the soft update coefficient of the exponential sliding average is set to 0.001;

[0033] Construct a combined reward function, which includes a sorting accuracy reward, a system throughput reward, and an energy efficiency reward. The sorting accuracy reward is in binary form, with a positive reward of ten for correct sorting and a negative penalty of ten for incorrect sorting. The system throughput reward is implemented by exponential function mapping, with a mapping coefficient of 0.2 for the exponential function. The value range of the system throughput reward is from zero to five. The energy efficiency reward takes into account the motion energy consumption of the actuator and the energy consumption of the cache unit, and the value range of the energy efficiency reward is from negative five to zero.

[0034] Calculate an empirical priority based on a timing differential error, where the empirical priority is equal to the sixth power of the sum of the absolute value of the timing differential error and a small positive number, and the sampling probability is proportional to the empirical priority;

[0035] The eps il on greedy strategy is used for exploration. The initial value of the eps il on greedy strategy is one, and it drops to 0.01 during the 500,000 training steps according to the exponential decay law. Each training round contains one thousand time steps, the batch size is 256, and the gradient clipping technology is used to limit the gradient norm to the range of negative ten to positive ten.

[0036] The priority score of the target is calculated according to the expected arrival time of the target, the unit time benefit and the complexity of the sorting action, and the multi-target sorting tasks are sorted based on the priority score, and the time axis is discretized into a time window of fixed length; in each time window, the action decision output by the strategy network is combined with the kinematic constraints, safety spacing requirements and energy consumption balance constraints of the sorting actuator to generate a local optimal scheduling solution, including:

[0037] Construct a target priority scoring model, wherein the target priority scoring model calculates the comprehensive priority score of the target based on the time urgency score, the unit time benefit score and the action complexity score; the time urgency score is calculated according to the time difference between the target estimated arrival time and the current time; the unit time benefit score is calculated according to the target value, the estimated sorting processing time and the mechanical movement time; the action complexity score is calculated according to the planned path length and the posture change amount;

[0038] The time axis is discretized into a time window of fixed length, the time window length is two seconds, and the overlap of adjacent time windows is fifty percent; in each time window, the objects to be sorted are sorted in descending order based on the comprehensive priority score, and an action decision sequence is generated according to the descending sorting result; the action decision sequence must simultaneously satisfy the speed constraint that the actuator speed is less than the maximum allowable speed, the acceleration constraint that the actuator acceleration is less than the maximum allowable acceleration, the safety distance constraint that the minimum distance between adjacent actuators is greater than the safety distance threshold, and the energy consumption balance constraint that the accumulated energy consumption in the time window is less than the energy consumption threshold; wherein the safety distance threshold is one and a half times the working radius of the actuator;

[0039] The scheduling scheme is updated in real time based on an event-driven mechanism. When a new sorting target arrives, a sorting task is completed, or an abnormal event occurs, re-planning is triggered; re-planning includes recalculating the comprehensive priority score based on the current system state, updating the descending order of the sorting targets, and generating an action decision sequence that meets the constraints;

[0040] At the end of each time window, the system status information is updated, the time window is slid to the next moment, and the sorting target sorting and action decision sequence generation based on the comprehensive priority score are repeated until all sorting tasks are completed.

[0041] A second aspect of an embodiment of the present invention provides a renewable resource intelligent sorting control system based on a YOLO framework, including:

[0042] The first unit is used to collect images of renewable resources on the conveyor belt in real time through a high-definition camera arranged above the renewable resource sorting conveyor belt, perform size normalization processing on the images of renewable resources to obtain standardized images, and input the standardized images into a pre-trained YOLOv5 deep neural network. The YOLOv5 deep neural network adopts an improved CSPDarknet53 as a backbone network, and introduces a dual attention module of a channel attention mechanism and a spatial attention mechanism in a feature extraction network;

[0043] The second unit is used to perform feature extraction and target detection on the input standardized image using the YOLOv5 deep neural network, obtain the position coordinate information and category probability information of the renewable resource target, determine the center point position, width and height of each renewable resource target according to the position coordinate information, classify the detected renewable resource targets into plastic, metal, glass, paper and other categories based on the category probability information, and calculate the confidence score of each target at the same time, retain the targets with confidence scores greater than a preset confidence threshold and generate target detection result data;

[0044] The third unit is used to calculate the estimated time for each renewable resource target to reach the sorting execution mechanism according to the position coordinate information and category information in the target detection result data, combined with the conveyor belt speed parameters, optimize the scheduling of multi-target parallel sorting tasks based on the deep reinforcement learning algorithm, generate the sorting execution sequence, and send the sorting control instruction to the pneumatic sorting actuator through the programmable logic controller, wherein the sorting control instruction includes the target category, execution time and pneumatic nozzle opening duration, so as to realize the accurate classification and delivery of renewable resources.

[0045] A third aspect of the embodiments of the present invention

[0046] An electronic device is provided, comprising:

[0047] processor;

[0048] a memory for storing processor-executable instructions;

[0049] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0050] A fourth aspect of the embodiments of the present invention is:

[0051] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0052] The beneficial effects of this application are as follows:

[0053] Improve the efficiency of recycling sorting. Use the YOLOv5 deep neural network to quickly and accurately identify and locate different types of recycled resources, and combine it with a deep reinforcement learning algorithm to optimize the scheduling of multi-objective parallel sorting tasks, thereby shortening the sorting time and improving the sorting efficiency.

[0054] Improve the accuracy of renewable resource sorting. Using the improved CSPDarknet53 backbone network and the YOLOv5 deep neural network with dual attention modules, the location and category of renewable resource targets can be detected more accurately. Combined with precise timing control and pneumatic sorting execution, the accurate classification and delivery of renewable resources can be achieved, reducing the sorting error rate.

[0055] The method realizes the automation and intelligence of the renewable resource sorting process, reduces manual intervention and reduces labor costs through real-time image acquisition by high-definition cameras, automatic identification and classification by deep learning algorithms, automatic scheduling of sorting tasks by deep reinforcement learning algorithms, and automatic control of sorting actuators by programmable logic controllers. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a flow chart of a renewable resource intelligent sorting control method based on the YOLO framework according to an embodiment of the present invention;

[0057] Figure 2 The figure is a schematic diagram of the structure of a renewable resource intelligent sorting control system based on the YOLO framework according to an embodiment of the present invention. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0059] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0060] Figure 1 FIG. 1 is a flow chart of a renewable resource intelligent sorting control method based on a YOLO framework according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0061] S101. The image of the renewable resources on the conveyor belt is collected in real time by a high-definition camera arranged above the renewable resource sorting conveyor belt, and the image of the renewable resources is normalized to obtain a standardized image, and the standardized image is input into a pre-trained YOLOv5 deep neural network, wherein the YOLOv5 deep neural network adopts an improved CSPDarknet53 as a backbone network, and introduces a dual attention module of a channel attention mechanism and a spatial attention mechanism in a feature extraction network;

[0062] S102. Use the YOLOv5 deep neural network to perform feature extraction and target detection on the input standardized image, obtain the position coordinate information and category probability information of the renewable resource target, determine the center point position, width and height of each renewable resource target according to the position coordinate information, and classify the detected renewable resource targets into plastic, metal, glass, paper and other categories based on the category probability information, and calculate the confidence score of each target at the same time, retain the targets with confidence scores greater than the preset confidence threshold and generate target detection result data;

[0063] S103. According to the position coordinate information and category information in the target detection result data, combined with the conveyor belt speed parameters, the estimated time for each renewable resource target to reach the sorting execution mechanism is calculated, and the multi-target parallel sorting tasks are optimized and scheduled based on the deep reinforcement learning algorithm, and the sorting execution sequence is generated. The sorting control instructions are sent to the pneumatic sorting actuator through the programmable logic controller, wherein the sorting control instructions include the target category, execution time and pneumatic nozzle opening duration, so as to realize the accurate classification and delivery of renewable resources.

[0064] For example, the YOLOv5-based renewable resource image detection and classification method first collects data through a high-definition camera. The high-definition camera is installed 2 meters above the sorting conveyor belt, using 4K resolution, a frame rate of 30fps, and a field of view of 90 degrees to ensure complete coverage of the conveyor belt working area. The collected image data is transmitted to the processing unit in real time via industrial Ethernet.

[0065] In the image preprocessing stage, the collected original images are standardized. First, the original images are uniformly scaled to 640×640 pixels and the bilinear interpolation algorithm is used for size transformation. Then the color space is converted to convert the RGB image to the YUV color space, and the brightness channel is extracted for histogram equalization to enhance the image contrast. Finally, the image is normalized to map the pixel value range to between 0 and 1.

[0066] The improved YOLOv5 network uses CSPDarknet53 as the backbone network. Based on the original CSPDarknet53, the dual attention module is introduced to enhance the feature extraction capability. The channel attention module obtains the global information of the channel dimension through global average pooling, learns the correlation between channels through two fully connected layers, and generates channel weights. The spatial attention module uses maximum pooling and average pooling to extract spatial features respectively, and then concatenates the two features and generates a spatial weight map through a 7×7 convolutional layer. The output feature map of the dual attention module is obtained by the weighted results of channel attention and spatial attention.

[0067] The target detection stage first extracts multi-scale features through the feature pyramid network FPN. FPN contains two feature fusion paths, top-down and bottom-up, to generate feature maps of five scales from P3 to P7. Each feature map passes through three prediction heads to predict the target position, category probability and confidence respectively. The position prediction uses relative coordinates to predict the offset of the target center point relative to the upper left corner of the grid and the aspect ratio. Category prediction outputs the probability distribution of five categories, including plastic, metal, glass, paper and other categories.

[0068] The post-processing of the prediction results includes coordinate conversion and non-maximum suppression. First, the relative coordinates are converted to absolute image coordinates to obtain the specific location and size information of each target. The prediction boxes with a confidence score lower than 0.5 are directly filtered out. Then the remaining prediction boxes are sorted in descending order of confidence, and the overlapping detection boxes are eliminated by the non-maximum suppression algorithm, and the IoU threshold is set to 0.45. The detection box that is finally retained is the final target detection result.

[0069] The result output contains the location information, category information, and confidence information of each detected target. The location information includes the coordinates of the upper left and lower right corners of the target box, the center point coordinates, width, and height. The category information includes the predicted category label and the corresponding probability value. The system saves the detection results in JSON format, including metadata such as timestamp and image ID, for subsequent analysis and processing.

[0070] Based on a large amount of experimental data verification, this method has achieved good results in common renewable resource target detection tasks. In the test data set, the average detection accuracy of plastic targets reached 92.3%, metal 91.8%, glass 90.5%, paper 93.1%, and the overall mAP reached 91.9%. In practical applications, the system can run stably and output detection results in real time, with an average processing delay of less than 50 milliseconds, meeting industrial sorting needs.

[0071] In an optional embodiment, using the YOLOv5 deep neural network to perform feature extraction and target detection on the input standardized image includes:

[0072] A dual attention mechanism is set after each stage network of the backbone network, and the dual attention mechanism includes a spatial attention module and a channel attention module, wherein the spatial attention module extracts the salient features of the target area and the global context features respectively through the maximum pooling branch and the average pooling branch, and the output features of the maximum pooling branch and the output features of the average pooling branch are input into a three-by-three convolutional layer and then fused to generate a spatial attention weight matrix, and the channel attention module calculates the importance weight of each channel using the global feature information, and the spatial attention weight matrix and the channel importance weight are adaptively fused through a learnable weight coefficient to obtain a first enhanced feature map;

[0073] Inputting the first enhanced feature map into a compressible excitation module in a feature pyramid network, the compressible excitation module comprising a global average pooling layer and a two-layer fully connected network, the global average pooling layer extracts channel dimension statistical features, the two-layer fully connected network learns the channel dependency relationship based on the channel dimension statistical features to obtain channel importance weights, and performing weighted fusion on different feature layers of the first enhanced feature map based on the channel importance weights to obtain a second enhanced feature map;

[0074] Inputting the second enhanced feature map into a dynamic anchor box optimization module in the detection head network, the dynamic anchor box optimization module includes an anchor box deformation network branch and an intersection-and-union prediction branch, wherein the anchor box deformation network branch predicts deformation parameters using the second enhanced feature map as input, deforms the initial anchor box based on the predicted deformation parameters to obtain a deformed anchor box, and the intersection-and-union prediction branch evaluates the matching degree between the deformed anchor box and the true target box and feeds back the matching degree information to the anchor box deformation network branch for iterative optimization to obtain an optimized detection box set;

[0075] The optimized detection frame set is post-processed, and the overlapping detection frames are screened using an improved soft non-maximum suppression algorithm, the improved soft non-maximum suppression algorithm extracts the depth feature vector in the overlapping detection frame, calculates the cosine distance of the depth feature vector to obtain the feature similarity, performs a weighted combination of the feature similarity and the geometric overlap of the detection frame to obtain a suppression score, and based on the suppression score, the overlapping detection frames in the optimized detection frame set are screened to obtain the final target detection result, wherein the weight coefficient of the suppression score is determined by optimizing the validation set.

[0076] The target detection method based on the YOLOv5 deep neural network improves the target detection accuracy by enhancing the feature extraction capability and optimizing the anchor frame mechanism. This method includes three main stages: image preprocessing, feature extraction and enhancement, and target detection and post-processing.

[0077] First, the input image is preprocessed. The image preprocessing stage scales the input image to a standard size and performs normalization operations, such as scaling the pixel value range to between 0 and 1, to improve the training efficiency and stability of the model. Assuming the input image size is 1920x1080, it is scaled to 640x640.

[0078] Subsequently, the improved YOLOv5 network is used for feature extraction and enhancement. The backbone network adopts the CSPDarknet53 structure, and a dual attention mechanism is added after each stage. Take the feature map output by a certain stage of the backbone network as an example, the size is 20x20x256. The spatial attention module first extracts the salient features of the target area and the global context features through the maximum pooling and average pooling operations of size 3x3 respectively. Then, the two feature maps are respectively input into a 3x3 convolutional layer, and the output feature maps are fused to generate the spatial attention weight matrix. The channel attention module uses the global average pooling operation to extract global feature information, and calculates the importance weight of each channel through a two-layer fully connected network. Finally, the spatial attention weight matrix and the channel importance weight are adaptively fused through the learnable weight coefficient to obtain the first enhanced feature map. Assuming that the learnable weight coefficient is 0.5, the final feature map is still 20x20x256. Then, the first enhanced feature map is input into the compressible excitation module (SE) in the feature pyramid network (FPN). The SE module first extracts channel dimension statistical features through global average pooling. Then, the dependency between channels is learned through a two-layer fully connected network to obtain the channel importance weights. Finally, the different feature layers of the first enhanced feature map are weighted fused based on the channel importance weights to obtain the second enhanced feature map.

[0079] Next, target detection is performed. The second enhanced feature map is input into the dynamic anchor box optimization module in the detection head network. The anchor box deformation network branch takes the second enhanced feature map as input and predicts the deformation parameters. Then, the initial anchor box is deformed based on the predicted deformation parameters to obtain the deformed anchor box. The intersection over union (IoU) prediction branch evaluates the matching degree between the deformed anchor box and the true target box, and feeds the matching degree information back to the anchor box deformation network branch for iterative optimization, and finally obtains the optimized detection box set. For example, the initial anchor box is (100, 100, 150, 150), and the deformation parameters are (0.1, 0.2, 0.05, 0.1), then the deformed anchor box is (110, 120, 157.5, 165).

[0080] Finally, post-processing is performed. The overlapping detection frames are screened using the improved soft non-maximum suppression (Soft-NMS) algorithm. The improved Soft-NMS algorithm extracts the deep feature vectors in the overlapping detection frames, and calculates the cosine distance of the deep feature vectors to obtain the feature similarity. Then, the feature similarity is weightedly combined with the geometric overlap (IoU) of the detection frames to obtain the suppression score. Assuming that the IoU of two overlapping detection frames is 0.8, the feature similarity is 0.9, and the weight coefficients are 0.7 and 0.3 respectively, the suppression score is 0.7*0.8+0.3*0.9=0.83. Finally, the overlapping detection frames in the optimized detection frame set are screened based on the suppression score to obtain the final target detection result. The weight coefficient is determined by optimizing the validation set, for example, by determining the optimal weight coefficient through grid search.

[0081] The beneficial effects of this method are reflected in three aspects:

[0082] First, improve the accuracy of target detection. The dual attention mechanism is used to enhance the feature expression ability, so that the network can focus on the target area and key features more effectively, thereby improving the accuracy of target detection.

[0083] Second, optimize the anchor frame mechanism. The dynamic anchor frame optimization module can adaptively adjust the anchor frame according to the shape and size of the target, improve the matching degree between the anchor frame and the target, and further improve the detection accuracy.

[0084] Third, improve detection efficiency. The improved Soft-NMS algorithm can effectively suppress redundant detection boxes and reduce the amount of calculation by considering feature similarity, thereby improving detection efficiency.

[0085] In an optional implementation, post-processing the optimized detection frame set and screening overlapping detection frames using an improved soft non-maximum suppression algorithm includes:

[0086] Arrange the overlapping detection frame set in descending order according to the confidence, select the overlapping detection frames with confidence greater than a preset confidence threshold as the candidate frame set, and use the ROI Align operation to uniformly sample the detection frames in the candidate frame set as feature blocks;

[0087] Input the feature block into a two-layer fully connected network to obtain a deep feature vector, calculate the cosine distance of the deep feature vector corresponding to the overlapping detection frame in the candidate frame set to obtain feature similarity, and calculate the generalized intersection-over-union ratio of the overlapping detection frame to obtain geometric overlap;

[0088] A dynamic weight factor is set based on the confidence of the detection frame in the candidate frame set, and the feature similarity and the geometric overlap are weighted based on the dynamic weight factor to obtain a suppression score;

[0089] Softly suppressing the candidate frame set according to the suppression score, suppressing the corresponding detection frame when the suppression score is greater than a first suppression threshold, linearly attenuating the confidence of the corresponding detection frame when the suppression score is between the first suppression threshold and a second suppression threshold, and retaining the corresponding detection frame when the suppression score is less than the second suppression threshold, to obtain a screened detection frame set;

[0090] The position parameters of adjacent detection frames with similar confidence levels in the filtered detection frame set are weighted averaged and fused to obtain the final target detection result.

[0091] In order to solve the problem of overlapping detection frames in target detection tasks, this embodiment provides an improved soft non-maximum suppression method. This method uses deep feature similarity and geometric overlap, combined with dynamic weight factors for weighting, to perform more precise screening of overlapping detection frames, effectively reducing false detections and improving detection accuracy.

[0092] First, obtain the detection box set output by the target detection model. Assume that the set contains five detection boxes with confidence levels of 0.9, 0.85, 0.8, 0.75, and 0.7, respectively, and the corresponding bounding box coordinates are (10, 10, 50, 50), (12, 12, 52, 52), (60, 60, 100, 100), (15, 15, 55, 55), and (62, 62, 102, 102).

[0093] Next, the detection box set is sorted in descending order of confidence. The confidence of the sorted detection boxes is 0.9, 0.85, 0.8, 0.75, and 0.7. The preset confidence threshold is set to 0.6. Since the confidence of all detection boxes is greater than 0.6, all detection boxes are used as candidate boxes.

[0094] Then, the ROI ligation operation is used to extract features from the candidate boxes. The image area corresponding to each candidate box is uniformly sampled into a feature block of size 7x7, for example.

[0095] The extracted feature blocks are input into a two-layer fully connected network to obtain a deep feature vector. Assume that the deep feature vector of the first detection box is [0.1, 0.2, 0.3, ...], the deep feature vector of the second detection box is [0.11, 0.22, 0.33, ...], and so on.

[0096] Calculate the cosine distance of the deep feature vectors corresponding to the overlapping detection frames in the candidate frame set to obtain the feature similarity. For example, the feature similarity calculation result of the first detection frame and the second detection frame is 0.98.

[0097] At the same time, the generalized intersection-over-union (GIOU) of the overlapping detection boxes is calculated to obtain the geometric overlap. For example, the geometric overlap between the first detection box and the second detection box is calculated to be 0.95.

[0098] Set dynamic weight factors based on the confidence of the detection boxes in the candidate box set. For example, the higher the confidence of the detection box, the larger its weight factor. Suppose the weight factor of the first detection box is 0.9, the weight factor of the second detection box is 0.85, and so on.

[0099] The feature similarity and geometric overlap are weighted based on the dynamic weight factor to obtain the suppression score. For example, the suppression score of the first detection box and the second detection box is 0.9*0.98+(1-0.9)*0.95=0.977.

[0100] Perform soft suppression on the candidate box set according to the suppression score. Set the first suppression threshold to 0.8 and the second suppression threshold to 0.6. Assuming that the suppression score of the first detection box and the second detection box is 0.977, which is greater than the first suppression threshold of 0.8, the second detection box is suppressed. Assuming that the suppression score of the third detection box and the fourth detection box is 0.7, which is between the first suppression threshold and the second suppression threshold, the confidence of the fourth detection box is linearly attenuated. Assuming that the suppression score of the third detection box and the fifth detection box is 0.5, which is less than the second suppression threshold of 0.6, the fifth detection box is retained.

[0101] Finally, the position parameters of adjacent detection frames with similar confidence in the filtered detection frame set are weighted averaged and fused. For example, if the confidence of the first detection frame and the fifth detection frame in the remaining detection frames are similar, their position parameters are weighted averaged to obtain the final target detection result.

[0102] The beneficial effects of this method are reflected in three aspects:

[0103] First, improve detection accuracy: By combining deep features and geometric information, the overlap degree of detection frames can be more accurately judged, redundant frames can be effectively suppressed, and false detections can be reduced, thereby improving detection accuracy.

[0104] Second, improve detection efficiency: The soft suppression mechanism avoids directly deleting the detection box and retains more potential correct detection results, especially in dense target scenarios, which can better balance precision and recall.

[0105] Third, enhanced robustness: The dynamic weight factor mechanism enables detection boxes with high confidence to have greater say in the suppression process, enhancing the robustness of the algorithm to noise and occlusion.

[0106] In an optional implementation, according to the position coordinate information and category information in the target detection result data, combined with the conveyor belt speed parameter, the estimated time for each renewable resource target to arrive at the sorting execution mechanism is calculated, and the multi-target parallel sorting tasks are optimized and scheduled based on the deep reinforcement learning algorithm, and the sorting execution sequence is generated, including:

[0107] Obtain pixel coordinates and category information corresponding to the position coordinate information of the target in the target detection result data, convert the pixel coordinates into physical coordinates through camera calibration parameters, and calculate the estimated time for the target to reach the sorting actuator based on the speed data collected by the conveyor speed sensor and the distance from the detection area to the sorting actuator;

[0108] Construct a reinforcement learning environment model, taking the target category information, physical coordinates, estimated arrival time, working status information and location information of the sorting execution mechanism, and buffer occupancy information as the state space, and taking the target allocation decision, execution timing arrangement, and buffer usage strategy of the sorting execution mechanism as the action space;

[0109] The reinforcement learning environment model is trained using a dual Q learning network structure, wherein the dual Q learning network includes a strategy network and a target network, and a reward function is constructed based on sorting accuracy, throughput, and energy efficiency;

[0110] The priority score of the target is calculated according to the expected arrival time of the target, the unit time benefit and the complexity of the sorting action, the multi-target sorting tasks are sorted based on the priority score, and the time axis is discretized into time windows of fixed length; within each time window, the action decision based on the output of the policy network is combined with the kinematic constraints, safety spacing requirements and energy balance constraints of the sorting actuator to generate a local optimal scheduling plan, and the sorting execution timing is obtained by sliding the time window.

[0111] A multi-objective parallel sorting method for renewable resources based on deep reinforcement learning aims to improve sorting efficiency and resource utilization. The core of this method is to dynamically generate the optimal sorting execution sequence using a reinforcement learning model based on information such as the target's estimated arrival time, category, and sorting mechanism status.

[0112] First, the detection result data of the renewable resource targets on the conveyor belt is obtained. These data include the pixel coordinates and category information of each target. For example, the pixel coordinates of target A are (100, 200), and the category is plastic bottle; the pixel coordinates of target B are (300, 400), and the category is glass bottle. The pixel coordinates are converted into actual physical coordinates through pre-calibrated camera parameters, such as focal length, principal point coordinates, etc. Assume that according to the camera calibration parameters, the physical coordinates of target A are (0.5, 1.0) meters, and the physical coordinates of target B are (1.5, 2.0) meters.

[0113] Next, the estimated time for the target to reach the sorting actuator is calculated in combination with the conveyor belt speed parameters. Assume that the speed measured by the conveyor belt speed sensor is 0.2 m / s and the distance from the detection area to the sorting actuator is 2 m. Then, the estimated time for target A to reach the sorting mechanism is (2-0.5) / 0.2=7.5 seconds, and the estimated time for target B is (2-1.5) / 0.2=2.5 seconds.

[0114] Then, a reinforcement learning environment model is constructed. The state space of the model contains the category information of the target (such as plastic bottles, glass bottles), physical coordinates, estimated arrival time, working status information (such as idle, busy) and location information of the sorting actuator, and occupancy information of the buffer area (such as full, not full). The action space includes which sorting agency the target is assigned to, the execution timing arrangement of each sorting agency, and the use strategy of the buffer area. For example, an action can be to assign target A to sorting agency No. 1, target B to sorting agency No. 2, sorting agency No. 1 starts sorting after 7 seconds, sorting agency No. 2 starts sorting after 2 seconds, and temporarily puts target A into the buffer area.

[0115] In order to train this reinforcement learning model, a dual Q-learning network structure is used. This network structure contains two neural networks: a policy network and a target network. By constantly interacting with the environment, the policy network learns how to choose the best action based on the current state, and the target network is used to evaluate the long-term value of these actions. During training, a reward function is used to guide the learning process. This reward function is designed based on sorting accuracy, throughput, and energy efficiency. For example, successfully sorting a target will receive a positive reward, and sorting errors or excessive energy consumption will receive a negative reward.

[0116] Next, the priority score of the target is calculated based on the target's estimated arrival time, unit time benefit (such as the benefit of sorting a certain category of targets), and sorting action complexity (such as the time required to sort different categories of targets). For example, assuming that sorting plastic bottles has higher benefits and the action of sorting glass bottles is more complex, target A may have a higher priority score than target B. Then, the multi-target sorting tasks are sorted based on the priority score.

[0117] Discretize the time axis into time windows of fixed length, for example, each time window is 1 second. In each time window, based on the action decision output by the policy network, combined with the kinematic constraints of the sorting actuator (such as maximum speed, acceleration), safety spacing requirements (such as the need to maintain a certain distance between different sorting mechanisms) and energy balance constraints (such as avoiding all sorting mechanisms working at the same time to cause excessive energy consumption), a local optimal scheduling plan is generated. For example, in a time window, it can be decided that sorting mechanism No. 1 sorts target A, and sorting mechanism No. 2 remains idle. By sliding the time window, the local optimal scheduling plans of each time window are connected, and finally a complete sorting execution sequence is obtained.

[0118] The beneficial effects of this method can be summarized in the following three aspects:

[0119] 1. Improve sorting efficiency: Through the optimized scheduling of the reinforcement learning model, the work of each sorting organization can be effectively arranged, reducing idle time and waiting time, thereby improving the overall sorting efficiency.

[0120] 2. Improve resource utilization: This method can dynamically adjust the sorting strategy and cache usage strategy based on information such as the target category, location, and arrival time, thereby maximizing the use of sorting resources and cache space.

[0121] 3. Reduce energy consumption: By considering energy consumption factors in the reward function and adding energy balance constraints in the scheduling process, the energy consumption in the sorting process can be effectively reduced.

[0122] In an optional implementation, a dual Q learning network structure is used to train the reinforcement learning environment model, wherein the dual Q learning network includes a strategy network and a target network, and constructing a reward function based on sorting accuracy, throughput, and energy efficiency includes:

[0123] Constructing a policy network and a target network, wherein the policy network and the target network have the same network architecture, the policy network is used for action selection and value function update, the target network is used for calculating the target Q value, the input layer dimensions of the policy network and the target network are consistent with the state vector dimensions, the policy network and the target network perform feature extraction through two hidden layers, and the output layer dimensions of the policy network and the target network correspond to the action space dimensions;

[0124] The parameters of the policy network are updated based on the temporal difference learning method. For the state and action, the target Q value is calculated. The target Q value is the sum of the product of the immediate reward and the discount factor and the smaller value of the two Q network prediction values. The discount factor is 0.99. The two Q network prediction values ​​are calculated based on the optimal action selected by the target network in the next state.

[0125] The parameters of the target network are updated by a soft update mechanism, the parameters of the target network are updated by an exponential sliding average of the policy network parameters, and the soft update coefficient of the exponential sliding average is set to 0.001;

[0126] Construct a combined reward function, which includes a sorting accuracy reward, a system throughput reward, and an energy efficiency reward. The sorting accuracy reward is in binary form, with a positive reward of ten for correct sorting and a negative penalty of ten for incorrect sorting. The system throughput reward is implemented by exponential function mapping, with a mapping coefficient of 0.2 for the exponential function. The value range of the system throughput reward is from zero to five. The energy efficiency reward takes into account the motion energy consumption of the actuator and the energy consumption of the cache unit, and the value range of the energy efficiency reward is from negative five to zero.

[0127] Calculate an empirical priority based on a timing differential error, where the empirical priority is equal to the sixth power of the sum of the absolute value of the timing differential error and a small positive number, and the sampling probability is proportional to the empirical priority;

[0128] The eps il on greedy strategy is used for exploration. The initial value of the eps il on greedy strategy is one, and it drops to 0.01 during the 500,000 training steps according to the exponential decay law. Each training round contains one thousand time steps, the batch size is 256, and the gradient clipping technology is used to limit the gradient norm to the range of negative ten to positive ten.

[0129] An intelligent sorting control method based on double Q-learning is designed to improve the accuracy, throughput and energy efficiency of the sorting system. The method uses a double Q-learning network structure to train the reinforcement learning environment model and combines the experience priority and eps il on greedy strategy for efficient exploration.

[0130] First, two neural networks with the same network architecture are constructed: the policy network and the target network. The input layer dimensions of these two networks are consistent with the dimensions of the state vector that describes the state of the sorting system. For example, the state vector can contain information such as the location, type, and buffer occupancy of the item. Both the policy network and the target network perform feature extraction through two hidden layers, each of which contains a certain number of neurons, for example, they can be set to 128 and 64 neurons respectively. The output layer dimension corresponds to the action space dimension. For example, if the sorting system has 5 possible actions (for example, moving items to different buffers), the output layer has 5 neurons. The policy network is used to select actions and update value functions, and the target network is used to calculate the target Q value.

[0131] Next, the parameters of the policy network are updated using a temporal difference learning method. For the current state and the action taken, a target Q value is calculated. The target Q value consists of two parts: one is the immediate reward, and the other is the product of a discount factor and the smaller of the two Q network predictions. The discount factor is used to balance the importance of current rewards and future rewards, and is usually set to a value close to 1, such as 0.99. The two Q network predictions are calculated based on the optimal action selected by the target network in the next state. By comparing the target Q value and the current Q value of the policy network, the temporal difference error is calculated and used to update the parameters of the policy network.

[0132] In order to stabilize the training process, a soft update mechanism is used to update the parameters of the target network. The parameters of the target network are updated by the exponential sliding average of the policy network parameters, and the soft update coefficient is set to a small value, such as 0.001. This means that the parameters of the target network will slowly approach the parameters of the policy network, thus avoiding drastic fluctuations.

[0133] In order to comprehensively consider multiple performance indicators of the sorting system, a combined reward function is constructed. The combined reward function consists of three parts: sorting accuracy reward, system throughput reward and energy efficiency reward. The sorting accuracy reward is in binary form: correct sorting gets a reward of 10, and incorrect sorting gets a penalty of -10. The system throughput reward is implemented through exponential function mapping. The mapping coefficient of the exponential function is 0.2, and the value range of the system throughput reward is 0 to 5. For example, when the throughput is 10, the reward value is 5*(1-exp(-0.2*10))≈4.32. The energy efficiency reward takes into account the motion energy consumption of the actuator and the energy consumption of the cache unit, and the value range is -5 to 0. For example, if the energy consumption is 100J, the reward value can be set to -5*(100 / maximum energy consumption).

[0134] In order to improve learning efficiency, the experience replay mechanism and priority sampling are adopted. The experience priority is calculated based on the temporal difference error, which is equal to the 0.6 power of the sum of the absolute value of the temporal difference error and a small positive number (such as 0.01). The sampling probability is proportional to the experience priority, which means that experience samples with higher temporal difference errors are more likely to be sampled for training.

[0135] In order to balance exploration and exploitation, the eps il on greedy strategy is used for exploration. The initial value of the eps il on greedy strategy is set to 1, and it decreases to 0.01 according to the exponential decay law during the 500,000 training steps. Each training round contains 1000 time steps and the batch size is 256. In order to prevent gradient explosion, the gradient clipping technique is used to limit the gradient norm to the range of -10 to 10. For example, if the calculated gradient norm is 15, it is clipped to 10.

[0136] Suppose in a sorting scenario, there are 5 types of items and 3 buffers. In the initial state, an item is located at the entrance, and the policy network chooses the action of moving the item to the first buffer according to the current state. If the sorting is correct, a reward of 10 is obtained, the throughput increases, the energy consumption increases, and finally a combined reward value is obtained. If the sorting is wrong, a penalty of -10 is obtained, the throughput remains unchanged, the energy consumption increases, and finally a negative combined reward value is obtained. This experience sample (state, action, reward, next state) is stored in the experience replay buffer. During the training process, a batch of experience samples are randomly sampled from the experience replay buffer to update the parameters of the policy network.

[0137] The intelligent sorting control method has the following beneficial effects:

[0138] 1. Improve sorting accuracy: Through the reinforcement learning algorithm, the control strategy is continuously optimized so that the sorting system can make the best decision based on the characteristics of the items and the system status, thereby improving the sorting accuracy.

[0139] 2. Improve system throughput: This method encourages the system to increase throughput and thus improve sorting efficiency through the design of reward functions.

[0140] 3. Reduce energy consumption: This method takes energy consumption as part of the reward function, encouraging the system to reduce energy consumption and improve energy efficiency while ensuring sorting accuracy and throughput.

[0141] In an optional implementation, the priority score of the target is calculated according to the estimated arrival time of the target, the unit time benefit and the complexity of the sorting action, the multi-target sorting tasks are sorted based on the priority score, and the time axis is discretized into a time window of fixed length; in each time window, the action decision output by the strategy network is combined with the kinematic constraints, safety spacing requirements and energy balance constraints of the sorting actuator to generate a local optimal scheduling solution, including:

[0142] Construct a target priority scoring model, wherein the target priority scoring model calculates the comprehensive priority score of the target based on the time urgency score, the unit time benefit score and the action complexity score; the time urgency score is calculated according to the time difference between the target estimated arrival time and the current time; the unit time benefit score is calculated according to the target value, the estimated sorting processing time and the mechanical movement time; the action complexity score is calculated according to the planned path length and the posture change amount;

[0143] The time axis is discretized into a time window of fixed length, the time window length is two seconds, and the overlap of adjacent time windows is fifty percent; in each time window, the objects to be sorted are sorted in descending order based on the comprehensive priority score, and an action decision sequence is generated according to the descending sorting result; the action decision sequence must simultaneously satisfy the speed constraint that the actuator speed is less than the maximum allowable speed, the acceleration constraint that the actuator acceleration is less than the maximum allowable acceleration, the safety distance constraint that the minimum distance between adjacent actuators is greater than the safety distance threshold, and the energy consumption balance constraint that the accumulated energy consumption in the time window is less than the energy consumption threshold; wherein the safety distance threshold is one and a half times the working radius of the actuator;

[0144] The scheduling scheme is updated in real time based on an event-driven mechanism. When a new sorting target arrives, a sorting task is completed, or an abnormal event occurs, re-planning is triggered; re-planning includes recalculating the comprehensive priority score based on the current system state, updating the descending order of the sorting targets, and generating an action decision sequence that meets the constraints;

[0145] At the end of each time window, the system status information is updated, the time window is slid to the next moment, and the sorting target sorting and action decision sequence generation based on the comprehensive priority score are repeated until all sorting tasks are completed.

[0146] The multi-objective sorting task scheduling method aims to optimize sorting efficiency and resource utilization. The core of this method is to calculate the priority score of the target according to the target's expected arrival time, unit time revenue and sorting action complexity, and sort the multi-objective sorting tasks based on this score.

[0147] First, a target priority scoring model is constructed. This model comprehensively considers three factors: time urgency, unit time benefit, and action complexity. The higher the time urgency score, the faster the target needs to be processed. For example, if the estimated arrival time of target A is 10 seconds later, the estimated arrival time of target B is 20 seconds later, and the current time is 0 seconds, then the time urgency score of target A is higher than that of target B. The higher the unit time benefit score, the higher the cost-effectiveness of processing the target. For example, target C is worth 10 yuan, the estimated processing time is 2 seconds, and the mechanical movement time is 1 second; target D is worth 5 yuan, the estimated processing time is 1 second, and the mechanical movement time is 1 second. Then the unit time benefit score of target C is higher than that of target D. The higher the action complexity score, the more difficult it is to operate the target. For example, the planned path length of target E is 10 meters, and the posture change is 90 degrees; the planned path length of target F is 5 meters, and the posture change is 45 degrees. Then the action complexity score of target E is higher than that of target F. The comprehensive priority score is calculated by weighting these three factors, and the weights can be adjusted according to actual conditions. For example, if time urgency is most important, the time urgency weight is set to be the highest.

[0148] Next, the time axis is discretized into fixed-length time windows, each of which is 2 seconds long and has a 50% overlap between adjacent time windows. For example, the first time window is 0-2 seconds, the second time window is 1-3 seconds, and so on. In each time window, the targets are sorted in descending order according to their comprehensive priority scores. For example, if there are targets G, H, and I in the time window, and their comprehensive priority scores are 9, 7, and 5, respectively, the sorting results are G, H, and I. An action decision sequence is generated based on the sorting results, that is, the actuators are arranged to process the targets in sequence.

[0149] When generating an action decision sequence, a series of constraints need to be met. Speed ​​constraint: The speed of the actuator cannot exceed the maximum allowed speed. For example, if the maximum allowed speed is 1 m / s, the speed of the actuator cannot exceed 1 m / s at any time. Acceleration constraint: The acceleration of the actuator cannot exceed the maximum allowed acceleration. For example, if the maximum allowed acceleration is 0.5 m / s2, the acceleration of the actuator cannot exceed 0.5 m / s2 at any time. Safety distance constraint: The minimum distance between adjacent actuators must be greater than the safety distance threshold, which is set to 1.5 times the working radius of the actuator. For example, if the working radius of the actuator is 0.5 m, the safety distance threshold is 0.75 m. Energy balance constraint: The cumulative energy consumption within the time window cannot exceed the energy consumption threshold. For example, if the energy consumption threshold is 100 joules, the total energy consumption of all actuators within the time window cannot exceed 100 joules.

[0150] The system uses an event-driven mechanism to update the scheduling plan in real time. When a new target arrives, a sorting task is completed, or an abnormal event occurs (such as a failure of an actuator), replanning is triggered. Replanning includes recalculating the comprehensive priority scores of all targets, updating the target ranking, and generating a new action decision sequence that meets the constraints.

[0151] At the end of each time window, the system status information is updated, such as the position, speed, and remaining energy of each actuator. Then the time window is slid to the next moment, and the target sorting and action decision sequence generation are repeated until all sorting tasks are completed. For example, if the current time window is 1-3 seconds, the next time window is 2-4 seconds.

[0152] The beneficial effects of this method are reflected in the following three aspects:

[0153] Improve sorting efficiency: By prioritizing high-value, time-critical targets, maximizing sorting revenue per unit time, thereby improving overall sorting efficiency.

[0154] Optimize resource utilization: By generating local optimal scheduling solutions based on the constraints, the processing power and energy of the actuators can be fully utilized to avoid wasting resources.

[0155] Enhance system robustness: Based on the event-driven mechanism, it can respond quickly to real-time changes in system status, adjust the scheduling plan in time, and enhance the adaptability and robustness of the system.

[0156] Figure 2 FIG. 1 is a schematic diagram of the structure of a renewable resource intelligent sorting control system based on the YOLO framework according to an embodiment of the present invention. Figure 2 As shown, the system comprises:

[0157] The first unit is used to collect images of renewable resources on the conveyor belt in real time through a high-definition camera arranged above the renewable resource sorting conveyor belt, perform size normalization processing on the images of renewable resources to obtain standardized images, and input the standardized images into a pre-trained YOLOv5 deep neural network. The YOLOv5 deep neural network adopts an improved CSPDarknet53 as a backbone network, and introduces a dual attention module of a channel attention mechanism and a spatial attention mechanism in a feature extraction network;

[0158] The second unit is used to perform feature extraction and target detection on the input standardized image using the YOLOv5 deep neural network, obtain the position coordinate information and category probability information of the renewable resource target, determine the center point position, width and height of each renewable resource target according to the position coordinate information, classify the detected renewable resource targets into plastic, metal, glass, paper and other categories based on the category probability information, and calculate the confidence score of each target at the same time, retain the targets with confidence scores greater than a preset confidence threshold and generate target detection result data;

[0159] The third unit is used to calculate the estimated time for each renewable resource target to reach the sorting execution mechanism according to the position coordinate information and category information in the target detection result data, combined with the conveyor belt speed parameters, optimize the scheduling of multi-target parallel sorting tasks based on the deep reinforcement learning algorithm, generate the sorting execution sequence, and send the sorting control instruction to the pneumatic sorting actuator through the programmable logic controller, wherein the sorting control instruction includes the target category, execution time and pneumatic nozzle opening duration, so as to realize the accurate classification and delivery of renewable resources.

[0160] According to a third aspect of the embodiments of the present invention,

[0161] An electronic device is provided, comprising:

[0162] processor;

[0163] a memory for storing processor-executable instructions;

[0164] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0165] According to a fourth aspect of the embodiments of the present invention,

[0166] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0167] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A renewable resource intelligent sorting control method based on the YOLO framework, characterized in that: include: The image of the renewable resources on the conveyor belt is collected in real time by a high-definition camera arranged above the renewable resources sorting conveyor belt, and the image of the renewable resources is normalized to obtain a standardized image, and the standardized image is input into a pre-trained YOLOv5 deep neural network. The YOLOv5 deep neural network adopts an improved CSPDarknet53 as the backbone network, and introduces a dual attention module of a channel attention mechanism and a spatial attention mechanism in the feature extraction network; Using the YOLOv5 deep neural network to perform feature extraction and target detection on the input standardized image, obtain the position coordinate information and category probability information of the renewable resource target, determine the center point position, width and height of each renewable resource target according to the position coordinate information, classify the detected renewable resource targets into plastic, metal, glass, paper and other categories based on the category probability information, and calculate the confidence score of each target at the same time, retain the targets with confidence scores greater than the preset confidence threshold and generate target detection result data; According to the position coordinate information and category information in the target detection result data, combined with the conveyor belt speed parameters, the estimated time for each renewable resource target to reach the sorting actuator is calculated, and the multi-target parallel sorting tasks are optimized and scheduled based on the deep reinforcement learning algorithm. The sorting execution sequence is generated, and the sorting control instructions are sent to the pneumatic sorting actuator through the programmable logic controller, wherein the sorting control instructions include the target category, execution time and pneumatic nozzle opening duration, so as to realize the accurate classification and delivery of renewable resources.

2. The method according to claim 1, characterized in that Using the YOLOv5 deep neural network to perform feature extraction and target detection on the input standardized image includes: A dual attention mechanism is set after each stage network of the backbone network, and the dual attention mechanism includes a spatial attention module and a channel attention module, wherein the spatial attention module extracts the salient features of the target area and the global context features respectively through the maximum pooling branch and the average pooling branch, and the output features of the maximum pooling branch and the output features of the average pooling branch are input into a three-by-three convolutional layer and then fused to generate a spatial attention weight matrix, and the channel attention module calculates the importance weight of each channel using the global feature information, and the spatial attention weight matrix and the channel importance weight are adaptively fused through a learnable weight coefficient to obtain a first enhanced feature map; Inputting the first enhanced feature map into a compressible excitation module in a feature pyramid network, the compressible excitation module comprising a global average pooling layer and a two-layer fully connected network, the global average pooling layer extracts channel dimension statistical features, the two-layer fully connected network learns the channel dependency relationship based on the channel dimension statistical features to obtain channel importance weights, and performing weighted fusion on different feature layers of the first enhanced feature map based on the channel importance weights to obtain a second enhanced feature map; Inputting the second enhanced feature map into a dynamic anchor box optimization module in the detection head network, the dynamic anchor box optimization module includes an anchor box deformation network branch and an intersection-and-union prediction branch, wherein the anchor box deformation network branch predicts deformation parameters using the second enhanced feature map as input, deforms the initial anchor box based on the predicted deformation parameters to obtain a deformed anchor box, and the intersection-and-union prediction branch evaluates the matching degree between the deformed anchor box and the true target box and feeds back the matching degree information to the anchor box deformation network branch for iterative optimization to obtain an optimized detection box set; The optimized detection frame set is post-processed, and the overlapping detection frames are screened using an improved soft non-maximum suppression algorithm, the improved soft non-maximum suppression algorithm extracts the depth feature vector in the overlapping detection frame, calculates the cosine distance of the depth feature vector to obtain the feature similarity, performs a weighted combination of the feature similarity and the geometric overlap of the detection frame to obtain a suppression score, and based on the suppression score, the overlapping detection frames in the optimized detection frame set are screened to obtain the final target detection result, wherein the weight coefficient of the suppression score is determined by optimizing the validation set.

3. The method according to claim 2, characterized in that Post-processing the optimized detection frame set and screening the overlapping detection frames using an improved soft non-maximum suppression algorithm includes: Arrange the overlapping detection frame set in descending order according to the confidence, select the overlapping detection frames whose confidence is greater than a preset confidence threshold as the candidate frame set, and use the ROI Align operation to uniformly sample the detection frames in the candidate frame set as feature blocks; Input the feature block into a two-layer fully connected network to obtain a deep feature vector, calculate the cosine distance of the deep feature vector corresponding to the overlapping detection frame in the candidate frame set to obtain feature similarity, and calculate the generalized intersection-over-union ratio of the overlapping detection frame to obtain geometric overlap; A dynamic weight factor is set based on the confidence of the detection frame in the candidate frame set, and the feature similarity and the geometric overlap are weighted based on the dynamic weight factor to obtain a suppression score; Softly suppressing the candidate frame set according to the suppression score, suppressing the corresponding detection frame when the suppression score is greater than a first suppression threshold, linearly attenuating the confidence of the corresponding detection frame when the suppression score is between the first suppression threshold and a second suppression threshold, and retaining the corresponding detection frame when the suppression score is less than the second suppression threshold, to obtain a screened detection frame set; The position parameters of adjacent detection frames with similar confidence levels in the filtered detection frame set are weighted averaged and fused to obtain the final target detection result.

4. The method according to claim 1, characterized in that: According to the position coordinate information and category information in the target detection result data, combined with the conveyor belt speed parameters, the estimated time for each renewable resource target to arrive at the sorting execution mechanism is calculated, and the multi-target parallel sorting tasks are optimized and scheduled based on the deep reinforcement learning algorithm. The sorting execution sequence is generated, including: Obtain pixel coordinates and category information corresponding to the position coordinate information of the target in the target detection result data, convert the pixel coordinates into physical coordinates through camera calibration parameters, and calculate the estimated time for the target to reach the sorting actuator based on the speed data collected by the conveyor speed sensor and the distance from the detection area to the sorting actuator; Construct a reinforcement learning environment model, taking the target category information, physical coordinates, estimated arrival time, working status information and location information of the sorting execution mechanism, and buffer occupancy information as the state space, and taking the target allocation decision, execution timing arrangement, and buffer usage strategy of the sorting execution mechanism as the action space; The reinforcement learning environment model is trained using a dual Q learning network structure, wherein the dual Q learning network includes a strategy network and a target network, and a reward function is constructed based on sorting accuracy, throughput, and energy efficiency; The priority score of the target is calculated according to the expected arrival time of the target, the unit time benefit and the complexity of the sorting action, the multi-target sorting tasks are sorted based on the priority score, and the time axis is discretized into time windows of fixed length; within each time window, the action decision based on the output of the policy network is combined with the kinematic constraints, safety spacing requirements and energy balance constraints of the sorting actuator to generate a local optimal scheduling plan, and the sorting execution timing is obtained by sliding the time window.

5. The method according to claim 4, characterized in that The reinforcement learning environment model is trained using a dual Q learning network structure, wherein the dual Q learning network includes a strategy network and a target network, and a reward function is constructed based on sorting accuracy, throughput, and energy efficiency, including: Constructing a policy network and a target network, wherein the policy network and the target network have the same network architecture, the policy network is used for action selection and value function update, the target network is used for calculating the target Q value, the input layer dimensions of the policy network and the target network are consistent with the state vector dimensions, the policy network and the target network perform feature extraction through two hidden layers, and the output layer dimensions of the policy network and the target network correspond to the action space dimensions; The parameters of the policy network are updated based on the temporal difference learning method. For the state and action, the target Q value is calculated. The target Q value is the sum of the product of the immediate reward and the discount factor and the smaller value of the two Q network prediction values. The discount factor is 0.

99. The two Q network prediction values ​​are calculated based on the optimal action selected by the target network in the next state. The parameters of the target network are updated by a soft update mechanism, the parameters of the target network are updated by an exponential sliding average of the policy network parameters, and the soft update coefficient of the exponential sliding average is set to 0.001; Construct a combined reward function, which includes a sorting accuracy reward, a system throughput reward, and an energy efficiency reward. The sorting accuracy reward is in binary form, with a positive reward of ten for correct sorting and a negative penalty of ten for incorrect sorting. The system throughput reward is implemented by exponential function mapping, with a mapping coefficient of 0.2 for the exponential function. The value range of the system throughput reward is from zero to five. The energy efficiency reward takes into account the motion energy consumption of the actuator and the energy consumption of the cache unit, and the value range of the energy efficiency reward is from negative five to zero. Calculate an empirical priority based on a timing differential error, where the empirical priority is equal to the sixth power of the sum of the absolute value of the timing differential error and a small positive number, and the sampling probability is proportional to the empirical priority; The epsilon greedy strategy is used for exploration. The initial value of the epsilon greedy strategy is one, and it decreases to 0.01 during the 500,000 training steps according to the exponential decay law. Each training round contains one thousand time steps, the batch size is 256, and the gradient clipping technology is used to limit the gradient norm to the range of negative ten to positive ten.

6. The method according to claim 1, characterized in that The priority score of the target is calculated according to the expected arrival time of the target, the unit time benefit and the complexity of the sorting action, and the multi-target sorting tasks are sorted based on the priority score, and the time axis is discretized into a time window of fixed length; in each time window, the action decision output by the strategy network is combined with the kinematic constraints, safety spacing requirements and energy consumption balance constraints of the sorting actuator to generate a local optimal scheduling solution, including: Construct a target priority scoring model, wherein the target priority scoring model calculates the comprehensive priority score of the target based on the time urgency score, the unit time benefit score and the action complexity score; the time urgency score is calculated according to the time difference between the target estimated arrival time and the current time; the unit time benefit score is calculated according to the target value, the estimated sorting processing time and the mechanical movement time; the action complexity score is calculated according to the planned path length and the posture change amount; The time axis is discretized into a time window of fixed length, the time window length is two seconds, and the overlap of adjacent time windows is fifty percent; in each time window, the objects to be sorted are sorted in descending order based on the comprehensive priority score, and an action decision sequence is generated according to the descending sorting result; the action decision sequence must simultaneously satisfy the speed constraint that the actuator speed is less than the maximum allowable speed, the acceleration constraint that the actuator acceleration is less than the maximum allowable acceleration, the safety distance constraint that the minimum distance between adjacent actuators is greater than the safety distance threshold, and the energy consumption balance constraint that the accumulated energy consumption in the time window is less than the energy consumption threshold; wherein the safety distance threshold is one and a half times the working radius of the actuator; The scheduling scheme is updated in real time based on an event-driven mechanism. When a new sorting target arrives, a sorting task is completed, or an abnormal event occurs, re-planning is triggered; re-planning includes recalculating the comprehensive priority score based on the current system state, updating the descending order of the sorting targets, and generating an action decision sequence that meets the constraints; At the end of each time window, the system status information is updated, the time window is slid to the next moment, and the sorting target sorting and action decision sequence generation based on the comprehensive priority score are repeated until all sorting tasks are completed.

7. A renewable resource intelligent sorting control system based on the YOLO framework, used to implement the method described in any one of claims 1 to 6, characterized in that: include: The first unit is used to collect images of renewable resources on the conveyor belt in real time through a high-definition camera arranged above the renewable resource sorting conveyor belt, perform size normalization processing on the images of renewable resources to obtain standardized images, and input the standardized images into a pre-trained YOLOv5 deep neural network. The YOLOv5 deep neural network adopts an improved CSPDarknet53 as a backbone network, and introduces a dual attention module of a channel attention mechanism and a spatial attention mechanism in a feature extraction network; The second unit is used to perform feature extraction and target detection on the input standardized image using the YOLOv5 deep neural network, obtain the position coordinate information and category probability information of the renewable resource target, determine the center point position, width and height of each renewable resource target according to the position coordinate information, classify the detected renewable resource targets into plastic, metal, glass, paper and other categories based on the category probability information, and calculate the confidence score of each target at the same time, retain the targets with confidence scores greater than a preset confidence threshold and generate target detection result data; The third unit is used to calculate the estimated time for each renewable resource target to reach the sorting execution mechanism according to the position coordinate information and category information in the target detection result data, combined with the conveyor belt speed parameters, optimize the scheduling of multi-target parallel sorting tasks based on the deep reinforcement learning algorithm, generate the sorting execution sequence, and send the sorting control instruction to the pneumatic sorting actuator through the programmable logic controller, wherein the sorting control instruction includes the target category, execution time and pneumatic nozzle opening duration, so as to realize the accurate classification and delivery of renewable resources.

8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Renewable resource sorting management system based on image analysis

    CN120411657A

  • Ship-borne intelligent swimming crab sorting model based on deep learning

    CN120635144A

  • Visual identification automatic sorting system for components

    CN120772155A

  • A component visual recognition automatic sorting system

    CN120772155B

  • Remote control method for waste sorting equipment based on wireless communication

    CN121364641A