Target detection and tracking method and system under storage environment based on improved YOLOv8n-MatchBox

By improving the YOLOv8n-MatchBox method, combining the SimAM attention mechanism and the ShuffleNetV2 lightweight structure, the accuracy and computational cost issues of target detection and tracking in warehousing environments are solved, and efficient target detection and tracking are achieved.

CN120635667AInactive Publication Date: 2025-09-12XINJIANG AGRI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510762729.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing target detection and tracking methods in warehousing environments have strong dependence on data labels, difficulty in cross-domain application, difficulty in detection in complex environments, and lack of three-dimensional coordinate guidance, resulting in low detection and tracking accuracy.

Method used

An improved YOLOv8n-MatchBox method is adopted. By integrating the SimAM attention mechanism into the Neck structure, combining the ShuffleNetV2 lightweight structure and the DeepSORT algorithm, the Hu invariant distance and occlusion judgment mechanism are introduced to build a custom model for the warehouse environment to achieve target detection and tracking.

Benefits of technology

The model's perception of small-scale targets has been enhanced, detection accuracy and tracking accuracy have been improved, computing costs have been reduced, and it is suitable for target detection and tracking in complex warehousing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635667A_ABST
    Figure CN120635667A_ABST
Patent Text Reader

Abstract

The invention provides an improved YOLOv8n-MatchBox-based target detection and tracking method and system in a storage environment. The method comprises the following steps: S1, obtaining video data of personnel and obstacles in the storage environment; s2, constructing a storage environment custom model; s3, inputting the acquired video data into the trained improved YOLOv8n target detection model, and outputting personnel and obstacle information; s4, inputting a movable person model and an obstacle model; s5, perfecting the storage environment custom model; s6, setting a target shielding threshold value; s7, inputting personnel and obstacle information output by the improved YOLOv8n target detection model into a KCF algorithm and an improved DeepSORT algorithm; according to the method, the detection result of the improved YOLOv8n model is used as the input of the KCF and the improved DeepSORT tracking model, a ShuffleNetV2 lightweight structure is introduced into the DeepSORT model, and the real three-dimensional coordinates are provided for the tracking system in cooperation with the customized environment model specified for the storage environment, so that the calculation cost is reduced while the tracking precision is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine vision technology, and in particular to a target detection and tracking method and system in a warehousing environment based on an improved YOLOv8n-MatchBox. Background Art

[0002] The main task of target detection is to locate the target of interest from the input image and then accurately determine the category of each target of interest. At the same time, target detection is the basis for other advanced vision problems such as scene segmentation, behavior understanding, motion recognition, and video content retrieval. Currently, target detection technology has been widely used in fields such as intelligent transportation, intelligent monitoring, drones, and military reconnaissance, and has a very broad prospect. In recent years, deep learning-based methods have dominated the field of target detection and have produced a series of excellent models, such as the R-CNN series and the YOLO series. However, existing target detection algorithms still have problems such as strong dependence on data labels, difficulty in cross-domain application, and difficulty in detection in complex environments. Traditional target detection and tracking methods typically rely on hand-crafted features and rules to identify and track targets. However, in complex conditions such as warehouse environments, targets often experience complex variations in scale, perspective, occlusion, and illumination, as well as background interference. This makes methods based on hand-crafted features and rules unable to effectively address these challenges. Furthermore, the lack of 3D coordinate guidance specific to warehouse environments makes traditional target detection and tracking methods less robust to complex target variations and background interference, resulting in low detection and tracking accuracy. In order to solve the above technical problems, a target detection and tracking method and system based on improved YOLOv8n-MatchBox in a warehousing environment are proposed. Summary of the Invention

[0003] In view of this, the present invention provides a target detection and tracking method and system in a warehousing environment based on an improved YOLOv8n-MatchBox to solve or alleviate the technical problems existing in the prior art and at least provide a beneficial option.

[0004] The technical solution of the present invention is implemented as follows: a target detection and tracking method in a warehouse environment based on an improved YOLOv8n-MatchBox, comprising the following steps: S1. Obtain video data of people and obstacles in the warehouse environment; S2. Build a custom model for the storage environment; S3. Input the acquired video data into the trained improved YOLOv8n target detection model and output the personnel and obstacle information. S4. Input movable personnel model and obstacle model; S5. Improve the custom model of the storage environment; S6. Set the target occlusion threshold. When the target is occluded, use the improved DeepSORT algorithm to track people and obstacles, calculate the Hu invariant distance to verify the position, and when the target is not occluded, use KCF tracking; S7. Input the personnel and obstacle information output by the improved YOLOv8n target detection model into the KCF algorithm and the improved DeepSORT algorithm, determine the target occlusion threshold, and select the target detection algorithm in MatchBox to continuously track all personnel and obstacles. S8. When the target finishes detecting the current frame, it automatically determines whether to read the next frame content.

[0005] Further preferably, in said S1, the personnel and obstacles in the storage environment are perceived in real time by a vehicle-mounted binocular camera to obtain video data of the personnel and obstacles in the storage environment.

[0006] Further preferably, in said S2, a virtual model of the storage environment is constructed based on the acquired video data, three-dimensional data is provided to the tracking system, and the personnel movement data in the storage environment is fed to the model so that it can predict the personnel movement path.

[0007] Further preferably, in the S3, an image dataset for detecting people and obstacles in a warehousing environment is constructed, all people and obstacles in the dataset are labeled, and a dataset for training the model is generated. The dataset is divided into three parts: a training set, a validation set, and a test set. The training set is used to train the improved YOLOv8n model, the validation set is used to adjust the model parameters and hyperparameters, and the test set is used to evaluate the final performance of the model, so that the trained model has stronger generalization ability while optimizing the model parameters and evaluating the model performance. The improved YOLOv8n model includes inserting a SimAM attention mechanism into the SPPCSPC module and the Cat structure of the Neck part of YOLOv8n.

[0008] Further preferably, in said S4, the personnel distribution data and the obstacle distribution data are entered into a custom model of the warehouse environment, and the possible movement paths of the personnel and some obstacles are identified.

[0009] Further preferably, in S6, the Hu invariant distance is calculated to verify the predicted position.

[0010] Further preferably, in the S7, each detection frame output by the improved YOLOv8n target detection model, including the position information, confidence and feature vector of people and obstacles, is used as input, and the feature training network of DeepSORT is improved, and the lightweight network ShuffleNetV2 is introduced to retrain the feature extraction model, thereby constructing a new multi-target tracking model DeepSORT-SNV2. ShufflNetV2 is a lightweight feature extraction model including a basic unit and a downsampling unit.

[0011] Further preferably, in S8, each frame of the video acquired by the binocular camera is detected and tracked, and its main content process includes: acquiring each frame of the video by the binocular camera and tracking it. For the case where people and obstacles are blocked in the video sequence, the improved DeepSORT algorithm can use the Kalman filter to predict the position of the target, and match it through appearance features when the target reappears.

[0012] A target detection and tracking system in a warehouse environment based on an improved YOLOv8n-MatchBox includes a binocular camera, a communication interface, a processor, and a memory. The binocular camera, the processor, the memory, and the communication interface communicate with each other. The memory stores program instructions executed by the processor, and the processor calls the program instructions.

[0013] Further preferably, the memory adopts a non-transitory computer-readable storage medium, and the non-transitory computer-readable storage medium stores computer instructions.

[0014] The embodiment of the present invention adopts the above technical solution, which has the following advantages: First, this paper integrates the SimAM attention mechanism into the Neck structure of the YOLOv8n model to effectively improve the model's perception of small-scale people and obstacles, thereby enhancing the detection performance of people and obstacles in complex warehousing scenarios. Then, the detection results of the improved YOLOv8n model are used as the input of the KCF and improved DeepSORT tracking models. The ShuffleNetV2 lightweight structure is introduced into the DeepSORT model, and combined with a custom environment model specified for the warehousing environment, it provides the tracking system with real three-dimensional coordinates, thereby ensuring tracking accuracy while reducing computational costs. 2. The present invention designs a tracking algorithm based on the occlusion judgment mechanism. By quantifying the relationship between the correlation quantity QROI and the occlusion state, the correlation quantity threshold is determined, and the tracking state is segmented according to the change of the correlation quantity. Based on the Kalman filter of the improved DeepSORT algorithm, the target position is predicted, and the authenticity of the predicted position is verified by using the Hu invariant distance method, which effectively improves the accuracy and speed of target detection and tracking, and provides technical support for the practical application of personnel and obstacle detection and tracking in warehousing environments.

[0015] The above summary is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments and features described above, further aspects, embodiments and features of the present invention will be readily apparent by reference to the accompanying drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0017] Figure 1 This is a flow chart of the target detection and tracking method in a warehouse environment based on the improved YOLOv8n-MatchBox of the present invention; Figure 2 This is a diagram of the target detection and tracking system in a warehouse environment based on the improved YOLOv8n-MatchBox of the present invention; Figure 3 This is a flow chart of the KCF algorithm in the present invention; Figure 4 This is the flow chart of the DeepSORT algorithm in the present invention. DETAILED DESCRIPTION

[0018] Hereinafter, only certain exemplary embodiments are briefly described. As will be appreciated by those skilled in the art, the described embodiments may be modified in various ways without departing from the spirit or scope of the present invention. Therefore, the drawings and description are to be considered as illustrative in nature and not restrictive.

[0019] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0020] like Figure 1 、 Figure 3 and Figure 4As shown, an embodiment of the present invention provides a target detection and tracking method in a warehouse environment based on an improved YOLOv8n-MatchBox, comprising the following steps: S1. Obtain video data of people and obstacles in the warehouse environment. Use the vehicle-mounted binocular camera to perceive people and obstacles in the warehouse environment in real time to obtain video data of people and obstacles in the warehouse environment. S2. Build a custom model of the warehouse environment. This virtual model is constructed based on the acquired video data, providing three-dimensional data for the tracking system. The model is fed with data on the movement of people in the warehouse environment, enabling it to predict their movement paths. S3. Input the acquired video data into the trained improved YOLOv8n target detection model, output personnel and obstacle information, construct an image dataset for personnel and obstacle detection in a warehouse environment, annotate all personnel and obstacles in the dataset, generate a dataset for training the model, and divide the dataset into three parts: training set, validation set, and test set. The training set is used to train the improved YOLOv8n model, the validation set is used to adjust the model parameters and hyperparameters, and the test set is used to evaluate the final performance of the model, so that the trained model has stronger generalization ability while optimizing the model parameters and evaluating the model performance. The improved YOLOv8n model includes inserting the SimAM attention mechanism into the SPPCSPC module and Cat structure of the Neck part of YOLOv8n; S4. Enter the movable personnel model and obstacle model, enter the personnel distribution data and obstacle distribution data into the custom model of the warehouse environment, and identify the possible movement paths of personnel and some obstacles; S5. Improve the custom model of the storage environment; S6. Set the target occlusion threshold. When the target is occluded, use the improved DeepSORT algorithm to track people and obstacles, and calculate the Hu invariant distance to verify the position. When the target is not occluded, use KCF tracking to calculate the Hu invariant distance to verify the predicted position. S7. Input the personnel and obstacle information output by the improved YOLOv8n target detection model into the KCF algorithm and the improved DeepSORT algorithm, and determine the target occlusion threshold. Select the target detection algorithm in MatchBox to continuously track all personnel and obstacles. Take each detection box output by the improved YOLOv8n target detection model, including the position information, confidence, and feature vector of the personnel and obstacles, as input. By improving the feature training network of DeepSORT, the lightweight network ShuffleNetV2 is introduced to retrain the feature extraction model, and then a new multi-target tracking model DeepSORT-SNV2 is constructed. ShufflNetV2 is a lightweight feature extraction model including basic units and downsampling units. S8. When the target finishes detecting the current frame, it automatically determines whether to read the next frame content, detects each frame of the video obtained by the binocular camera and tracks it. The main content process includes: tracking each frame of the video obtained by the binocular camera. In the case where people and obstacles are blocked in the video sequence, the improved DeepSORT algorithm can use the Kalman filter to predict the position of the target, and match it based on appearance features when the target reappears.

[0021] In one embodiment, in S3, SimAM is a simple and effective parameter-free three-dimensional attention module. By adjusting the spatial attention of the input features, the module can not only improve the model's ability to pay attention to information at different positions, but also find the importance of each neuron without increasing the original network parameters, and infer the three-dimensional attention weight for the feature map. Embedding the SimAM attention mechanism in the SPPCSPC module can enhance the model's adaptability to human targets of different scales, proportions, directions, etc., thereby improving the accuracy of human and obstacle target detection; introducing the SimAM attention mechanism in the Cat structure can reduce the interference of complex environments on target detection in an end-to-end manner, so that the network can better extract more critical visual features of people and obstacles, further improving model performance. In addition, in order to make the model both fast and accurate, data enhancement methods are used in the network training process to enrich sample features.

[0022] In one embodiment, the Hu invariant distance is calculated to verify the predicted position, and the specific steps include: Assume that the coordinates of a two-dimensional image are (x, y), the size is M×N, and the grayscale value is f(x, y). In the continuous case, the definition of the p+q order geometric moment of the image is as shown in formula (1): (1) The P+Q order center distance of the image is defined as shown in formula (2): (2) Where (x, (y) represents the center of gravity of the image, and the geometric distance of the discrete digital image is shown in formula (3): (3) Where M and N represent the image width and image height respectively, and the normalized center distance is defined as shown in formula (4): (4) Where ρ = (p + q) / 2 + 1, using the second-order normalization and third-order normalization center distance, we can derive the seven Hu invariant distances of the image, which remain unchanged when the image is rotated, translated, and scaled, as shown in Equation (5): (5) For the seven Hu invariant distance normalized cross-correlation operations, the correlation value of the region of interest ROI of the initial frame and the current frame of the image in the video sequence is calculated, which is called QROI. The calculation formula is shown in formula (6); then this value is used to segment the learning rate; (6) Wherein, NP(i) is the i-th invariant distance feature of the target object in the current frame, and is the mean of the seven invariant distance features of the target object in the current frame; NR(i) is the i-th invariant distance feature of the target object in the initial frame, and is the mean of the seven invariant distance features of the object in the initial frame. The larger the QROI value, the greater the correlation between the initial frame and the current frame, and vice versa. QROI∈[0,1]. Due to the invariance characteristics of the seven Hu invariant distances, the QROI value remains unchanged when the image is rotated, translated or scaled, but the relevant amount of the target object will change under occlusion. Therefore, the change of QROI is used to judge whether the target is occluded, and then an occlusion judgment mechanism is constructed.

[0023] In one embodiment, establishing an occlusion judgment mechanism specifically includes: during the video detection process, judging whether the target is occluded by the change of the relevant quantity QROI value. Therefore, the learning rate in the KCF algorithm is segmented using the relevant quantity, and an occlusion judgment mechanism is established to achieve adaptive update of the target model, as shown in formula (7): (7) Where Q represents the correlation quantity, φ represents the threshold of the correlation quantity, and β is the learning rate in the KCF algorithm. When the Q value is greater than or equal to the threshold of the correlation quantity, it means that the target object is not occluded. At this time, the learning rate of the KCF algorithm is set to 0.02. When the Q value is less than the threshold of the correlation quantity, it means that the target object is occluded. At this time, the learning rate of the KCF algorithm is set to 0.

[0024] In one embodiment, in S7, feature extraction: extracting feature vectors for tracking from the detection results of the improved YOLOv8n; preliminary screening: filtering out low-confidence detection results according to the confidence threshold to reduce false detection; initialization tracking: for each remaining detection frame, using the KCF algorithm for preliminary target tracking; occlusion judgment: calculating the occlusion degree of each detection frame, and judging whether it exceeds the preset occlusion threshold; target tracking: when the target is not occluded, the KCF algorithm is used to track the target normally; when the judgment mechanism determines that the target is occluded, the KCF tracking is stopped and switched to the improved DeepSORT algorithm to predict the target position, and QROI is used to judge the authenticity of the target that appears again. When the target reappears after occlusion, the KCF algorithm is continued to be used for tracking until the end of the last frame content.

[0025] In one embodiment, in S7, the basic unit first introduces a channel separation operation to divide the input feature map channel into two independent branches with equal number of channels. The left branch is not operated, and the right branch performs three consecutive convolution layer operations, specifically including one depth-separable convolution layer and two ordinary 1×1 convolution layers that fuse feature information between channels. Then, the outputs of the left and right branches are spliced ​​together through the Concat operation, aiming to improve the calculation speed while ensuring that the number of output channels is the same as the number of input channels; then, a channel reorganization operation is performed on the Concat processing result to ensure information exchange between channels. The downsampling unit does not use the channel separation operation, but directly copies the two branches and performs downsampling with a step size of 2. The downsampling process is completed by performing similar operations and then splicing them together. The spatial size of the resulting feature map is halved and the number of channels is twice that of the input. Compared with traditional deep neural network structures such as ResNet and VGG, ShuffleNetV2 introduces a channel reorganization operation. This operation can reduce the demand for computing and storage while maintaining model performance, resulting in a lighter design. This also enables it to be better deployed on embedded systems, mobile devices or edge computing devices. The DeepSORT-SNV2 network has higher computing efficiency and fewer model parameters without losing network calculation accuracy, and better balances the relationship between speed and accuracy.

[0026] In one embodiment, in S8, based on the KCF algorithm, an occlusion judgment mechanism is set using a constant distance. When the target is not occluded, the KCF algorithm is used to track the target normally. When the judgment mechanism determines that the target is occluded, KCF tracking is stopped and the improved DeepSORT algorithm is switched to predict the target position. At the same time, QROI is used to judge the authenticity of the target that appears again. When the target appears again after being occluded, the KCF algorithm is continued to be used for tracking until the end of the last frame content.

[0027] like Figure 2 As shown, an embodiment of the present invention also provides a target detection and tracking system in a warehouse environment based on the improved YOLOv8n-MatchBox, including a binocular camera, a communication interface, a processor and a memory. The binocular camera, the processor, the memory and the communication interface communicate with each other. The memory stores program instructions executed by the processor, and the processor calls the program instructions. The memory adopts a non-transitory computer-readable storage medium, and the non-transitory computer-readable storage medium stores computer instructions.

[0028] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various modifications and substitutions within the technical scope disclosed in the present invention, and such modifications and substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A target detection and tracking method in a warehouse environment based on improved YOLOv8n-MatchBox, characterized in that: The following steps are involved: S1. Obtain video data of people and obstacles in the warehouse environment; S2. Build a custom model for the storage environment; S3. Input the acquired video data into the trained improved YOLOv8n target detection model and output the personnel and obstacle information. S4. Input movable personnel model and obstacle model; S5. Improve the custom model of the storage environment; S6. Set the target occlusion threshold. When the target is occluded, use the improved DeepSORT algorithm to track people and obstacles, calculate the Hu invariant distance to verify the position, and when the target is not occluded, use KCF tracking; S7. Input the personnel and obstacle information output by the improved YOLOv8n target detection model into the KCF algorithm and the improved DeepSORT algorithm, determine the target occlusion threshold, and select the target detection algorithm in MatchBox to continuously track all personnel and obstacles. S8. When the target finishes detecting the current frame, it automatically determines whether to read the next frame content.

2. The target detection and tracking method in a warehouse environment based on the improved YOLOv8n-MatchBox according to claim 1 is characterized in that: In S1, the vehicle-mounted binocular camera is used to perceive people and obstacles in the storage environment in real time to obtain video data of people and obstacles in the storage environment.

3. The target detection and tracking method in a warehouse environment based on the improved YOLOv8n-MatchBox according to claim 1 is characterized in that: In said S2, a virtual model of the warehouse environment is constructed based on the acquired video data, providing three-dimensional data for the tracking system, and feeding the model with personnel movement data in the warehouse environment so that it can predict the personnel movement path.

4. The target detection and tracking method in a warehouse environment based on the improved YOLOv8n-MatchBox according to claim 1, characterized in that: In the S3, an image dataset for detecting people and obstacles in a warehouse environment is constructed, all people and obstacles in the dataset are labeled, and a dataset for training the model is generated. The dataset is divided into three parts: a training set, a validation set, and a test set. The training set is used to train the improved YOLOv8n model, the validation set is used to adjust the model parameters and hyperparameters, and the test set is used to evaluate the final performance of the model, so that the trained model has stronger generalization ability while optimizing the model parameters and evaluating the model performance. The improved YOLOv8n model includes inserting the SimAM attention mechanism into the SPPCSPC module of the Neck part of YOLOv8n and the Cat structure.

5. The target detection and tracking method in a warehouse environment based on the improved YOLOv8n-MatchBox according to claim 1, characterized in that: In said S4, the personnel distribution data and the obstacle distribution data are entered into the custom model of the warehouse environment, and the possible movement paths of the personnel and some obstacles are identified.

6. The target detection and tracking method in a warehouse environment based on the improved YOLOv8n-MatchBox according to claim 1, characterized in that: In S6 , the Hu invariant distance is calculated to verify the predicted position.

7. The target detection and tracking method in a warehouse environment based on the improved YOLOv8n-MatchBox according to claim 1, characterized in that: In the S7, each detection frame output by the improved YOLOv8n target detection model, including the position information, confidence and feature vector of people and obstacles, is used as input. By improving the feature training network of DeepSORT, the lightweight network ShuffleNetV2 is introduced to retrain the feature extraction model, and then a new multi-target tracking model DeepSORT-SNV2 is constructed. ShufflNetV2 is a lightweight feature extraction model including a basic unit and a downsampling unit.

8. The target detection and tracking method in a warehouse environment based on the improved YOLOv8n-MatchBox according to claim 1, characterized in that: In S8, each frame of the video captured by the binocular camera is detected and tracked. The main content process includes: each frame of the video captured by the binocular camera is tracked. In the case where people and obstacles are blocked in the video sequence, the improved DeepSORT algorithm can use the Kalman filter to predict the position of the target and match it based on appearance features when the target reappears.

9. A target detection and tracking system in a warehouse environment based on an improved YOLOv8n-MatchBox, characterized by: The system comprises a binocular camera, a communication interface, a processor and a memory. The binocular camera, the processor, the memory and the communication interface communicate with each other. The memory stores program instructions executed by the processor, and the processor calls the program instructions.

10. The target detection and tracking system in a warehouse environment based on the improved YOLOv8n-MatchBox according to claim 9, characterized in that: The memory employs a non-transitory computer-readable storage medium that stores computer instructions.