Anomaly recognition method for sorting center based on video detection technology
By applying video detection technology and LSTM network in the sorting center, combined with multi-object tracking algorithm, automatic identification and management of item abnormalities is achieved, and the problems of difficulty and low efficiency of manual identification in the existing technology are solved, and transportation safety and service quality of logistics enterprises are improved.
Patent Information
- Application Number
- CN202210272042.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-18
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-03-18
AI Technical Summary
Existing logistics companies use manual inspection and video search methods in sorting centers to find out the abnormalities of items, resulting in loss of items, cassette problems, etc., affecting service quality and corporate reputation.
The sorting center abnormal recognition method based on video detection technology is adopted, combined with long and short-term memory network (LSTM) and multi-object tracking algorithm, and automatic identification and management of item abnormalities through video object detection, trajectory prediction and abnormal identification.
It improves the item detection and tracking accuracy and abnormal identification capabilities of the sorting center, reduces manual intervention, improves transportation safety and efficiency, reduces item loss rate, and improves the service quality and credibility of logistics companies.
Smart Images

Figure CN114663808B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of video detection, and in particular relates to a method for identifying anomalies in a sorting center based on video detection technology. Background Art
[0002] With the improvement of residents' consumption level and the rise of online shopping in China, the logistics industry has been rapidly developed. Many new enterprises have entered the logistics industry to seize the market, and established enterprises have maintained their competitive advantages through price advantages and improved service quality. Under the fierce competition in the industry, safety and speed have become important evaluation indicators for logistics enterprises.
[0003] Regarding the safety of goods, the focus is on the two processes of sorting and "last mile" delivery of transported goods. Because there are too many factors involved in delivery problems caused by regions, corporate systems, social culture, etc., this article will not consider them for the time being. The focus is on ensuring the safety of transported goods and transportation processes during the sorting process.
[0004] According to the survey, existing logistics companies mostly use manual inspection and video search to solve the safety problems of items in sorting centers, which is always unsatisfactory for some small items. At the same time, when faced with problems such as item jamming and item loss, manual intervention methods are backward, lack timeliness, waste resources and fail to achieve good results. The loss of high-priced items will greatly reduce the service quality of logistics companies and even cause the loss of corporate reputation of logistics companies, which will have a very bad impact on the company.
[0005] The multi-target tracking problem uses a multi-target tracking algorithm to match the existing target trajectory based on the target detection results in each frame of the image; for newly appearing targets, new targets need to be generated; for targets that have left the camera's field of view, the trajectory tracking needs to be terminated. In this process, the matching of targets and detections can be regarded as target re-identification.
[0006] Recurrent Neural Network (RNN), also known as Time Recursive Neural Network, is different from a fully connected network in that the value of the hidden layer in a fully connected network depends only on the input layer, and the signal of the neuron can only propagate to the upper layer, while the state of the hidden layer in the RNN depends not only on the input layer, but also on the previous state, that is, the output of the neuron can also be fed back to itself at the next moment. This feature of RNN enables RNN to focus on multiple input values in advance, allowing RNN to learn and display the temporal dynamic behavior of time series. The purpose of designing RNN is to give it the ability to remember like a human, so RNN is usually used to handle tasks of time-varying systems.
[0007] Long short-term memory (LSTM) is a special RNN that is mainly used to solve the gradient vanishing and gradient exploding problems in the long sequence training process. Compared with ordinary RNN, LSTM can perform better in longer sequences and can remember historical information for a longer period of time. LSTM has a chain structure composed of repeated modules like RNN. The difference is that the repeated modules of LSTM add some gate structures on the basis of the repeated modules of RNN. These gates are some control parameters used to control the flow of information in and between networks. Summary of the invention
[0008] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a sorting center anomaly recognition method based on video detection technology. By using video detection technology and big data processing, a long short-term memory network is introduced to improve the multi-target algorithm, and video targets in the logistics sorting center environment can be detected and tracked. The RS-Loss loss function is used to simplify the neural network model and improve the running speed.
[0009] The present invention provides a method for identifying abnormalities in a sorting center based on video detection technology, comprising the following steps:
[0010] Step S1. When the system starts, the first frame of the video is detected, and three frames are detected continuously;
[0011] Step S2. Run the object detector to perform video object detection to obtain the bounding box of the object;
[0012] Step S3. Build a target state prediction model based on LSTM to predict the target trajectory;
[0013] Step S4. Construct a TWP tracker to implement trajectory similarity calculation and data association, and assign a digital ID to each object;
[0014] Step S5. Create a track token to manage the status of the tracking target that has been assigned a digital ID;
[0015] Step S6: Establish a target anomaly identifier to identify tracking targets with abnormal status.
[0016] As a further technical solution of the present invention, in step S2, the specific process of running the object detector to perform video target detection includes the following steps:
[0017] Step S21. Perform a convolution operation on the original image through the image preprocessing module in the Musk convolutional neural network model to obtain a common feature map as the input of the RPN network. The size of the common feature map is set to N×H×W pixels, where H is the height of the original image and W is the bottom of the original image.
[0018] Step S22. Perform a 3×3 convolution on the common feature map to obtain a feature map of 256×H×W pixels, that is, there are H×W 256-dimensional feature vectors in the map;
[0019] Step S23. Predetermine anchor points and standard boxes. Each point on the feature map corresponds to an area on the original image. The area size is m×m, where m is the ratio of the original image to the feature map. The center of each area is set as the anchor point. Each anchor point determines K standard boxes. Therefore, H×W points in the feature map correspond to H×W×K boxes in the original image. The aspect ratio of the standard box is 1, the sizes are 64xp; 128xp; 256x0p, K=3×3=9, and three ratios are set for each standard frame.
[0020] Step S24. The feature map contains H×W vectors, each of which is 256-dimensional. Two full-connection operations are performed on each feature vector to obtain two sub-feature maps, namely, the target background feature map and the original image offset feature map;
[0021] The size of the target background feature map is 4K×H×W pixels, K corresponds to the K anchored canonical boxes, 4 represents 4 different states, H is the height of the feature map, W is the bottom of the feature map,
[0022] Set 4 components to represent different states of the target and background:
[0023] a,b,c,d∈{1,0},
[0024] a, c are target quantities,
[0025] b, d are displacements,
[0026] For backgrounds that cannot be moved,
[0027] The background moving within the corresponding time period,
[0028] For temporarily stationary targets,
[0029] Corresponding to the stationary target in the video image,
[0030] The non-movable background is a feature area that has no changes or movements for 20 consecutive frames;
[0031] The size of the original image offset feature map is 4K×H×W pixels, K is the K anchored standard boxes, 4 represents the 4 coordinates of the standard box (x, y, h, w), (x, y) represents the coordinates of the anchor point, that is, the center point of the standard box, (h, w) is the width and height of the standard box, each anchor point corresponds to K standard boxes, marking K×H×W standard boxes;
[0032] Set an offset for the canonical box and a threshold α. If the offset is less than α, it is marked as a positive sample in training. If the offset is greater than α, it is marked as a negative sample.
[0033] Assume that the center points of the canonical frame A and its corresponding real frame B are (x i ,y i ) and (x j ,y j ); the width of A is w i , the width of B is w j ; The height of A is H i , B's is H j ,
[0034] The offset is:
[0035]
[0036] The parameter μ x , μ y , μ w , μ h ∈(0, 0.1),σ x =σ y =0.1,σ h =σ w =0.2;
[0037] Step S25. Obtain candidate boxes through the Musk convolutional neural network model based on the anchor points and feature maps. The loss function of the model training is
[0038]
[0039] Among them, i is the i-th anchor point participating in the training in the feature map, there are N in total, p i is the predicted probability, which represents the probability that the original image defined by the anchor point is the target. is the cross entropy of the binary classification; is the regression of the canonical box.
[0040] Furthermore, in step S2, the setting of the ratio size of the K standard boxes of the anchor points in the RPN network is specifically as follows:
[0041] Let the length, width and height of the carton be (a, b, c), let the width of the outer matrix of the carton in the image video be W, and the height be H,
[0042] The linear mapping formula from matrix to image matrix is
[0043]
[0044]
[0045] Among them, the parameter β is the distance between the carton and the camera, β∈(0,1), the parameter γ is the wide angle between the carton and the camera, γ∈(0.3,1), and the parameter μ is the depression angle between the carton and the camera, μ∈(0.2,1).
[0046] Further, in step S3, the target state prediction model adopts the loss function ME-Loss;
[0047] Object detection and instance segmentation are multiple tasks. The loss function of multiple tasks adopts the weighted summation of the loss functions of multiple subtasks. The formula is in, is the loss function of the t-th task in the ith stage, is a hyperparameter used to determine the weight of each task in each stage;
[0048] Using an integrated loss function, the loss function is defined as
[0049] Given a set of logits and the output of the final fully connected layer, the loss function is calculated as follows:
[0050] Step S31. Calculate the difference x between the sum ij =s j -s i ,
[0051] Step S32. Based on the difference values of and, the difference term obtained for each pair of samples can be expressed as a basic term
[0052]
[0053] Step S33. Normalize and sum L to obtain the final loss function:
[0054]
[0055] Furthermore, in step S3, the target state prediction model based on LSTM uses time series to obtain target motion information, input: x t (x,y,z),z t (r,a,e); output: x t+1 (x,y,z),z t+1 (r,a,e);
[0056] Among them, x t is the state vector of the target at time t, x, y, z are the positions in the three-dimensional coordinate system with the center of the video as the far point; z t is the measurement state of the target at time t, r, a, e represent the distance, azimuth and angle in the polar coordinate system with the camera as the origin respectively;
[0057] Based on the LSTM neuron structure, x t ,z t As the input of the model, X represents the input target set {x0,x1,...,x t , x t+1}, W represents the learnable parameters of the fully connected layer, and σ represents the sigmoid activation function tanh represents the tanh function Concat represents a vector concatenation function, which is used to merge two vectors into a longer vector. h t represents the output set of the hidden layer {h0,h1,......,h t ,h t+1},c t Represents the cell state set of LSTM {c0,c1,......c t ,c t+1 ,}; The specific process of the model from input to output is:
[0058] Step S31a. Input data x t (x,y,z) and z t (r,a,e), merged into a vector through the concat function,
[0059]
[0060] Step S3b. The hidden layer h obtained from the previous node t According to the weight matrix w xh Calculate and get the output of the middle hidden layer:
[0061] Step S3c. Use the sigmoid activation function σ(x) and the tanh function to obtain the cell state c of this node. t+1 ,
[0062] Step S3d. Calculate h t+1 , and the output x t+1 ,
[0063]
[0064] x t+1 =w to ·h t+1 +b o ;
[0065] Among them, b i , b h , b o , is the offset.
[0066] Furthermore, in step S4, the TWP tracking model is divided into two stages. In the first stage, the target detection is taken as a node in the graph to build a network flow model; then the model is solved to obtain a preliminary tracking trajectory; wherein, the data association cost is calculated by the LSTM motion model and the appearance model based on the color histogram; in the second stage, in order to reduce the missed targets and target IDs, the trajectories are clustered and optimized, and in the clustering, the similarity of the targets is calculated by the siamese network and the LSTM motion model.
[0067] Furthermore, in the system in step S1, an abnormality index is added to form a new evaluation index system. The evaluation index includes three factors: tracking accuracy, tracking trajectory consistency and abnormal situation identification rate, as follows:
[0068] Use ID scores to judge tracking accuracy.
[0069] Identification precision:
[0070] Identification recall:
[0071] Identification FI:
[0072] Use SQE to judge the consistency of tracking trajectory.
[0073] SQE estimates the transformation of ID{three digits} by analyzing the feature distance between targets with the same identity and targets with different identities; the Gaussian mixture model is used to measure the distance distribution, and the distance measurement mode uses the Euclidean distance:
[0074] The Gaussian mixture model and Euclidean distance formula of the standard features are used to obtain the distance distribution model, and the feature distance is standardized to obey the chi-square distribution
[0075] The evaluation index formula is:
[0076] Among them, n is the number of trajectories, L is the average length of the trajectory, dif is the degree of identity change within the trajectory, and the mean distance of the Gaussian mixture model is calculated to determine whether there are multiple identities;
[0077] Use Get strange to judge the identification rate of abnormal situations.
[0078]
[0079]
[0080] Among them, FN is False Negative, which is the sum of the number of false negatives detected and tracked in the entire video;
[0081] FP is False Positve, the sum of the number of false positives in the entire video detection and tracking;
[0082] TP is TruePositive correctly marked, the sum of correctly marked targets detected and tracked in the entire video;
[0083] M is the environmental parameter, which is determined according to the number and movement speed of the targets in the video;
[0084] Gs is the anomaly capture rate, which indicates the detection rate of true anomaly targets detected in the video targets;
[0085] TPG is the sum of correctly detected and tracked targets in the entire video taking into account Gs;
[0086] FNG is the sum of the number of missed detections and target tracking in the entire video when Gs is taken into account;
[0087] FPG is the sum of the number of false alarms in the entire video detection and tracking when Gs is taken into account;
[0088] α is the weight of TPG, and the value of α is adjusted according to the recognition accuracy of the anomaly detector.
[0089] Furthermore, in step S5, a track Token is constructed to manage the status of the tracking target to which the digital ID has been assigned. The specific steps are as follows:
[0090] Step S51. Activate the tracking state and assign a tracker to the target detected in three consecutive frames;
[0091] Step S52: when the apparent information of the target is greatly disturbed, the apparent feature matching cannot reach the threshold, and the tracking state is converted to a waiting state;
[0092] Step S53: When the target is activated by the detector again, the features of the latest three frames stored in the feature library are matched with the features of the newly appeared target. If the association is successful, the target is switched to the tracking state again.
[0093] Step 4: If there is no match, it means that a new tracking target has appeared, and a series of initialization operations are performed on the new target; for the successfully tracked target, the features in the target feature library and the various parameters in the target motion system must be updated in real time to complete the update of the system status.
[0094] Furthermore, in step S5, a two-stage tracker is established using a motion model based on the LSTM network to track and predict the trajectory.
[0095] Further, in step S6, in the target management database of the system, the target state is classified into three categories according to the motion state: stationary, uniform motion, and variable speed falling motion. The abnormal conditions of express shipments transported by the sorting center are classified into express shipment jamming, express shipment dropping, and express shipment overlapping. By combining the position information and state information of the express shipment, it is judged and identified whether it is in an abnormal state.
[0096] Assume that the storage attributes of the express in the database are [ID, Place, V, Status],<ID,Place,V,Status> ;
[0097] ID is a digital ID assigned by the system to each target express through the TWP tracker. The digital ID has three states:
[0098]
[0099] Place is the location information of the express shipment. There are three types of Place:
[0100]
[0101] V is the speed information of the express, and there are three levels of V:
[0102]
[0103] Status is the status of the shipment, one is normal and three are abnormal:
[0104]
[0105] Status=f(ID,Place,V)=ID×Place×V.
[0106] The advantages of the present invention are:
[0107] 1. Classical tracking filtering algorithms such as Kalman filtering and UKF need to make assumptions based on different motion modes and the target environment, that is, they require a priori target motion model, noise distribution, etc. The algorithm proposed in this invention does not require any prior knowledge at all; compared with the classical algorithm, it has higher estimation accuracy; it is suitable for estimating different maneuvering target states, and does not require complex parameter adjustment process during estimation.
[0108] 2. Compared with the traditional multi-target tracking evaluation index system, which can only evaluate the detection accuracy and tracking accuracy, the MOTG index proposed in the present invention takes into account both the detection accuracy and the tracking accuracy. At the same time, it introduces the capture rate Gs of abnormal targets, adapts to the specific environment of the sorting center, and achieves a better evaluation effect.
[0109] 3. Aiming at the logistics environment and combining the actual research on logistics item packaging, the specific parameter K was improved in the RPN network, so that the accuracy and speed of the algorithm in deriving candidate boxes were further improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0110] Figure 1 is a system flow chart of the present invention;
[0111] Figure 2 This is a flow chart of the Token-based digital ID management method of the present invention;
[0112] Figure 3 is a flow chart of the track evaluation system of the present invention;
[0113] Figure 4 It is a schematic diagram of the LSTM-based prediction model of the present invention. DETAILED DESCRIPTION
[0114] See also Figure 1 This embodiment provides a method for identifying abnormalities in a sorting center based on video detection technology. The specific process is as follows:
[0115] Step 1: Initialization, the system integrates the video resources. At the beginning, it detects the first frame of the video and detects three frames in succession.
[0116] Step 2: Use the improved RNP network to process the initial image to obtain a feature map of unified pixels, and then process it through two convolutional layers and one pooling layer. The specific steps are as follows:
[0117] Step 1: Through the image preprocessing module in the Musk convolutional neural network model, a series of convolution operations are performed on the original image to obtain a common feature map as the input of the RPN network. The size of the common feature map is set to N×H×W pixels, where H is the height of the original image and W is the bottom of the original image.
[0118] Step 2: Perform a 3×3 convolution on the common feature map to obtain a feature map of 256×H×W pixels, that is, there are H×W 256-dimensional feature vectors in the map.
[0119] Step 3: Predetermine anchor points and standard boxes
[0120] Since there is a one-to-one mapping relationship between the feature map and the original image, each point on the feature map corresponds to an area on the original image. This area is generally very small. Let the size of this area be m×m, where m is the ratio of the original image to the feature map. However, the shape of this area is uncertain and needs to be adjusted according to the detected target and background.
[0121] The center of each area is set as the anchor point. Through this anchor point, it can be determined whether it is the target area. Each anchor point determines K different standard boxes. Therefore, the feature map has H×W points, corresponding to H×W×K different boxes in the original image. The size of the box here is set manually.
[0122] Here, the ratio of K value to the specification box is generally a standard value. In order to achieve more accurate target selection, this article summarizes the most suitable K value and specification box ratio and size in the logistics warehousing environment by combining the express delivery size in the actual logistics market (see the derivation method of right protection point 3), as shown in the following figure:
[0123] Standard box ratio (aspect ratio): 1;
[0124] Standard frame size: 64xp; 128xp; 256x0p;
[0125] K=3×3=9, and three ratios are set for each size of the standard box.
[0126] Step 4: The feature map has H×W vectors, each of which is 256-dimensional. Two full-connection operations are performed on each feature vector to obtain two sub-feature maps:
[0127] 1. Target background feature map - classification target and background
[0128] The size of the feature map is 4K×H×W pixels, K corresponds to the K anchored canonical boxes, 4 corresponds to 4 different states, H is the height of the feature map, and W is the bottom of the feature map.
[0129] Set 4 components to represent different states of the target and background:
[0130]
[0131] a,b,c,d∈{1,0};
[0132] a, c are target quantities,
[0133] b, d are displacements,
[0134] ——Corresponding background, background that cannot be moved (such as shelves, conveyor belts, pillars, walls, etc.)
[0135] ——Corresponding background, background that can move in certain time periods (such as forklifts, stacked goods, etc.)
[0136] ——Corresponding targets, temporarily stationary targets (such as stuck express mail, express mail in temporary storage area, sorting personnel, etc.)
[0137] ——Corresponding target, moving target in the video image (such as express mail on the conveyor belt, sorting personnel, etc.)
[0138] For the background that cannot move, it is set as a feature area that has no changes or movements for 20 consecutive frames. Such fixed background areas will be stripped out in the next feature scan, reducing the recognition of such areas and greatly reducing the time for target detection.
[0139] 2. Original image offset feature map - mark each standard box
[0140] The size of the feature map is 4K×H×W pixels, K corresponds to the K anchored canonical boxes, 4 represents the 4 coordinates of the canonical box (x, y, h, w), (x, y) represents the coordinates of the anchor point, that is, the center point of the canonical box, and (h, w) represents the width and height of the canonical box. Since each pixel of the feature map corresponds to a small area of the original image, each area has an anchor point, and each anchor point corresponds to K canonical boxes, it is necessary to mark K×H×W canonical boxes. Set an offset for the canonical box, discard the canonical box that is larger than the true canonical box, reduce the training amount of the CNN model, set a threshold α, if the offset is less than α, it is marked as a positive sample in training, if the offset is greater than α, it is marked as a negative sample. Perform a special transformation on the offset variable to make the distribution of the offset more uniform and easier to fit.
[0141] Assume that the center points of the canonical frame A and its corresponding real frame B are (x i ,y i ) and (x j ,y j ); the width of A is w i , the width of B is w j ; The height of A is H i , B's is H j .
[0142] The offsets are marked as:
[0143]
[0144] The parameter μ x , μ y , μ w , μ h ∈(0, 0.1),σ x =σ y =0.1,σ h =σ w =0.2.
[0145] Step 5: Combine the anchor points set in step 2 with the K×H×W feature maps obtained in step 4, and obtain the candidate boxes through post-processing of the Musk convolutional neural network model. The loss function of the model training is as follows:
[0146]
[0147] i is the i-th anchor point in the feature map that participates in the training. There are N anchor points in total. i Represents the predicted probability, which represents the probability (0 or 1) that the original image defined by the anchor point is the target. This term calculates the cross entropy for binary classification; This term computes the regression of the canonical box.
[0148] Step 3: Build a target state prediction model based on LSTM to predict the target trajectory. Figure 3 and Figure 4 To conduct specific process analysis;
[0149] Step 1: Input data x t (x,y,z) and z t (r,a,e) are combined into a vector using the concat function.
[0150]
[0151] Step 2: The hidden layer h obtained from the previous node t According to the weight matrix w xh Calculate and get the output of the middle hidden layer:
[0152]
[0153] Step 3: Use the sigmoid activation function σ(x) and the tanh function to obtain the cell state c of this node t+1
[0154]
[0155] Step 4: Calculate h t+1 , the output x t+1
[0156]
[0157] x t+1 =w to ·h t+1 +b o ;
[0158] where b i , b h , b o , is the bias in the formula, which is set manually.
[0159] Step 5: Model training:
[0160] The small batch PMSpro algorithm that combines the adaptive gradient descent method with the Momentum algorithm is used to reduce the swing in the gradient descent, allowing learning at a larger learning rate and accelerating the descent process. The algorithm formula is as follows:
[0161]
[0162] γ←ργ+(1-ρ)·g·g
[0163]
[0164] θ←θ+φ
[0165] Loss function:
[0166] The MSE function is used as the loss function to minimize the MSE of the predicted value and the true value. The formula is as follows:
[0167]
[0168] in, represents the true state of the target at time t, x t represents the predicted state of the target at time t, and D represents the dimension of the target state.
[0169] Step 4: Build a TWP tracker to implement trajectory similarity calculation and data association, and assign a digital ID to each object.
[0170] Step 1: Data association in the first stage
[0171] The data association problem is considered as a maximum a posteriori problem (MAP), where D = {di} is a set of detection responses. Where di = (pi, si, ai, ti), where di is the position of the object, si is the size of the object, ai is the appearance feature of the object, and t is the index frame of the object in the video. Define j = {dj1, dj2, ..., djn} as a trajectory composed of detection responses based on time series. Therefore, the purpose of data association is to maximize the posterior probability of generating a trajectory yt = {tn} under given detection response conditions. The following model can be obtained:
[0172]
[0173] tT i ∩T j =Φ,Ψi≠j
[0174] Among them, the conditional probability of detection is established using Bernoulli distribution, indicating whether the detection response is a true detection result or a false negative.
[0175] For the probability of the trajectory, a Markov chain is used to model it, which consists of three parts. Specifically, it represents the probability of the trajectory starting, ending, and being associated with the detection. The final MAP optimization problem can be viewed as an integer linear programming problem. The formula is as follows:
[0176]
[0177] stf en,i ,f i,j ,f ex,i ,f i ∈{0,1}
[0178] Step 1: Data association in the second stage
[0179] First, as mentioned above, the detection response is treated as a node in the network flow for data association. The cost between two detections consists of two parts. On the one hand, the appearance model is built using the color histogram; it does not take into account and has a certain rotation invariance. On the other hand, the IoU between the detection and prediction generated by the LSTM motion model is used to measure the similarity. Finally, the two are combined to complete the data association task in the second stage. Among them, di is the position of the i-th object and pi is the prediction.
[0180] cost(d i ,d j )=αIoU(p i ,d j )+(1-α)A1(d i ,d j )
[0181] Step 5: Create a track token to manage the status of the tracking target that has been assigned a digital ID. The specific steps are explained in conjunction with the figure. The specific steps are as follows:
[0182] Step 1: Activate the tracking state and assign a tracker for the target detected in three consecutive frames;
[0183] Step 2: When the target's appearance information is greatly disturbed due to occlusion, illumination, motion, etc., the appearance feature matching cannot reach the threshold and the tracking state is converted to the waiting state;
[0184] Step 3: When the target is activated by the detector again, the features of the latest three frames stored in the feature library are matched with the features of the newly appeared target. If the association is successful, the target is switched to the tracking state again.
[0185] Step 4: If there is no match, it means that a new tracking target has appeared, and a series of initialization operations are performed on the new target. For the successfully tracked target, the features in the target feature library and the various parameters in the target motion system should be updated in real time to complete the update of the system status.
[0186] Step 6: Establish a target anomaly detector to identify and specially mark the tracking targets with abnormal status.
[0187] In the target management database of the system, the target state is divided into three categories according to the motion state: static, uniform motion, and variable speed falling motion;
[0188] Abnormal situations of express shipments transported by the sorting center are divided into the following situations:
[0189] Express delivery jams - express delivery jams on the conveyor belt due to corners and congestion;
[0190] Parcel dropped - parcel dropped due to collision or acceleration on the conveyor belt or during machine transportation;
[0191] Overlapping of parcels - on the conveyor belt, collisions and other issues may cause parcels to overlap and pile up, resulting in sorting errors and the disappearance of parcels.
[0192] The system combines the location information and status information of the express to judge and identify whether it is in an abnormal state. Assume that the storage attributes of the express in the database are [ID, Place, V, Status]:
[0193] (ID, Place, V, Status>
[0194] ID is a digital ID assigned by the system to each target express through the TWP tracker. The digital ID has three states:
[0195]
[0196] Place indicates the location of the express shipment. There are three types of Place:
[0197]
[0198] V indicates the speed information of the express. There are three levels of V:
[0199]
[0200] Status indicates the status of the shipment, one is normal and three are abnormal:
[0201]
[0202] Status=f(ID,Place,V)=ID×Place×V
[0203] The basic principles, main features and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above specific embodiments. The above specific embodiments and the description in the specification are only for further illustrating the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of the present invention to be protected is defined by the claims and their equivalents.
Claims
1. A method for identifying abnormalities in a sorting center based on video detection technology, characterized in that: The method comprises the following steps: Step S1. When the system starts, the first frame of the video is detected, and three frames are detected continuously; Step S2. Run the object detector to perform video object detection to obtain the bounding box of the object; Step S3. Build a target state prediction model based on LSTM to predict the target trajectory; Step S4. Construct a TWP tracker to implement trajectory similarity calculation and data association, and assign a digital ID to each object; Step S5. Create a track token to manage the status of the tracking target that has been assigned a digital ID; Step S6. Establish a target anomaly discriminator to identify the tracking target with abnormal status; In step S3, based on the LSTM target state prediction model, the target motion information is obtained by using the time series, and the input is: t (x,y,z),z t (r,a,e); output: x t+1 (x,y,z),z t+1 (r,a,e); Among them, x t is the state vector of the target at time t, x, y, z are the positions in the three-dimensional coordinate system with the center of the video as the origin; z t is the measurement state of the target at time t, r, a, e represent the distance, orientation and angle in the polar coordinate system with the camera as the origin respectively; Based on the LSTM neuron structure, x t ,z t As the input of the model, X represents the input target set {x0,x1,...,x t , x t+1 }, W represents the learnable parameters of the fully connected layer, and σ represents the sigmoid activation function tanh represents the tanh function Concat represents a vector concatenation function, which is used to merge two vectors into a longer vector. h t represents the output set of the hidden layer {h0,h1,......,h t ,h t+1 }, c t Represents the cell state set of LSTM {c0,c1,...c t ,c t+1 ,}; The specific process of the model from input to output is: Step S31a. Input data x t (x,y,z) and z t (r,a,e), merged into a vector through the concat function, Step S3b. The hidden layer h obtained from the previous node t According to the weight matrix w xh Calculate and get the middle hidden layer Output: Step S3c. Use the sigmoid activation function σ(x) and the tanh function to obtain the cell state c of this node. t+1 , Step S3d. Calculate h t+1 , and the output x t+1 , Among them, b i , b h , b o , is the offset.
2. The method for identifying abnormalities in a sorting center based on video detection technology according to claim 1 is characterized in that: In step S2, the specific process of running the object detector to perform video target detection includes the following steps: Step S21. Perform a convolution operation on the original image through the image preprocessing module in the Musk convolutional neural network model to obtain a common feature map as the input of the RPN network. The size of the common feature map is set to N×H×W pixels, where H is the height of the original image and W is the bottom of the original image. Step S22. Perform a 3×3 convolution on the common feature map to obtain a feature map of 256×H×W pixels, that is, there are H×W 256-dimensional feature vectors in the map; Step S23. Predetermine anchor points and standard boxes. Each point on the feature map corresponds to an area on the original image. The area size is m×m, where m is the ratio of the original image to the feature map. The center of each area is set as the anchor point. Each anchor point determines K standard boxes. Therefore, H×W points in the feature map correspond to H×W×K boxes in the original image. The aspect ratio of the standard box is 1, the size is 64xp; 128xp; 256x0p, K = 3 × 3 = 9, and three ratios are set for each standard frame; Step S24. The feature map contains H×W vectors, each of which is 256-dimensional. Two full-connection operations are performed on each feature vector to obtain two sub-feature maps, namely the target background feature map and the original image offset feature map. The size of the target background feature map is 4K×H×W pixels, K corresponds to the K anchored standard boxes, 4 represents 4 different states, H is the height of the feature map, W is the bottom of the feature map, Set 4 components to represent different states of the target and background: a,b,c,d∈{1,0}, a, c are target quantities, b, d are displacements, For backgrounds that cannot be moved, The background moving within the corresponding time period, For temporarily stationary targets, Corresponding to the stationary target in the video image, The non-movable background is a feature area that has no changes or movements for 20 consecutive frames; The size of the original image offset feature map is 4K×H×W pixels, K is the K anchored standard boxes, 4 represents the 4 coordinates of the standard box (x, y, h, w), (x, y) represents the coordinates of the anchor point, that is, the center point of the standard box, (h, w) is the width and height of the standard box, each anchor point corresponds to K standard boxes, marking K×H×W standard boxes; Set an offset for the canonical box and a threshold α. If the offset is less than α, it is marked as a positive sample in training. If the offset is greater than α, it is marked as a negative sample. Assume that the center points of the canonical frame A and its corresponding real frame B are (x i ,y i ) and (x j ,y j ); the width of A is w i , the width of B is w j ; The height of A is H i ,B's is H j , The offset is: among which parameters x ,m y ,m w ,m h ∈(0,0.1),σ x =s y =0.1,σ h =s w =0.2; Step S25. Obtain candidate boxes through the Musk convolutional neural network model based on the anchor points and feature maps. The loss function of the model training is Among them, i is the i-th anchor point participating in the training in the feature map, there are N in total, p i is the predicted probability, which represents the probability that the original image defined by the anchor point is the target. is the cross entropy of the binary classification; is the regression of the canonical box.
3. The method for identifying abnormalities in a sorting center based on video detection technology according to claim 1 is characterized in that: In step S2, the setting of the proportion size of the K standard boxes of the anchor points in the RPN network is specifically as follows: the length, width and height of the carton are set to (a, b, c), the width of the peripheral matrix of the carton in the image video is set to W, and the height is set to H, The linear mapping formula from matrix to image matrix is Among them, the parameter β is the distance between the carton and the camera, β∈(0,1), the parameter γ is the wide angle between the carton and the camera, γ∈(0.3,1), and the parameter μ is the depression angle between the carton and the camera, μ∈(0.2,1).
4. The method for identifying abnormalities in a sorting center based on video detection technology according to claim 1, characterized in that: In step S3, the target state prediction model adopts the loss function ME-Loss; Object detection and instance segmentation are multiple tasks. The loss function of multiple tasks adopts the weighted summation of the loss functions of multiple subtasks. The formula is in, is the loss function of the t-th task in the ith stage, is a hyperparameter used to determine the weight of each task in each stage; Using an integrated loss function, the loss function is defined as Given a set of logits and the output of the final fully connected layer, the loss function is calculated as follows: Step S31. Calculate the difference x between the sum ij =s j -s i , Step S32. Based on the difference values of and, the difference term obtained for each pair of samples can be expressed as a basic term Step S33. Normalize and sum L to obtain the final loss function:
5. The method for identifying abnormalities in a sorting center based on video detection technology according to claim 1, characterized in that: In step S4, the TWP tracking model is divided into two stages. In the first stage, the target detection is taken as a node in the graph to build a network flow model; then the model is solved to obtain a preliminary tracking trajectory; wherein the data association cost is calculated by the LSTM motion model and the appearance model based on the color histogram; in the second stage, in order to reduce missed targets and target IDs, the trajectories are clustered and optimized, and in the clustering, the similarity of the targets is calculated by the siamese network and the LSTM motion model.
6. The method for identifying abnormalities in a sorting center based on video detection technology according to claim 1, characterized in that: In the system in step S1, an abnormality index is added to form a new evaluation index system. The evaluation index includes three factors: tracking accuracy, tracking trajectory consistency and abnormal situation identification rate, which are as follows: Use ID scores to judge tracking accuracy. Identification precision: Identification recall: Identification FI: Use SQE to judge the consistency of tracking trajectory. SQE estimates the transformation of ID{three digits} by analyzing the feature distance between targets with the same identity and targets with different identities; The Gaussian mixture model is used to measure the distance distribution, and the distance measurement mode uses the Euclidean distance: The Gaussian mixture model and Euclidean distance formula of the standard features are used to obtain the distance distribution model, and the feature distance is standardized to obey the chi-square distribution The evaluation index formula is: Among them, n is the number of trajectories, L is the average length of the trajectory, dif is the degree of identity change within the trajectory, and the mean distance of the Gaussian mixture model is calculated to determine whether there are multiple identities; Use Get strange to judge the identification rate of abnormal situations. Among them, FN is False Negative, which is the sum of the number of false negatives detected and tracked in the entire video; FP is False Positve, the sum of the number of false positives in the entire video detection and tracking; TP is TruePositive correctly marked, the sum of correctly marked targets detected and tracked in the entire video; M is the environmental parameter, which is determined according to the number and movement speed of the targets in the video; Gs is the anomaly capture rate, which indicates the detection rate of true anomaly targets detected in the video targets; TPG is the sum of correctly detected and tracked targets in the entire video taking into account Gs; FNG is the sum of the number of missed detections and target tracking in the entire video when Gs is taken into account; FPG is the sum of the number of false alarms in the entire video detection and tracking when Gs is taken into account; α is the weight of TPG, and the value of α is adjusted according to the recognition accuracy of the anomaly detector.
7. The method for identifying abnormalities in a sorting center based on video detection technology according to claim 1, characterized in that: In step S5, a track Token is constructed to manage the status of the tracking target to which a digital ID has been assigned. The specific steps are: Step S51. Activate the tracking state and assign a tracker to the target detected in three consecutive frames; Step S52: When the apparent information of the target is greatly disturbed, the apparent feature matching cannot reach the threshold and the tracking state is converted to a waiting state; Step S53: When the target is activated by the detector again, the features of the latest three frames stored in the feature library are matched with the features of the newly appeared target. If the association is successful, the target is switched to the tracking state again. Step 4: If there is no match, it means that a new tracking target has appeared, and a series of initialization operations are performed on the new target; for the successfully tracked target, the features in the target feature library and the various parameters in the target motion system must be updated in real time to complete the update of the system status.
8. The method for identifying abnormalities in a sorting center based on video detection technology according to claim 1, characterized in that: In step S5, a two-stage tracker is established using a motion model based on an LSTM network to track and predict the trajectory.
9. The method for identifying abnormalities in a sorting center based on video detection technology according to claim 1, characterized in that: In step S6, in the target management database of the system, the target state is divided into three categories according to the motion state: static, uniform motion, and variable speed falling motion. The abnormal conditions of express shipments transported by the sorting center are divided into express shipment jamming, express shipment dropping, and express shipment overlapping. By combining the position information and state information of the express shipment, it is judged and identified whether it is in an abnormal state. Assume that the storage attributes of the express in the database are [ID, Place, V, Status],<ID,Place.V,Status> ; ID is a digital ID assigned by the system to each target express through the TWP tracker. The digital ID has three states: Place is the location information of the express shipment. There are three types of Place: V is the speed information of the express, and there are three levels of V: Status is the status of the shipment, one is normal and three are abnormal: Status=f(ID,Place,V)=ID×Place×V.
Citation Information
Patent Citations
Multi-target tracking method based on depth track prediction
CN110135314A
Multi-target tracking method combined with video scene feature perception
CN110660083A