Parcel waybill number identification method based on target tracking

By adopting a target tracking method in the sorting center, combining the CRN collaborative representation learning network model and multi-task learning framework, the problems of low recall and identification errors of package waybill number recognition in the prior art are solved, and the waybill number recognition effect with high accuracy and high recall is achieved.

CN120088501AActive Publication Date: 2025-06-03HARBIN INST OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510159761.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-03
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

In the prior art, when identifying parcel waybill numbers in sorting centers, the recall rate is low and it is prone to identify errors due to packages with the same appearance, and the time information of the parcels are not effectively utilized.

Method used

A target tracking-based method is adopted, and a CRN collaborative representation learning network model is designed, combined with a multi-task learning framework, the feature representation of object detection and ReID tasks is optimized, and the inter-frame data correlation is used using Kalman filtering and Hungarian algorithms, and the appearance characteristics of wrapped ReID are introduced to improve the accuracy and robustness of the association.

Benefits of technology

It has achieved high accuracy and high recall rate of waybill number recognition, with an accuracy rate of 99.9%, and a recall rate of 99%. It has significantly alleviated the ID jump problem in multi-target tracking, and improved the real-time and accuracy of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088501A_ABST
    Figure CN120088501A_ABST
Patent Text Reader

Abstract

The invention provides a parcel waybill number identification method based on target tracking. The method comprises the steps of designing and training a proposed CRN cooperative representation learning network model on the basis of traditional target detection and ReID models to improve the ability of the model to extract target space and appearance features; then, according to the IOU of the detection frame, the similarity of the appearance features, and a Kalman filtering and Hungary matching algorithm for depth classification, target tracks between the frames are associated; and finally, associating the waybill numbers subjected to six-side scanning near the parcel falling time with the track to obtain the waybill number corresponding to the parcel in each frame. By introducing the method, parcel waybill number matching errors caused by overlapping or motion blur can be relieved, the sorting efficiency is improved, the operation cost is reduced, and the intelligent level of the whole logistics system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and particularly to a method for identifying a package waybill number based on object tracking. Background Art

[0002] In the logistics field, the problem of identifying the package waybill number in the narrow-band grid of the sorting center, as a special image retrieval problem, has been widely studied. The goal of package waybill number identification is: when a user gives a query for a certain package on the grid, the system returns what the waybill number of this package is. The candidate waybill numbers come from the set of waybill numbers of all packages that have undergone six-sided scanning in a period of time before the package falls into the grid.

[0003] Currently, there are two common methods for identifying the package waybill number on the sorting grid. One is to install a barcode scanner above the grid to identify the package waybill number by scanning the label on the surface of the package. Since the posture of the package when it falls into the grid is not fixed, the surface where the label is located may not be within the field of view of the barcode scanner and cannot be read by the barcode scanner. Although the accuracy of this method is close to 100%, the recall rate is relatively low.

[0004] The other is to compare the image of the package on the grid with the images of packages that have undergone six-sided scanning through the ReID method to find the corresponding waybill number. Six-sided scanning is also called a six-sided code reading system. There are five barcode scanners installed above and around it, and there is also a bottom line scanning and reading module at the bottom. It can stably read the package waybill number regardless of the orientation of the package label. Six-sided scanning is very expensive. Generally, multiple narrow-bands share one six-sided scanner, so the ReID method still needs to be used to retrieve the waybill number corresponding to the package on the grid. The drawback of this method is that once packages with the same appearance appear continuously, it is very easy to identify wrongly. Therefore, the logistics industry needs a better method for identifying package waybill numbers.

[0005] Both of the two methods mentioned above only focus on visual information and do not use the time information when the package falls into the grid. The present invention will use the object tracking method combined with the time when the package falls into the grid to replace the traditional ReID method to implement a waybill number identification algorithm with high accuracy and high recall rate. Summary of the Invention

[0006] The purpose of the present invention is to solve the problems in the prior art, and a method for identifying a package waybill number based on object tracking is proposed.

[0007] The present invention is realized through the following technical solutions. The present invention proposes a method for identifying a package waybill number based on object tracking, and the method includes the following steps:

[0008] Step 1, design and train a CRN collaborative representation learning network model. The CRN collaborative representation learning network model jointly optimizes the feature representations of object detection and ReID tasks by introducing a multi-task learning framework;

[0009] In Step 1, a multi-task learning dataset is constructed: The object detection dataset contains 15,000 labeled images, with a total of 50,000 different packages, serving as Dataset A; The ReID task dataset has 100,000 images, including images of 30,000 different packages, with image samples of each package under multiple perspectives and scenarios, serving as Dataset B;

[0010] The CRN collaborative representation learning network model is specifically as follows: A reciprocal network REN is adopted to perform collaborative learning between different tasks; The reciprocal network separates the feature maps of object detection and ID embedding extraction into two different task-driven branches, and in this way, task-dependent representations can be learned; Specifically, given the shared features, a new structure combining self-relation and cross-relation is designed to enhance the feature representation. The self-relation prompts the hidden nodes to learn task-dependent features, while the cross-relation aims to improve the collaborative learning of the two tasks; At the same time, a scale-aware attention network SAAN is adopted to improve the alignment of ID embedding extraction and enhance the model's adaptability to scale changes; The output of SAAN is a feature tensor containing rich semantics of all objects in all resolutions;

[0011] In Step 2, first, the Kalman filter algorithm is used to predict the state of the object in the next frame, then the Hungarian algorithm is used to match the detection results of the current frame with the predicted state, and the appearance features of package ReID are further introduced to improve the accuracy and robustness of the association;

[0012] In Step 3, an algorithm that matches the trajectory with the waybill number is used to obtain the final matching result.

[0013] Furthermore, in Step 1, during the training process of the CRN collaborative representation learning network model, it is necessary to optimize both the object detection and ReID tasks simultaneously; The training objective of the CRN collaborative representation learning network model is to minimize the joint loss of the object detection and ReID tasks, and thus a multi-task loss function is defined The loss function consists of the object detection loss and the ReID loss and is balanced by a trade-off parameter λ to balance the contributions of the two tasks:

[0014]

[0015] The object detection task adopts an anchor-based detection framework, uses cross-entropy loss to classify the object categories, and uses smooth L1 loss to regress the position of the bounding box; Therefore, the object detection loss is expressed as:

[0016]

[0017] Among them, y i is the true label of the anchor, p i is the class probability predicted by the model, t i and are the predicted bounding box parameters and the true bounding box parameters respectively;

[0018] The goal of the ReID task is to learn a discriminative feature representation so that different images of the same package are as close as possible in the feature space, while images of different packages are as far away as possible. Therefore, the triplet loss is used to optimize the ReID task; given an anchor sample e a , a positive sample e p , which is in the same package as the anchor, and a negative sample e n , which is in a different package from the anchor, the triplet loss is expressed as:

[0019]

[0020] Among them, α is the margin parameter, used to control the distance between the positive and negative samples, and [·]+ means taking the positive value.

[0021] Furthermore, in step two, the CRN collaborative representation learning network model is inferred. During the inference process, the CRN collaborative representation learning network model receives an input image I, extracts multi-scale features F = Backbone(I) through the shared backbone network. Subsequently, the model is divided into two branches: the object detection branch outputs the bounding boxes {b 1 , b 2 , …, b n} of the object through B = Detector(F), where each b i contains location and class information; the ReID branch extracts the appearance feature vectors E = {e 1 , e 2 , …, e n} of each object through E = ReID(F), where is a d-dimensional feature vector; finally, the CRN collaborative representation learning network model outputs the object detection results and the corresponding ReID features simultaneously, realizing the joint inference of detection and re-identification.

[0022] Furthermore, in step two, the Kalman filter algorithm is used to predict the state of the package in the next frame; assuming the state vector of the object is where (x, y) represents the location of the object, represents the speed of the object; the Kalman filter predicts the state of the object in the next frame through the motion model: x k|k-1 = F k x k-1|k-1 ; where Fk is the state transition matrix, and x k|k-1 is the predicted state, and x k-1|k-1 is the state estimate of the previous frame. Then, the prediction of the covariance matrix is:

[0023]

[0024] where P k|k-1 is the predicted covariance matrix, and Q k is the process noise covariance matrix; when a new detection box observation value z k arrives, the Kalman filter updates the state estimate through the following formula:

[0025]

[0026] Update the state estimate and covariance matrix:

[0027] x k|k = x k|k-1 + K k (z k - H k x k|k-1 )

[0028] P k|k = (I - K k H k )P k|k-1 .

[0029] Furthermore, in step two, the Hungarian algorithm is used to match the detection results of the current frame with the predicted target states; calculate the matching cost between each detection box z k and each predicted state x k|k-1 ; each element c i,j of the cost matrix C represents the matching cost between the i-th detection box and the j-th target trajectory. Based on the position information and appearance information, the Mahalanobis distance is used to measure the position difference between the detection box and the predicted state:

[0030]

[0031] where is the covariance matrix of the observation residual;

[0032] The cosine similarity is used to measure the similarity between the appearance feature vectors e i and e j of the detection box and the target trajectory:

[0033]

[0034] Finally, the matching cost c i,j is the weighted sum of the position information and appearance information:

[0035]

[0036] Use the Hungarian algorithm to solve the cost matrix C and find the optimal matching scheme.

[0037] Further, in step three, construct a time window. Assume the time when the package drops into the grid is t g , and the time when the system receives the waybill number is t s , then the time window W is defined as:

[0038] W = [t g - Δt, t g + Δt]

[0039] where Δt is an adjustable time tolerance parameter used to control the size of the time window; through the time window, the system can search for the package trajectory that matches the waybill number within W.

[0040] Further, in step three, establish a probability matching model between the trajectory and the waybill number. Assume there are N candidate package trajectories T = {T 1 , T 2 , …, T N} within the time window W. Each trajectory T i contains a series of state information {x i , y i , t i , e i}, where x i , y i represent the position of the package, t i represents the timestamp, and e i represents the appearance feature vector of the package;

[0041] For each candidate trajectory T i , calculate its matching probability P(T i | S) with the waybill number, where S represents the relevant information of the waybill number. The matching probability P(T i | S) can be decomposed into the following parts:

[0042] P(T i | S) = P(t i | t s ) · P(e i | e s )

[0043]

[0044] where P(t i | t s ) represents the trajectory Ti Timestamp t i Time of receiving the waybill number t s The matching degree of P(e i |e s ) represents the trajectory T i Appearance features i The appearance characteristics of the package corresponding to the waybill number when scanned on six sides s The cosine similarity is used to measure the similarity of appearance features.

[0045] Furthermore, in step 3, after calculating the matching probabilities of all candidate trajectories, the maximum a posteriori probability criterion is used to select the optimal matching result, that is, the trajectory T with the largest matching probability is selected. * As the final matching result:

[0046]

[0047] The present invention also proposes an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method for identifying a parcel waybill number based on target tracking are implemented.

[0048] The present invention also proposes a computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the steps of the method for identifying a parcel waybill number based on target tracking.

[0049] Compared with the prior art, the present invention has the following beneficial effects:

[0050] The present invention effectively solves the competition problem between target detection and ReID tasks by designing a CRN collaborative representation learning network model and combining it with a multi-task learning framework, thereby improving the robustness and generalization ability of the model. At the same time, the Kalman filter and the Hungarian algorithm are used for inter-frame data association, and the appearance features of the package ReID are introduced, which significantly alleviates the ID jump problem in multi-target tracking and enhances the tracking stability. In addition, the time window mechanism and the probability matching model are used to solve the matching problem caused by the delay in the sending of package information, thereby improving the real-time performance and accuracy of the system. Finally, the present invention realizes efficient and stable parcel waybill number recognition, with an accuracy rate of 99.9% and a recall rate of 99%. It has strong adaptive capabilities and wide application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0052] Figure 1 Flowchart of the package waybill number recognition method based on target tracking according to the present invention;

[0053] Figure 2 Schematic diagram of the overall network structure of CRN;

[0054] Figure 3 Schematic diagram of the result example of target detection;

[0055] Figure 4 Schematic diagram of the result example of ReID clustering;

[0056] Figure 5 Schematic diagram of the result example of package waybill number recognition;

[0057] Figure 6 Schematic diagram of the recall rate of applying the package waybill number recognition technology in a certain sorting center;

[0058] Figure 7 Schematic diagram of the accuracy rate of applying the package waybill number recognition technology in a certain sorting center. Specific implementation manners

[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0060] The goals that the present invention aims to achieve include:

[0061] Real-time problem: In order to make the time for processing each frame of the video as short as possible, it is hoped to use a small model to simultaneously complete the target detection and ReID tasks, and then use a rule-based method to associate the data between frames to obtain the package trajectory.

[0062] Over - competition problems between object detection and ReID in multi - task learning: 1) Competition in object representation learning. In traditional multi - task learning, object class confidence, object scale, and ID information are obtained by processing a shared feature tensor. Although this method is efficient, it ignores the inherent differences between different tasks. 2) Competition in semantic level assignment. Modern object detection models always introduce a feature pyramid structure to improve the localization accuracy of objects with different sizes. In this framework, objects of different scales will be assigned to features with different resolutions. However, this assignment is not suitable for ReID because it will lead to misalignment of the semantic levels.

[0063] ID jump problem in multi - object tracking: Due to factors such as mutual occlusion and overlap between packages, the ID of the same package often changes during the tracking process, and the ID cannot be directly and simply matched with the waybill number. Therefore, an inter - frame data association algorithm that can keep the ID stable needs to be designed.

[0064] Problem of delayed package information sending: Sometimes the package has already fallen into the grid, but after a few seconds, the candidate waybill number updates the new information of falling into the grid. Therefore, a module needs to be designed to handle the problems caused by the delayed sending of package information.

[0065] Therefore, the present invention proposes a method for identifying the waybill number of packages based on object tracking, which is specifically aimed at the task of identifying the waybill number of packages in the grids of a logistics sorting center. The method includes designing and training the proposed CRN collaborative representation learning network model on the basis of traditional object detection and ReID models to improve the model's ability to extract object space and appearance features; then, based on the detection box IOU, appearance feature similarity, and depth classification, using the Kalman filter and Hungarian matching algorithm to associate the object trajectories between frames; finally, associating the waybill numbers scanned on six sides near the package falling time with the trajectories to obtain the waybill number corresponding to the package in each frame. By introducing this method, it is possible to alleviate the mis - matching of package waybill numbers caused by overlap or motion blur, improve the sorting efficiency, reduce the operating cost, and enhance the intelligent level of the overall logistics system.

[0066] Specifically, in combination with Figures 1 - 7 , the present invention proposes a method for identifying the waybill number of packages based on object tracking, and the method includes the following steps:

[0067] Step 1, design and train the CRN collaborative representation learning network model, aiming to improve the model's ability to extract object space and appearance features; the CRN collaborative representation learning network model enhances the robustness and generalization ability of the model in complex scenarios by introducing a multi - task learning framework and jointly optimizing the feature representations of object detection and ReID tasks;

[0068] In step one, a multi-task learning dataset is constructed: In order to minimize the time taken to process each frame of the video, the present invention trains a small model to simultaneously complete object detection and ReID tasks. The object detection dataset of the present invention contains 15,000 labeled images, totaling 50,000 different packages, which is used as dataset A; the ReID task dataset has 100,000 images, containing images of 30,000 different packages, with image samples of each package in multiple perspectives and scenarios, which is used as dataset B.

[0069] The CRN collaborative representation learning network model is specifically as follows: To address the problem that the performance degradation of multi-task learning mainly stems from the excessive competition between object detection and ReID tasks, the present invention fully considers the essential differences between the two tasks and designs a CRN model, which contains two sub-networks to solve the above problems. The present invention uses a reciprocal network REN for collaborative learning between different tasks; the reciprocal network separates the feature maps of object detection and ID embedding extraction into two different task-driven branches, and in this way, it can learn task-dependent representations; specifically, given the shared features, a new structure that combines self-relation and cross-relation is designed to enhance the feature representation. The self-relation prompts the hidden nodes to learn task-dependent features, while the cross-relation aims to improve the collaborative learning of the two tasks; at the same time, a scale-aware attention network SAAN is used to improve the alignment of ID embedding extraction and enhance the model's adaptability to scale changes. This architecture enhances the influence of object-related regions and channels; the output of SAAN is a feature tensor containing rich semantics of all objects at all resolutions; according to the above design concept, the ID embedding based on aggregated features helps prevent semantic misalignment during the matching process.

[0070] In step two, first, the Kalman filter algorithm is used to predict the state of the object in the next frame, and then the Hungarian algorithm is used to match the detection results of the current frame with the predicted state, and the appearance features of package ReID are further introduced to improve the accuracy and robustness of the association.

[0071] In step three, an algorithm that matches the trajectory with the waybill number is used to obtain the final matching result, which can greatly alleviate the impact caused by the difference between the time when the package actually lands and the time when the tracking system receives a new candidate waybill number.

[0072] In step one, during the training process of the CRN collaborative representation learning network model, it is necessary to optimize both object detection and ReID tasks simultaneously; to ensure that the two tasks can work together rather than compete with each other, the present invention designs a multi-task loss function and combines the characteristics of the reciprocal network (REN) and the scale-aware attention network (SAAN). The training objective of the CRN collaborative representation learning network model is to minimize the joint loss of object detection and ReID tasks, and thus a multi-task loss function is defined. The loss function consists of an object detection loss and a ReID loss and is balanced by a trade-off parameter λ for the contributions of the two tasks:

[0073]

[0074] The object detection task adopts an anchor-based detection framework, uses cross-entropy loss to classify object categories, and uses smooth L1 loss to regress the positions of bounding boxes; thus, the object detection loss is expressed as:

[0075]

[0076] where y i is the ground truth label of the anchor, p i is the class probability predicted by the model, t i and are the predicted bounding box parameters and the ground truth bounding box parameters respectively;

[0077] The goal of the ReID task is to learn a discriminative feature representation such that different images of the same package are as close as possible in the feature space, while images of different packages are as far away as possible. Therefore, triplet loss is adopted to optimize the ReID task; given an anchor sample e a , a positive sample e p that is of the same package as the anchor, and a negative sample e n that is of a different package from the anchor, the triplet loss is expressed as:

[0078]

[0079] where α is the margin parameter used to control the distance between positive and negative samples, and [·]+ denotes taking the positive value.

[0080] In step two, CRN collaborative representation learning network model inference is performed. During the inference process, the CRN collaborative representation learning network model receives an input image I, extracts multi-scale features F = Backbone(I) through a shared backbone network. Subsequently, the model is divided into two branches: the object detection branch outputs the bounding boxes {b 1 ,b 2 ,…,b n} of the objects through B = Detector(F), where each b i contains position and class information; the ReID branch extracts the appearance feature vectors E = {e 1 ,e 2 ,…,en}, where is a d-dimensional feature vector; finally, the CRN collaborative representation learning network model outputs the object detection result and the corresponding ReID feature at the same time, realizing the joint inference of detection and re-identification.

[0081] In step two, the Kalman filter algorithm is used to predict the state of the package in the next frame; assume that the state vector of the target is where (x, y) represents the position of the target, represents the speed of the target; the Kalman filter predicts the state of the target in the next frame through the motion model: x k|k-1 = F k x k-1|k-1 ; where F k is the state transition matrix, x k|k-1 is the predicted state, x k-1|k-1 is the state estimate of the previous frame, then the prediction of the covariance matrix is:

[0082]

[0083] where P k|k-1 is the predicted covariance matrix, Q k is the process noise covariance matrix; when the new detection box observation value z k arrives, the Kalman filter updates the state estimate through the following formula:

[0084]

[0085] Update the state estimate and covariance matrix:

[0086] x k|k = x k|k-1 + K k (z k - H k x k|k-1 )

[0087] P k|k = (I - K k H k )P k|k-1 .

[0088] In step two, the Hungarian algorithm is used to match the detection result of the current frame with the predicted target state; calculate the matching cost between each detection box z k and each predicted state x k|k-1 ; each element c i,j of the cost matrix C represents the matching cost between the i-th detection box and the j-th target trajectory. Based on the position information and appearance information, the Mahalanobis distance is used to measure the position difference between the detection box and the predicted state:

[0089]

[0090] where is the covariance matrix of the observation residuals;

[0091] The cosine similarity is used to measure the similarity between the detection box and the appearance feature vector e i and e j :

[0092]

[0093] Finally, the matching cost c i,j is the weighted sum of the position information and the appearance information:

[0094]

[0095] The Hungarian algorithm is used to solve the cost matrix C to find the optimal matching scheme. The matching result will determine which detection boxes are associated with the existing target trajectories and which detection boxes may represent newly emerging targets.

[0096] In step three, a time window is constructed. To handle the inconsistency between the package dropping time and the time when the system receives the waybill number, the present invention introduces a time window mechanism. Assume that the package dropping time is t g , and the time when the system receives the waybill number is t s , then the time window W is defined as:

[0097] W = [t g - Δt, t g + Δt]

[0098] where Δt is an adjustable time tolerance parameter used to control the size of the time window; through the time window, the system can search for the package trajectory that matches the waybill number within W.

[0099] In step three, a probabilistic matching model between the trajectory and the waybill number is established. Assume that there are N candidate package trajectories T = {T 1 , T 2 , …, T N} within the time window W. Each trajectory T i contains a series of state information {x i , y i , t i , e i}, where x i , y i represent the position of the package, t i represents the timestamp, and e i represents the appearance feature vector of the package;

[0100] For each candidate trajectory T i , calculate its matching probability P(T i |S), where S represents the relevant information of the waybill number, such as the appearance characteristics of the package. The matching probability P(T i |S) can be decomposed into the following parts:

[0101] P(T i |S) = P(t i |t s )·P(e i |e s )

[0102]

[0103] Among them, P(t i |t s ) represents the matching degree between the timestamp t i of the trajectory T i and the receiving time t s of the waybill number; it is assumed that the time difference follows a Gaussian distribution; P(e i |e s ) represents the matching degree between the appearance feature e i of the trajectory T i and the appearance feature e s of the package corresponding to the waybill number in the six-sided scan, and the cosine similarity is used to measure the similarity of the appearance features.

[0104] In step three, after calculating the matching probabilities of all candidate trajectories, the maximum a posteriori probability criterion is used to select the optimal matching result, that is, select the trajectory T * with the largest matching probability as the final matching result:

[0105]

[0106] To further improve the adaptive ability of the system, the present invention also designs a feedback and update mechanism for the matching result. After each matching result is generated, the system updates the matching model according to the actual grid dropping situation. Specifically, the system adjusts the time window W and the parameters in the matching probability model according to the actual grid dropping time and the matching result T * so that the model can better adapt to the changes in the actual scenario.

[0107] The following combines specific sorting center grid video data to illustrate the specific implementation manner of the present invention.

[0108] Execute step one: The designed CRN is as Figure 2As shown in the figure. The CRN is trained using the AdamW optimizer, and then object detection and ReID are performed on the divided test set. The obtained results are compared with the ground truth. Figure 3 is an example of the object detection result. Figure 4 is an example of the ReID clustering result. Compared with other neural networks that simultaneously implement object detection and ReID, the average values of the evaluation metrics of the CRN are shown in Table 1.

[0109] Table 1

[0110]

[0111] Execute Step 2: In the inference stage, the CRN model simultaneously outputs the object detection result and the corresponding ReID feature through joint inference, realizing the seamless combination of detection and re-identification. Specifically, the model first extracts multi-scale features through a shared backbone network, and then outputs the bounding box information and appearance feature vector of the object through the object detection branch and the ReID branch respectively. Then, the Kalman filter algorithm is used to predict the state of the object in the next frame, including position and speed information, and the Hungarian algorithm is used to match the detection result of the current frame with the predicted state. In the matching process, a cost matrix is constructed and the optimal matching scheme is solved by combining the position information and appearance feature of the object, so as to determine the association relationship between the detection box and the existing object trajectory or the appearance of a new object. This process effectively alleviates the ID jump problem in multi-object tracking and improves the stability and accuracy of tracking.

[0112] Execute Step 3: Through the above steps, the system can accurately obtain the waybill number corresponding to each package in each frame of the image, as Figure 5 shown. In order to evaluate the performance of the entire waybill number recognition technology, the system conducts statistical analysis on the accuracy rate and recall rate. The experimental results show that the accuracy rate of this technology is as high as 99.9%, and the recall rate also reaches 99%, as Figure 6 and Figure 7 shown. This result fully proves the high efficiency and reliability of the present invention in practical applications, and can effectively handle complex scenarios such as package information transmission delay and target occlusion, providing strong technical support for package sorting and logistics management.

[0113] The present invention also proposes an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method for recognizing the waybill number of a package based on object tracking are realized.

[0114] The present invention also proposes a computer-readable storage medium for storing computer instructions, and when the computer instructions are executed by a processor, the steps of the method for recognizing the waybill number of a package based on object tracking are realized.

[0115] The memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DRRAM). It should be noted that the memory of the method described in the present invention is intended to include but not limited to these and any other suitable types of memory.

[0116] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a high-definition digital video disc (DVD)), or a semiconductor medium (such as a solid state disc (SSD)), etc.

[0117] In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware processor or completed by the combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0118] It should be noted that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0119] The above has introduced in detail a method for identifying a package waybill number based on target tracking proposed by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for identifying a parcel waybill number based on target tracking, characterized in that: The method comprises the following steps: Step 1: Design and train the CRN collaborative representation learning network model. The CRN collaborative representation learning network model introduces a multi-task learning framework to jointly optimize the feature representation of the target detection and ReID tasks. In step 1, a multi-task learning dataset is constructed: the object detection dataset contains 15,000 labeled images, including 50,000 different packages in total, as dataset A; the ReID task dataset has 100,000 images, including 30,000 images of different packages, each package has image samples from multiple perspectives and scenes, as dataset B; The CRN collaborative representation learning network model is specifically as follows: a reciprocal network REN is used to perform collaborative learning between different tasks; the reciprocal network separates the feature maps of object detection and ID embedding extraction into two different task-driven branches, in this way, the representation of dependent tasks can be learned; specifically, given shared features, a new structure combining self-relationships and cross-relationships is designed to enhance feature representation, the self-relationships prompt hidden nodes to learn the features of dependent tasks, and the cross-relationships are intended to improve the collaborative learning of the two tasks; at the same time, a scale-aware attention network SAAN is used to improve the alignment of ID embedding extraction and enhance the model's adaptability to scale changes; the output of SAAN is a feature tensor containing rich semantics of all targets at all resolutions; Step 2: The Kalman filter algorithm is first used to predict the state of the target in the next frame, and then the Hungarian algorithm is used to match the detection result of the current frame with the predicted state. The appearance features of the package ReID are further introduced to improve the accuracy and robustness of the association. Step 3: Use the algorithm for matching the track with the waybill number to obtain the final matching result.

2. The method according to claim 1, characterized in that In step 1, during the training process of the CRN collaborative representation learning network model, it is necessary to optimize both the target detection and ReID tasks simultaneously; the training goal of the CRN collaborative representation learning network model is to minimize the joint loss of the target detection and ReID tasks, thereby defining a multi-task loss function The loss function is composed of the target detection loss and ReID loss The contribution of the two tasks is balanced by a trade-off parameter λ: The target detection task adopts an anchor-based detection framework and uses cross entropy loss To classify the target category, and use smooth L1 loss to regress the position of the bounding box; therefore, the target detection loss is expressed as: Among them, y i is the real label of the anchor, p i is the class probability predicted by the model, t i and are the predicted bounding box parameters and the true bounding box parameters respectively; The goal of the ReID task is to learn a discriminative feature representation so that different images of the same package are as close as possible in the feature space, while images of different packages are as far away as possible. Therefore, the triplet loss is used to optimize the ReID task. Given an anchor sample e a , a positive sample e p , which is in the same package as the anchor point, and a negative sample e n , which is wrapped differently from the anchor point, the triplet loss is expressed as: Among them, α is the interval parameter used to control the distance between positive and negative samples, and [·]+ indicates a positive value.

3. The method according to claim 1, characterized in that: In step 2, the CRN collaborative representation learning network model is inferred. During the inference process, the CRN collaborative representation learning network model receives an input image I and extracts multi-scale features F = Backbone (I) through a shared backbone network. Subsequently, the model is divided into two branches: the target detection branch outputs the target bounding box {b1, b2, …, b n }, where each b i Contains location and category information; the ReID branch extracts the appearance feature vector E = {e1, e2, …, e n },in is a d-dimensional feature vector; finally, the CRN collaborative representation learning network model simultaneously outputs the target detection results and the corresponding ReID features, realizing the joint reasoning of detection and re-identification.

4. The method according to claim 3, characterized in that In step 2, the Kalman filter algorithm is used to predict the state of the next frame of the package; assuming that the state vector of the target is Where (x, y) represents the location of the target. Indicates the speed of the target; Kalman filter predicts the state of the target in the next frame through the motion model: x k|k-1 =F k x k-1|k-1 ; where F k is the state transition matrix, x k|k-1 is the predicted state, x k-1|k-1 is the state estimate of the previous frame, then the prediction of the covariance matrix is: Where P k|k-1 is the prediction covariance matrix, Q k is the process noise covariance matrix; when the new detection box observation z k When it arrives, the Kalman filter updates the state estimate by the following formula: Update the state estimate and covariance matrix: x k|k =x k|k-1 +K k (z k -H k x k|k-1 ) P k|k =(I-K k H k )P k|k-1 。 5. The method according to claim 4, characterized in that In step 2, the Hungarian algorithm is used to match the detection result of the current frame with the predicted target state; each detection box z is calculated k With each predicted state x k|k-1 The matching cost between them; each element c of the cost matrix C i,j Represents the matching cost of the i-th detection box and the j-th target track. Based on the position information and appearance information, the Mahalanobis distance is used to measure the position difference between the detection box and the predicted state: in is the covariance matrix of the observed residuals; Use cosine similarity to measure the appearance feature vector e of the detection box and the target trajectory i and e j Similarities between: Finally, the matching cost c i,j is the weighted sum of location information and appearance information: Use the Hungarian algorithm to solve the cost matrix C and find the optimal matching solution.

6. The method according to claim 1, characterized in that In step 3, the time window is constructed. Assume that the time for the package to be dropped is t g , the time when the system receives the waybill number is t s , then the time window W is defined as: W=[t g -Δt,t g +Δt] Among them, Δt is an adjustable time tolerance parameter used to control the size of the time window; through the time window, the system can search for the package trajectory matching the waybill number within W.

7. The method according to claim 6, characterized in that In step 3, a probability matching model between trajectory and waybill number is established. Assume that there are N candidate package trajectories T = {T1, T2, …, T N }, each trajectory T i Contains a series of status information {x i ,y i ,t i ,e i }, where x i ,y i Indicates the location of the package, t i Indicates the timestamp, e i A feature vector representing the appearance of the package; For each candidate trajectory T i , calculate the matching probability P(T i |S), where S represents the relevant information of the waybill number, and the matching probability P(T i |S) can be broken down into the following parts: P(T i |S)=P(t i |t s )·P(e i |e s ) Among them, P(t i |t s ) represents the trajectory T i Timestamp t i Time of receiving the waybill number t s The matching degree of P(e i |e s ) represents the trajectory T i Appearance features i The appearance characteristics of the package corresponding to the waybill number when scanned on six sides s The cosine similarity is used to measure the similarity of appearance features.

8. The method according to claim 7, characterized in that In step 3, after calculating the matching probabilities of all candidate trajectories, the maximum a posteriori probability criterion is used to select the optimal matching result, that is, the trajectory T with the largest matching probability is selected. * As the final matching result:

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium for storing computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Pedestrian re-identification algorithm implementation method based on HSV and SDALF

    CN107679467A

  • Multi-target tracking method suitable for vehicle motion characteristics

    CN116152297A

  • A system and method for understanding and explaining spoken interactions using speech acoustic and linguistic markers

    GB202311309D0

  • System and method for attention-aware relation mixer for person search

    US20240153308A1