Video-based contraband tracking in packages method, apparatus, device, and medium

By extracting the minimum bounding rectangle and angular features of contraband items, and combining inter-frame matching detection and motion state, the problem of multi-frame image association was solved, enabling full-process trajectory tracking of contraband items and improving security inspection efficiency and accuracy.

CN122244105APending Publication Date: 2026-06-19HUNAN SUKE INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUNAN SUKE INTELLIGENT TECH CO LTD
Filing Date
2026-05-25
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing technologies cannot correlate the same package target in multiple frames of images, resulting in fragmented detection results. This makes it impossible to track the entire process of prohibited items from their appearance to removal from the security scanner, thus limiting the accuracy of security checks and the efficiency of rapid response.

Method used

By acquiring images from security inspection videos, the minimum bounding rectangle and angular features of the target contraband are extracted. Combined with inter-frame matching detection and motion state, the offset prediction value and detection value are calculated to achieve accurate tracking of contraband in packages.

Benefits of technology

It has improved the efficiency and accuracy of tracking prohibited items in packages, realized the full-process trajectory tracking of prohibited items, and enhanced the precision control and rapid handling efficiency of the security inspection process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122244105A_ABST
    Figure CN122244105A_ABST
Patent Text Reader

Abstract

This application relates to the field of video image processing technology, and discloses a method, apparatus, device, and medium for tracking contraband in packages based on video. The method includes: acquiring security inspection video and obtaining the corresponding image; extracting target features determined by length and angle features from the previous frame image; performing matching detection in the current frame image based on the target features; determining the current offset prediction value and the current offset detection value; and outputting the actual position of the target contraband in the target package in the current frame image based on the current offset prediction value or fusion result, thereby completing the video-based tracking of contraband in packages. The apparatus, device, and storage medium all correspond to this method. This application solves the technical problem that the inability to associate the same package target in multiple frames of images leads to fragmented detection results, making it impossible to achieve full-process trajectory tracking of contraband from its appearance to its removal from the security scanner, thus limiting the precise control and rapid handling efficiency of the security inspection process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video image processing technology, specifically a video-based method, apparatus, device, and medium for tracking contraband in packages. Background Technology

[0002] With the rapid development of my country's civil aviation and rail transit industries, more and more people are choosing to travel by air, high-speed rail, and subway. To ensure safety, passengers' carry-on bags must undergo security checks upon entering stations. However, due to the rapid increase in passenger flow and parcel volume at transportation hubs, the contradiction between continuous tracking of dynamic targets and tracing the source of prohibited items during security checks is becoming increasingly apparent. Traditional security screening machines rely on manual image interpretation, requiring security personnel to identify prohibited items only after the parcel has completely passed through and a complete image has been generated. During this time, the movement trajectory of the parcel on the conveyor belt cannot be correlated in real time. Once prohibited items are detected, it is difficult to quickly locate the parcel and its transmission position, making it unsuitable for high-frequency, high-flow security screening scenarios such as high-speed rail and airports.

[0003] In recent years, X-ray imaging-based security inspection technology has gradually combined with artificial intelligence. Deep learning-driven target detection algorithms, such as YOLO and Faster R-CNN, have begun to be applied to the identification of contraband. These algorithms automatically analyze scanned images to assist in image interpretation, reducing the workload of manual inspection to some extent. However, current technologies still rely on single-frame static detection and lack the ability to continuously track dynamic packages. When a package partially appears, moves, or overlaps with other packages during transmission, the algorithm cannot associate the same package target in multiple frames, resulting in fragmented detection results. This makes it impossible to track the entire trajectory of contraband from its appearance to its removal from the security scanner, limiting the accuracy and speed of security inspection. Summary of the Invention

[0004] The purpose of this application is to provide a video-based method, apparatus, device, and medium for tracking contraband in packages, in order to solve the technical problem in the prior art that the same package target in multiple frames of images cannot be associated, resulting in fragmented detection results and the inability to achieve full-process trajectory tracking of contraband from appearance to removal in the security inspection device, thus limiting the precise control and rapid handling efficiency of the security inspection process.

[0005] To achieve the above objectives, this application provides a video-based method for tracking contraband in packages, comprising: Acquire security inspection video and obtain corresponding images, including at least the current frame image and the previous frame image corresponding to the current frame image; for the current frame image, identify at least one target package and at least one target contraband item in the target package; Extract target features determined by length and angle features from the previous frame image. The length feature is determined based on the line segment defined by the minimum bounding rectangle of the target contraband, and the angle feature is determined based on the line segment corresponding to the length feature. Based on the target features, perform matching detection in the current frame image. For the current frame image where the matching detection is successful, the current offset prediction value is determined by combining the corresponding motion state, and the current offset detection value is calculated based on the preset selection rules. Based on the current offset prediction value or the fusion result, the actual position of the target contraband in the target package in the current frame image is output to complete the video-based tracking of contraband in the package. The fusion result is determined by fusing the current offset prediction value and the current offset detection value.

[0006] Preferably, the extraction of the target features includes: Find the minimum bounding rectangle corresponding to a target contraband item; Within the preset image coordinate system: Calculate the minimum bounding rectangle. The variance of pixel values ​​at each point in each column on the axis; based on the minimum bounding rectangle in The projection on the axis is combined with the preset division rules to divide the region; in each region after division, the column with the largest pixel value variance is determined, and the midpoint corresponding to each column with the largest pixel value variance is found. The midpoints are connected in sequence to obtain multiple line segments; the length corresponding to the line segment is defined as the length feature, the angle formed by adjacent line segments is defined as the angle feature, and the obtained length feature and angle feature are defined as the target feature.

[0007] Preferably, the matching detection based on target features in the current frame image includes: Obtain the target features from the previous frame and define them as the first target features; Based on the first target feature, the corresponding target feature is extracted from the current frame image and defined as the second target feature; Based on the comparison results of the first target feature and the second target feature, and combined with the preset comparison threshold, when the comparison result meets the comparison threshold, it is determined that the matching detection is successful; wherein, the comparison result includes the overall length difference rate and the overall angle difference rate.

[0008] Preferably, the current offset prediction value is determined based on a preset frame rate and the movement speed and acceleration corresponding to the previous frame image.

[0009] Preferably, the calculation of the current offset detection value based on a preset selection rule specifically involves initiating the calculation of the current offset detection value based on the selection rule; initiating the calculation of the current offset detection value includes: For a single contraband item in a target package, the sliding distance is determined based on the current offset prediction value and the number of slides corresponding to the preset sliding operation, thus determining the sliding distance corresponding to the target contraband item; when the matching detection is successful, the sliding distance is defined as the first offset detection value corresponding to the target contraband item. Based on preset grouping rules, for a target package, the contraband items with a first offset detection value are grouped into groups; for each group, the group variance and group mean corresponding to all first offset detection values ​​are calculated; when the group variance of a group meets a preset group variance threshold, the group is determined to be a stable group; the mean of the group means corresponding to all stable groups is calculated and defined as the second offset detection value. Calculate the mean of all second offset detection values ​​in the current frame image and define it as the current offset detection value corresponding to the current frame image.

[0010] Preferably, the movement speed corresponding to the current frame image is updated based on the preset frame rate and the current offset detection value, and the acceleration corresponding to the current frame image is updated accordingly; the updated movement speed and acceleration corresponding to the current frame image are used to optimize the current offset prediction value when the next frame image is the current frame image.

[0011] Preferably, the output of the actual position includes: When there is no current offset detection value, the current offset prediction value is used as the tracking offset of the current frame image. The position of the target contraband in the target package is updated based on the tracking offset to complete the video-based tracking of contraband in the package. When a current offset detection value exists, the fusion result is used as the tracking offset of the current frame image. The position of the target contraband in the target package is updated based on this tracking offset to complete the video-based tracking of contraband in the package. The fusion of the current offset prediction value and the current offset detection value corresponding to the fusion result is based on the weighting factor of the current offset prediction value. This weighting factor is determined based on the number of current offset detection values, the current offset detection value, the mean of the current offset detection values, and the current offset prediction value that exist throughout the entire tracking process for a target contraband.

[0012] To achieve the above objectives, this application also provides a video-based package contraband tracking device, which applies the video-based package contraband tracking method described above, including: The target determination module is used to acquire security inspection videos and obtain corresponding images, including at least the current frame image and the previous frame image corresponding to the current frame image; for the current frame image, it determines at least one target package and at least one target contraband item in the target package; The feature and matching module is used to extract target features determined by length and angle features from the previous frame image. The length feature is determined based on the line segment defined by the minimum bounding rectangle of the target contraband, and the angle feature is determined based on the line segment corresponding to the length feature. Based on the target features, matching detection is performed in the current frame image. The tracking optimization module is used to determine the current offset prediction value for the current frame image that has been successfully matched and detected, combined with the corresponding motion state, and calculate the current offset detection value based on the preset selection rules. Based on the current offset prediction value or the fusion result, it outputs the actual position of the target contraband in the target package in the current frame image, so as to complete the video-based tracking of contraband in the package. The fusion result is determined based on the fusion of the current offset prediction value and the current offset detection value.

[0013] To achieve the above objectives, this application also provides a video-based contraband tracking device for packages, including at least one processor, at least one memory, and a data bus; The processor and the memory communicate with each other via the data bus; The memory stores program instructions that can be executed by the processor, which invokes the program instructions to execute the video-based method for tracking contraband in packages as described above.

[0014] To achieve the above objectives, this application also provides a storage medium storing a computer program thereon, which, when executed by a processor, implements the video-based method for tracking contraband in packages as described above.

[0015] Beneficial Effects: The video-based method, apparatus, equipment, and medium for tracking contraband in packages disclosed in this application improve the overall efficiency and accuracy of contraband tracking by employing a continuous prediction and correction approach. Specifically, a state transition matrix method is used for prediction, and a feature vector-based method is used for correction. This solves the technical problem in existing technologies where it is impossible to correlate the same package target in multiple frames of images, leading to fragmented detection results and the inability to achieve full-process trajectory tracking of contraband from its appearance to its removal from the security scanner, thus limiting the precise control and rapid handling efficiency of the security inspection process. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1A flowchart illustrating a video-based method for tracking contraband in packages, as provided in an embodiment of this application. Figure 2 A schematic diagram illustrating the application of the video-based package contraband tracking method provided in this application embodiment; Figure 3 A structural block diagram of a video-based package contraband tracking device provided in an embodiment of this application; in the figure: 10, target determination module; 20, feature and matching module; 30, tracking optimization module.

[0018] The implementation, functional features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0019] The technical solutions in the embodiments of this application will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0020] In this document, the term "comprising" is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0021] In video image processing for tracking contraband in parcels, current technical solutions primarily rely on a combination of X-ray image target detection and multi-target tracking. This approach has been widely applied in security inspections and logistics. In practical applications, mainstream algorithm frameworks are typically based on classic deep learning models such as the YOLO (You Only Look Once) series and Faster R-CNN (Region-based Convolutional Neural Networks). These algorithms are used for efficient detection of contraband. Simultaneously, to achieve continuous localization of multiple targets, advanced tracking frameworks such as DeepSORT (Simple Online and Realtime Tracking with Deep Association Metric) or ByteTrack are often used to ensure accurate capture and continuous tracking of the target's trajectory during movement. Specifically, the workflow of these algorithms can be divided into two main stages: The first stage involves extracting feature information of the contraband through convolutional neural networks. This process requires deep analysis of the input X-ray image to identify potentially dangerous items, such as knives, lithium batteries, and flammable and explosive materials, generating corresponding bounding boxes for these targets, and determining their specific categories. The second stage utilizes inter-frame feature matching detection technology and target motion state prediction methods to associate the same target in different frames, thereby forming a complete target motion trajectory. This method can effectively handle the rapid movement of packages on a conveyor line and overcome interference problems caused by occlusion or changes in lighting, ultimately achieving accurate positioning and continuous tracking of contraband inside packages. However, the existing technical solutions mentioned above have significant technical shortcomings when performing multi-target tracking tasks, including: First, when implementing multi-target tracking, the DeepSORT algorithm primarily relies on matching high-confidence bounding boxes with target appearance features for data association. However, this mechanism is prone to association errors when targets are severely occluded or when multiple targets have highly similar appearances. For example, when parts of two targets are occluded by other objects, or when their appearance features such as color and texture are very similar, the algorithm may fail to correctly match the bounding boxes with the corresponding targets, resulting in unstable tracking results and poor overall tracking robustness.

[0022] Secondly, while the ByteTrack algorithm optimizes its strategy for utilizing low-confidence detection boxes, improving tracking performance by introducing more low-confidence detection information, its core still heavily relies on the quality of the detection results. Especially in scenarios with dense package occlusion, such as multiple packages closely arranged and obscuring each other on a logistics conveyor line, the algorithm is prone to trajectory breaks and target ID switching issues. This leads to frequent changes in the target's identity during tracking, making it difficult to maintain stable and continuous tracking, severely impacting reliability in practical applications.

[0023] Furthermore, both DeepSORT and ByteTrack require simultaneous target detection and tracking during operation, resulting in significant computational overhead. On one hand, the detection module needs to locate the target in each frame, a computationally intensive task; on the other hand, the tracking module needs to perform target matching and trajectory prediction, further increasing the system's burden. Therefore, in practical applications, these algorithms often struggle to maintain high positioning accuracy while ensuring real-time performance, especially in high-speed conveyor belt scenarios where the requirement for stable and continuous tracking of contraband is difficult to fully meet. This performance bottleneck limits the widespread application of existing technologies in industrial settings.

[0024] To address the aforementioned technical deficiencies, this embodiment discloses a video-based method, apparatus, device, and medium for tracking contraband in packages. In summary, this embodiment designs a technical solution for rapid video tracking of contraband in packages. When a contraband item enters the video frame and is detected, it can be tracked quickly and accurately until it leaves the video frame, achieving full-process trajectory tracking. The entire process is smooth and judder-free, providing security personnel with precise real-time location guidance, effectively reducing the time required to locate contraband, and improving security response speed and processing efficiency, thus meeting the urgent needs of modern transportation hubs for intelligent and efficient security checks.

[0025] The present embodiment will now provide a detailed description of the video-based method, apparatus, equipment, and medium for tracking contraband in packages.

[0026] Reference Figure 1 , Figure 1 A flowchart illustrating a video-based method for tracking contraband in packages, as provided in this application embodiment.

[0027] Firstly, such as Figure 1 As shown, this embodiment discloses a video-based method for tracking contraband in packages, including: S10: Acquire security inspection video and obtain corresponding images, including at least the current frame image and the previous frame image corresponding to the current frame image; for the current frame image, identify at least one target package and at least one target contraband in the target package.

[0028] In this specific application, S10 completes the real-time acquisition of video images from the security inspection machine and the detection of the current frame image, that is, detecting new packages and their contraband in the current frame image. In specific implementation, S10 is specifically manifested as follows: Real-time acquisition of each frame of video image, recording its sequence number. Obtain the direction of movement of the package in the video, for example, to the right. In this embodiment, the rightward direction is used as an example for explanation; other directions are similar. Obtain the preset frame rate of the video. Exemplary Obtain the horizontal movement speed of the preset video's 0th frame. The value is 0, acceleration The value is 0. It should be noted that, numerically, the image number... Starting from 0 and incrementing by 1 each time; in the security inspection field, the video movement direction of security inspection machines is generally horizontal scrolling, with no vertical direction. The preset initial values ​​can all be 0, or other empirical values. Based on this, real-time acquisition of security inspection machine video images is achieved.

[0029] Furthermore, the image performs package detection, including identifying and acquiring information on contraband, such as category and location, from new packages that were not present in the previous frame, and determining whether there are any historical packages corresponding to the previous frame. Based on this, the detection of the current frame is completed.

[0030] Through S10, this embodiment completes the selection of packages that require video-based tracking of contraband, providing corresponding targets for contraband tracking.

[0031] S20: Extract target features determined by length and angle features from the previous frame image, where the length feature is determined based on the line segment defined by the minimum bounding rectangle of the target contraband, and the angle feature is determined based on the line segment corresponding to the length feature; perform matching detection in the current frame image based on the target features.

[0032] Analysis of existing X-ray security inspection package contraband tracking algorithms reveals that current technologies largely rely on deep, full-domain image semantic feature extraction. This results in cumbersome computational chains and high feature redundancy, significantly consuming local edge computing resources of security equipment and easily causing excessive real-time tracking latency. Furthermore, the complex feature matching logic is difficult to adapt to high-speed, high-volume security inspection scenarios. Therefore, this embodiment optimizes the target features for video-based package contraband tracking.

[0033] Specifically, the extraction of target features includes: Find the minimum bounding rectangle corresponding to a target contraband item; Within the preset image coordinate system: Calculate the minimum bounding rectangle. The variance of pixel values ​​at each point in each column on the axis; based on the minimum bounding rectangle in The projection on the axis is combined with the preset division rules to divide the region; in each region after division, the column with the largest pixel value variance is determined, and the midpoint corresponding to each column with the largest pixel value variance is found. The midpoints are connected in sequence to obtain multiple line segments; the length corresponding to the line segment is defined as the length feature, the angle formed by adjacent line segments is defined as the angle feature, and the obtained length feature and angle feature are defined as the target feature.

[0034] In summary, this embodiment targets the minimum bounding rectangle of contraband items and extracts length and angle structured geometric features based on pixel column variance distribution. While ensuring the accuracy of package tracking and identification, it simplifies the feature calculation process, reduces local computing power overhead, and adapts to the low latency and high real-time operation requirements of security inspection edge terminals.

[0035] Real-time acquisition of each frame of video image and extraction of target features, specifically manifested as follows: Obtain the minimum bounding rectangle (MBR) of a contraband item; calculate the minimum bounding rectangle of this MBR. The variance of pixel values ​​in each column on the axis; based on this MBR... Projecting onto the axis yields several regions; identifying the column with the largest pixel value variance in each region; calculating the midpoints of these columns; connecting each midpoint sequentially to obtain the corresponding line segments; and obtaining the corresponding length and angle features based on these line segments.

[0036] Reference Figure 2 , Figure 2 This is a schematic diagram illustrating the application of the video-based method for tracking contraband in packages provided in this embodiment of the application.

[0037] like Figure 2 As shown, in one specific application of this embodiment, the package moves from left to right, with the left side representing the previous frame and the right side representing the current frame. Taking the previous frame as an example, the extraction of target features is explained. The MBR of the contraband in the image is obtained; the MBR is calculated in... The variance of pixel values ​​in each column on the axis; based on this MBR... The projection on the axis is divided into equal parts. The four regions; find the column with the largest variance in pixel values ​​for each region, i.e. Figure 2 The left side of the middle section shows the vertical red line; calculate the midpoint of each of these columns, i.e. , , and Connect the midpoints in sequence to obtain the corresponding line segments. line segments and line segments Based on these line segments, the corresponding length and angle characteristics are obtained, i.e., line segments. line segments and line segments The corresponding length values ​​and line segments respectively line segments and line segments The angle formed and The corresponding angle values. It should be noted that: in cases such as... Figure 2 In the example shown, the region is divided using an average division rule. However, in practical applications, in response to computing power allocation requirements and / or detection accuracy requirements, the division rule can also be a dynamic algorithm corresponding to the computing power allocation requirements and / or detection accuracy requirements. For example, an adaptive solution algorithm is constructed to determine the number of regions to be divided and the computing power allocation requirements and / or detection accuracy requirements, thereby achieving real-time optimization of the region division.

[0038] In tracking packages between security inspection frames, existing technologies mostly employ methods such as global pixel association matching and multi-dimensional deep feature comparison. These methods have high inter-frame computational complexity and large feature redundancy, severely consuming the local computing power of edge devices and making them unsuitable for the continuous frame real-time tracking requirements of high-speed security inspection pipelines. Therefore, this embodiment optimizes the matching detection in the current frame image based on the simple but high-quality target features representing the length and angle structure of contraband extracted earlier in S20.

[0039] Specifically, based on target features, matching detection is performed in the current frame image, including: Obtain the target features from the previous frame and define them as the first target features; Based on the first target feature, the corresponding target feature is extracted from the current frame image and defined as the second target feature; Based on the comparison results of the first target feature and the second target feature, and combined with the preset comparison threshold, when the comparison result meets the comparison threshold, it is determined that the matching detection is successful; wherein, the comparison result includes the overall length difference rate and the overall angle difference rate.

[0040] Combination Figure 2 The following details the implementation of a feature-based matching detection method in the current frame image. For the current frame image on the right, the MBR of the contraband in the image is also obtained; the MBR is calculated in... The variance of pixel values ​​in each column on the axis; based on this MBR... The projection on the axis is divided into equal parts. The four regions; find the column with the largest variance in pixel values ​​for each region, i.e. Figure 2 The right side of the middle section shows the vertical red line; calculate the midpoint of each of these columns, i.e. , , and Connect the midpoints in sequence to obtain the corresponding line segments. line segments and line segments Based on these line segments, the corresponding length and angle characteristics are obtained, i.e., line segments. line segments and line segments The corresponding length values ​​and line segments respectively line segments and line segments The angle formed and The corresponding angle values.

[0041] Based on this, the first target feature obtained is a line segment. line segments and line segments The corresponding length values ​​and line segments respectively line segments and line segments The angle formed and The corresponding angle values ​​yielded the second target feature as a line segment. line segments and line segments The corresponding length values ​​and line segments respectively line segments and line segments The angle formed and The corresponding angle values.

[0042] The specific application of the overall length difference rate in the comparison results is as follows: Calculate the line segments separately and line segments Length difference rate line segments and line segments Length difference rate and line segments and line segments Length difference rate The mathematical expression corresponding to the calculation method is as follows: in, It is the absolute value operator; Based on length difference rate Length difference rate and length difference rate Calculate the overall length difference rate The mathematical expression corresponding to the calculation method is as follows: The specific application of the overall angular difference rate in the comparison results is as follows: Calculate the angles separately and Angular difference rate , and Angular difference rate The mathematical expression corresponding to the calculation method is as follows: in, It is the absolute value operator; Based on angle difference rate and angle difference rate Calculate the overall angular difference rate The mathematical expression corresponding to the calculation method is as follows: Based on this, we obtain the overall length difference rate. and overall angle difference rate The comparison results.

[0043] In a specific application of this embodiment, for example, the following comparison threshold can be set: when Less than the preset threshold ,and Less than the preset threshold When the target feature of the contraband at that location is successfully matched, the preset threshold value can be... , .

[0044] Based on the above, the structured target features corresponding to the length and angle of the contraband extracted by S20 are used to extract the feature parameters of the two frames across frames. The inter-frame matching is completed by judging the difference rate threshold of length and angle, which greatly simplifies the inter-frame association operation logic. While ensuring the tracking accuracy, it reduces the local computing power consumption of the device and realizes low-latency continuous and stable tracking of security inspection targets.

[0045] Based on the above S10 and S20, an efficient and high-quality data foundation is provided for video-based tracking of contraband in packages. However, in practical applications, it has been found that existing security inspection frame-to-frame contraband tracking methods mostly rely solely on inter-frame matching results to locate target offsets, without considering the continuous movement sequence of the package for predictive correction. This makes them prone to position jumps and tracking drifts during high-speed movement on the assembly line and changes in target posture, which reduces positioning accuracy and consumes additional edge computing power. Therefore, this embodiment further optimizes the real-time positioning of packages based on the successful inter-frame feature matching determination in S20.

[0046] S30: For the current frame image where the matching detection is successful, determine the current offset prediction value in combination with the corresponding motion state, and calculate the current offset detection value based on the preset selection rules; based on the current offset prediction value or fusion result, output the actual position of the target contraband in the target package in the current frame image to complete the video-based tracking of contraband in the package; wherein, the fusion result is determined based on the fusion of the current offset prediction value and the current offset detection value.

[0047] Specifically, the current offset prediction value is determined based on the preset frame rate and the movement speed and acceleration corresponding to the previous frame image.

[0048] In the specific application of this embodiment, the mathematical expression for calculating the current offset prediction value is as follows: Calculate the current frame image Current offset prediction value : in: This represents the movement speed corresponding to the previous frame. This is the acceleration corresponding to the previous frame image; , The frame rate is the preset value mentioned above. It should be noted that the length concepts such as the current offset prediction value and the current offset detection value in this embodiment are all based on the unit of one pixel.

[0049] Furthermore, the following calculations are performed simultaneously: Calculate the current frame image Corresponding movement speed : Calculate the current frame image corresponding acceleration : It should be noted that the current offset prediction value calculated here is only a preliminary prediction, and will be corrected based on the detection results to improve accuracy; the motion speed and acceleration calculation corresponding to the current frame image here provide the data basis for subsequent corrections.

[0050] Specifically, the current offset detection value is calculated based on preset selection rules, which means that the calculation of the current offset detection value is started based on the selection rules.

[0051] In the specific application of this embodiment, to further control the generation of calculations, a prerequisite design was implemented for the calculation of the current offset detection value. In a feasible implementation, the following two conditions are set: first, if Less than For example, such as That is, the first 5 frames; secondly, it can be... Divisible, for example, such as When either of the above two conditions is met, the calculation of the current offset detection value is initiated.

[0052] When initiating the calculation of the current offset detection value, the following is included: For a target contraband in a target package, the sliding distance is determined based on the current offset prediction value and the number of sliding operations corresponding to the preset sliding operation, and the sliding distance corresponding to the target contraband is determined; when the matching detection is successful, the sliding distance is defined as the first offset detection value corresponding to the target contraband.

[0053] In one specific application of this embodiment, the following sliding operation is performed on the current frame image to perform target feature matching: Get the preset maximum number of swipes Exemplary Each swipe moves one pixel; iterate through all swipe counts and calculates the length of the swipe for the ... The sliding distance of the second slide The sliding distance The mathematical expression for the calculation method is as follows: Among them: when At that time, that is, at the first slide ;when At that time, that is, during the second slide, That is, the length of the slide of one pixel; when At that time, that is, during the third slide, That is, the length of the slide of two pixels; that is, the number of times. (from 1 to (positive integers) and the formula used for calculation (i.e., 0 to) The relationship between integers is ,and The unit is relative to the current offset prediction value. Correspondingly, that is The unit is also the length of a pixel. Indicates to The value is rounded up, for example: during the second slide, hour, During the third slide, hour, During the fourth slide, hour, . Used to control the direction of the slide, in this embodiment, for example, a positive number represents right and a negative number represents left; a further example is: in the second slide ( )hour, hour, , On the second slide, it slides one pixel to the right; on the third slide... )hour, hour, , On the third slide, it slides one pixel to the left. Therefore, in reality, in When sliding to the left still yields a positive result, the current offset prediction value is... Swipe left and right alternately nearby for a higher hit rate; When sliding to the left results in a negative number, the result is directly changed. This ensures the result is positive.

[0054] Therefore, the rightward sliding distance of the contraband MBR can be obtained. At this point, the above-mentioned operation of matching and detecting targets in the current frame image is performed. When the matching and detection is successful, the sliding distance is... Define the first offset detection value corresponding to the target contraband, and then end the sliding.

[0055] Based on preset grouping rules, for a target package, the prohibited items with the first offset detection value are grouped into groups. For each group, the group variance and group mean corresponding to all first offset detection values ​​are calculated. When the group variance of a group meets the preset group variance threshold, the group is determined to be a stable group. The mean of the group means corresponding to all stable groups is calculated and defined as the second offset detection value.

[0056] In the specific application of this embodiment, the set of contraband items within a target package that successfully matches the target features is positioned as... . Set All prohibited items are grouped according to material category, for example, such as metals, glass, organic materials, and liquids. The following evaluation and screening operations are performed on all prohibited item groups within the target package: Calculate the group variance corresponding to the first offset detection value of all samples within the group. and group mean If the group variance Less than the preset value Exemplary If the matching detection results of prohibited items in this group are considered stable and reliable, the group is classified as a stable group, and the deviation of this group is considered to be the group mean. Calculate the group mean for all stable groups. mean The mean Defined as the second offset detection value.

[0057] Calculate the mean of all second offset detection values ​​in the current frame image and define it as the current offset detection value corresponding to the current frame image.

[0058] In the specific application of this embodiment, a second offset detection value is detected in the current frame image. Calculate the average of all target packages The mean Defined as the current offset detection value corresponding to the current frame image. At this point, the movement speed and acceleration corresponding to the current frame image are updated simultaneously.

[0059] Specifically, the movement speed corresponding to the current frame image is updated based on the preset frame rate and the current offset detection value, and the acceleration corresponding to the current frame image is updated accordingly; the updated movement speed and acceleration corresponding to the current frame image are used to optimize the current offset prediction value when the next frame image is the current frame image.

[0060] In the specific application of this embodiment, based on the current offset detection value The correction of the motion state is completed, and the corresponding mathematical expression is as follows: Calculate the updated current frame image Corresponding movement speed : Calculate the updated current frame image corresponding acceleration : The updated movement speed corresponding to the current frame image and acceleration Used to optimize the current offset prediction value when the next frame image is the current frame image.

[0061] Therefore, based on the successful inter-frame feature matching determination obtained in S20 of this embodiment, the offset prediction is calculated by combining the package's temporal motion state with the measured values. The real-time pose of the target is calibrated by dual data fusion, thus locking the dynamic position of the contraband inside the package and ensuring stable and reliable continuous video tracking with lightweight computing power.

[0062] Based on the above optimizations, this embodiment provides a complete and accurate data foundation for the optimized output of the actual location.

[0063] Specifically, the output of the actual location includes: When there is no current offset detection value, the current offset prediction value is used as the tracking offset of the current frame image. The position of the target contraband in the target package is updated based on the tracking offset to complete the video-based tracking of contraband in the package.

[0064] In this specific application, to strictly control computation and ensure stable computing power and tracking quality, when the current offset detection value is not activated, the current offset prediction value is used as the tracking offset of the current frame image; that is, the current offset prediction value and the current offset detection value are not fused. However, thanks to the aforementioned update of the movement speed and acceleration corresponding to the current frame image, this embodiment can still maintain accurate video-based tracking of contraband in packages even without fusion. Correspondingly, the current offset prediction value is used as the tracking offset of the current frame image. The mathematical expression is: When a current offset detection value exists, the fusion result is used as the tracking offset of the current frame image. The position of the target contraband in the target package is updated based on this tracking offset to complete the video-based tracking of contraband in the package. The fusion of the current offset prediction value and the current offset detection value corresponding to the fusion result is based on the weighting factor of the current offset prediction value. This weighting factor is determined based on the number of current offset detection values, the current offset detection value, the mean of the current offset detection values, and the current offset prediction value that exist throughout the entire tracking process for a target contraband.

[0065] In the specific application of this embodiment, combined with the aforementioned selection rules, the periodic calculation of the current offset detection value is realized, specifically as follows: When a current offset detection value exists, the current offset detection value and the current offset prediction value are fused. The specific fusion method is as follows: Calculate the weighting factor for the current offset prediction value of the current frame image. Each time the current offset detection value is calculated, a sample is generated, numbered from 1 to... For the current frame image as the first Each sample, its weighting factor The mathematical expression for the calculation method is as follows: in: The total number of samples; For the first The current offset detection value of each sample, i.e., the current frame image. of ; For the first The current offset prediction value of each sample, i.e., the current frame image. of ; This is the average of the current offset detection values ​​generated during the current tracking. Therefore, the fusion result is used as the tracking offset of the current frame image. The corresponding mathematical expression is: Based on the tracking offset of the output In the wrapping motion taking a horizontal rightward direction as an example, the optimized output of the actual position is specifically manifested as follows: When it is determined that the package is appearing for the first time in the current frame, the recognition result is directly used as its position in the current frame; when it is determined that the package is not appearing for the first time in the current frame, the position of the package in the previous frame is obtained. Based on the tracking offset of the output The actual location of the package in the current frame image is ,in, This represents the horizontal coordinate (pixel position) of the top-left corner of the package detection box. This represents the top-left corner (vertical pixel position) of the package detection bounding box. This represents the width of the package detection bounding box (in pixels horizontally). This indicates the height of the package detection box (vertical pixel length).

[0066] In summary, the video-based package contraband tracking method of this embodiment has at least the following technical innovations and corresponding technical effects: A novel method for extracting the target features of contraband in packages is proposed, which solves the problems of slow speed and instability of existing extraction methods; A novel matching detection method for contraband targets in packages is proposed, which solves the problems of slow speed and inaccuracy of existing matching detection methods. A novel method for predicting image horizontal offset is proposed to address the slow speed of existing prediction methods. A novel fusion method for image horizontal offset prediction and detection is proposed to address the problems of slow speed and inaccuracy in existing fusion methods.

[0067] Based on the above, the video-based package contraband tracking method of this embodiment improves the overall efficiency and accuracy of package contraband tracking by employing a continuous prediction and correction approach. Specifically, a state transition matrix method is used for prediction, and a feature vector-based method is used for correction. This solves the technical problem in existing technologies where it is impossible to associate the same package target in multiple frames of images, resulting in fragmented detection results and the inability to achieve full-process trajectory tracking of contraband from its appearance to its removal from the security scanner, thus limiting the precise control and rapid handling efficiency of the security inspection process.

[0068] Reference Figure 3 , Figure 3 A structural block diagram of a video-based package contraband tracking device provided in an embodiment of this application; in the figure: 10, target determination module; 20, feature and matching module; 30, tracking optimization module.

[0069] Secondly, such as Figure 3 As shown, this embodiment also discloses a video-based contraband tracking device for packages, which applies the video-based contraband tracking method for packages described above, including: The target determination module 10 is used to acquire security inspection video and obtain corresponding images, including at least the current frame image and the previous frame image corresponding to the current frame image; for the current frame image, it determines at least one target package and at least one target contraband in the target package; The feature and matching module 20 is used to extract target features determined by length features and angle features in the previous frame image, wherein the length feature is determined based on the line segment determined by the minimum bounding rectangle of a target contraband, and the angle feature is determined based on the line segment corresponding to the length feature; and based on the target features, matching detection is performed in the current frame image. The tracking optimization module 30 is used to determine the current offset prediction value for the current frame image that has been successfully matched and detected, in combination with the corresponding motion state, and calculate the current offset detection value based on the preset selection rules; based on the current offset prediction value or the fusion result, it outputs the actual position of the target contraband in the target package in the current frame image, so as to complete the tracking of contraband in the package based on video; wherein, the fusion result is determined based on the fusion of the current offset prediction value and the current offset detection value.

[0070] Thirdly, this embodiment also discloses a video-based contraband tracking device in a package, including at least one processor, at least one memory, and a data bus; The processor and memory communicate with each other via a data bus; The memory stores program instructions that can be executed by the processor, which calls the program instructions to execute the video-based method for tracking contraband in packages as described above.

[0071] Fourthly, this embodiment also discloses a storage medium storing a computer program, which, when executed by a processor, implements the video-based method for tracking contraband in packages as described above.

[0072] It should be noted that the video-based package contraband tracking device, equipment, and storage medium of this embodiment correspond to the aforementioned video-based package contraband tracking method. Therefore, any content not specifically described in the video-based package contraband tracking device, equipment, and storage medium of this embodiment, including but not limited to functional definitions, working principles, and technical effects, can be referred to the description in the aforementioned video-based package contraband tracking method, and will not be repeated here.

[0073] In the embodiments provided in this application, it should be understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, code, or any suitable combination thereof. For hardware implementation, the processor may be implemented in one or more of the following: application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, other electronic units designed to implement the functions described herein, or combinations thereof. For software implementation, some or all of the processes of the embodiments may be performed by a computer program instructing the associated hardware. During implementation, the program may be stored in a computer-readable storage medium or transmitted as one or more instructions or code on a computer-readable storage medium. Computer-readable storage media include computer storage media and communication media, wherein communication media include any medium that facilitates the transmission of a computer program from one place to another. Storage media may be any available medium accessible to a computer. Computer-readable storage media may include, but are not limited to, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code having the form of instructions or data structures and accessible to a computer.

[0074] Finally, it should be noted that the above description is only a preferred embodiment of this application and is not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A video-based contraband tracking method in a package, characterized by, include: Acquire security inspection video and obtain corresponding images, including at least the current frame image and the previous frame image corresponding to the current frame image; For the current frame image, identify at least one target package and at least one target contraband item within the target package; Extract target features determined by length and angle features from the previous frame image. The length feature is determined based on the line segment defined by the minimum bounding rectangle of the target contraband, and the angle feature is determined based on the line segment corresponding to the length feature. Based on the target features, perform matching detection in the current frame image. For the current frame image where the matching detection is successful, the current offset prediction value is determined by combining the corresponding motion state, and the current offset detection value is calculated based on the preset selection rules. Based on the current offset prediction value or the fusion result, the actual position of the target contraband in the target package in the current frame image is output to complete the video-based tracking of contraband in the package. The fusion result is determined by fusing the current offset prediction value and the current offset detection value.

2. The video-based contraband tracking in packages method of claim 1, wherein, The extraction of the target features includes: Find the minimum bounding rectangle corresponding to a target contraband item; Within the preset image coordinate system: Calculate the minimum bounding rectangle. The variance of pixel values ​​at each point in each column on the axis; based on the minimum bounding rectangle in The projection on the axis is combined with the preset division rules to divide the region; in each region after division, the column with the largest pixel value variance is determined, and the midpoint corresponding to each column with the largest pixel value variance is found. The midpoints are connected in sequence to obtain multiple line segments; the length corresponding to the line segment is defined as the length feature, the angle formed by adjacent line segments is defined as the angle feature, and the obtained length feature and angle feature are defined as the target feature.

3. The video-based contraband tracking in packages method of claim 1, wherein, The aforementioned matching detection based on target features in the current frame image includes: Obtain the target features from the previous frame and define them as the first target features; Based on the first target feature, the corresponding target feature is extracted from the current frame image and defined as the second target feature; Based on the comparison results of the first target feature and the second target feature, and combined with the preset comparison threshold, when the comparison result meets the comparison threshold, it is determined that the matching detection is successful; wherein, the comparison result includes the overall length difference rate and the overall angle difference rate.

4. The video-based contraband tracking in packages method of claim 1, wherein, The current offset prediction value is determined based on the preset frame rate and the movement speed and acceleration corresponding to the previous frame image.

5. The video-based contraband tracking in packages method of claim 1, wherein, The calculation of the current offset detection value based on the preset selection rules specifically involves starting the calculation of the current offset detection value based on the selection rules. When initiating the calculation of the current offset detection value, the following is included: For a single contraband item in a target package, the sliding distance is determined based on the current offset prediction value and the number of slides corresponding to the preset sliding operation, thus determining the sliding distance corresponding to the target contraband item; when the matching detection is successful, the sliding distance is defined as the first offset detection value corresponding to the target contraband item. Based on preset grouping rules, for a target package, the contraband items with a first offset detection value are grouped into groups; for each group, the group variance and group mean corresponding to all first offset detection values ​​are calculated; when the group variance of a group meets a preset group variance threshold, the group is determined to be a stable group; the mean of the group means corresponding to all stable groups is calculated and defined as the second offset detection value. Calculate the mean of all second offset detection values ​​in the current frame image and define it as the current offset detection value corresponding to the current frame image.

6. The video-based contraband tracking in packages method of claim 5, wherein, The movement speed corresponding to the current frame image is updated based on the preset frame rate and the current offset detection value, and the acceleration corresponding to the current frame image is updated accordingly. The updated movement speed and acceleration corresponding to the current frame image are used to optimize the current offset prediction value when the next frame image is the current frame image.

7. The video-based contraband tracking in packages method of claim 1, wherein, The output of the actual position includes: When there is no current offset detection value, the current offset prediction value is used as the tracking offset of the current frame image. The position of the target contraband in the target package is updated based on the tracking offset to complete the video-based tracking of contraband in the package. When a current offset detection value exists, the fusion result is used as the tracking offset of the current frame image. The position of the target contraband in the target package is updated based on this tracking offset to complete the video-based tracking of contraband in the package. The fusion of the current offset prediction value and the current offset detection value corresponding to the fusion result is based on the weighting factor of the current offset prediction value. This weighting factor is determined based on the number of current offset detection values, the current offset detection value, the mean of the current offset detection values, and the current offset prediction value that exist throughout the entire tracking process for a target contraband.

8. A video-based contraband tracking apparatus for packages, which applies the video-based contraband tracking method according to any one of claims 1 to 7, characterized by The device includes: The target determination module is used to acquire security inspection videos and obtain corresponding images, including at least the current frame image and the previous frame image corresponding to the current frame image; for the current frame image, it determines at least one target package and at least one target contraband item in the target package; The feature and matching module is used to extract target features determined by length and angle features from the previous frame image. The length feature is determined based on the line segment defined by the minimum bounding rectangle of the target contraband, and the angle feature is determined based on the line segment corresponding to the length feature. Based on the target features, matching detection is performed in the current frame image. The tracking optimization module is used to determine the current offset prediction value for the current frame image that has been successfully matched and detected, combined with the corresponding motion state, and calculate the current offset detection value based on the preset selection rules. Based on the current offset prediction value or the fusion result, it outputs the actual position of the target contraband in the target package in the current frame image, so as to complete the video-based tracking of contraband in the package. The fusion result is determined based on the fusion of the current offset prediction value and the current offset detection value.

9. A video-based contraband tracking device in a package, comprising: Includes at least one processor, at least one memory, and a data bus; The processor and the memory communicate with each other via the data bus; The memory stores program instructions that can be executed by the processor, which invokes the program instructions to execute the video-based method for tracking contraband in packages as described in any one of claims 1 to 7.

10. A medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the video-based method for tracking contraband in packages as described in any one of claims 1 to 7.