Parcel detection method and device and storage medium
By obtaining continuous frame images of the monitoring area on the embedded device, extracting the area of interest and performing background elimination and contour extraction, combining the target tracking algorithm and privacy protection design, the problems of high false alarm rate, high computing resource consumption and high deployment cost in package detection are solved, and accurate and low-cost package status detection is achieved.
Patent Information
- Application Number
- CN202510547383.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-19
AI Technical Summary
The existing parcel detection technology has high false alarm rate, high computing resource consumption, high deployment cost and privacy leakage risks, making it difficult to achieve real-time and accurate parcel status detection in embedded devices.
By acquiring the original images of the continuous frames of the monitoring area, extracting the area of interest, performing background elimination and contour extraction, using the target tracking algorithm to construct motion trajectories, combining spatial position and time persistence to determine the arrival and departure status of the package, using a tilt-mounted camera and privacy protection design to reduce computing resource consumption and false alarm rate.
It realizes the precise detection of package status under low computing power conditions on embedded devices, reduces false alarm rates, adapts to multi-scenario environments, protects user privacy, and reduces hardware costs and maintenance complexity.
Smart Images

Figure CN120510185A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of security monitoring, and in particular to a package detection method, device and storage medium. Background Art
[0002] Traditional package detection technologies primarily rely on physical sensors (such as pressure and infrared sensors) or video surveillance combined with deep learning solutions, which present significant drawbacks. Sensor solutions are susceptible to environmental interference, resulting in high false alarm rates. For example, windblown debris or temperature fluctuations can trigger false alarms, and they struggle to distinguish packages from other objects. Deep learning-based video surveillance, while capable of target recognition, requires high-performance computing resources, making it difficult to run in real time on embedded devices and posing the risk of user privacy breaches. Furthermore, while RFID technology can accurately identify packages, it requires pre-installed tags and is expensive to deploy, limiting its application.
[0003] Therefore, there is an urgent need for a package detection method that does not rely on physical sensors or preset tags and can simultaneously reduce computing resource consumption and false alarm rate. Summary of the Invention
[0004] In view of this, the present invention provides a package detection method, device, and storage medium that do not rely on physical sensors or preset tags and can simultaneously reduce computing resource consumption and false alarm rate. The technical solution is as follows.
[0005] In a first aspect, the present invention provides a package detection method, the method comprising:
[0006] Obtain continuous frame original images of the monitoring area and extract the region of interest from each frame original image;
[0007] Perform background removal and contour extraction on the image of the region of interest to obtain candidate targets;
[0008] The target tracking algorithm is used to associate candidate targets in different frames into the motion trajectory of the same target, and the motion trajectory of the valid target is obtained;
[0009] According to the movement trajectory of the valid target, determine whether the valid target leaves or enters the area of interest for the first time;
[0010] When a valid target enters the area of interest for the first time and its presence time exceeds the set threshold, the package is confirmed to have arrived; when a valid target leaves the area of interest, the package is confirmed to have left.
[0011] The package detection method provided by the present invention achieves multiple technological breakthroughs in determining the arrival and departure status of packages through continuous frame image dynamic analysis technology. First, by acquiring continuous raw image frames of the monitored area and extracting regions of interest (ROIs), computing resources are focused on key areas, reducing the amount of data processing compared to full-frame image processing solutions. Combined with background removal and contour extraction techniques, the proposed method improves target segmentation accuracy. Second, a motion trajectory continuity model is constructed based on the target tracking algorithm. By correlating candidate targets across frames, the proposed method accurately distinguishes between brief disturbances and actual package stops. A package arrival is detected when a target first enters the ROI and its presence exceeds a set threshold, while a departure is detected when the target leaves the ROI. This dual verification mechanism, based on spatial location and temporal continuity, significantly reduces the false alarm rate caused by environmental interference. Furthermore, this method achieves full visual detection through a non-dependent architecture (without the need for physical sensors or RFID tags as in traditional solutions), overcoming the deployment limitations of traditional solutions caused by insufficient sensor sensitivity (e.g., infrared cannot detect non-heat-generating objects) or tag dependency. It can adapt to packages of different materials and access control scenarios of various sizes, maintaining a high target detection rate even in backlit and low-contrast environments. Furthermore, a trajectory continuity recovery mechanism significantly reduces the false alarm rate for package departures.
[0012] In an optional embodiment, the original image is acquired by an image acquisition device that is installed obliquely on the door body of the access control area, and the installation angle of the image acquisition device is determined according to the door body structure of the access control area.
[0013] In the package detection method provided by this invention, the image acquisition device is tilted and mounted on the door body, with the angle adjusted according to the door structure. By adjusting the camera's tilt angle, the coverage area outside the door is expanded, solving the limited field of view of traditional vertically mounted cameras and improving the completeness of package detection in the region of interest (ROI).
[0014] In an optional embodiment, extracting the region of interest includes:
[0015] Extract the image of the region of interest from each frame of the original image, and determine whether the image of the region of interest in the current frame is valid;
[0016] If it is valid, the image of the region of interest of the current frame is converted into a grayscale image, and the grayscale image is enhanced by the contrast-limited adaptive histogram equalization algorithm;
[0017] If not valid, the processing of the image of the region of interest in the current frame is skipped, and it is determined whether the image of the region of interest in the next frame is valid.
[0018] In an optional implementation, determining whether the region of interest image is valid includes:
[0019] When the image of the region of interest meets any invalidity condition, confirming that the image of the region of interest is invalid;
[0020] The invalid conditions include: the boundary of the image of the region of interest exceeds the valid range of the original image, the area of the region of interest is smaller than a preset threshold, or the pixel information in the region of interest is abnormal.
[0021] The package detection method provided by the present invention skips invalid frame processing and reduces invalid calculation amount through ROI boundary, area and pixel anomaly detection (such as exceeding the image range or area being too small).
[0022] In an optional embodiment, background removal and contour extraction are performed on an image of a region of interest to obtain a candidate target, including:
[0023] The background of the image of the region of interest is removed by using a mixed Gaussian model to obtain a foreground mask;
[0024] Optimize the foreground mask through opening and closing operations;
[0025] The optimized foreground mask is contour extracted, and interference targets are filtered according to the area threshold to obtain candidate targets.
[0026] In an optional embodiment, the target tracking algorithm is used to associate candidate targets in different frames with the motion trajectory of the same target to obtain the motion trajectory of the valid target, including:
[0027] The regional overlap between the candidate target in the current frame and the historical tracking target is calculated by intersection-over-union matching, and the target association is performed according to the preset matching threshold;
[0028] Update the motion trajectory of the target that is continuously matched successfully and assign a unique tracking identifier to the same target;
[0029] Generate new tracking identifiers for unsuccessful matching candidates.
[0030] The package detection method provided by the present invention associates cross-frame targets through the intersection-over-union (IOU) matching threshold and combines it with a unique identifier mechanism to achieve continuous tracking of the package's motion trajectory.
[0031] In an optional embodiment, judging whether the valid target leaves or enters the region of interest for the first time based on the motion trajectory of the valid target includes:
[0032] Analyze the motion trajectory of valid targets through difference analysis algorithm;
[0033] When a valid target enters a preset boundary range of the region of interest for the first time, it is confirmed that the valid target enters the region of interest for the first time;
[0034] When the overlap between the valid target and the region of interest is lower than a preset threshold and no valid target is detected in consecutive target frames, it is determined that the valid target leaves the region of interest;
[0035] The preset threshold is that the continuous target frames detect that the valid target exists in the region of interest, and the position change amplitude of the valid target is less than the preset range.
[0036] The package detection method provided by the present invention distinguishes between temporary storage and permanent departure of packages based on an overlap threshold (e.g., the amplitude of the position change of consecutive frames is less than a preset range) and statistics on the number of disappeared frames, thus avoiding the misjudgment problem caused by temporary occlusion of the target (e.g., pedestrians passing by) in traditional solutions.
[0037] In an optional embodiment, the method further includes:
[0038] The preset privacy area in the original image is blurred and the original clarity of the non-privacy area is retained to generate a privacy-preserving image.
[0039] The package detection method provided by the present invention desensitizes sensitive areas inside the door (such as indoor scenes) through background blurring technology (such as Gaussian blur), thereby ensuring the accuracy of package detection while meeting the requirements of privacy protection.
[0040] In summary, the package detection method provided by the present invention is based on the continuous frame ROI processing mechanism (acquiring image → extracting ROI → background elimination → target tracking → trajectory determination), combined with the optimization of the camera tilt installation angle (adapting to the door structure to expand the monitoring field of view), focusing computing resources on key areas, and reducing the amount of data; through the mixed Gaussian model dynamic background modeling and morphological optimization, it effectively suppresses the interference of rain, snow, and sudden changes in illumination, and improves the target segmentation accuracy; ROI validity verification (boundary, area, pixel anomaly detection) and CLAHE grayscale enhancement work together to filter invalid frames and improve the detection rate of low-contrast scene contours; the intersection-to-union matching algorithm and the unique identifier allocation mechanism realize independent tracking of multiple targets and restore them through trajectory continuity; the entry and exit area judgment logic (overlap threshold + continuous frame disappearance detection) is combined with the time and space dual-dimensional verification to eliminate false triggering caused by short-term interference; the privacy area fuzzy processing (Gaussian blurring of sensitive areas) significantly reduces the risk of privacy leakage while ensuring detection accuracy, meeting privacy protection requirements.
[0041] In a second aspect, the present invention provides a package detection device, comprising:
[0042] The region extraction module is used to obtain the continuous frame original image of the monitoring area and extract the region of interest from each frame original image;
[0043] The candidate target acquisition module is used to perform background removal and contour extraction on the image of the region of interest to obtain candidate targets;
[0044] The tracking association module is used to associate candidate targets in different frames into the motion trajectory of the same target through the target tracking algorithm to obtain the motion trajectory of the valid target;
[0045] A judgment module is used to judge whether a valid target enters or leaves a region of interest based on its motion trajectory;
[0046] The confirmation module is used to confirm the arrival of a package when a valid target enters the area of interest for the first time and its continuous presence time exceeds a set threshold; and to confirm the departure of a package when a valid target leaves the area of interest.
[0047] In a third aspect, the present invention provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the package detection method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.
[0048] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to cause a computer to execute the package detection method of the first aspect or any corresponding embodiment thereof.
[0049] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions, wherein the computer instructions are used to enable a computer to execute the package detection method of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0051] Figure 1 is a schematic flow chart of a package detection method according to an embodiment of the present invention;
[0052] Figure 2 is a flowchart of a package detection method according to a specific example of an embodiment of the present invention;
[0053] Figure 3 is a schematic flow chart of the ByteTrack algorithm according to an embodiment of the present invention;
[0054] Figure 4 is a structural block diagram of a package detection device according to an embodiment of the present invention;
[0055] Figure 5 Schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0056] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.
[0057] Package detection technology, a crucial component of intelligent security, is primarily used in access control systems and logistics management, aiming to automatically identify the arrival and departure status of packages. Traditional solutions rely on three main approaches: video surveillance combined with deep learning models, physical sensors (such as pressure / infrared sensors), and RFID tag recognition. However, all of these solutions suffer from significant technical drawbacks.
[0058] In the field of video surveillance, although deep learning-based solutions (such as YOLO and Faster R-CNN) can achieve target detection, they rely on high-performance computing devices (such as GPUs) for real-time processing, resulting in high hardware costs and difficulty in deployment in embedded systems. In addition, such solutions require complete acquisition of surveillance footage, which poses a risk of user privacy leakage, especially in home or office access control scenarios. Although physical sensor solutions (such as pressure sensors to detect changes in package weight) can reduce computing requirements, they are susceptible to environmental interference (such as wind-blown debris triggering false alarms) and lack sensitivity, making it difficult to distinguish packages from other objects (such as garbage bags or temporarily placed items). RFID technology requires pre-tags on packages, which not only increases deployment costs, but is also limited by the tag's recognition range and cannot detect untagged packages, limiting its actual application scenarios.
[0059] The core contradiction of the above technical solutions lies in the inability to simultaneously optimize computing resource consumption, environmental robustness, and deployment costs. For example, while deep learning solutions offer high detection accuracy, the hardware cost and privacy risks are difficult to balance; while sensor solutions offer low costs, their false alarm rates and adaptability to specific scenarios are insufficient. Existing technologies lack a solution that balances low false alarm rates, low computing power requirements, and non-invasive detection. Therefore, a technical approach based on dynamic image analysis is urgently needed. This approach, which uses lightweight algorithms to accurately determine package status while avoiding the privacy, cost, and environmental adaptability shortcomings of traditional solutions, could provide a cost-effective, practical solution for the smart security sector.
[0060] Therefore, in order to overcome the shortcomings of the above-mentioned existing technical solutions, the embodiment of the present invention provides a package detection method. Figure 1 FIG. 1 is a flow chart of a package detection method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0061] S101 , obtaining successive frames of original images of a monitoring area, and extracting a region of interest from each frame of the original image.
[0062] Specifically, an image acquisition device (such as a camera) captures continuous frame images of the access control monitoring area, generating multiple frames of raw image data in a time series. Each frame corresponds to the monitoring screen at a specific moment, ensuring that the movement of dynamic targets (such as moving packages) is fully recorded. Continuous frame processing provides the temporal data foundation for subsequent motion trajectory analysis. Continuous frame raw images are uncompressed, unprocessed image sequences acquired continuously in chronological order, preserving the original pixel information of the scene.
[0063] A region of interest (ROI) is defined in each image frame. This represents the core monitoring area for package detection (e.g., the 1.5m x 1m package placement area in front of a door). Cropping or masking operations are used to retain only the image data within the ROI, eliminating interference from irrelevant areas (e.g., the scene inside the door or the sky background), thus reducing the computational complexity of subsequent processing. A region of interest (ROI) is a predefined rectangular or polygonal area within an image that represents the physical space to be analyzed. Extraction is the process of segmenting an ROI subimage from a full-frame image, typically achieved through coordinate cropping or a mask matrix.
[0064] S102: performing background removal and contour extraction on the image of the region of interest to obtain a candidate target.
[0065] Specifically, in step S102, background elimination is a technique for removing static or slowly changing backgrounds from an image, retaining only moving targets. Background elimination is a technique for separating the static background and dynamic foreground in an ROI image through dynamic background modeling techniques (such as inter-frame differences and statistical models). Static background refers to fixed scenes (such as the ground and walls), and dynamic foreground refers to moving targets (packages and pedestrians). After eliminating the background, only foreground pixels are retained to highlight the target. Contour extraction is a process of identifying the outer boundaries of a target through edge detection algorithms (such as the Canny operator). Contour extraction is the process of performing edge detection (such as gradient calculation and threshold segmentation) on foreground pixels to extract the outer contour of the target. Noise (such as fallen leaves and flying insects) is filtered out through geometric shape analysis (such as area and aspect ratio), and candidate targets that meet the package characteristics are retained. The candidate target is a set of pixels that are preliminarily determined to be potential packages after background elimination and contour extraction.
[0066] S103 , associating candidate targets in different frames into motion trajectories of the same target through a target tracking algorithm, and obtaining the motion trajectory of a valid target.
[0067] Specifically, in the above step S103, the candidate targets in multiple consecutive frames are associated across frames through the target tracking algorithm, the position changes of the same package in different frames are determined, and its motion trajectory is constructed. The target tracking algorithm is an algorithm for matching targets across frames and recording their motion paths. The core is the spatial position and feature similarity measurement. The motion trajectory is a sequence of positions of the target in consecutive frames, which characterizes its moving direction, speed and stay state. The valid target is a target instance that is determined to be a real package after trajectory association verification. For example, if the candidate target A of the nth frame and the candidate target B of the n+1th frame are highly matched in spatial position and shape features, they are determined to be the same target and their movement paths are recorded. Through trajectory continuity analysis, short-lived interference targets (such as small animals passing through the ROI) are excluded, and only valid targets that persist are retained.
[0068] S104: Determine, based on the movement trajectory of the valid target, whether the valid target leaves or enters the region of interest for the first time.
[0069] Specifically, the first entry refers to the state where the target appears in the ROI for the first time during the monitoring period. Leaving refers to the state where the target completely moves out of the ROI or has no overlap with the ROI. The set threshold is the minimum time or space overlap condition used to trigger the state change (such as 30 seconds, overlapping area <10%). The first entry into the ROI is judged as if the target has not appeared in the ROI in the historical frame, and the current frame appears in the ROI boundary for the first time, it is marked as the first entry. Leaving the ROI is judged as if the target gradually moves out of the ROI boundary in multiple consecutive frames, and the overlapping area with the ROI is lower than the threshold (such as <10%), it is marked as leaving. Time persistence verification is that the target entering for the first time must continue to exist in the ROI for more than the set time threshold (such as 30 seconds) to avoid misjudgment of the scenario where the courier temporarily places the package and then retrieves it immediately.
[0070] S105. When a valid target enters the area of interest for the first time and its continuous presence time exceeds a set threshold, the package is confirmed to have arrived; when a valid target leaves the area of interest, the package is confirmed to have left.
[0071] Specifically, the duration of existence refers to the time span from the first entry to the departure (or continued stay) of the target. Event confirmation is the final judgment of the package status change based on preset logic (time and space conditions). Package arrival confirmation is when the target enters the ROI for the first time and the duration of its existence exceeds the threshold. The system generates a package arrival event and triggers subsequent responses (such as push notifications and video recording). Package departure confirmation is when the target leaves the ROI. The system generates a package departure event, marks the package as taken or removed, and triggers reverse processing (such as data archiving).
[0072] Optionally, in the above step S101, the original image is acquired by an image acquisition device that is installed obliquely on the door body of the access control area, and the installation angle of the image acquisition device is determined according to the door body structure of the access control area.
[0073] Extracting a region of interest includes: extracting an image of the region of interest from each frame of the original image, and determining whether the image of the region of interest in the current frame is valid; if valid, converting the image of the region of interest in the current frame into a grayscale image, and performing image enhancement on the grayscale image using a contrast-limited adaptive histogram equalization algorithm; if invalid, skipping processing the image of the region of interest in the current frame, and determining whether the image of the region of interest in the next frame is valid. The image of the region of interest is determined to be invalid when any invalid condition is met; the invalid condition includes: the boundary of the image of the region of interest exceeds the valid range of the original image, the area of the region of interest is less than a preset threshold, or the pixel information in the region of interest is abnormal.
[0074] Optionally, in the above step S102, background removal and contour extraction are performed on the image of the region of interest to obtain candidate targets, including: background removal of the image of the region of interest through a mixed Gaussian model to obtain a foreground mask; optimizing the foreground mask through opening and closing operations; contour extraction of the optimized foreground mask, and filtering interference targets according to the area threshold to obtain candidate targets.
[0075] Optionally, in the above step S103, the candidate targets in different frames are associated as the motion trajectory of the same target through the target tracking algorithm to obtain the motion trajectory of the valid target, including: calculating the area overlap between the candidate target in the current frame and the historical tracking target through intersection-over-union matching, and associating the targets according to a preset matching threshold; updating the motion trajectory of the target that is continuously matched successfully, and assigning a unique tracking identifier to the same target; generating a new tracking identifier for the candidate target that is not successfully matched.
[0076] Optionally, in the above step S104, based on the motion trajectory of the valid target, it is determined whether the valid target leaves or enters the area of interest for the first time, including: analyzing the motion trajectory of the valid target through a difference analysis algorithm; when the valid target enters the preset boundary range of the area of interest for the first time, confirming that the valid target enters the area of interest for the first time; when the overlap between the valid target and the area of interest is lower than a preset threshold and the continuous target frame does not detect the valid target, determining that the valid target leaves the area of interest; wherein the preset threshold is that the continuous target frame detects that the valid target exists in the area of interest and the position change amplitude of the valid target is less than the preset range.
[0077] Optionally, the above method further includes: blurring a preset privacy area in the original image and retaining the original clarity of the non-privacy area to generate a privacy-preserving image.
[0078] In summary, through the synergistic integration of multi-level technical features, an efficient, stable, and privacy-safe package detection solution has been constructed. Based on continuous frame image analysis, it dynamically extracts key areas within the monitoring area, focusing computing resources on this core area, significantly reducing the data processing burden and ensuring rapid response on embedded devices. By intelligently adjusting the camera's mounting angle, it adaptively expands the monitoring field of view, accurately covering access control areas and avoiding the blind spots caused by traditional solutions due to limited viewing angles. Combining dynamic background modeling with morphological optimization techniques effectively overcomes complex environmental interference such as sudden changes in illumination and rainy and snowy conditions, reliably distinguishing foreground objects from background noise and improving object segmentation accuracy. A multi-stage verification mechanism automatically filters invalid or abnormal image frames, avoiding redundant computation. Grayscale enhancement and contrast optimization enhance the ability to recognize object outlines in low-light and low-contrast scenes. A cross-frame object tracking algorithm establishes an independent motion trajectory for each package, supporting concurrent detection and accurate tracking of multiple targets. It maintains trajectory continuity even in brief occlusions or when objects intersect, significantly reducing false positives and missed detections. The entry and exit zone determination mechanism integrates dual verification of spatial location and temporal continuity to effectively distinguish between real package events and short-term interference, ensuring the reliability of state change triggering. The privacy protection design uses intelligent fuzzy processing of sensitive areas to avoid the leakage of user privacy information while maintaining detection accuracy, meeting strict data security regulations. The modular algorithm architecture is compatible with standardized development interfaces, supports seamless upgrades and functional expansion, and reserves technical space for the future integration of lightweight deep learning models. The overall solution is centered on non-dependent visual detection, breaking through the limitations of traditional sensors in deployment costs, environmental adaptability, and privacy risks, significantly reducing hardware investment and maintenance complexity. It is suitable for diverse scenarios such as community access control and logistics sites, and provides a cost-effective and robust technical practice example for the field of smart security.
[0079] For example, the package detection method provided in the above embodiment will be described in detail below using a specific example.
[0080] This example builds an access control security system based on classic computer vision algorithms, primarily implementing dynamic package monitoring and intelligent early warning capabilities. The system uses cameras to capture real-time surveillance footage outside the door. When a package is detected within the preset warning area, a three-step response mechanism is immediately triggered: a package arrival event message is generated and pushed to the user's mobile device via the cloud backend; an intelligent video capture mechanism is activated, automatically saving 15 seconds of dynamic footage after the target appears; and key video clips are encrypted and uploaded to cloud storage. When the target leaves the monitored area, the system simultaneously executes the reverse process: a package departure notification is sent, video evidence is captured 15 seconds before and after the disappearance, and data is archived. The entire solution utilizes feature extraction and pattern recognition technologies to automate the entire process, ensuring real-time performance while ensuring the integrity of event traceability. Finally, a background blur feature is implemented to protect user privacy.
[0081] The process of this example is as follows Figure 2 As shown, the specific process includes the following.
[0082] First, we check for a valid ROI (Region of Interest). This step ensures that the region of interest is valid and suitable for further processing. If the ROI is invalid, processing of that region is skipped. After extracting a valid ROI, we convert the image to grayscale to reduce color interference and improve processing efficiency. Next, we apply the Contrast Limited Adaptive Histogram Equalization (CLAHE) algorithm to enhance image contrast and detail, making subsequent foreground extraction clearer.
[0083] The expression of the CLAHE algorithm is:
[0084]
[0085] Where L is the maximum grayscale level, hi is the number of pixels at grayscale level i, and N is the total number of pixels. If the histogram value of a block exceeds the preset threshold C, the excess is cropped and evenly distributed across all grayscale levels to ensure that the total number of pixels remains unchanged.
[0086] Background subtraction is performed using a Gaussian mixture model (GMM) to generate a foreground mask, facilitating subsequent moving target extraction. Morphological operations (such as opening and closing) are then performed to optimize the mask, remove noise, and fill potential holes. Next, contours are detected and small objects are filtered out, ultimately extracting valid moving regions. These moving regions are considered potential target areas for further tracking.
[0087] The expression of GMM is:
[0088]
[0089] Where πk is the mixing weight (∑πk=1), μk and Σk are the mean and covariance matrix of the kth Gaussian distribution, respectively.
[0090] The expression of the opening operation is:
[0091]
[0092] The expression of the closing operation is:
[0093]
[0094] in, represents the corrosion operation, Represents the dilation operation, A is the input image, and B is the structural element.
[0095] The BYTETracker tracking algorithm is used to track each moving target and update the target tracking results. BYTETracker can effectively match new targets with existing targets by calculating the overlap (IOU) between targets and maintain tracking stability during motion.
[0096] The process of bytetrack algorithm is as follows Figure 3 As shown, it includes the following processes.
[0097] Target detection and confidence segmentation: The detection boxes of all targets in the current frame are obtained through a detector (such as YOLOX), and are divided into high-confidence targets (completely visible packages) and low-confidence targets (packages that may be occluded or partially visible) according to the confidence threshold.
[0098] Initial matching and trajectory prediction: Use a motion model (such as Kalman filter) to predict the current frame position based on the historical trajectory, prioritize the intersection-over-union (IOU) matching of high-confidence detection boxes with the predicted position, and update the trajectory of the tracked target.
[0099] Low-resolution frame secondary matching: The unmatched low-confidence detection frame is matched with the remaining track twice to restore the target that is temporarily occluded (such as the package is blocked by the courier's body for 0.5 seconds) to avoid misjudgment of departure.
[0100] Track lifecycle management: continuously unmatched tracks are deleted after a certain timeout period, and new tracks are generated for unmatched high-confidence detection boxes, ultimately outputting the motion tracks of all active targets.
[0101] Check the validity of existing targets to confirm their validity. Use adaptive thresholding and morphological processing to further verify the target's existence. Also, use difference analysis to determine if the target has disappeared. If so, update the template or mark the target as deleted based on the difference analysis results.
[0102] The newly detected target is matched with the existing target through the IOU (Intersection over Union) matching strategy. If the IOU value is higher than the set threshold, the existing target is updated; if it is lower than the threshold, it is considered a new target and is created and tracked.
[0103] While tracking the target, overlapping boxes are merged and non-maximum suppression (NMS) is used to further filter out the most reliable target boxes and remove redundant boxes with excessive overlap. Finally, the local coordinates of the ROI are converted to global image coordinates to ensure that the final result is consistent with the original image.
[0104] NMS is used to suppress redundant detection frames. The core steps include:
[0105] Sort the detection boxes by confidence and select the box with the highest score Bmax.
[0106] Calculate the intersection over union (IoU) of other boxes with Bmax;
[0107] IoU=Area(Bi∪Bmax) / Area(Bi∩Bmax);
[0108] If IoU>threshold, Bi is removed.
[0109] For background blurring, users specify a rectangular area to ensure that it remains clear during the blurring process. By blurring the image horizontally and vertically, and then overlaying the original content of the specified area onto the blurred image, the desired visual effect is achieved while protecting user privacy.
[0110] Through these steps, the entire target detection and tracking process can work stably and effectively, adapting to the common target detection needs in application scenarios such as security doors.
[0111] This embodiment also provides a package detection device for implementing the aforementioned embodiments and preferred implementations. Details already described will not be repeated. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. While the devices described in the following embodiments are preferably implemented using software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0112] This embodiment provides a package detection device, such as Figure 4 Shown, including:
[0113] The region extraction module 401 is used to obtain continuous frames of original images of the monitoring area and extract the region of interest from each frame of the original image;
[0114] The candidate target acquisition module 402 is used to perform background removal and contour extraction on the image of the region of interest to obtain candidate targets;
[0115] The tracking association module 403 is used to associate candidate targets in different frames into motion trajectories of the same target through a target tracking algorithm to obtain the motion trajectory of a valid target;
[0116] A determination module 404 is configured to determine whether a valid target enters or leaves a region of interest based on the motion trajectory of the valid target;
[0117] The confirmation module 405 is used to confirm the arrival of a package when a valid target enters the area of interest for the first time and its continuous presence time exceeds a set threshold; and to confirm the departure of a package when a valid target leaves the area of interest.
[0118] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0119] The package inspection device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0120] The embodiment of the present invention also provides a computer device having the above Figure 4 The package detection device shown.
[0121] See also Figure 5 , Figure 5 is a structural diagram of a computer device provided by an optional embodiment of the present invention, such as Figure 5As shown, the computer device includes: one or more processors 10, memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in the memory or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 5 A processor 10 is taken as an example.
[0122] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.
[0123] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0124] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0125] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0126] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0127] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.
[0128] A portion of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium that can be accessed by the computer.
[0129] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A package detection method, characterized in that: The method comprises: Obtain continuous frame original images of the monitoring area and extract the region of interest from each frame original image; Performing background removal and contour extraction on the image of the region of interest to obtain a candidate target; The target tracking algorithm is used to associate candidate targets in different frames into the motion trajectory of the same target, and the motion trajectory of the valid target is obtained; According to the motion trajectory of the valid target, determining whether the valid target leaves or enters the region of interest for the first time; When the valid target enters the area of interest for the first time and its presence time exceeds a set threshold, the package is confirmed to have arrived; when the valid target leaves the area of interest, the package is confirmed to have left.
2. The method according to claim 1, characterized in that The original image is acquired by an image acquisition device that is installed obliquely on a door body of an access control area. The installation angle of the image acquisition device is determined according to the door body structure of the access control area.
3. The method according to claim 2, characterized in that Extract regions of interest, including: Extracting an image of a region of interest from each frame of the original image, and determining whether the image of the region of interest in the current frame is valid; If valid, the current frame region of interest image is converted into a grayscale image, and the grayscale image is enhanced by a limited contrast adaptive histogram equalization algorithm; If not valid, the processing of the image of the region of interest in the current frame is skipped, and it is determined whether the image of the region of interest in the next frame is valid.
4. The method according to claim 3, characterized in that The determining whether the region of interest image is valid includes: When the image of the region of interest meets any invalidity condition, confirming that the image of the region of interest is invalid; The invalid conditions include: the boundary of the image of the region of interest exceeds the valid range of the original image, the area of the region of interest is smaller than a preset threshold, or the pixel information in the region of interest is abnormal.
5. The method according to claim 3, characterized in that The performing background removal and contour extraction on the image of the region of interest to obtain a candidate target includes: Performing background elimination on the image of the region of interest using a mixed Gaussian model to obtain a foreground mask; Optimizing the foreground mask by opening and closing operations; The optimized foreground mask is contour extracted, and interference targets are filtered according to the area threshold to obtain candidate targets.
6. The method according to claim 5, characterized in that The target tracking algorithm is used to associate candidate targets in different frames with the motion trajectory of the same target to obtain the motion trajectory of the valid target, including: The regional overlap between the candidate target in the current frame and the historical tracking target is calculated by intersection-over-union matching, and the target association is performed according to the preset matching threshold; Update the motion trajectory of the target that is continuously matched successfully and assign a unique tracking identifier to the same target; Generate new tracking identifiers for unsuccessful matching candidates.
7. The method according to claim 6, characterized in that The determining, based on the motion trajectory of the valid target, whether the valid target leaves or enters the region of interest for the first time includes: Analyze the motion trajectory of valid targets through difference analysis algorithm; When a valid target enters a preset boundary range of the region of interest for the first time, confirming that the valid target enters the region of interest for the first time; When the overlap between the valid target and the region of interest is lower than a preset threshold and the valid target is not detected in consecutive target frames, determining that the valid target leaves the region of interest; The preset threshold is that the valid target is detected in the region of interest by consecutive target frames, and the position change amplitude of the valid target is less than a preset range.
8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: The privacy area preset in the original image is blurred, and the original clarity of the non-privacy area is retained to generate a privacy-preserving image.
9. A package detection device, characterized in that: The device comprises: The region extraction module is used to obtain the continuous frame original image of the monitoring area and extract the region of interest from each frame original image; A candidate target acquisition module is used to perform background removal and contour extraction on the image of the region of interest to obtain a candidate target; The tracking association module is used to associate candidate targets in different frames into the motion trajectory of the same target through the target tracking algorithm to obtain the motion trajectory of the valid target; A judgment module, configured to judge whether the valid target enters or leaves the region of interest according to the motion trajectory of the valid target; The confirmation module is used to confirm the arrival of a package when the valid target enters the area of interest for the first time and its continuous presence time exceeds a set threshold; and to confirm the departure of a package when the valid target leaves the area of interest.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the package detection method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Abnormal intrusion detection method based on motion template
CN103577833A
Suspicious target detection tracking and recognition method based on dual-camera cooperation
CN104301669A
Parcel detection method and device, equipment and storage medium
CN113468918A
Parcel detection method and device, computer readable medium and electronic equipment
CN116416191A
Parcel tracking method based on security check image
CN118196382A
Cited By
Real-time detection and early warning method and device for wall surface of safety channel and medium
CN120786038A
Video-based method and device for tracking contraband in parcel, equipment and medium
CN122244105A