Article monitoring method and device, electronic equipment and storage medium

By monitoring images to detect the handling and placement of items by target personnel, identifying and recording behaviors, the problem of data inconsistency in the sales management of specific items is solved, and management efficiency is improved.

CN115497038BActive Publication Date: 2026-04-17QINGDAO INTELLIFUSION TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QINGDAO INTELLIFUSION TECH CO LTD
Filing Date
2022-08-30
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In the sales management of specific items, the reliance on manual operation for inbound and outbound management leads to inconsistencies between sales data and actual inventory, resulting in problems such as taking too much, taking too little, misplacing, and missing records, which are difficult to trace and lead to low management efficiency.

Method used

By monitoring the target area through surveillance images, the system detects the handling and placement of target personnel and items, identifies and monitors behaviors, and records behavioral information to trace the actual movement of items.

Benefits of technology

It improves the efficiency of sales management for specific items, ensures that sales data is consistent with actual inventory, and facilitates the tracking of item transactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115497038B_ABST
    Figure CN115497038B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides an article monitoring method, the method comprises the following steps: obtaining a monitoring image of a target area, the monitoring image comprising a target person and a target article; detecting the state of the target person and the target article based on the monitoring image to obtain taking and placing state information between the target person and the target article; identifying the behavior of the target person based on the taking and placing state information to obtain a behavior identification result of the target person; and monitoring the target article through the behavior identification result of the target person. The state of the target person and the target article is detected through the monitoring image of the target area to obtain the taking and placing state information between the target person and the target article, the behavior of the target person is identified according to the taking and placing state information between the target person and the target article, the target article can be monitored through the behavior of the target person, the behavior of the target person on the target article is facilitated to be tracked, and therefore the sales management efficiency of specific articles is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an object monitoring method, device, electronic device, and storage medium. Background Technology

[0002] In the sales management of specific items, these items are generally allocated and sold by authorized and qualified personnel. For example, prescription drugs in pharmacies can only be dispensed, purchased, and used with a prescription from a licensed physician or licensed assistant physician. The sale of specific items typically requires unified inventory management. However, because inventory management relies on manual handling, personnel may take too much, take too little, misplace, or omit information, leading to discrepancies between sales data and actual inventory. This makes subsequent investigation difficult and results in low efficiency in the sales management of specific items. Summary of the Invention

[0003] This invention provides a method for monitoring items. By monitoring images of a target area, the method detects the status of target personnel and target items, obtains information on the handling and placement status between the target personnel and the target items, and identifies the behavior of the target personnel towards the target items based on this information. This allows for the monitoring of target items by tracking the behavior of the target personnel towards the target items, thereby improving the sales management efficiency of specific items.

[0004] In a first aspect, embodiments of the present invention provide a method for monitoring items, the method comprising:

[0005] Acquire surveillance images of the target area, the surveillance images including target personnel and target items;

[0006] Based on the surveillance image, the status of the target person and the target item is detected to obtain the handling and placement status information between the target person and the target item.

[0007] Based on the picking and placing status information, the behavior of the target person is identified to obtain the behavior identification result of the target person;

[0008] The target items are monitored based on the behavioral recognition results of the target personnel.

[0009] Optionally, the monitoring image includes consecutive frame images, and the step of performing state detection on the target person and the target item based on the monitoring image to obtain the handling and placement state information between the target person and the target item includes:

[0010] For each frame of the image, state detection is performed on the target person and the target item to obtain the handling and placement attributes between the target person and the target item in each frame of the image.

[0011] The picking and placing attributes are sorted according to the time sequence of the corresponding frame images to obtain the picking and placing status information between the target person and the target item.

[0012] Optionally, the step of performing state detection on the target person and the target item in each frame of the image to obtain the handling and placement attributes between the target person and the target item in each frame of the image includes:

[0013] Target detection is performed on each frame of the image to obtain the hand detection result of the target person;

[0014] The hand detection results are used to detect the attributes of the target item to obtain the handling attributes between the target person and the target item in each frame of the image.

[0015] Optionally, the step of performing target detection on each frame of the image to obtain the hand detection result of the target person includes:

[0016] Perform a first target detection on each frame of the image to obtain a human body detection box;

[0017] Perform a second target detection on each frame of the image to obtain a hand detection bounding box;

[0018] The human body detection box and the hand detection box in each frame image are associated to obtain the association relationship between the human body detection box and the hand detection box in each frame image;

[0019] Based on the correlation between the human body detection box and the hand detection box in each frame of the image, and whether the hand detection box in each frame of the image includes the target item, the hand detection result of the target person is obtained.

[0020] Optionally, associating the human detection box and the hand detection box in each frame of the image to obtain the association relationship between the human detection box and the hand detection box in each frame of the image includes:

[0021] Calculate the detection box distance between the human body detection box and the hand detection box in each frame image;

[0022] Based on the detection box distance, the association between the human body detection box and the hand detection box in each frame image is obtained.

[0023] Optionally, sorting the pick-up and place attributes according to the time sequence of the corresponding frame images to obtain the pick-up and place status information between the target person and the target item includes:

[0024] Assign a first identifier to human body detection boxes belonging to the same target person in each frame image, and assign a second identifier to hand detection boxes belonging to the same target person in adjacent frame images;

[0025] Based on the first identifier and the second identifier, the picking and placing attributes corresponding to each target person are sorted according to the time sequence of the corresponding frame images to obtain the picking and placing status information between each target person and the corresponding target item.

[0026] Optionally, the step of performing behavior recognition on the target person based on the picking and placing status information to obtain the behavior recognition result of the target person includes:

[0027] The pick-up and put-down status information is filtered to obtain filtered pick-up and put-down status information;

[0028] Based on the filtered pick-up and put-down status information, the target person's behavior is identified to obtain the target person's behavior identification result.

[0029] Secondly, embodiments of the present invention provide an item monitoring device, the device comprising:

[0030] The acquisition module is used to acquire surveillance images of the target area, the surveillance images including target personnel and target items;

[0031] The detection module is used to perform state detection on the target person and the target item based on the monitoring image, and obtain the handling and placement state information between the target person and the target item;

[0032] The identification module is used to identify the behavior of the target person based on the picking and placing status information, and obtain the behavior identification result of the target person;

[0033] The monitoring module is used to monitor the target item based on the behavior recognition results of the target person.

[0034] Thirdly, embodiments of the present invention provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps in the item monitoring method provided in embodiments of the present invention.

[0035] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the article monitoring method provided in the embodiments of the present invention.

[0036] In this embodiment of the invention, a monitoring image of a target area is acquired, the monitoring image including target personnel and target items; based on the monitoring image, state detection is performed on the target personnel and the target items to obtain handling and placement status information between the target personnel and the target items; based on the handling and placement status information, behavior recognition is performed on the target personnel to obtain the behavior recognition result; and the target items are monitored based on the behavior recognition result. By performing state detection on the target personnel and target items through the monitoring image of the target area to obtain handling and placement status information between the target personnel and the target items, and by recognizing the target personnel's behavior towards the target items based on the handling and placement status information, the target items can be monitored through the target personnel's behavior towards the target items, facilitating the tracking of the target personnel's behavior towards the target items, thereby improving the sales management efficiency of specific items. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 This is a flowchart of an article monitoring method provided in an embodiment of the present invention;

[0039] Figure 2 This is a schematic diagram of the structure of an item monitoring device provided in an embodiment of the present invention;

[0040] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0042] Please see Figure 1 , Figure 1 This is a flowchart of an article monitoring method provided by an embodiment of the present invention, such as... Figure 1 As shown, the method for monitoring this item includes the following steps:

[0043] 101. Obtain surveillance images of the target area.

[0044] In this embodiment of the invention, the target area can be an area for storing target items. For example, if the target item is a prescription drug, the target area is a shelf or counter for storing prescription drugs. Of course, the target item can also be other specific items subject to sales restrictions, such as antipyretics for special periods or food containing specific ingredients.

[0045] An image monitoring device can be installed above the target area to monitor it and obtain a monitoring image of the target area, which includes the target personnel and target items. The image monitoring device can be activated when personnel are detected entering the target area to monitor it and obtain a monitoring image including the target personnel and target items.

[0046] The aforementioned target personnel can be sales or management personnel of the target items, possessing the authority to retrieve or store the target items. There can be one or more target personnel. During the monitoring of the target area by the image surveillance equipment, when at least one person is detected in the monitoring screen, facial recognition is performed on that person to obtain their identity information. Based on this identity information, it is determined whether the person has the authority to retrieve or store the target items. If the person does not have the authority, an alarm is triggered, and the person is prompted to leave the target area. If the person has the authority, they are identified as a target personnel, and continuous tracking and monitoring are conducted until they leave the target area, resulting in a monitoring image containing both the target personnel and the target items.

[0047] 102. Based on the surveillance images, perform status detection on the target personnel and target items to obtain the handling and placement status information between the target personnel and the target items.

[0048] In this embodiment of the invention, the aforementioned state detection is used to detect the state between the target person and the target item. The aforementioned pick-up and put-down state information may include the target person being in a state of picking up the target item, and the target person being in a state of not picking up the target item.

[0049] Specifically, whether a target person is in the state of picking up a target item can be determined by the positional relationship between the target person's hand and the target item, or by the trajectory of the target person's hand and the trajectory of the target item. For example, if the target person's hand and the target item are in the same position, it can be determined that the target person is in the state of picking up the target item. Or, if the trajectory of the target person's hand and the trajectory of the target item are the same, it can be determined that the target person is in the state of picking up the target item.

[0050] 103. Based on the picking and putting status information, perform behavior recognition on the target personnel to obtain the behavior recognition results of the target personnel.

[0051] In this embodiment of the invention, the aforementioned behavior recognition result includes the behavior of accessing the target item and the behavior of not accessing the target item. The aforementioned pick-up and drop-off status information can indicate whether the target person has picked up the target item. When the target person is in the state of picking up the target item, the target person's behavior can be identified as the behavior of accessing the target item. When the target person is in the state of not picking up the item, the target person's behavior can be identified as the behavior of not accessing the target item.

[0052] 104. Monitor target items based on the behavioral identification results of target personnel.

[0053] In this embodiment of the invention, the aforementioned behavior recognition result includes access behavior and non-access behavior. Access behavior refers to the act of accessing or storing the target item, while non-access behavior refers to the act of not accessing or storing the target item. When the behavior recognition result is access behavior, the target person's behavior is recorded; when the behavior recognition result is non-access behavior, the target person's behavior is not recorded. Specifically, the aforementioned access behavior includes the act of storing the target item and the act of taking the target item, while the aforementioned non-access behavior indicates the act of not storing or taking the target item.

[0054] By monitoring the behavior records of target personnel and the target items, discrepancies between sales records and actual inventory can be traced back by retrieving the behavior records of each target personnel. Specifically, the aforementioned behavior records include the time of the behavior and the quantity of target items accessed. The time of the behavior can be the monitoring time of the surveillance image, and the quantity of target items accessed can be obtained by quantity detection from the surveillance image.

[0055] In this embodiment of the invention, a monitoring image of a target area is acquired, the monitoring image including target personnel and target items; based on the monitoring image, state detection is performed on the target personnel and the target items to obtain handling and placement status information between the target personnel and the target items; based on the handling and placement status information, behavior recognition is performed on the target personnel to obtain the behavior recognition result; and the target items are monitored based on the behavior recognition result. By performing state detection on the target personnel and target items through the monitoring image of the target area to obtain handling and placement status information between the target personnel and the target items, and by recognizing the target personnel's behavior towards the target items based on the handling and placement status information, the target items can be monitored through the target personnel's behavior towards the target items, facilitating the tracking of the target personnel's behavior towards the target items, thereby improving the sales management efficiency of specific items.

[0056] Optionally, in the step of detecting the status of the target personnel and target items based on the monitoring images to obtain the handling and placement status information between the target personnel and target items, where the monitoring images include continuous frame images, the status of the target personnel and target items in each frame image can be detected to obtain the handling and placement attributes between the target personnel and target items in each frame image; the handling and placement attributes are sorted according to the time order of the corresponding frame images to obtain the handling and placement status information between the target personnel and target items.

[0057] In this embodiment of the invention, consecutive frame images refer to frame images that are consecutive in time, and each frame image in the aforementioned consecutive frame images includes both the target person and the target object. The aforementioned consecutive frame images can also be referred to as a video stream.

[0058] Specifically, for each frame of the image, the handling and placement attributes between the target person and the target item are detected. These handling and placement attributes can include the picking attribute and the not picking attribute. The picking attribute refers to the target person picking up the target item, and the not picking attribute refers to the target person not picking up the target item.

[0059] After obtaining the handling and placement attributes between the target person and the target item in each frame, the handling and placement attributes can be sorted according to the time sequence of the corresponding frame images to obtain the handling and placement attribute sequence as the aforementioned handling and placement state information. The aforementioned handling and placement state attribute sequence can be represented as an = [a1, a2, ..., ai, ..., an], where ai represents the handling and placement attributes between the target person and the target item in the i-th frame image, and n represents the total number of frame images.

[0060] The behavior of the target person can be identified based on the aforementioned pick-up and drop-down attribute sequence 'an', thus obtaining the behavior identification result. Specifically, the behavior of the target person can be identified based on the number of picked-up and unpicked attributes in the pick-up and drop-down attribute sequence 'an'. When the number of picked-up attributes in the pick-up and drop-down attribute sequence exceeds a preset number, the target person's behavior can be identified as a pick-up / drop-down behavior; when the number of picked-up attributes in the pick-up and drop-down attribute sequence does not exceed the preset number, the target person's behavior can be identified as a non-pick-up / drop-down behavior.

[0061] Optionally, in the step of performing state detection on the target person and target item in each frame of the image to obtain the handling attributes between the target person and target item in each frame of the image, target detection can be performed on each frame of the image to obtain the hand detection result of the target person; the attribute detection of the target item can be performed on the hand detection result to obtain the handling attributes between the target person and target item in each frame of the image.

[0062] In this embodiment of the invention, for each frame of a surveillance image, target detection can be performed on the hands of the target person. Specifically, a human target detector can be used to detect the target person in each frame of the image, and the detection result of the target person, including the position of the target person, can be determined. Then, a hand target detector can be used to detect the hands in each frame of the image, and the detection result of the hands, including the position of the hands, can be obtained. Matching the position of the target person with the position of the hands, the hand detection result of the target person is obtained.

[0063] In one possible embodiment, a first hand target detector can be used to directly detect whether a hand is holding a target item, and match the hand holding the target item or the hand not holding the target item with the target person to determine the holding and placing attributes between the target person and the target item.

[0064] Taking prescription drugs as an example, human object detector A detects each frame of the image and outputs a bounding box for the category of human in each frame. A first hand object detector B detects each frame of the image and outputs bounding boxes for the categories of "hand holding prescription drugs" and "hand not holding prescription drugs" in each frame. For each frame, when human object detector A detects a human bounding box R... A The first hand target detector B detected a hand detection box R corresponding to the "hand holding prescription drugs". B At that time, output the human detection box R. A The hand detection bounding box R corresponding to "the hand holding the prescription drug" B And the "grabbing attribute" of holding prescription drugs; when the human target detector A detects the human detection box R AThe first hand target detector B detected the hand detection box R corresponding to "the hand that is not holding the prescription drug". B At that time, output the human detection box R. A The hand detection box R corresponding to "hand without prescription medication" B The "unclaimed" attribute indicates that the person is not holding the prescription medication; when the human target detector A detects the human bounding box R... A When the first hand target detector B does not detect any results, the human body detection box RA, the hand detection box is unknown, and the "handling and placing attribute unknown" of the hand holding the prescription drug is output; otherwise, no output is output for this frame.

[0065] In another possible embodiment, the detection results of the hands in each frame of the image can be detected by a second hand target detector, and the hand detection results can be matched with the corresponding target person to determine the target person's hand detection results. The target item classifier can be used to classify the target person's hand detection results to determine whether there is a target item in the target person's hand detection results, and then determine the handling attribute between the target person and the target item.

[0066] Taking prescription drugs as an example, a human object detector A detects each frame of the image and outputs a bounding box representing the target category in each frame. A second hand object detector C detects each frame of the video image and outputs a bounding box representing the target category as a hand. A target item classifier D classifies the hand bounding boxes detected by the second hand object detector C and outputs one of two categories: "Prescription drug available" or "No prescription drug available." For each frame, when human object detector A detects a human bounding box R... A The second hand target detector C detected the hand detection box R. C And use the target item classifier D to classify R C When the classification result is "prescription drugs available", output the human detection box R. A Hand detection frame R C And the "grabbing attribute" of holding prescription drugs; when the human target detector A detects the human bounding box R A The second hand target detector C detected the hand detection box R. C And use the target item classifier D to classify R C When the classification result is "no prescription drugs available", output the human detection box R. A Hand detection frame R C The "unclaimed attribute" indicates that the person is not holding the prescription medication; when the human target detector A detects the human bounding box R... A When the second hand target detector C does not detect a result, the human body detection box R is output. AIf the hand detection bounding box is unknown and the "handling / placing attribute" of the hand holding prescription drugs is unknown, the frame will not be output in other cases.

[0067] Optionally, in the step of performing target detection on each frame of the image to obtain the hand detection result of the target person, a first target detection can be performed on each frame of the image to obtain a human body detection box; a second target detection can be performed on each frame of the image to obtain a hand detection box; the human body detection box and the hand detection box in each frame of the image can be associated to obtain the association relationship between the human body detection box and the hand detection box in each frame of the image; and the hand detection result of the target person can be obtained based on the association relationship between the human body detection box and the hand detection box in each frame of the image and whether the hand detection box in each frame of the image includes the target object.

[0068] In this embodiment of the invention, a human target detector can be used to perform a first target detection on each frame of the image to obtain the human detection box of each target person in each frame of the image; a first hand target detector can be used to perform a second detection on each frame of the image to obtain the hand detection box of the hand holding the target item or the hand detection box of the hand not holding the target item in each frame of the image; in each frame of the image, the hand detection box of the hand holding the target item or the hand detection box of the hand not holding the target item is matched with the human detection box of the target person; the successfully matched human detection box is associated with the hand detection box of the hand holding the target item or the hand detection box of the hand not holding the target item, and the corresponding association relationship is recorded; the identity information of the target person is identified based on the human detection box, thereby obtaining the hand detection result of each target person in each frame of the image.

[0069] Taking prescription drugs as an example, human object detector A detects each frame of the image and outputs a bounding box for the category of human in each frame. A first hand object detector B detects each frame of the image and outputs bounding boxes for the categories of "hand holding prescription drugs" and "hand not holding prescription drugs" in each frame. For each frame, when human object detector A detects a human bounding box R... A The first hand target detector B detected a hand detection box R corresponding to the "hand holding prescription drugs". B At that time, output the human detection box R. A The hand detection bounding box R corresponding to "the hand holding the prescription drug" B And the "take-up attribute" of holding prescription drugs, and the human detection box R A The hand detection bounding box R corresponding to "the hand holding the prescription drug" B Perform matching, and select the successfully matched human detection bounding box R. A The hand detection box R corresponding to "the hand holding the prescription drug" B Perform associations and record the corresponding association relationships based on the human detection bounding box R.A The identification information of the target personnel is obtained, thereby obtaining the hand detection results for each target personnel; when the human target detector A detects the human detection box R... A The first hand target detector B detected the hand detection box R corresponding to "the hand that is not holding the prescription drug". B At that time, output the human detection box R. A The hand detection box R corresponding to "hand without prescription medication" B The "unclaimed" attribute indicates that the person did not have the prescription medication, and the human detection frame is set to R. A The hand detection bounding box R corresponding to "hand without prescription medication" B Perform matching, and match the successfully matched human body detection box with the hand detection box R corresponding to "hand without prescription drugs". B The system performs association and records the corresponding relationships. Based on the human detection bounding box, it identifies the target person's identity information, thus obtaining the hand detection results for each target person. When human target detector A detects human detection bounding box R... A When the first hand target detector B does not detect any results, it outputs the human body detection box RA, the hand detection box is unknown, and the "handling and placing attribute unknown" of the hand holding the prescription drug, and does not perform matching; otherwise, this frame is not output.

[0070] In one possible embodiment, a human target detector can be used to perform a first target detection on each frame of the image to obtain the human detection box of each target person in each frame of the image; a second hand target detector can be used to detect the hand detection box in each frame of the image. In each frame of the image, the hand detection box is matched with the human detection box of the target person, the successfully matched human detection box is associated with the hand detection box, and the corresponding association relationship is recorded. The identity information of the target person is identified based on the human detection box, thereby obtaining the hand detection result of each target person.

[0071] Taking prescription drugs as an example, a human object detector A detects each frame of the image and outputs a bounding box representing the target category in each frame. A second hand object detector C detects each frame of the video image and outputs a bounding box representing the target category as a hand. A target item classifier D classifies the hand bounding boxes detected by the second hand object detector C and outputs one of two categories: "Prescription drug available" or "No prescription drug available." For each frame, when human object detector A detects a human bounding box R... A The second hand target detector C detected the hand detection box R. C and the human detection box R A and hand detection frame R C Perform matching, and select the successfully matched human detection bounding box R. A With hand detection frame RC Perform associations and record the corresponding association relationships based on the human detection bounding box R. A The identification information of the target personnel is obtained, thereby obtaining the hand detection results of each target personnel; when the target item classifier D is used to classify R... C When the classification result is "prescription drugs available", output the human detection box R. A Hand detection frame R C And the "take-up attribute" of holding prescription drugs; when using the target item classifier D to R C When the classification result is "no prescription drugs available", output the human detection box R. A Hand detection frame R C The "unclaimed attribute" indicates that the person is not holding the prescription medication; when the human target detector A detects the human bounding box R... A When the second hand target detector C does not detect a result, the human body detection box R is output. A If the hand detection bounding box is unknown and the "handling / placing attribute" of the hand holding prescription drugs is unknown, the frame will not be output in other cases.

[0072] Optionally, in the step of associating the human detection box and the hand detection box in each frame of an image to obtain the association relationship between the human detection box and the hand detection box in each frame of an image, the detection box distance between the human detection box and the hand detection box in each frame of an image can be calculated; based on the detection box distance, the human detection box and the hand detection box can be associated to obtain the association relationship between the human detection box and the hand detection box in each frame of an image.

[0073] In this embodiment of the invention, the human detection box can be represented as (x1, y1, w1, h1, ε1), where (x1, y1) represents the center point position of the human detection box, and the hand detection box can be represented as (x2, y2, w2, h2, ε2), where (x2, y2) represents the center point position of the hand detection box. The aforementioned hand detection boxes can be the hand detection boxes output by the first hand target detector or the second hand target detector. In each frame of the image, the distance between (x1, y1) and (x2, y2) can be calculated as the detection box distance. Human detection boxes and hand detection boxes with a detection box distance less than a preset value are associated, and the association relationship between the human detection boxes and hand detection boxes in each frame of the image is recorded.

[0074] Optionally, in the step of sorting the pick-up and put-down attributes according to the time sequence of the corresponding frame images to obtain the pick-up and put-down status information between the target person and the target item, a first identifier can be assigned to the human body detection boxes belonging to the same target person in each frame image, and a second identifier can be assigned to the hand detection boxes belonging to the same target person in adjacent frame images; based on the first identifier and the second identifier, the pick-up and put-down attributes corresponding to each target person are sorted according to the time sequence of the corresponding frame images to obtain the pick-up and put-down status information between each target person and the corresponding target item.

[0075] In this embodiment of the invention, target tracking can be performed on each frame of the image to obtain the tracking results between each target person and the corresponding target item. Based on the tracking results and the handling attributes between each target person and the corresponding target item, the handling status information between each target person and the corresponding target item can be obtained. The above target tracking can be based on target tracking algorithms such as LK optical flow and KCF.

[0076] Specifically, after outputting human bounding boxes in each frame of an image through a human target detector and hand bounding boxes in each frame of an image through a first hand target detector or a second hand target detector, a first identifier is assigned to the human bounding box and a second identifier is assigned to the hand bounding box within the same frame. The first identifier can track an ID, and the second identifier can also track an ID. In different frames of images, the similarity between human bounding boxes in different frames is calculated. Human bounding boxes with a similarity higher than a preset value are identified as belonging to the same target person and assigned the same first identifier P. k,i Where i is the i-th frame image, k is the k-th human body detection box in the i-th frame image, and the human body detection box tracking sequence P for the k-th target person is obtained. k Similarly, in different frame images, the similarity between hand detection boxes in different frame images is calculated, and hand detection boxes with a similarity higher than a preset value are identified as the same target person's hand detection boxes and assigned the same second identifier H. k,i Where i is the i-th frame image, k is the k-th hand detection box in the i-th frame image, and the hand detection box tracking sequence H of the k-th target person is obtained. k Then, by using the correlation between the human detection bounding box and the hand detection bounding box in the same frame image, the tracking sequence P of the human detection bounding box is... k With hand detection box tracking sequence H j Perform association and based on the tracking sequence P k or H k Sort the pick-up and put-down attributes of each target person to obtain the pick-up and put-down status information an between each target person and the corresponding target item.

[0077] In one possible embodiment, the human body detection box and the hand detection box can also be matched according to the Hungarian matching algorithm to obtain the human body detection box tracking sequence P of the k-th target person. k The hand detection box tracking sequence H of the corresponding k-th target person k And according to the tracking sequence P k or H k Sort the pick-up and put-down attributes of each target person to obtain the pick-up and put-down status information an between each target person and the corresponding target item.

[0078] Optionally, in the step of performing behavior recognition on the target person based on the state information to obtain the behavior recognition result of the target person, the picking and placing state information can be filtered to obtain filtered picking and placing state information; and the target person can be identified based on the filtered picking and placing state information to obtain the behavior recognition result of the target person.

[0079] In this embodiment of the invention, the picking and placing status information *an* of each target person can be filtered to obtain filtered picking and placing status information *bn*. This filters out noise caused by false detections, thereby improving the accuracy of behavior recognition. The filtered picking and placing status information *bn* can be input into the behavior recognition model, and the behavior recognition model outputs the behavior recognition result of the corresponding target person.

[0080] Specifically, the aforementioned behavior recognition model can be a time-series model based on recurrent neural networks or long short-term memory networks, and this model is pre-trained. A dataset is obtained by collecting sample state information `cn` and corresponding behavior labels. This dataset is then used to train the behavior recognition model, resulting in a pre-trained model. The sample state information `cn` is obtained using the same method as the pick-up / placement state information `an`.

[0081] Taking prescription drug sales management as an example, the target item is prescription drugs, the target personnel are sales staff or managers, and the handling status information is a sequence of human hands handling prescription drugs. This sequence includes three handling attributes: "hands picked up" corresponds to taking the medication, "hands not picked up" corresponds to not taking the medication, and "hands not picked up" corresponds to "unknown." After the person with the tracking ID leaves the video frame, the analysis can filter the sequence of human hands handling prescription drugs to remove noise caused by false detections during target detection. Analyzing the filtered sequence reveals the behavior of the tracking ID in handling prescription drugs during the monitoring image. This is illustrated in Table 1 below.

[0082] Table 1

[0083]

[0084] In Table 1, during frames 12-16, the human target's hand was seen taking medication in the monitoring frame; during frames 17-25, the hand left the monitoring frame without taking medication. Therefore, it can be determined that this tracking ID performed a suspected medication replenishment operation, suggesting that the corresponding sales or management personnel may have placed prescription drugs in the prescription drug cabinet.

[0085] As shown in Table 2 below:

[0086] Table 2

[0087]

[0088] In Table 2, during frame 50-57, the human target's hand appeared on the monitoring screen without holding the medication; during frame 58-63, the hand left the monitoring screen with the medication. Therefore, it can be determined that the tracking ID performed a suspected medication retrieval operation, suggesting that the corresponding sales or management personnel removed the prescription medication from the prescription drug cabinet.

[0089] In this embodiment of the invention, the time when the hand leaves the monitoring screen after taking the medicine is determined to be the time when the corresponding sales or management personnel's behavior occurs, and the behavior of the sales or management personnel and the time of occurrence of the behavior are recorded. When there is a difference between the inventory and the sales record, the corresponding behavior is found based on the sales time and the time of occurrence of the behavior, thereby improving the sales management efficiency of prescription drugs.

[0090] It should be noted that the item monitoring method provided in this embodiment of the invention can be applied to devices such as smart cameras, smartphones, computers, and servers that can monitor items.

[0091] Optional, please see Figure 2 , Figure 2 This is a schematic diagram of the structure of an item monitoring device provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the device includes:

[0092] The acquisition module 201 is used to acquire surveillance images of the target area, wherein the surveillance images include target personnel and target items;

[0093] The detection module 202 is used to perform state detection on the target person and the target item based on the monitoring image, and obtain the handling and placement state information between the target person and the target item;

[0094] The identification module 203 is used to identify the behavior of the target person based on the picking and placing status information, and obtain the behavior identification result of the target person;

[0095] The monitoring module 204 is used to monitor the target item based on the behavior recognition results of the target person.

[0096] Optionally, the monitored image includes consecutive frame images, and the detection module 202 includes:

[0097] The detection submodule is used to perform state detection on the target personnel and the target items in the monitoring image to obtain the handling and placement attributes between the target personnel and the target items in each frame of the image;

[0098] The sorting submodule is used to sort the pick-up and put-down attributes according to the time sequence of the corresponding frame images to obtain the pick-up and put-down status information between the target person and the target item.

[0099] Optionally, the detection submodule includes:

[0100] The first detection unit is used to perform target detection on each frame of the image to obtain the hand detection result of the target person.

[0101] The second detection unit is used to perform attribute detection on the hand detection results to obtain the handling attributes between the target person and the target item in each frame of the image.

[0102] Optionally, the first detection unit includes:

[0103] The first detection subunit is used to perform a first target detection on each frame of the image to obtain a human body detection box;

[0104] The second detection subunit is used to perform second target detection on each frame of the image to obtain a hand detection box.

[0105] The association subunit is used to associate the human body detection box and the hand detection box in each frame image to obtain the association relationship between the human body detection box and the hand detection box in each frame image;

[0106] The third detection subunit is used to obtain the hand detection result of the target person in each frame image based on the correlation between the human body detection box and the hand detection box in each frame image and whether the hand detection box includes the target item.

[0107] Optionally, the association subunit is further configured to calculate the detection box distance between the human body detection box and the hand detection box in each frame image; and to associate the human body detection box and the hand detection box according to the detection box distance to obtain the association relationship between the human body detection box and the hand detection box in each frame image.

[0108] Optionally, the sorting submodule includes:

[0109] The identifier allocation unit is used to assign a first identifier to human body detection boxes belonging to the same target person in each frame image, and to assign a second identifier to hand detection boxes belonging to the same target person in adjacent frame images.

[0110] The sorting unit is used to sort the pick-up and put-down attributes corresponding to each target person according to the first identifier and the second identifier in the time sequence of the corresponding frame images, so as to obtain the pick-up and put-down status information between each target person and the corresponding target item.

[0111] Optionally, the identification module 203 includes:

[0112] The filtering submodule is used to filter the pick-up and put-down status information to obtain filtered pick-up and put-down status information.

[0113] The identification submodule is used to identify the behavior of the target person based on the filtered pick-up and put-down status information, and obtain the behavior identification result of the target person.

[0114] It should be noted that the item monitoring device provided in this embodiment of the invention can be applied to devices such as smart cameras, smartphones, computers, and servers that can monitor items.

[0115] The item monitoring device provided in this embodiment of the invention can implement all the processes implemented by the item monitoring method in the above-described method embodiments, and can achieve the same beneficial effects. To avoid repetition, it will not be described again here.

[0116] See Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 3 As shown, it includes: a memory 302, a processor 301, and a computer program for an item monitoring method stored in the memory 302 and executable on the processor 301, wherein:

[0117] The processor 301 is used to call the computer program stored in the memory 302 and perform the following steps:

[0118] Acquire surveillance images of the target area, the surveillance images including target personnel and target items;

[0119] Based on the surveillance image, the status of the target person and the target item is detected to obtain the handling and placement status information between the target person and the target item.

[0120] Based on the picking and placing status information, the behavior of the target person is identified to obtain the behavior identification result of the target person;

[0121] The target items are monitored based on the behavioral recognition results of the target personnel.

[0122] Optionally, the monitoring image includes consecutive frame images, and the processor 301 executes the step of performing state detection on the target person and the target item based on the monitoring image to obtain the handling and placement state information between the target person and the target item, including:

[0123] For each frame of the image, state detection is performed on the target person and the target item to obtain the handling and placement attributes between the target person and the target item in each frame of the image.

[0124] The picking and placing attributes are sorted according to the time sequence of the corresponding frame images to obtain the picking and placing status information between the target person and the target item.

[0125] Optionally, the processor 301 performs state detection on the target person and the target item in each frame of the image to obtain the handling and placement attributes between the target person and the target item in each frame of the image, including:

[0126] Target detection is performed on each frame of the image to obtain the hand detection result of the target person;

[0127] The hand detection results are used to detect the attributes of the target item to obtain the handling attributes between the target person and the target item in each frame of the image.

[0128] Optionally, the process of performing target detection on each frame of the image to obtain the hand detection result of the target person, executed by processor 301, includes:

[0129] Perform a first target detection on each frame of the image to obtain a human body detection box;

[0130] Perform a second target detection on each frame of the image to obtain a hand detection bounding box;

[0131] The human body detection box and the hand detection box in each frame image are associated to obtain the association relationship between the human body detection box and the hand detection box in each frame image;

[0132] Based on the correlation between the human body detection box and the hand detection box in each frame of the image, and whether the hand detection box in each frame of the image includes the target item, the hand detection result of the target person is obtained.

[0133] Optionally, the process executed by processor 301 to associate the human detection box and the hand detection box in each frame of the image to obtain the association relationship between the human detection box and the hand detection box in each frame of the image includes:

[0134] Calculate the detection box distance between the human body detection box and the hand detection box in each frame image;

[0135] Based on the detection box distance, the association between the human body detection box and the hand detection box in each frame image is obtained.

[0136] Optionally, the step of processor 301 sorting the pick-up and place attributes according to the time sequence of the corresponding frame images to obtain the pick-up and place status information between the target person and the target item includes:

[0137] Assign a first identifier to human body detection boxes belonging to the same target person in each frame image, and assign a second identifier to hand detection boxes belonging to the same target person in adjacent frame images;

[0138] Based on the first identifier and the second identifier, the picking and placing attributes corresponding to each target person are sorted according to the time sequence of the corresponding frame images to obtain the picking and placing status information between each target person and the corresponding target item.

[0139] Optionally, the process executed by processor 301 to perform behavior recognition on the target person based on the pick-up and put-down state information, and obtain the behavior recognition result of the target person, includes:

[0140] The pick-up and put-down status information is filtered to obtain filtered pick-up and put-down status information;

[0141] Based on the filtered pick-up and put-down status information, the target person's behavior is identified to obtain the target person's behavior identification result.

[0142] The electronic device provided in this embodiment of the invention can implement all the processes of the item monitoring method in the above-described method embodiments, and can achieve the same beneficial effects. To avoid repetition, further details are omitted here.

[0143] This invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the item monitoring method or the application-side item monitoring method provided in this invention, and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0144] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0145] The above description discloses only preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. Therefore, equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.

Claims

1. An article monitoring method characterized by, Includes the following steps: Acquire surveillance images of the target area, the surveillance images including target personnel and target objects, and the surveillance images including consecutive frame images; Based on the surveillance images, state detection is performed on the target personnel and the target items to obtain the handling and placement status information between the target personnel and the target items. This includes: performing a first target detection on each frame of the image to obtain a human body detection box; performing a second target detection on each frame of the image to obtain a hand detection box; associating the human body detection box and the hand detection box in each frame of the image to obtain the association relationship between the human body detection box and the hand detection box in each frame of the image; obtaining the hand detection result of the target personnel based on the association relationship between the human body detection box and the hand detection box in each frame of the image and whether the hand detection box in each frame of the image includes the target item; identifying the identity information of the target personnel based on the human body detection box to obtain the hand detection result of each target personnel; performing attribute detection on the target item based on the hand detection result to obtain the handling and placement attributes between the target personnel and the target item in each frame of the image; and sorting the handling and placement attributes according to the time order of the corresponding frame images to obtain the handling and placement status information between the target personnel and the target item. Based on the picking and placing status information, the behavior of the target person is identified to obtain the behavior identification result of the target person; The target items are monitored based on the behavioral recognition results of the target personnel.

2. The article monitoring method as described in claim 1, characterized in that, The step of associating the human detection box and the hand detection box in each frame of the image to obtain the association relationship between the human detection box and the hand detection box in each frame of the image includes: Calculate the detection box distance between the human body detection box and the hand detection box in each frame image; Based on the detection box distance, the association between the human body detection box and the hand detection box in each frame image is obtained.

3. The article monitoring method as described in claim 2, characterized in that, The step of sorting the pick-up and put-down attributes according to the time sequence of the corresponding frame images to obtain the pick-up and put-down status information between the target person and the target item includes: Assign a first identifier to human body detection boxes belonging to the same target person in each frame image, and assign a second identifier to hand detection boxes belonging to the same target person in adjacent frame images; Based on the first identifier and the second identifier, the picking and placing attributes corresponding to each target person are sorted according to the time sequence of the corresponding frame images to obtain the picking and placing status information between each target person and the corresponding target item.

4. The article monitoring method according to any one of claims 1 to 3, characterized in that, The step of performing behavior recognition on the target person based on the picking and placing status information to obtain the behavior recognition result of the target person includes: The pick-up and put-down status information is filtered to obtain filtered pick-up and put-down status information; Based on the filtered pick-up and put-down status information, the target person's behavior is identified to obtain the target person's behavior identification result.

5. An item monitoring device, characterized in that, The device includes: The acquisition module is used to acquire surveillance images of the target area, the surveillance images including target personnel and target items, and the surveillance images including continuous frame images; The detection module is used to perform state detection on the target personnel and the target items based on the monitoring images to obtain the handling and placement state information between the target personnel and the target items. This includes: performing a first target detection on each frame of the image to obtain a human body detection box; performing a second target detection on each frame of the image to obtain a hand detection box; associating the human body detection box and the hand detection box in each frame of the image to obtain the association relationship between the human body detection box and the hand detection box in each frame of the image; obtaining the hand detection result of the target personnel based on the association relationship between the human body detection box and the hand detection box in each frame of the image and whether the hand detection box in each frame of the image includes the target item; identifying the identity information of the target personnel based on the human body detection box to obtain the hand detection result of each target personnel; performing attribute detection on the target item based on the hand detection result to obtain the handling and placement attributes between the target personnel and the target item in each frame of the image; and sorting the handling and placement attributes according to the time sequence of the corresponding frame images to obtain the handling and placement state information between the target personnel and the target item. The identification module is used to identify the behavior of the target person based on the picking and placing status information, and obtain the behavior identification result of the target person; The monitoring module is used to monitor the target item based on the behavior recognition results of the target person.

6. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the article monitoring method as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the article monitoring method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Dynamic commodity recognition method, unmanned vending cabinet and vending method thereof

    CN112907168A