Drug sales monitoring method and system

CN115620186BActive Publication Date: 2026-09-29QINGDAO INTELLIFUSION TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211051299.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2026-09-29
Estimated Expiration
2042-08-30

AI Technical Summary

Benefits of technology

[0013]本发明实施例提供一种药品售卖监控方法及系统,该药品售卖监控方法通过在监控到药品存放区域有第一对象进入时,监控第一对象在药品存放区域内的行为是否为取药行为,从而在监控到第一对象在药品存放区域内的行为是取药行为时,生成取药行为信息,同时在监控到药品结账区域有第二对象进入时,监控第二对象在药品结账区域内的行为是否支付行为,从而在监控到第二对象在药品结账区域内的行为是支付行为时,生成支付行为信息,这样在生成的取药行为信息的预设时长内查询到生成的支付行为信息时,生成药品售卖事件,以使得药店人员无法对药品售卖进行弄虚作假,极大地提高了药品售卖情况的真实性和准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115620186B_ABST
    Figure CN115620186B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a kind of drug selling monitoring method and system, belong to data processing field.The method comprises: when it is determined that the first target area has first object to enter according to the first monitoring video of the first target area corresponding, it is determined whether the behavior of first object in the first target area is taking medicine behavior according to the first monitoring video;When it is determined that the behavior of first object in the first target area is taking medicine behavior, generate taking medicine behavior information;When it is determined that the second target area has second object to enter according to the second monitoring video of the second target area corresponding, it is determined whether the behavior of second object in the second target area is payment behavior according to the second monitoring video;When it is determined that the behavior of second object in the second target area is payment behavior, generate payment behavior information;When generated payment behavior information is inquired within the preset time length of taking medicine behavior information, generate drug selling event.The method improves the authenticity of drug selling situation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method and system for monitoring drug sales. Background Technology

[0002] Medicines are an indispensable part of people's daily lives, and medicine safety is related to people's health and safety. Pharmacy staff need to sell medicines in accordance with national or regional regulations. Currently, the main method is to use a point-of-sale system to scan the barcodes of medicines to record the sales information of pharmacy staff. However, there are cases of pharmacy staff falsifying information, which cannot guarantee the authenticity and accuracy of the sales information. Summary of the Invention

[0003] This invention provides a method and system for monitoring drug sales, aiming to improve the authenticity and accuracy of drug sales information in pharmacies.

[0004] In a first aspect, embodiments of the present invention provide a method for monitoring drug sales, comprising: when it is determined, based on a first surveillance video corresponding to a first target area, that a first object has entered the first target area, determining, based on the first surveillance video, whether the behavior of the first object in the first target area is a drug retrieval behavior, wherein the first target area includes a drug storage area; when it is determined that the behavior of the first object in the first target area is a drug retrieval behavior, generating drug retrieval behavior information, wherein the drug retrieval behavior information is used to indicate that the behavior of the first object in the first target area is a drug retrieval behavior; when it is determined, based on a second surveillance video corresponding to a second target area, that a second object has entered the second target area, determining, based on the second surveillance video, whether the behavior of the second object in the second target area is a payment behavior, wherein the second target area includes a drug checkout area; when it is determined that the behavior of the second object in the second target area is a payment behavior, generating payment behavior information, wherein the payment behavior information is used to indicate that the behavior of the second object in the second target area is a payment behavior; and when the generated payment behavior information is found within a preset time period for generating the drug retrieval behavior information, generating a drug sales event.

[0005] In one embodiment, determining whether the behavior of the first object in the first target area constitutes medication retrieval based on the first surveillance video includes: inputting multiple frames of images from the first surveillance video into a preset target detection model for processing to obtain target detection results for each frame of the multiple frames; tracking the first object based on the target detection results for each frame of the images; and determining whether the behavior of the first object in the first target area constitutes medication retrieval based on the target detection results for each frame of the images when the first object leaves the first target area.

[0006] In one embodiment, the multi-frame images include at least a first image and a second image adjacent to the first image. The step of tracking the first object based on the target detection results of each frame of the images includes: calculating the intersection-union ratio (IUR) between the head and shoulder detection boxes of the first object in the first image and the head and shoulder detection boxes of each object to be detected in the second image; determining the largest IUR among the multiple IURs as the target IUR, and determining whether the target IUR is greater than or equal to a preset IUR threshold; when the target IUR is greater than or equal to the preset IUR threshold, marking the object to be detected in the second image corresponding to the target IUR as the first object.

[0007] In one embodiment, the target detection result includes a hand status label of the first object. Determining whether the first object's behavior within the first target area constitutes medication retrieval based on the target detection results of each frame of the image includes: counting the number of images corresponding to the hand status label being a medication placement status label to obtain a first image count, where the medication placement status label describes the first object's hand being in a medication placement state; determining the percentage of the first image count relative to the total number of the multiple frames of images; and determining that the first object's behavior within the first target area constitutes medication retrieval when the percentage is greater than or equal to a preset percentage threshold.

[0008] In one embodiment, the target detection result includes a hand state label of the first object. Determining whether the first object's behavior within the first target area constitutes medication retrieval based on the target detection results of each frame of the image includes: determining a first target image, where the first target image is the image in the multi-frame image where the hand state label is the first detected one corresponding to a medication placement label, the medication placement label describing that the first object's hand is in a medication placement state; counting the number of second target images where the hand state label corresponds to the medication placement label, the second target images being images located after the first target image in the multi-frame image set; incrementing the number of second target images by 1 to obtain a second image count; and determining that the first object's behavior within the first target area constitutes medication retrieval when the second image count is greater than or equal to a preset image count threshold.

[0009] In one embodiment, the drug collection behavior information includes a drug collection behavior identifier and the body feature information of the first object. The body feature information includes at least one of facial feature information and head and shoulder feature information. The drug sales monitoring method further includes: calculating a first similarity between the body feature information of the first object and preset body feature information; and generating a drug sales event when the first similarity is greater than or equal to a preset similarity threshold.

[0010] In one embodiment, the payment behavior information includes a payment behavior identifier and the body feature information of the second object. After calculating the first similarity between the body feature information of the first object and the preset body feature information, the method further includes: when the first similarity is less than a preset similarity threshold, calculating the second similarity between the body feature information of the first object and the body feature information of the second object; and when the second similarity is greater than or equal to the similarity threshold, generating a drug sales event.

[0011] Secondly, embodiments of the present invention also provide a drug sales monitoring system, comprising a first monitoring device, a second monitoring device, and a server, wherein the first monitoring device and the second monitoring device are respectively communicatively connected to the server; the first monitoring device is used to monitor a first target area and obtain a first monitoring video, the first target area including a drug storage area; when it is determined from the first monitoring video that a first object has entered the first target area, it is determined from the first monitoring video whether the behavior of the first object in the first target area is a drug retrieval behavior; when it is determined that the behavior of the first object in the first target area is a drug retrieval behavior, drug retrieval behavior information is generated and sent to the server, the drug retrieval behavior information indicating that the behavior of the first object in the first target area is a drug retrieval behavior; The second monitoring device is used to monitor a second target area and obtain a second monitoring video, the second target area including a drug checkout area; when it is determined from the second monitoring video that a second object has entered the second target area, it is determined from the second monitoring video whether the behavior of the second object in the second target area is a payment behavior; when it is determined that the behavior of the second object in the second target area is a payment behavior, payment behavior information is generated and sent to the server, the payment behavior information being used to indicate that the behavior of the second object in the drug checkout area is a payment behavior; the server is used to determine whether it has received the payment behavior information sent by the second monitoring device within a preset time period when it receives the drug collection behavior information; when it receives the payment behavior information sent by the second monitoring device within the preset time period, a drug sales event is generated.

[0012] Thirdly, embodiments of the present invention also provide a drug sales monitoring system, comprising a first monitoring device, a second monitoring device, and a server, wherein the first monitoring device and the second monitoring device are respectively communicatively connected to the server; the first monitoring device is used to monitor a first target area, obtain a first monitoring video, and send the first monitoring video to the server, the first target area including a drug storage area; the second monitoring device is used to monitor a second target area, obtain a second monitoring video, and send the second monitoring video to the server, the second target area including a drug checkout area; the server is used to acquire the first monitoring video sent by the first monitoring device and the second monitoring video sent by the second monitoring device; when it is determined from the first monitoring video that a first object has entered the first target area, according to the... The system determines whether the behavior of the first object in the first target area constitutes medication retrieval based on the first surveillance video. If the behavior is determined to be medication retrieval, medication retrieval behavior information is generated, indicating that the first object's behavior in the first target area constitutes medication retrieval. If the system determines that a second object has entered the second target area based on the second surveillance video, it determines whether the second object's behavior in the second target area constitutes payment. If the behavior is determined to be payment, payment behavior information is generated, indicating that the second object's behavior in the second target area constitutes payment. If the generated payment behavior information is found within a preset time period after the medication retrieval behavior information is generated, a medication sale event is generated.

[0013] This invention provides a method and system for monitoring drug sales. The method monitors whether a first person's behavior in the drug storage area constitutes drug retrieval when the first person enters the storage area. If the first person's behavior is confirmed to be drug retrieval, drug retrieval information is generated. Simultaneously, when a second person enters the drug checkout area, the method monitors whether the second person's behavior is confirmed to be payment. If the second person's behavior is confirmed to be payment, payment information is generated. When payment information is retrieved within a preset timeframe of the drug retrieval information, a drug sales event is generated. This prevents pharmacy staff from falsifying drug sales, significantly improving the authenticity and accuracy of drug sales data. Attached Figure Description

[0014] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a schematic diagram of a scenario for implementing the drug sales monitoring method provided in the embodiments of the present invention;

[0016] Figure 2 This is a flowchart illustrating a drug sales monitoring method provided in an embodiment of the present invention;

[0017] Figure 3 This is a schematic diagram of a frame from the first surveillance video in an embodiment of the present invention;

[0018] Figure 4 yes Figure 2 A flowchart illustrating the sub-steps of the drug sales monitoring method in China;

[0019] Figure 5 This is a schematic diagram of a network layer of the preset target detection model in an embodiment of the present invention;

[0020] Figure 6 This is a schematic diagram of the target detection results of multiple frames of images in an embodiment of the present invention;

[0021] Figure 7 This is another schematic diagram of the target detection results of multiple frames of images in an embodiment of the present invention;

[0022] Figure 8 This is a schematic block diagram of a drug sales monitoring system provided in an embodiment of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0025] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0026] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0027] Please see Figure 1 , Figure 1 This is a schematic diagram of a scenario for implementing the drug sales monitoring method provided in the embodiments of the present invention.

[0028] like Figure 1 As shown, the scenario includes a first monitoring device 100, a second monitoring device 200, and a server 300. The first monitoring device 100 and the second monitoring device 200 are respectively connected to the server 300. The first monitoring device 100 can be installed on the top or side wall of the first target area of ​​the pharmacy, so that the first target area is within the monitoring range of the first monitoring device 100. The first target area includes the medicine storage area. The second monitoring device 200 can be installed on the top or side wall of the second target area, so that the medicine checkout area is within the monitoring range of the second monitoring device 200. The second target area includes the medicine checkout area, which is the area where customers move around when checking out near the cashier.

[0029] In one embodiment, a first monitoring device 100 monitors a first target area and obtains a first monitoring video. The first target area includes a medicine storage area. When it is determined from the first monitoring video that a first object has entered the first target area, it is determined from the first monitoring video whether the first object's behavior in the first target area is a medicine retrieval behavior. When it is determined that the first object's behavior in the first target area is a medicine retrieval behavior, medicine retrieval behavior information is generated and sent to a server 300. The medicine retrieval behavior information is used to indicate that the first object's behavior in the first target area is a medicine retrieval behavior. A second monitoring device 200 monitors a second target area and obtains a second monitoring video. The second target area includes a medicine checkout area. When a second object is determined to have entered the second target area based on the second surveillance video, the system determines whether the second object's behavior within the second target area constitutes a payment behavior. If the second object's behavior within the second target area is determined to be a payment behavior, payment behavior information is generated and sent to server 300. This payment behavior information indicates that the second object's behavior within the drug checkout area constitutes a payment behavior. When server 300 receives drug collection behavior information from first monitoring device 100, it determines whether it has received payment behavior information from second monitoring device 200 within a preset time period. If it has received payment behavior information from second monitoring device 200 within the preset time period, a drug sales event is generated.

[0030] In one embodiment, a first monitoring device 100 monitors a first target area, obtains a first monitoring video, and sends the first monitoring video to a server 300. The first target area includes a medicine storage area. A second monitoring device 200 monitors a second target area, obtains a second monitoring video, and sends the second monitoring video to the server 300. The second target area includes a medicine checkout area. The server 300 acquires the first monitoring video sent by the first monitoring device 100 and the second monitoring video sent by the second monitoring device 200. When it is determined from the first monitoring video that a first object has entered the first target area, the server determines from the first monitoring video whether the first object's behavior in the first target area is... Medication pickup behavior; when it is determined that the first object's behavior in the first target area is a medication pickup behavior, medication pickup behavior information is generated, which is used to indicate that the first object's behavior in the first target area is a medication pickup behavior; when it is determined from the second surveillance video that a second object has entered the second target area, it is determined from the second surveillance video whether the second object's behavior in the second target area is a payment behavior; when it is determined that the second object's behavior in the second target area is a payment behavior, payment behavior information is generated, which is used to indicate that the second object's behavior in the second target area is a payment behavior; when the generated payment behavior information is found within the preset time period for generating medication pickup behavior information, a drug sales event is generated.

[0031] The first monitoring device 100 or the second monitoring device 200 may include one or more shooting devices and chips. The server 300 may be an independent server, a server cluster composed of multiple servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0032] The following will combine Figure 1 The following scenario provides a detailed description of the drug sales monitoring method provided by embodiments of the present invention. It should be noted that... Figure 1 The scenarios described are only used to explain the drug sales monitoring method provided in the embodiments of the present invention, but do not constitute a limitation on the application scenarios of the drug sales monitoring method provided in the embodiments of the present invention.

[0033] Please see Figure 2 , Figure 2 This is a flowchart illustrating a drug sales monitoring method provided in an embodiment of the present invention.

[0034] like Figure 2 As shown, the drug sales monitoring method includes steps S101 to S105.

[0035] Step S101: When it is determined from the first surveillance video corresponding to the first target area that a first object has entered the first target area, determine from the first surveillance video whether the behavior of the first object in the first target area is a drug-taking behavior.

[0036] In this embodiment of the invention, the first target area includes a drug storage area, the first object may include a pharmacist or a customer of a pharmacy, the drug storage area may include a first drug storage area and / or a second drug storage area, the first drug storage area is used to store drugs that need to be registered before they can be sold, and drugs that need to be registered before they can be sold may include prescription drugs, cough medicines, antipyretics, antiviral drugs, antibiotics, etc., the second drug storage area is used to store drugs that can be sold without registration.

[0037] In one embodiment, multiple frames of images are acquired from a first surveillance video, and a first background image corresponding to a first target region is acquired. The pixel differences between each frame of the multiple frames and the first background image are calculated. If there is an image in each frame with a pixel difference greater than or equal to a preset pixel difference threshold, it is determined that a first object has entered the first target region. If the pixel difference corresponding to each frame is less than the pixel difference threshold, it is determined that no first object has entered the first target region. The first background image is the image captured by the first monitoring device when there is no first object in the first target region. For example, such as... Figure 3 As shown, the region of interest 41 in a frame of the first surveillance video includes the first object 42, and it can be determined that the first object has entered the first target area.

[0038] In one embodiment, multiple frames of images are acquired from a first surveillance video, and each frame is input into a preset target detection model for processing to obtain the target detection results for each frame. If each frame contains an image whose target detection result includes a head-shoulder detection box or a face detection box, it is determined that a first object has entered the first target region. Conversely, if the target detection result for each frame does not include a head-shoulder detection box or a face detection box, it is determined that no first object has entered the first target region. By using the target detection results of the images, it is possible to accurately determine whether a first object has entered the first target region, thereby improving the authenticity and accuracy of drug sales events.

[0039] In one embodiment, multiple frames of images are acquired from a first surveillance video, and a first background image corresponding to a first target region is acquired. The pixel difference between each frame of the multiple images and the first background image is calculated, and the image corresponding to the pixel difference being greater than or equal to a preset pixel difference threshold is acquired from the multiple frames to obtain a target image. The target image is then input into a preset target detection model for processing to obtain a target detection result for the target image. If the target detection result of the target image includes a head and shoulder detection box or a face detection box, it is determined that a first object has entered the first target region. If the target detection result of the target image does not include a head and shoulder detection box or a face detection box, it is determined that no first object has entered the first target region. By comprehensively considering the target detection result of the target image and the pixel difference between the target image and the background image, the accuracy of detecting whether a first object has entered the first target region can be further improved, thereby improving the authenticity and accuracy of drug sales information.

[0040] In one embodiment, such as Figure 4 As shown, step S101 includes sub-steps S1011 to S1013.

[0041] Sub-step S1011: Input the multi-frame images from the first monitoring video into the preset target detection model for processing to obtain the target detection results of each frame in the multi-frame images.

[0042] In this embodiment of the invention, the target detection result includes a head and shoulder detection frame of the first object or the object to be detected, and a hand status label of the first object. The hand status label includes a drug placement status label or an idle status label. The drug placement status label is used to describe that the first object's hand is in a drug placement state, and the idle status label is used to describe that the first object's hand is in an idle state.

[0043] In one embodiment, the preset target detection model includes a lightweight backbone network, a bottleneck network, and a detection head network. The bottleneck network includes transposed convolutional layers and a path aggregation network (PAN). Using a lightweight backbone network can reduce the amount of data in the target detection model and improve its running speed. Furthermore, the path aggregation network in the bottleneck network uses transposed convolutional layers, which can reduce the accuracy loss during the transformation and quantization process of the target detection model, thus ensuring the accuracy of the target detection model deployed on monitoring devices or servers.

[0044] In one embodiment, the image is normalized, and features are extracted from the normalized image using a lightweight backbone network to obtain multiple first feature maps of different sizes. The multiple first feature maps are then processed by a transposed convolutional layer and a path aggregation network to obtain multiple second feature maps. The multiple second feature maps are then processed by a detection head network to obtain multiple target feature maps. Based on the multiple target feature maps, the target detection result of the first object is determined.

[0045] In one embodiment, the lightweight backbone network includes a convolutional network and a Cross-Stage Partial Network (CSP). Existing lightweight networks like YOLOv5s use a backbone network consisting of a Focus network and a CSP. However, the Focus network requires downsampling and slicing the image, and frequent downsampling and slicing operations can severely impact the cache of monitoring devices or servers. Compared to existing lightweight networks like YOLOv5s, the object detection model in this invention uses a convolutional network and a CSP, eliminating the need for frequent downsampling and slicing operations. This reduces the cache usage of monitoring devices or servers and improves the deployment flexibility of the object detection model.

[0046] In one embodiment, the lightweight backbone network includes ShuffleNetV2. Compared to the existing lightweight network YOLOv5s, the target detection model in this embodiment uses ShuffleNetV2 as its backbone network. This eliminates the need for frequent downsampling and slicing of images, reducing cache usage on monitoring devices or servers and improving the deployment flexibility of the target detection model. Furthermore, ShuffleNetV2 has a smaller data volume, thus reducing the overall data volume of the target detection model.

[0047] In one embodiment, a convolutional network is used to convolve the normalized image to obtain a convolutional image, and a cross-stage local network is used to extract features from the convolutional image to obtain multiple first feature maps; alternatively, ShuffleNetV2 is used to extract features from the normalized image to obtain multiple first feature maps. The following explanation uses a lightweight backbone network including a convolutional network and a cross-stage local network, and a bottleneck network including transposed convolutional layers and a path aggregation network, as an example to illustrate the preset object detection model.

[0048] like Figure 5 As shown, the preset target detection network includes a lightweight backbone network 10, a bottleneck network 20, and a detection head network 30. The lightweight backbone network 10 includes a convolutional network 11, a first CSP layer 12, a second CSP layer 13, and a third CSP layer 14. The convolutional network 11 is connected to the first CSP layer 12, the first CSP layer 12 is connected to the second CSP layer 13, and the second CSP layer 13 is connected to the third CSP layer 14. The bottleneck network 20 includes a first transposed convolutional layer 211, a second transposed convolutional layer 212, a first splicing layer 213, a second splicing layer 214, a fourth CSP layer 215, a first convolutional layer 216, a third splicing layer 217, a fifth CSP layer 218, a second convolutional layer 219, a fourth splicing layer 220, and a sixth CSP layer 221. The detection head network 30 includes a third convolutional layer 31, a fourth convolutional layer 32, and a fifth convolutional layer 33.

[0049] The first transposed convolutional layer 211 is connected to the third CSP layer 14 and the first splicing layer 213. The second transposed convolutional layer 212 is connected to the first splicing layer 213 and the second splicing layer 214. The first splicing layer 213 is connected to the second CSP layer 13. The second splicing layer 214 is connected to the first CSP layer 12 and the fourth CSP layer 215. The fourth CSP layer 215 is connected to the first convolutional layer 216. The first convolutional layer 216 is connected to the third splicing layer 217. The third splicing layer 217 is connected to the first splicing layer 213. The fifth CSP layer 218 is connected to the second convolutional layer 219. The second convolutional layer 219 is connected to the fourth splicing layer 220. The fourth splicing layer 220 is connected to the third CSP layer 14 and the sixth CSP layer 221. The third convolutional layer 31 is connected to the fourth CSP layer 215, the fourth convolutional layer 32 is connected to the fifth CSP layer 218, and the fifth convolutional layer 33 is connected to the sixth CSP layer 221.

[0050] In one embodiment, the image is normalized, and then convolved by a convolutional network 11 to obtain a convolutional image. Features are extracted from the convolutional image through a first CSP layer 12 to obtain feature map A1, then through a second CSP layer 13 to obtain feature map A2, and finally through a third CSP layer 14 to obtain feature map A3. Feature map A3 is then transposed and convolved by a first transposed convolutional layer 211 to obtain a feature map. Figure B1; Feature map B2 is obtained by concatenating feature map A2 and feature map B1 through the first concatenation layer 213; feature map B3 is obtained by transposing and convolving feature map B2 through the second transposed convolutional layer 212; feature map B4 is obtained by concatenating feature map A1 and feature map B3 through the second concatenation layer 214; feature map C1 is obtained by processing feature map B4 through the fourth CSP layer 215; and target feature map D1 is obtained by convolving feature map C1 through the third convolutional layer 31.

[0051] Feature map B5 is obtained by convolving feature map C1 through the first convolutional layer 216; feature map B6 is obtained by concatenating feature map B2 and feature map B5 through the third concatenation layer 217; feature map C2 is obtained by processing feature map B6 through the fifth CSP layer; and feature map D2 is obtained by convolving feature map C2 through the fourth convolutional layer 32. Feature map B7 is obtained by convolving feature map C2 through the second convolutional layer 219; feature map B8 is obtained by concatenating feature map A3 and feature map B7 through the fourth concatenation layer 220; feature map C3 is obtained by processing feature map B8 through the sixth CSP layer; and feature map D3 is obtained by convolving feature map C3 through the fifth convolutional layer 33.

[0052] Sub-step S1012: Track the first object based on the target detection results of each frame image.

[0053] In this embodiment of the invention, the multi-frame images include a first image and a second image adjacent to the first image. The tracking of the first object based on the target detection results of each frame can be achieved by: calculating the intersection-union ratio (IUR) between the head and shoulder detection boxes of the first object in the first image and the head and shoulder detection boxes of each object to be detected in the second image; determining the largest IUR among multiple IURs as the target IUR, and determining whether the target IUR is greater than or equal to a preset IUR threshold; when the target IUR is greater than or equal to the preset IUR threshold, marking the object to be detected in the second image corresponding to the target IUR as the first object. The IUR threshold can be set based on actual conditions, and this embodiment of the invention does not specifically limit it. Since the size of a person's head and shoulders is relatively fixed, and a person's head and shoulders are usually not occluded, the head and shoulders can be accurately identified, thereby accurately calculating the IUR between the head and shoulder detection boxes, which can improve the accuracy of tracking the first object. Furthermore, the calculation of the IUR between the head and shoulder detection boxes is small, which can improve the efficiency of tracking the first object.

[0054] In one embodiment, when the target cross-union ratio (CUNR) is less than a CUNR threshold, the object to be detected in the second image corresponding to the target CUNR is marked as a new first object. For example, let the area of ​​the head and shoulder detection bounding box of the first object A in the first image be S. A In the second image, the areas of the head and shoulder detection boxes for objects B1, B2, and B3 are S, respectively. B1 S B2 and S B3 Then, the intersection-union ratios (IoU) between the head and shoulder detection bounding boxes of the first object A and the head and shoulder detection bounding boxes of objects B1, B2, and B3 are respectively (S). A ∩S B1 ) / (S A ∪S B1 ), (S A ∩S B2 ) / (S A ∪S B2 ) and (S A ∩S B3 ) / (S A ∪S B3 ), because (S A ∩S B2 ) / (S A ∪S B2 If (S) is the largest, then (S) A ∩S B2 ) / (S A ∪S B2If the crossover ratio (CRR) is greater than or equal to the preset threshold, then the object to be detected, B2, is marked as the first object, A. If (S...) A ∩S B2 ) / (S A ∪S B2 If the crossover ratio (CRO) is less than or equal to the preset crossover ratio threshold, then the object to be detected, B2, will be marked as the new first object, C.

[0055] Sub-step S1013: When the first object leaves the first target area, determine whether the first object’s behavior in the first target area is a drug-taking behavior based on the target detection results of each frame image.

[0056] Through sub-steps S1011 to S1013, when a first object is detected entering the drug storage area, a target detection model can be used to perform target detection on the collected image and track the first object based on the target detection results. In this way, when the first object leaves the drug storage area, it is possible to determine in real time and accurately whether the first object's behavior in the drug storage area is a drug retrieval behavior based on the target detection results.

[0057] In this embodiment of the invention, the target detection result includes a hand status label of the first object. The hand status label includes a drug-holding status label or an idle status label. The drug-holding status label is used to describe that the first object's hand is in a drug-holding state, and the idle status label is used to describe that the first object's hand is in an idle state. The first object's hand being in a drug-holding state means that the first object's hand is holding a drug, while the first object's hand being in an idle state means that the first object's hand is not holding a drug.

[0058] In one embodiment, the number of images corresponding to hand status tags that are also drug placement status tags is counted to obtain a first number of images; the percentage of the first number of images relative to the total number of multi-frame images is determined; when this percentage is greater than or equal to a preset percentage threshold, the behavior of the first object in the first target area is determined to be a drug retrieval behavior; when this percentage is less than the preset percentage threshold, the behavior of the first object in the first target area is determined not to be a drug retrieval behavior. The total number of multi-frame images and the preset percentage threshold can be set based on actual conditions, and this embodiment of the invention does not impose specific limitations on them. For example, the total number of multi-frame images is 10 frames, and the preset percentage threshold is 70% or 50%. By comprehensively considering the hand tag status in the target detection results of multi-frame images, the interference caused by misidentification by the target detection model can be reduced, and the recognition accuracy of the first object's behavior in the first target area can be improved.

[0059] For example, the total number of images is 10, and the hand state label of the first object in the target detection results of these 10 images is as follows: Figure 6 As shown, Figure 6In the text, "1" represents the drug placement status label, indicating that the first object's hand is in a drug placement state, and "0" represents the idle state label, indicating that the first object's hand is in an idle state or is a misidentification. Statistically, the number of images corresponding to the hand status label "1" is 7 frames. The percentage of images corresponding to the hand status label "1" in the total number of multiple frames is calculated to be 70%, while the percentage threshold is 50%. Therefore, it can be determined that the first object's behavior in the first target area is a drug retrieval behavior.

[0060] In one embodiment, a first target image is determined. The first target image is the image in a multi-frame image where the first detected hand state label is a drug placement state label, and the drug placement state label describes that the first object's hand is in a drug placement state. The number of second target images corresponding to the hand state label being a drug placement state label is counted. The second target images are images in the multi-frame image that are located after the first target image. The number of second target images is incremented by 1 to obtain the number of second images. When the number of second images is greater than or equal to a preset image number threshold, it is determined that the first object's behavior in the first target area is a drug retrieval behavior. The image number threshold can be set based on actual conditions, and this embodiment of the invention does not specifically limit it. For example, the image number threshold can be 3 or 5. By comprehensively considering the hand label states in the target detection results of multiple frames, the interference caused by misidentification of the target detection model can be reduced, and the recognition accuracy of the first object's behavior in the first target area can be improved.

[0061] For example, the total number of images is 10, and the hand state label of the first object in the target detection results of these 10 images is as follows: Figure 7 As shown, Figure 7 In the code, "1" represents the medication placement status label, indicating that the first object's hand has been detected as being in a medication placement state; "0" represents the idle status label, indicating that the first object's hand has been detected as being in an idle state or that the detection was incorrect. Figure 7 It can be seen that the image with the first hand state label being the drug placement state label is image 51, and there are 5 images after image 51 where the hand state label in the target detection results is "1". Therefore, the number of second target images corresponding to the hand state label being the drug placement state label is 5, so the number of second images is 6, and the image number threshold is 5. Since 6 is greater than 5, it can be determined that the behavior of the first object in the first target area is the drug retrieval behavior.

[0062] In one embodiment, the target detection results corresponding to the last few frames of a multi-frame image are obtained, and the number of images corresponding to head and shoulder detection boxes that do not contain the first object in the target detection results is counted to obtain the third image count. The percentage of the third image count to the total number of the last few frames is calculated, and when the percentage is greater than or equal to a preset percentage, it is determined that the tracked first object has left the first target region. By considering the target detection results of the multi-frame images, it is possible to accurately determine whether the tracked first object has left the first target region.

[0063] The total number of the last few frames and the preset percentage can be set based on actual conditions, and this embodiment of the invention does not impose specific limitations on this. For example, the total number of the last few frames is 4, and the preset percentage is 50%. The target detection results for the last 4 frames out of 10 consecutive images are obtained. If the target detection results for more than 2 of these 4 frames do not contain the head and shoulder detection boxes of the first object, it can be determined that the tracked first object has left the first target region.

[0064] Step S102: When it is determined that the behavior of the first object in the first target area is a drug retrieval behavior, drug retrieval behavior information is generated.

[0065] In this embodiment of the invention, the drug retrieval behavior information is used to indicate that the behavior of the first object in the first target area is a drug retrieval behavior. The drug retrieval behavior information includes a drug retrieval behavior identifier, or the drug retrieval behavior information includes a drug retrieval behavior identifier and the body feature information of the first object. The body feature information of the first object can be obtained by the first monitoring device identifying the first object. The body feature information may include facial feature information or head and shoulder feature information.

[0066] Step S103: When it is determined from the second surveillance video corresponding to the second target area that a second object has entered the second target area, determine from the second surveillance video whether the behavior of the second object in the second target area is a payment behavior.

[0067] In this embodiment of the invention, the second target area includes a drug checkout area, which is the area where customers are active when checking out near the cashier.

[0068] In one embodiment, multiple frames of images are acquired from a second surveillance video, and a second background image corresponding to the second target area is acquired. The pixel differences between each frame of the multiple frames and the second background image are calculated. If there is an image in each frame with a pixel difference greater than or equal to a preset pixel difference threshold, it is determined that a second object has entered the second target area. If the pixel difference in each frame is less than the pixel difference threshold, it is determined that no second object has entered the second target area. The second background image is the image captured by the second surveillance device when there is no second object in the second target area.

[0069] In one embodiment, multiple frames of images are acquired from a second surveillance video, and each frame is input into a preset target detection model for processing to obtain the target detection results for each frame. If each frame contains an image whose target detection result includes a head-shoulder detection box or a face detection box, it is determined that a second object has entered the second target region. Conversely, if the target detection result for each frame does not include a head-shoulder detection box or a face detection box, it is determined that no second object has entered the second target region. By using the target detection results of the images, it is possible to accurately determine whether a second object has entered the second target region, thereby improving the authenticity and accuracy of drug sales events.

[0070] In one embodiment, multiple frames of images are acquired from a second surveillance video, and a second background image corresponding to the second target region is acquired. The pixel difference between each frame of the multiple images and the second background image is calculated, and the image corresponding to the pixel difference being greater than or equal to a preset pixel difference threshold is acquired from the multiple frames to obtain the target image. The target image is then input into a preset target detection model for processing to obtain the target detection result of the target image. When the target detection result of the target image includes a head and shoulder detection box or a face detection box, it is determined that a second object has entered the second target region. When the target detection result of the target image does not include a head and shoulder detection box or a face detection box, it is determined that no second object has entered the second target region. By comprehensively considering the target detection result of the target image and the pixel difference between the target image and the background image, the detection accuracy of whether a second object has entered the second target region can be further improved, thereby improving the authenticity and accuracy of drug sales information.

[0071] In one embodiment, determining whether the behavior of the second object within the second target area constitutes a payment activity based on the second surveillance video can be achieved by: determining whether the duration of the second object's presence in the second target area exceeds a preset duration threshold based on the second surveillance video; and determining that the second object's behavior within the second target area exceeds the preset duration threshold when the duration exceeds the preset duration threshold. The preset duration threshold can be set based on actual circumstances, and this embodiment of the invention does not impose specific limitations on it. For example, the preset duration threshold could be 3 minutes.

[0072] In one embodiment, determining whether the behavior of a second object within a second target area constitutes a payment behavior based on the second surveillance video can be achieved by: inputting multiple frames of images from the second surveillance video into a payment behavior recognition model for processing to obtain action behavior labels for each frame; counting the number of images corresponding to payment behavior labels to obtain the number of target images; determining the percentage of the number of target images relative to the total number of frames in the second surveillance video; and determining that the behavior of the second object within the second target area constitutes a payment behavior when this percentage is greater than or equal to a preset percentage threshold. By comprehensively considering the action behavior labels of multiple frames, interference caused by misidentification of payment behavior labels can be reduced, thereby improving the accuracy of identifying the behavior of the second object within the second target area.

[0073] In this embodiment of the invention, the payment behavior recognition model is obtained by iteratively training a neural network model based on multiple positive training samples and multiple negative training samples. The positive training samples include positive sample images and labeled payment behavior tags. The action of the object in the positive sample image is to use an electronic device to scan a code for payment. The negative training samples include negative sample images and labeled non-payment behavior tags. The action of the object in the negative sample image does not involve using an electronic device to scan a code for payment, such as simply playing with an electronic device or taking a selfie with an electronic device.

[0074] Step S104: When it is determined that the behavior of the second object in the second target area is a payment behavior, payment behavior information is generated.

[0075] In this embodiment of the invention, payment behavior information is used to indicate that the behavior of the second object in the second target area is a payment behavior. The payment behavior information includes a payment behavior identifier, or the payment behavior information includes a payment behavior identifier and the identity feature information of the second object. The identity feature information of the second object can be obtained by the second monitoring device identifying the second object. The identity feature information may include facial feature information and head and shoulder feature information.

[0076] Step S105: When the generated payment behavior information is found within the preset time period for generating drug collection behavior information, a drug sales event is generated.

[0077] In this embodiment of the invention, the preset duration can be set based on actual conditions, and this embodiment of the invention does not impose specific limitations on it. For example, if the preset duration is 5 minutes, then if the generated payment behavior information is found within 5 minutes of generating the medication pickup behavior information, a medication sale event is generated. In other words, if payment behavior information is generated within 5 minutes after the medication pickup behavior information is generated, a medication sale event is generated.

[0078] In one embodiment, if no new drug registration record is found within a preset time period after the drug sales event is generated, a third surveillance video collected within the preset time period for the generated drug sales event is acquired. The third surveillance video includes surveillance videos obtained from a first target area and surveillance videos obtained from a second target area. The violation verification results for the third surveillance video are obtained, and based on the violation verification results, a prompt message indicating whether the store clerk has engaged in any violation is output. By identifying the suspected violation of drug sales by a pharmacy clerk when no new drug registration record is found within the preset time period after the drug sales event is generated, the violation verification results for the third surveillance video can be used to further determine whether the clerk has engaged in any violation, greatly improving the efficiency and accuracy of detecting whether the clerk has engaged in any violation of drug sales regulations.

[0079] When customers purchase prescription drugs, cough medicines, fever reducers, antiviral drugs, antibiotics, or other medications that require registration before sale, they need to scan a QR code with their mobile phones to enter their drug registration information. The mobile phone then uploads the drug registration information to the server via the internet. Upon receiving the drug registration information, the server generates and stores a drug registration record based on the received information. The drug registration information may include the customer's ID number, contact number, and the type of drug purchased.

[0080] It is understood that the violation verification result can be entered by the verification personnel through watching a third surveillance video, or it can be obtained by analyzing the third surveillance video. This embodiment of the invention does not specifically limit this. The violation verification result includes whether the employee has committed a violation or not. When the violation verification result indicates that the employee has committed a violation, a first prompt message is output to the associated electronic device to indicate that the employee has committed a violation. Conversely, when the violation verification result indicates that the employee has not committed a violation, a second prompt message is output to the associated electronic device to indicate that the employee has not committed a violation.

[0081] In one embodiment, the method for analyzing the third surveillance video to obtain the violation verification result can be as follows: if, based on the third surveillance video, it is determined that the behavior of the first object in the first target area is medication collection, and the behavior of the second object in the second target area is payment, a video segment containing the first target area is obtained from the third surveillance video; text recognition is performed on each frame of the video segment to obtain text information; if the text information contains at least one preset keyword, the violation verification result is determined to be that the store clerk has violated regulations; if the text information does not contain the preset keyword, the violation verification result is determined to be that the store clerk has not violated regulations. The preset keywords may include prescription drugs, cough medicine, antipyretics, antiviral drugs, antibiotics, etc.

[0082] In one embodiment, if payment behavior information is found within a preset time period for generating medication pickup behavior information, a first similarity is calculated between the body feature information of the first object and preset body feature information; if the first similarity is greater than or equal to a preset similarity threshold, a medication sale event is generated. The body feature information may include facial feature information and / or head and shoulder feature information, and the preset body feature information is pre-recorded pharmacist body feature information. By further determining whether a pharmacist in a pharmacy has sold medication based on the similarity between the body feature information of the first object and the preset body feature information when payment behavior information is found within the preset time period for generating medication pickup behavior information, the accuracy of determining medication sale events can be improved, thereby further enhancing the authenticity and accuracy of medication sales information in pharmacies.

[0083] In one embodiment, if payment behavior information is found within a preset time period for generating medication pickup behavior information, a first similarity is calculated between the body feature information of a first object and preset body feature information; if the first similarity is less than a preset similarity threshold, a second similarity is calculated between the body feature information of the first object and the body feature information of a second object; if the second similarity is greater than or equal to the similarity threshold, a medication sale event is generated. By determining that the object picking up the medication is a customer when the first similarity is less than the preset similarity threshold, and by determining whether the customer picking up the medication and the customer making the payment are the same person through the second similarity between the body feature information of the first object and the body feature information of the second object, it can be determined whether the customer picking up the medication and the customer making the payment are the same person. This enables customers to pick up and pay for medication themselves, and also improves the authenticity and accuracy of the pharmacy's medication sales information.

[0084] It is understood that the drug storage area may include a first drug storage area and / or a second drug storage area. The first drug storage area is used to store drugs that require registration before they can be sold, and the second area is used to store drugs that can be sold without registration. Customers cannot go to the first drug storage area to retrieve drugs, but pharmacy staff can go to the first drug storage area to retrieve or place drugs. Both customers and pharmacy staff can go to the second drug storage area to retrieve or place drugs.

[0085] For example, in a scenario where the first drug storage area is within the monitoring range of a first monitoring device, and the drug checkout area is within the monitoring range of a second monitoring device, the first monitoring device can monitor the head, shoulders, and hand positions of the pharmacy staff within the first drug storage area. This allows the first monitoring device to track the staff by their head and shoulders and to identify whether their hand movements constitute drug retrieval. When this is identified, the first monitoring device sends a drug retrieval identifier to the server. The second monitoring device can monitor whether the customer's behavior in the drug checkout area constitutes payment. When this is identified, the second monitoring device sends a payment identifier to the server. The server can then use the drug retrieval identifier from the first monitoring device and the payment identifier from the second monitoring device to determine if a drug sale event has occurred. Upon confirming a drug sale event, the server can generate and store a drug sale monitoring record.

[0086] For example, in a scenario where the second drug storage area is within the monitoring range of the first monitoring device and the drug checkout area is within the monitoring range of the second monitoring device, the first monitoring device can monitor the head, shoulders, face, and hand status of pharmacy staff or customers in the second drug storage area. In this way, the first monitoring device can track the staff through the staff's or customers' head, shoulders, or face, and identify whether the staff or customers' behavior in the second drug storage area is a drug-taking behavior through the staff's or customers' hand status. When it is identified that the pharmacy staff or customers' behavior in the second drug storage area is a drug-taking behavior, the first monitoring device sends a drug-taking behavior identifier and the staff's or customers' facial feature information to the server.

[0087] The second monitoring device can monitor whether a customer's behavior in the drug checkout area constitutes a payment. When the second monitoring device identifies a customer's behavior as a payment, it sends a payment behavior identifier and the customer's facial feature information to the server. The server can then use the drug collection behavior identifier and facial feature information sent by the first monitoring device, as well as the payment behavior identifier and facial feature information sent by the second monitoring device, to determine whether a drug sale event has occurred. Specifically, if the facial feature information sent by the first monitoring device matches the pre-entered facial feature information of the store clerk, it can be determined that the person collecting the drug is a store clerk. If the facial feature information sent by the first monitoring device matches the facial feature information sent by the second monitoring device, it can be determined that the person collecting the drug is a customer. When it is determined that the person collecting the drug is either a customer or a store clerk, a drug sale event can be confirmed, and the server generates and stores a drug sale monitoring record.

[0088] It is understood that steps S101 and S102 can be executed by the first monitoring device, steps S103 and S104 can be executed by the second monitoring device, and step S105 can be executed by the server. Alternatively, the first monitoring device sends the first monitoring video to the server, the second monitoring device sends the second monitoring video to the server, and steps S101 to S105 are all executed by the server. This embodiment of the invention does not specifically limit the executing entity for steps S101 to S105.

[0089] The drug sales monitoring method provided in the above embodiments monitors whether the behavior of a first person entering the drug storage area constitutes drug retrieval when such behavior is detected. If the first person's behavior is confirmed to be drug retrieval, drug retrieval information is generated. Simultaneously, when a second person enters the drug checkout area, the method monitors whether their behavior is confirmed to be payment. If the second person's behavior is confirmed to be payment, payment information is generated. Thus, when payment information is retrieved within a preset timeframe of the generated drug retrieval information, a drug sales event is generated. This prevents pharmacy staff from falsifying drug sales, greatly improving the authenticity and accuracy of drug sales data.

[0090] Please see Figure 8 , Figure 8 This is a schematic block diagram of a drug sales monitoring system provided in an embodiment of the present invention.

[0091] like Figure 8 As shown, the drug sales monitoring system 400 includes a first monitoring device 410, a second monitoring device 420, and a server 430. The first monitoring device 410 and the second monitoring device 420 are respectively connected to the server 430.

[0092] In one embodiment, the first monitoring device 410 is used to monitor a first target area and obtain a first monitoring video, the first target area including a medicine storage area; when it is determined from the first monitoring video that a first object has entered the first target area, it is determined from the first monitoring video whether the behavior of the first object in the first target area is a medicine retrieval behavior; when it is determined that the behavior of the first object in the first target area is a medicine retrieval behavior, medicine retrieval behavior information is generated and sent to the server, the medicine retrieval behavior information being used to indicate that the behavior of the first object in the first target area is a medicine retrieval behavior;

[0093] The second monitoring device 420 is used to monitor a second target area and obtain a second monitoring video. The second target area includes a drug checkout area. When it is determined from the second monitoring video that a second object has entered the second target area, it is determined from the second monitoring video whether the behavior of the second object in the second target area is a payment behavior. When it is determined that the behavior of the second object in the second target area is a payment behavior, payment behavior information is generated and sent to the server. The payment behavior information is used to indicate that the behavior of the second object in the drug checkout area is a payment behavior.

[0094] The server 430 is configured to determine, upon receiving the medication collection information, whether it has received the payment information sent by the second monitoring device within a preset time period; and to generate a medication sales event upon receiving the payment information sent by the second monitoring device within the preset time period.

[0095] In one embodiment, the first monitoring device 410 is used to monitor a first target area, obtain a first monitoring video, and send the first monitoring video to a server. The first target area includes a medicine storage area.

[0096] The second monitoring device 420 is used to monitor the second target area, obtain the second monitoring video, and send the second monitoring video to the server. The second target area includes the drug checkout area.

[0097] The server 430 is configured to acquire a first monitoring video sent by the first monitoring device and a second monitoring video sent by the second monitoring device; when it is determined from the first monitoring video that a first object has entered the first target area, it determines whether the behavior of the first object in the first target area is a medication retrieval behavior; when it is determined that the behavior of the first object in the first target area is a medication retrieval behavior, it generates medication retrieval behavior information, which indicates that the behavior of the first object in the first target area is a medication retrieval behavior; when it is determined from the second monitoring video that a second object has entered the second target area, it determines whether the behavior of the second object in the second target area is a payment behavior; when it is determined that the behavior of the second object in the second target area is a payment behavior, it generates payment behavior information, which indicates that the behavior of the second object in the second target area is a payment behavior; and when the generated payment behavior information is found within a preset time period for generating the medication retrieval behavior information, a drug sales event is generated.

[0098] In one embodiment, the second monitoring device 420 or the server 430 is further configured to determine, based on the second monitoring video, whether the duration of the second object in the second target area is greater than a preset duration threshold; and when the duration is determined to be greater than the preset duration threshold, determine that the behavior of the second object in the second target area is a payment behavior.

[0099] In one embodiment, the server 430 is further configured to, when no new drug registration record is found within a preset time period after the drug sales event is generated, acquire a third monitoring video collected within a preset time period for the generated drug sales event, the third monitoring video including monitoring video obtained from monitoring the first target area and monitoring video obtained from monitoring the second target area; acquire violation verification results for the third monitoring video, and output a prompt message indicating whether the store clerk has engaged in any violation based on the violation verification results.

[0100] In one embodiment, the first monitoring device 410 or the server 430 is further configured to input multiple frames of images from the first monitoring video into a preset target detection model for processing to obtain target detection results for each frame of the multiple frames; track the first object based on the target detection results for each frame of the images; and when the first object leaves the first target area, determine whether the behavior of the first object in the first target area is a drug-taking behavior based on the target detection results for each frame of the images.

[0101] In one embodiment, the multi-frame images include at least a first image and a second image adjacent to the first image. The first monitoring device 410 or the server 430 is further configured to calculate the intersection-over-union ratio (IoU) between the head and shoulder detection bounding boxes of the first object in the first image and the head and shoulder detection bounding boxes of each object to be detected in the second image; determine the largest IoU among the plurality of IoU ratios as the target IoU ratio, and determine whether the target IoU ratio is greater than or equal to a preset IoU ratio threshold; when the target IoU ratio is greater than or equal to the preset IoU ratio threshold, mark the object to be detected in the second image corresponding to the target IoU ratio as the first object.

[0102] In one embodiment, the target detection result includes a hand status tag of the first object, the first monitoring device 410 or the server 430, and is further used to count the number of images corresponding to the hand status tag being a drug placement status tag, to obtain a first image count, wherein the drug placement status tag describes that the first object's hand is in a drug placement state; determine the percentage of the first image count to the total number of the multi-frame images; and when the percentage is greater than or equal to a preset percentage threshold, determine that the first object's behavior in the first target area is a drug retrieval behavior.

[0103] In one embodiment, the target detection result includes a hand status tag of the first object, the first monitoring device 410 or the server 430, and is further used to determine a first target image, the first target image being the image in the multi-frame image where the hand status tag is the first detected one corresponding to a drug placement status tag, the drug placement status tag describing that the first object's hand is in a drug placement state; counting the number of second target images where the hand status tag is the drug placement status tag, the second target image being the image in the multi-frame image located after the first target image; incrementing the number of second target images by 1 to obtain the number of second images; when the number of second images is greater than or equal to a preset image number threshold, determining that the first object's behavior in the first target area is a drug retrieval behavior.

[0104] In one embodiment, the drug-taking behavior information includes a drug-taking behavior identifier and the body feature information of the first object. The body feature information includes at least one of facial feature information and head and shoulder feature information. The server 430 is further configured to calculate a first similarity between the body feature information of the first object and preset body feature information; when the first similarity is greater than or equal to a preset similarity threshold, a drug sales event is generated.

[0105] In one embodiment, the payment behavior information includes a payment behavior identifier and the body feature information of the second object. The server 430 is further configured to calculate a second similarity between the body feature information of the first object and the body feature information of the second object when the first similarity is less than a preset similarity threshold; and to generate a drug sales event when the second similarity is greater than or equal to the similarity threshold.

[0106] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the drug sales monitoring system described above can be referred to the corresponding process in the aforementioned drug sales monitoring method embodiments, and will not be repeated here.

[0107] This invention also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs, which can be executed by one or more processors to implement any of the drug sales monitoring methods provided in the specification of this invention.

[0108] The storage medium can be the internal storage unit of the first monitoring device, the second monitoring device, or the server described in the foregoing embodiments, such as the hard drive or memory of the first monitoring device, the second monitoring device, or the server. The storage medium can also be an external storage device of the first monitoring device, the second monitoring device, or the server, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, or a Flash Card.

[0109] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware embodiments, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0110] It should be understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. It should be noted that, herein, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0111] The sequence numbers of the above embodiments of the present invention are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The above descriptions are only specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for monitoring drug sales, characterized in that, include: When it is determined from the first surveillance video corresponding to the first target area that a first object has entered the first target area, it is determined from the first surveillance video whether the behavior of the first object in the first target area is a drug picking behavior. The first target area is used to store drugs that need to be registered before they can be sold. When it is determined that the behavior of the first object in the first target area is a drug retrieval behavior, drug retrieval behavior information is generated, which is used to indicate that the behavior of the first object in the first target area is a drug retrieval behavior; When it is determined from the second surveillance video corresponding to the second target area that a second object has entered the second target area, it is determined from the second surveillance video whether the behavior of the second object in the second target area is a payment behavior, and the second target area includes the drug checkout area; When it is determined that the behavior of the second object in the second target area is a payment behavior, payment behavior information is generated, which is used to indicate that the behavior of the second object in the second target area is a payment behavior; When the generated payment behavior information is found within a preset time period after the generation of the medication collection behavior information, a medication sale event is generated. If no new drug registration record is found within a preset time period after the drug sales event is generated, a third monitoring video collected within a preset time period for the generated drug sales event is obtained. The third monitoring video includes the monitoring video obtained by monitoring the first target area and the monitoring video obtained by monitoring the second target area. Obtain the violation verification results for the third surveillance video, and based on the violation verification results, output a prompt message indicating whether the store employee has engaged in any violation; The step of obtaining the violation verification result for the third surveillance video includes: if, based on the third surveillance video, it is determined that the behavior of the first object in the first target area is medication collection and the behavior of the second object in the second target area is payment, then a video segment containing the first target area is obtained from the third surveillance video; text recognition is performed on each frame of the video segment to obtain text information; if the text information contains at least one preset keyword, the violation verification result is determined to be that the store clerk has violated regulations; if the text information does not contain the preset keyword, the violation verification result is determined to be that the store clerk has not violated regulations.

2. The drug sales monitoring method according to claim 1, characterized in that, The step of determining whether the behavior of the second object in the second target area constitutes a payment behavior based on the second surveillance video includes: Based on the second surveillance video, determine whether the duration for which the second object is located in the second target area is greater than a preset duration threshold; When the duration is determined to be greater than a preset duration threshold, the behavior of the second object within the second target area is determined to be a payment behavior.

3. The drug sales monitoring method according to claim 1 or 2, characterized in that, The step of determining whether the behavior of the first object in the first target area constitutes medication collection based on the first surveillance video includes: The first surveillance video is input into a preset target detection model for processing to obtain the target detection results of each frame of the multi-frame image. Based on the target detection results of each frame of the image, the first object is tracked; When the first object is detected leaving the first target area, the target detection results of each frame of the image are used to determine whether the first object’s behavior in the first target area is a drug-taking behavior.

4. The drug sales monitoring method according to claim 3, characterized in that, The multi-frame images include at least a first image and a second image adjacent to the first image. Tracking the first object based on the target detection results of each frame of the images includes: Calculate the intersection-over-union ratio (IoU) between the head and shoulder detection bounding boxes of the first object in the first image and the head and shoulder detection bounding boxes of each object to be detected in the second image; The largest cross-union ratio among the multiple cross-union ratios is determined as the target cross-union ratio, and it is determined whether the target cross-union ratio is greater than or equal to a preset cross-union ratio threshold. When the target cross-union ratio is greater than or equal to a preset cross-union ratio threshold, the object to be detected in the second image corresponding to the target cross-union ratio is marked as the first object.

5. The drug sales monitoring method according to claim 3, characterized in that, The target detection result includes the hand state label of the first object. Determining whether the first object's behavior within the first target area constitutes medication-taking behavior based on the target detection results of each frame of the image includes: The number of images corresponding to the hand status label and the drug placement status label is counted to obtain the first image count. The drug placement status label describes that the hand of the first object is in the drug placement state. Determine the percentage of the first image count relative to the total number of the multi-frame images; When the percentage is greater than or equal to a preset percentage threshold, the behavior of the first object within the first target area is determined to be a drug-taking behavior.

6. The drug sales monitoring method according to claim 3, characterized in that, The target detection result includes the hand state label of the first object. Determining whether the first object's behavior within the first target area constitutes medication-taking behavior based on the target detection results of each frame of the image includes: A first target image is determined, which is the image in the multi-frame image where the first detected hand state label is the drug placement state label, and the drug placement state label describes that the hand of the first object is in a drug placement state; The number of second target images corresponding to the hand status label and the drug placement status label is counted, where the second target image is the image located after the first target image in the multi-frame image; Add 1 to the number of the second target images to get the number of the second images; When the number of the second images is greater than or equal to a preset image number threshold, the behavior of the first object in the first target area is determined to be a drug-taking behavior.

7. The drug sales monitoring method according to claim 1 or 2, characterized in that, The medication dispensing behavior information includes a medication dispensing behavior identifier and the body characteristic information of the first object. The body characteristic information includes at least one of facial feature information and head and shoulder feature information. The medication sales monitoring method further includes: Calculate the first similarity between the body feature information of the first object and the preset body feature information; When the first similarity is greater than or equal to a preset similarity threshold, a drug sales event is generated.

8. The drug sales monitoring method according to claim 7, characterized in that, The payment behavior information includes a payment behavior identifier and the body feature information of the second object. After calculating the first similarity between the body feature information of the first object and the preset body feature information, the method further includes: When the first similarity is less than a preset similarity threshold, a second similarity is calculated between the body feature information of the first object and the body feature information of the second object; When the second similarity is greater than or equal to the similarity threshold, a drug sale event is generated.

9. A drug sales monitoring system, characterized in that, The drug sales monitoring system includes a first monitoring device, a second monitoring device, and a server, wherein the first monitoring device and the second monitoring device are respectively communicatively connected to the server; The first monitoring device is used to monitor the first target area and obtain the first monitoring video. The first target area is used to store medicines that need to be registered before they can be sold. When it is determined from the first surveillance video that a first object has entered the first target area, it is determined from the first surveillance video whether the behavior of the first object in the first target area is a drug-taking behavior. When it is determined that the behavior of the first object in the first target area is a drug retrieval behavior, drug retrieval behavior information is generated and sent to the server. The drug retrieval behavior information is used to indicate that the behavior of the first object in the first target area is a drug retrieval behavior. The second monitoring device is used to monitor a second target area and obtain a second monitoring video. The second target area includes a drug checkout area. When it is determined from the second monitoring video that a second object has entered the second target area, the device determines whether the behavior of the second object in the second target area constitutes a payment behavior. When it is determined that the behavior of the second object in the second target area is a payment behavior, payment behavior information is generated and sent to the server. The payment behavior information indicates that the behavior of the second object in the drug checkout area is a payment behavior. The server is used to determine, upon receiving the medication collection information, whether it has received the payment information sent by the second monitoring device within a preset time period. When the payment behavior information sent by the second monitoring device is received within a preset time period, a drug sales event is generated; The server is also used to acquire a third monitoring video collected within a preset time period for the generated drug sales event when no new drug registration record is found within a preset time period after the drug sales event is generated. The third monitoring video includes monitoring video obtained from monitoring the first target area and monitoring video obtained from monitoring the second target area. Obtain the violation verification result for the third surveillance video, and based on the violation verification result, output a prompt message indicating whether the store clerk has engaged in any violation. Obtaining the violation verification result for the third surveillance video includes: if, based on the third surveillance video, the behavior of the first object in the first target area is determined to be medication collection, and the behavior of the second object in the second target area is determined to be payment, obtain a video segment containing the first target area from the third surveillance video; perform text recognition on each frame of the video segment to obtain text information; if the text information contains at least one preset keyword, determine that the violation verification result indicates the store clerk has engaged in any violation; if the text information does not contain the preset keyword, determine that the violation verification result indicates the store clerk has not engaged in any violation.

10. A drug sales monitoring system, characterized in that, The drug sales monitoring system includes a first monitoring device, a second monitoring device, and a server, wherein the first monitoring device and the second monitoring device are respectively communicatively connected to the server; The first monitoring device is used to monitor the first target area, obtain the first monitoring video, and send the first monitoring video to the server. The first target area is used to store medicines that need to be registered before they can be sold. The second monitoring device is used to monitor the second target area, obtain the second monitoring video, and send the second monitoring video to the server. The second target area includes the drug checkout area. The server is used to acquire the first monitoring video sent by the first monitoring device and the second monitoring video sent by the second monitoring device. When it is determined from the first surveillance video that a first object has entered the first target area, it is determined from the first surveillance video whether the behavior of the first object in the first target area is a drug-taking behavior. When it is determined that the behavior of the first object in the first target area is a drug retrieval behavior, drug retrieval behavior information is generated, which is used to indicate that the behavior of the first object in the first target area is a drug retrieval behavior; When it is determined from the second surveillance video that a second object has entered the second target area, it is determined from the second surveillance video whether the behavior of the second object in the second target area is a payment behavior; when it is determined that the behavior of the second object in the second target area is a payment behavior, payment behavior information is generated, which is used to indicate that the behavior of the second object in the second target area is a payment behavior; when the generated payment behavior information is found within a preset time period for generating the drug collection behavior information, a drug sales event is generated; If no new drug registration record is found within a preset time period after the drug sales event is generated, a third monitoring video collected within a preset time period for the generated drug sales event is obtained. The third monitoring video includes the monitoring video obtained by monitoring the first target area and the monitoring video obtained by monitoring the second target area. Obtain the violation verification result for the third surveillance video, and based on the violation verification result, output a prompt message indicating whether the store clerk has engaged in any violation. Obtaining the violation verification result for the third surveillance video includes: if, based on the third surveillance video, the behavior of the first object in the first target area is determined to be medication collection, and the behavior of the second object in the second target area is determined to be payment, obtain a video segment containing the first target area from the third surveillance video; perform text recognition on each frame of the video segment to obtain text information; if the text information contains at least one preset keyword, determine that the violation verification result indicates the store clerk has engaged in any violation; if the text information does not contain the preset keyword, determine that the violation verification result indicates the store clerk has not engaged in any violation.

Citation Information

Patent Citations

  • Detection method and device for user operation

    CN108921081A

  • Unmanned shelf payment detection method, device and system

    CN110647783A

  • Pedestrian crossing road guardrail detection method and device and storage medium

    CN112434627A