Method, system, device, storage medium and program product for identifying behavior information
By identifying the packaging posture, handheld items behavior and bag information in the surveillance video frame in the supermarket, the problem of difficulty in detecting packaging behavior in time is solved, and efficient identification and reduction of misjudgment is achieved.
Patent Information
- Application Number
- CN202411621282.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-11-14
AI Technical Summary
In supermarkets, it is difficult for the loss prevention officer to discover and stop packaging in a timely and accurately manner, resulting in loss of goods.
By identifying the packaging posture of the human object in the monitoring video frame, the behavior of the handheld item, and the preset posture of the target packaging information, we judge whether the packaging behavior occurs and issue a prompt message.
It improves the efficiency of identification of packaging behavior, reduces the possibility of misjudgment, ensures that packaging behavior is stopped in a timely manner, and reduces product losses.
Smart Images

Figure CN119152579B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information processing technologies, and in particular, to a method, system, device, storage medium, and program product for identifying behavior information. Background Art
[0002] With the development of retail business, supermarkets no longer require customers to store their bags. Instead, customers put the goods into their own bags (hereinafter referred to as "bagging behavior") and take them out of the supermarket, resulting in losses of goods. To detect and prevent bagging behavior in a timely manner, supermarkets generally have loss prevention staff who manually patrol the supermarket to check for possible bagging behavior.
[0003] However, the entire process of bagging behavior generally takes a short time. Supermarkets are large in area and have a large number of people, making it difficult for loss prevention staff to detect and prevent bagging behavior in a timely and accurate manner, resulting in low inspection efficiency. Summary of the Invention
[0004] The main objective of the embodiments of this application is to provide a method, system, device, storage medium, and program product for identifying behavior information, which can accurately determine whether a human object has committed a bagging behavior by combining the bagging posture, the behavior of holding an object, and the preset posture corresponding to the target bag information, reducing the possibility of misjudgment and improving the recognition efficiency of bagging behavior.
[0005] In a first aspect, the embodiments of this application provide a method for identifying behavior information, including: obtaining a monitoring video frame within a specified area; identifying target bag information existing in the human object area in the monitoring video frame; detecting the bagging posture and the behavior of holding an object committed by the human object; if the bagging posture is the preset posture corresponding to the target bag information, and the human object has the behavior of holding an object before the bagging posture occurs, determining that the human object has committed a bagging behavior within the specified area and sending a prompt message.
[0006] In one embodiment, the identifying target bag information existing in the human object area in the monitoring video frame includes: identifying the human trajectory information in the monitoring video frame; intercepting a human image from the monitoring video frame according to the human trajectory information; inputting the human image into a preset bag detection model so that the preset bag detection model outputs the target bag information existing in the human image, and the target bag information includes the bag position and / or the bag type.
[0007] In one embodiment, before inputting the human image into the preset bag detection model, it further includes: obtaining a bag sample image set, where the positions and types of different bag objects are marked in the bag sample images; training a preset detector with the bag sample image set to obtain the preset bag detection model.
[0008] In one embodiment, detecting the bag-packing posture of the human object includes: identifying the human trajectory information in the monitoring video frame; intercepting a human image from the monitoring video frame according to the human trajectory information; inputting the human image into a preset posture detection model so that the preset posture detection model outputs the bag-packing postures existing in the human image, the types of bag bags for which the bag-packing postures occur, and the occurrence time of the bag-packing postures.
[0009] In one embodiment, before inputting the human image into the preset posture detection model, it further includes: obtaining a bag-packing sample image set, in which various bag-packing postures and non-bag-packing postures corresponding to different target bag bag information are marked; training a preset detector with the bag-packing sample image set to obtain the preset posture detection model.
[0010] In one embodiment, detecting the behavior of the human object holding an item includes: identifying the human trajectory information in the monitoring video frame; intercepting a human image from the monitoring video frame according to the human trajectory information; inputting the human image into a preset behavior detection model so that the preset behavior detection model outputs the hand of the human object in the human image and the target item held by the hand; if the distance between the position of the target item and the position of the hand of the human object is less than a preset threshold, determining that the human object has performed the behavior of holding an item, and determining the occurrence time of the behavior of holding an item.
[0011] In one embodiment, before inputting the human image into the preset behavior detection model, it further includes: obtaining a sample image set of holding an item, in which items with different external features are marked; training a preset detector with the sample image set of holding an item to obtain the preset behavior detection model.
[0012] In one embodiment, the types of bag bags include one or more of a backpack carried on both shoulders of a human, a single-shoulder bag with the bag bag above the waist of the human, a single-shoulder bag with the bag bag below the waist of the human, a bag carried on the forearm of the human, and a bag carried by the hand of the human.
[0013] In a second aspect, an embodiment of the present application provides a supermarket anti-theft detection method, including: obtaining a monitoring video frame in a supermarket; identifying target bag bag information existing in a human object area in the monitoring video frame; detecting the bag-packing posture and the behavior of holding a commodity of the human object; if the bag-packing posture is a preset posture corresponding to the target bag bag information, and the human object has the behavior of holding a commodity before the occurrence of the bag-packing posture, determining that the human object has performed the behavior of packing the commodity in the supermarket, and sending out an anti-theft prompt message.
[0014] In a third aspect, an apparatus for identifying behavior information provided by an embodiment of the present application includes:
[0015] An acquisition module, configured to acquire surveillance video frames within a specified area;
[0016] An identification module, configured to identify target bag information existing in a human object area in the surveillance video frames;
[0017] A detection module, configured to detect a bagging posture and a behavior of holding an item performed by the human object;
[0018] A determination module, configured to determine that the human object performs a bagging behavior in the specified area and issue a prompt message if the bagging posture is a preset posture corresponding to the target bag information and the human object has a behavior of holding an item before the bagging posture occurs.
[0019] In an embodiment, the identification module is configured to identify human trajectory information in the surveillance video frames; intercept a human body image from the surveillance video frames according to the human trajectory information; and input the human body image into a preset bag detection model, so that the preset bag detection model outputs the target bag information existing in the human body image, where the target bag information includes a bag position and / or a bag type.
[0020] In an embodiment, it further includes: a first training module, configured to, before inputting the human body image into the preset bag detection model, acquire a bag sample image set, where positions and types of different bag objects are marked in the bag sample images; and train a preset detector by using the bag sample image set to obtain the preset bag detection model.
[0021] In an embodiment, the detection module is configured to identify human trajectory information in the surveillance video frames; intercept a human body image from the surveillance video frames according to the human trajectory information; and input the human body image into a preset posture detection model, so that the preset posture detection model outputs a bagging posture existing in the human body image, a bag type for which the bagging posture occurs, and a time when the bagging posture occurs.
[0022] In an embodiment, it further includes: a second training module, configured to, before inputting the human body image into the preset posture detection model, acquire a bagging sample image set, where multiple bagging postures and non-bagging postures corresponding to different target bag information are marked in the bagging sample images; and train a preset detector by using the bagging sample image set to obtain the preset posture detection model.
[0023] In one embodiment, a detection module is configured to identify human trajectory information in the monitored video frame; intercept a human image from the monitored video frame according to the human trajectory information; input the human image into a preset behavior detection model so that the preset behavior detection model outputs the hand of the human object in the human image and the target item held by the hand; if the distance between the position of the target item and the position of the hand of the human object is less than a preset threshold, determine that the human object has performed the behavior of holding an item, and determine the occurrence time of the behavior of holding an item.
[0024] In one embodiment, it further includes: a third training module, configured to, before inputting the human image into the preset behavior detection model, obtain a sample image set of holding an item, where different-shaped items are marked in the sample image of holding an item; train a preset detector with the sample image set of holding an item to obtain the preset behavior detection model.
[0025] In one embodiment, the types of bags include one or more of a backpack carried on both shoulders of a human, a single-shoulder bag with the bag above the waist of the human, a single-shoulder bag with the bag below the waist of the human, a bag carried on the forearm of the human, and a bag carried by the hand of the human.
[0026] In a fourth aspect, an embodiment of the present application provides a recognition system for behavior information, including:
[0027] A bag detection module, configured to obtain a monitored video frame in a specified area and identify target bag information existing in the human object area in the monitored video frame;
[0028] A hand-held item detection module, configured to detect the behavior of a human object holding an item;
[0029] A bag-packing posture detection module, configured to detect the bag-packing posture of the human object;
[0030] A bag-packing behavior decision module, respectively connected to the bag detection module, the hand-held item detection module, and the bag-packing posture detection module, configured to, if the bag-packing posture is a preset posture corresponding to the target bag information and the human object has the behavior of holding an item before the bag-packing posture occurs, determine that the human object has performed a bag-packing behavior in the specified area and send a prompt message.
[0031] In a fifth aspect, an embodiment of the present application provides an electronic device, including:
[0032] At least one processor; and
[0033] A memory communicatively connected to the at least one processor;
[0034] Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the electronic device to execute the method described in any of the above aspects.
[0035] In a sixth aspect, an embodiment of the present application provides a cloud device, including:
[0036] At least one processor; and
[0037] A memory communicatively connected to the at least one processor;
[0038] Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the cloud device to execute the method described in any of the above aspects.
[0039] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when a processor executes the computer-executable instructions, the method described in any of the above aspects is implemented.
[0040] In an eighth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any of the above aspects is implemented.
[0041] The method, system, device, storage medium and program product for identifying behavior information provided by the embodiments of the present application identify the target bag information in the human object area of the monitored video frame, as well as the bag-packing posture and the behavior of holding an item by the human object, and match the bag-packing posture of the human object with the target bag information. If the bag-packing posture is the preset posture corresponding to the target bag information, and the human object has the behavior of holding an item before the bag-packing posture occurs, it is determined that the human object has performed a bag-packing behavior in the specified area. At this time, a prompt message can be sent to timely stop the bag-packing behavior. In this way, by combining the bag-packing posture, the behavior of holding an item, and the preset posture corresponding to the target bag information, it is possible to accurately determine whether the human object has performed a bag-packing behavior, reduce the possibility of misjudgment, and improve the recognition efficiency of the bag-packing behavior. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0043] Figure 1A schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0044] Figure 2 A schematic diagram of an application scenario of an identification system for behavior information provided by an embodiment of the present application;
[0045] Figure 3 A schematic framework diagram of an identification system for behavior information provided by an embodiment of the present application;
[0046] Figure 4 A schematic flowchart of a method for identifying behavior information provided by an embodiment of the present application;
[0047] Figure 5 A schematic diagram of the packing posture of a single-shoulder bag provided by an embodiment of the present application;
[0048] Figure 6 A schematic flowchart of a method for identifying behavior information provided by an embodiment of the present application;
[0049] Figure 7 A schematic flowchart of a supermarket anti-theft detection method provided by an embodiment of the present application;
[0050] Figure 8 A schematic structural diagram of an identification device for behavior information provided by an embodiment of the present application;
[0051] Figure 9 A schematic structural diagram of a cloud device provided by an embodiment of the present application.
[0052] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0053] Here, exemplary embodiments will be described in detail, and their examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application.
[0054] The terms "first", "second", "third", "fourth", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances. For example, without departing from the scope of this article, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information.
[0055] Depending on the context, as used herein, the word "if" may be interpreted as "when" or "while" or "in response to determining".
[0056] Furthermore, as used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise.
[0057] It should be further understood that the terms "comprising", "including" indicate the presence of features, steps, operations, elements, components, items, kinds, and / or groups, but do not preclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, kinds, and / or groups.
[0058] The term "and / or" herein is used to describe the associated relationship of associated objects and specifically represents three possible relationships. For example, A and / or B can represent: A exists alone, both A and B exist simultaneously, and B exists alone.
[0059] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.
[0060] To clearly describe the technical solutions of the embodiments of this application, the nouns involved in this application are first defined as follows:
[0061] Bag: Various bags that can hold items, including but not limited to backpacks, shoulder bags, cross-body bags, briefcases, canvas bags, plastic bags, paper bags, etc.
[0062] Bagging behavior: The act of putting items into a self-owned bag, such as putting supermarket goods into a self-owned bag.
[0063] POS: Point of Sales, generally referring to a cashier terminal device. In the embodiments of this application, it may refer to a self-service scanning and cashier integrated device provided by a supermarket for customers.
[0064] ID: Identification, a unique identifier.
[0065] Yolov5: A deep learning model for object detection, belonging to the "You Only Look Once" (YOLO) series.
[0066] ResNet: Residual Network, also known as Residual Network.
[0067] Resnet-18: It is a deep convolutional neural network architecture and a member of the ResNet family.
[0068] The watermark information embedding method of the embodiments of the present application can be applied to any field involving watermark processing, such as scenarios like AIGC model distribution traceability and AIGC synthetic map marking protection.
[0069] With the development of the retail business, supermarkets no longer require customers to store their bags. Instead, customers put the goods into their own bags (hereinafter referred to as "bagging behavior") and take them out of the supermarket, resulting in losses of goods. To detect and stop the bagging behavior in a timely manner, supermarkets generally have loss prevention staff who check for possible bagging behavior in the supermarket through manual patrols.
[0070] However, the entire process of the bagging behavior generally takes a short time. The supermarket is large and crowded with people, making it difficult for loss prevention staff to detect and stop the bagging behavior in a timely and accurate manner, resulting in low inspection efficiency.
[0071] In related technologies, for the problem of supermarket loss prevention, many manufacturers have developed products in fields such as self-checkout machines and unmanned supermarkets. For example, for the needs of supermarket loss prevention, there are generally solutions such as loss prevention at the POS machine end and loss prevention based on pick-and-place recognition. Among them, loss prevention at the POS end is for self-checkout equipment. However, people who bag goods in the supermarket may not pass through the checkout area and directly leave through the non-shopping passage, resulting in the bagged goods not being taken out in front of the self-checkout machine. In such cases, the self-checkout machine cannot detect the risk. Therefore, it can be complementary to the solution of the embodiments of the present application. The solution of loss prevention based on pick-and-place recognition requires a long link. For example, it is necessary to confirm that the customer takes the goods from the shelf and compare them with the goods taken when checking out, involving multiple links such as pick-and-place recognition, commodity recognition, personnel tracking, and connection with billing data, which is difficult to implement or has a high cost.
[0072] In an optional embodiment, for the bagging recognition method, the bagging behavior can be recognized based on a video-based action recognition method. However, this method requires a larger computing device to support the calculation of inter-frame information. In addition, there are complexities in obtaining training samples and the generalization ability of the model.
[0073] To solve at least one of the above problems, an embodiment of the present application provides a recognition scheme for behavior information. By recognizing the target bag information in the human object area of the surveillance video frame, as well as the bag-packing posture and the behavior of holding an item by the human object, and matching the bag-packing posture of the human object with the target bag information, if the bag-packing posture is the preset posture corresponding to the target bag information, and there is a behavior of holding an item by the human object before the bag-packing posture occurs, it is determined that the human object has performed a bag-packing behavior in the specified area. At this time, a prompt message can be sent to timely stop the bag-packing behavior. In this way, by combining the bag-packing posture, the behavior of holding an item, and the preset posture corresponding to the target bag information, it is possible to accurately determine whether the human object has performed a bag-packing behavior, reduce the possibility of misjudgment, and improve the recognition efficiency of the bag-packing behavior.
[0074] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict between the embodiments, the embodiments and the features in the embodiments can be combined with each other. In addition, the step sequence in the following method embodiments is only an example and is not strictly limited.
[0075] As Figure 1 shown, this embodiment provides an electronic device 1, including: at least one processor 11 and a memory 12. Figure 1 Taking one processor as an example. The processor 11 and the memory 12 are connected through a bus 10. The memory 12 stores instructions executable by the processor 11. The instructions are executed by the processor 11 so that the electronic device 1 can execute all or part of the processes of the methods in the following embodiments, so as to realize that by combining the bag-packing posture, the behavior of holding an item, and the preset posture corresponding to the target bag information, it is possible to accurately determine whether the human object has performed a bag-packing behavior, reduce the possibility of misjudgment, and improve the recognition efficiency of the bag-packing behavior.
[0076] In one embodiment, the electronic device 1 may be a mobile phone, a tablet computer, a notebook computer, a desktop computer, or a large computing system composed of multiple computers.
[0077] Figure 2 This is a schematic diagram of an application scenario 200 of a recognition system for behavior information provided by an embodiment of the present application. As Figure 2 shown, the system includes: a server 210 and a terminal 220, where:
[0078] The server 210 may be a data platform that provides a recognition service for behavior information, such as a surveillance service platform of a supermarket. In an actual scenario, a surveillance service platform may have multiple servers 210. Figure 2 Taking 1 server 210 as an example.
[0079] The terminal 220 can be devices such as a computer, a mobile phone, or a tablet used by a user to log in to the monitoring service platform. There can also be multiple terminals 220. Figure 2 For illustration, two terminals 220 are taken as an example.
[0080] Information can be transmitted between the terminal 220 and the server 210 via the Internet so that the terminal 220 can access the data on the server 210. The above-mentioned terminal 220 and / or server 210 can both be implemented by the electronic device 1.
[0081] The recognition scheme for behavior information in the embodiments of this application can be deployed on the server 210, or on the terminal 220, or partially on the server 210 and partially on the terminal 220. In the actual scenario, it can be selected based on actual needs, and this embodiment does not make a limitation.
[0082] When the recognition scheme for behavior information is fully or partially deployed on the server 210, a call interface can be opened to the terminal 220 to provide algorithm support for the terminal 220.
[0083] The method provided in the embodiments of this application can be implemented by the electronic device 1 executing corresponding software code and realizing it through data interaction with the server. Among them, the electronic device 1 can be a local terminal device. When this method runs on the server, this method can be implemented and executed based on a cloud interaction system, where the cloud interaction system includes a server and a client device.
[0084] In a possible implementation manner, the method provided in the embodiments of this application provides a graphical user interface through a terminal device, where the terminal device can be the aforementioned local terminal device or the client device in the aforementioned cloud interaction system.
[0085] As Figure 3 shown, it is a framework schematic diagram of a recognition system for behavior information in an embodiment of this application. The system includes a bag detection module, a hand-held item detection module, a bag-packing posture detection module, and a bag-packing behavior decision module, where:
[0086] The bag detection module is used to obtain monitoring video frames in a specified area and identify the target bag information existing in the human object area in the monitoring video frames.
[0087] The hand-held item detection module is used to detect the hand-held item behavior of the human object.
[0088] The bag-packing posture detection module is used to detect the bag-packing posture of the human object.
[0089] The packing behavior decision-making module is respectively connected to the bag detection module, the hand-held item detection module, and the packing posture detection module, and is used to determine that the human object performs a packing behavior in the specified area and send out a prompt message if the packing posture is the preset posture corresponding to the target bag information and the human object has a hand-held item behavior before the packing posture occurs.
[0090] In the system according to the embodiment of the present application, by identifying the human body trajectory in the monitoring video frame, the target bag information in the human object area is determined by the bag detection module according to the human body trajectory, the packing posture of the human object is recognized by the packing posture detection module according to the monitoring video frame, and the hand-held item behavior of the human object is recognized by the hand-held item detection module according to the monitoring video frame. Then, the packing behavior decision-making module matches the packing posture of the human object with the target bag information. If the packing posture is the preset posture corresponding to the target bag information and the human object has a hand-held item behavior before the packing posture occurs, it is determined that the human object performs a packing behavior in the specified area. At this time, a prompt message can be sent out, such as outputting the packing behavior decision conclusion, the information of the person packing, or the video segment where the packing behavior occurs, so as to stop the packing behavior in time. In this way, by combining the packing posture, the hand-held item behavior, and the preset posture corresponding to the target bag information, it is possible to accurately determine whether the human object performs a packing behavior, reduce the possibility of misjudgment, and improve the recognition efficiency of the packing behavior.
[0091] Optionally, the system may further include a human body trajectory extraction module, which is connected to the bag detection module and is used to extract human body trajectory information from the monitoring video frame to form a structured description of the human object. Specifically, for a captured frame of an image, the positions of all human objects are detected from the image using a human body detection algorithm, and then the detected human objects are associated with the same human objects in the previous frame using a human body tracking algorithm. Finally, the trajectory information of each human object in the monitoring video frame is formed. The human body trajectory information includes but is not limited to the identity ID of the human object and the position coordinates of the human object in each frame. The human body trajectory information can be used as the prerequisite information for packing behavior recognition and is sent to the subsequent bag detection module together with the corresponding monitoring video frame.
[0092] Please refer to Figure 4 , which is the method for recognizing behavior information according to an embodiment of the present application. This method can be executed by the Figure 1 shown electronic device 1 and can be applied to the Figures 2 - 3 shown supermarket monitoring application scenario to achieve that by combining the packing posture, the hand-held item behavior, and the preset posture corresponding to the target bag information, it is possible to accurately determine whether the human object performs a packing behavior, reduce the possibility of misjudgment, and improve the recognition efficiency of the packing behavior. In this embodiment, taking the terminal 220 as the execution end as an example, the method includes the following steps:
[0093] Step 401: Obtain the monitored video frames within the specified area.
[0094] In this step, the specified area can refer to the area covered by the monitoring, such as the monitoring coverage area of a supermarket or a shopping mall. The monitored video frames can be collected in real time from a monitoring camera, or can be the monitored videos stored locally or remotely. There is no limitation on the obtaining method of the monitored video frames.
[0095] Step 402: Identify the target bag information existing in the human object area in the monitored video frames.
[0096] In this step, in order to accurately identify whether the bagging behavior occurs in the specified area, the target bag information existing in the human object area in the monitored video frames can be identified first, so as to determine whether the human object is carrying a bag and what kind of bag it is. The target bag information includes but is not limited to information such as the type, size, and appearance of the target bag carried by the human object. Here, the bag types include but are not limited to types such as backpacks, shoulder bags, handbags, and cross-body bags. By identifying the target bag information, a reliable data basis can be provided for subsequent bagging behavior decision-making.
[0097] In an embodiment, identifying the target bag information existing in the human object area in the monitored video frames includes: identifying the human trajectory information in the monitored video frames. Intercepting the human image from the monitored video frames according to the human trajectory information. Inputting the human image into a preset bag detection model so that the preset bag detection model outputs the target bag information existing in the human image, and the target bag information includes the bag position and / or the bag type.
[0098] In this embodiment, by identifying the human trajectory information and intercepting the human image, the human object area in the video frames can be effectively located. The identification method of the human trajectory information can refer to the description of the human trajectory extraction module in the foregoing embodiment. By inputting the human image into the preset bag detection model, the target bag information existing in the human image can be accurately identified. The target bag information includes but is not limited to the position and type of the target bag. In this way, "bag detection" and "bag classification" are implemented through the same model. This model takes the human image as the input and gives two outputs, namely the position of the bag and the type of the bag, through reasoning, thereby improving the efficiency and accuracy of bag information identification. The preset model has been trained and optimized, and can process a large number of video frames in a short time, improving the processing efficiency of the system. On the other hand, by using the human trajectory information, the human object in the video can be dynamically tracked, and the continuous identification and monitoring of the target bag can be maintained even in a complex scene. This dynamic tracking ability ensures the continuity and stability of the identification process.
[0099] In one embodiment, before inputting a human body image into a preset bag detection model, it further includes: obtaining a bag sample image set, where the positions and types of different bag objects are marked in the bag sample images. Training a preset detector using the bag sample image set to obtain the preset bag detection model.
[0100] In this embodiment, by pre-obtaining and marking the bag sample image set and then training the preset detector using these marked sample image sets, the accuracy and robustness of the bag detection model can be significantly improved. To balance speed and accuracy, the preset detector can be implemented by a Yolov5 detector, and other detectors that can perform similar functions are also applicable. The positions and types of different bag objects are marked in the sample images, enabling the training process to better capture the features and variability of the bags, thereby improving the detection performance of the model in practical applications. So that the bag detection model can adapt to different types and styles of bags, improving the generalization ability of the model and enabling it to maintain an efficient detection effect in diverse scenarios. Through this preprocessing and training step, the finally obtained preset bag detection model can more accurately identify and classify the bag objects in the human body image, enhancing the recognition efficiency and reliability of the overall system.
[0101] In one embodiment, the bag types include one or more of a backpack carried on both shoulders of a human body, a single-shoulder bag carried above the waist of a human body, a single-shoulder bag carried below the waist of a human body, a bag carried on the forearm of a human body, and a bag carried in the hand of a human body.
[0102] In this embodiment, the bag types can be classified according to the "bag-carrying posture" of a person. The bag types include but are not limited to: a backpack carried on both shoulders, a single-shoulder bag (the bag is above the waist), a single-shoulder bag (the bag is near and below the waist), a bag carried on the forearm, and a bag carried in the hand. The above five bag types respectively correspond to different "bag-packing postures", and are subsequently matched with the conclusions given by the posture detection model to achieve higher alarm accuracy. Other bag types can also be configured according to actual scenario requirements. During the training process of the bag detection model, the bag sample image set can contain as many different bag types as possible, so that the trained bag detection model can identify a variety of bags.
[0103] Step 403: Detect the bag-packing posture and the behavior of holding an item of the human object.
[0104] In this step, the bag-packing posture refers to the posture of the human hand or an object in contact with the bag, and at least includes the posture of the human body object putting an item into their own bag, such as the posture of a user putting the goods on the shelf into their own bag in a supermarket scenario. In an actual scenario, if a bag-packing behavior occurs, there must be a bag-packing posture. Therefore, the bag-packing posture can be used as reference information for subsequent bag-packing behavior decisions. The hand-holding behavior of an object refers to the behavior of the human body object's hand holding an item, such as the behavior of a user holding a commodity in a supermarket scenario. Generally, the human body object uses their hand to put the item into the bag. Therefore, the hand-holding behavior of an object can be used as reference information for subsequent bag-packing behavior decisions. By combining the bag-packing posture and the hand-holding behavior of an object, the accuracy of bag-packing behavior decisions can be improved.
[0105] In one embodiment, detecting the hand-holding behavior of a human body object in step 403 includes: identifying the human body trajectory information in the monitored video frame. Intercepting the human body image from the monitored video frame according to the human body trajectory information. Inputting the human body image into a preset behavior detection model so that the preset behavior detection model outputs the hand of the human body object and the target item held by the hand in the human body image. If the distance between the position of the target item and the position of the hand of the human body object is less than a preset threshold, it is determined that the human body object has a hand-holding behavior, and the occurrence time of the hand-holding behavior is determined.
[0106] In this embodiment, the hand-holding behavior in the human body image can be identified through a pre-trained behavior detection model. By identifying the human body trajectory information in the video frame, the relevant human body image can be accurately intercepted from the video to ensure the accuracy of subsequent detection. After inputting the intercepted human body image into the preset behavior detection model, the hand of the human body object and the target item held by the hand in the human body image can be identified. By calculating the distance between the position of the target item and the position of the human body object's hand and comparing it with the preset threshold, it can be accurately determined whether a hand-holding behavior has occurred, reducing misjudgment, and the occurrence time of this behavior can be accurately recorded. The workload of manual monitoring is reduced, the detection efficiency and accuracy are improved, and it can be applied to multiple fields such as security monitoring and retail analysis, with significant application value. The preset threshold can be set according to the actual scenario requirements. For example, the preset threshold can be the length of the target item. The position of the target item can be represented by the center point coordinates of the target item, and the hand position can be represented by the key point coordinates of the hand.
[0107] Taking the supermarket scenario as an example, the target item is the target commodity held by the human hand. The behavior detection model takes the human body image as input and gives the position of the target commodity held by the human hand in the human body image. To avoid misdetection of the commodities in the background, the detected commodity position is matched with the hand position. If the distance between the center point coordinates of the commodity and the key point coordinates of the hand is less than the threshold (such as the length of the commodity), it is determined that the detected commodity is the target commodity held by the human body.
[0108] In one embodiment, before inputting the human body image into the preset behavior detection model, it further includes: obtaining a sample image set of handheld items, where the items with different shape features are marked in the sample images of handheld items. Training the preset detector with the sample image set of handheld items to obtain the preset behavior detection model.
[0109] In this embodiment, the sample image of handheld items refers to the image of a human hand holding an item. Generally, different types of items have different shape features. For example, different commodities have different outer packaging features. To improve the recognition accuracy of the model, it is possible to collect as many sample images of a human hand holding different types of items as possible to form a sample image set of handheld items, and mark the items with different shape features held by the human hand in these sample images, so as to provide rich training data for the model. Since the sample images cover the appearance features of various items, the model can learn the recognition features of different items during the training process. By training the preset detector with these sample image sets, the finally obtained preset behavior detection model has strong generalization ability and recognition accuracy, ensuring that the model can more accurately identify the items held by the human object in the surveillance video frame. It improves the detection efficiency and accuracy of the model in practical applications, reduces the situation of false detection and missed detection, and can be applied to various scenarios that require accurate identification of the behavior of holding items, such as security monitoring, retail management and other fields.
[0110] The preset detector can be implemented by the Yolov5 detector, and other detectors are also applicable.
[0111] Taking the supermarket scenario as an example, in order to improve the accuracy of commodity detection and distinguish it from non-commodities, the output of the behavior detection model can be configured to include at least 5 categories: boxed commodities, bottled commodities, bagged commodities, mobile phones and other items, where mobile phones and other items are non-commodities.
[0112] In one embodiment, detecting the packing posture of the human object in step 403 includes: identifying the human trajectory information in the surveillance video frame. Intercepting the human body image from the surveillance video frame according to the human trajectory information. Inputting the human body image into the preset posture detection model, so that the preset posture detection model outputs the packing postures existing in the human body image, the types of bags where the packing postures occur, and the occurrence time of the packing postures.
[0113] In this embodiment, by identifying the human body trajectory information in the video frame, a clear human body image can be effectively intercepted from the video, ensuring the accuracy of subsequent pose detection. Subsequently, by inputting the intercepted human body image into a preset pose detection model, the packing poses existing in the image can be automatically identified. It can not only identify the packing poses, but also further determine the types of bags involved in the packing poses and the specific occurrence time. The occurrence time of the packing pose can be a time point or a continuous period of time. In this way, the efficiency and accuracy of pose recognition are improved, the need for manual intervention is reduced, and it can be applied to multiple fields such as security monitoring and behavior analysis, with broad application prospects.
[0114] In one embodiment, before inputting the human body image into the preset pose detection model, it further includes: obtaining a set of packing sample images, in which various packing poses and non-packing poses corresponding to different target bag information are marked. Using the set of packing sample images to train a preset detector to obtain a preset pose detection model.
[0115] In this embodiment, the packing sample image refers to an image containing a packing pose, such as an image of the pose where a human object puts an item into a bag. Different types of bags and various packing poses and non-packing poses related thereto may be marked in the packing sample image. The set of packing sample images usually includes various different types of bags, such as backpacks, handbags, suitcases, etc. These images show various poses of people interacting with these bags in different scenarios. The sample image set not only includes packing poses, but also may include non-packing poses (such as the actions of a human object standing or walking without involving a bag) to help the model distinguish packing poses from other non-packing poses. Marking information can be attached to each sample image, indicating the type of bag appearing in the image and the corresponding packing poses and non-packing poses. These marking information provide supervision signals, enabling the model to learn the features of various packing actions and non-packing actions during the training process.
[0116] Optionally, the preset detector here can be a Resnet-18 classifier, and other types of classifiers are also applicable.
[0117] In an actual scenario, the same type of bag has similar "packing poses", while different "bag types" have different "packing poses", and they are separable from the poses of not packing. Taking the supermarket scenario as an example, assume that the bag types include a backpack carried on the human shoulders, a single-shoulder bag with the bag above the human waist, a single-shoulder bag with the bag below the human waist, a bag carried on the human forearm, and a bag carried in the human hand, these 5 types.
[0118] The pose detection model takes the human body image as input and can give 6 categories, including a "non-packing pose" and the packing poses corresponding to these 5 bag types.
[0119] As shown Figure 5 in the figure, it is a schematic diagram of the bag-packing posture of a single-shoulder bag provided by an embodiment of the present application. In this figure, a human object is carrying a bag on one shoulder, the bag is located at and above the human waist, and the human object holds a commodity in contact with the bag, and a bag-packing posture occurs.
[0120] Step 404: If the bag-packing posture is a preset posture corresponding to the target bag information, and there is a behavior of holding an item by the human object before the bag-packing posture occurs, it is determined that the human object has a bag-packing behavior in the specified area, and a prompt message is sent.
[0121] In this step, the bag-packing posture of the human object detected in step 403 is matched with the target bag information. If the bag-packing posture is a preset posture corresponding to the target bag information, and there is a behavior of holding an item by the human object before the bag-packing posture occurs, it is determined that the human object has a bag-packing behavior in the specified area. At this time, a prompt message can be sent to stop the bag-packing behavior in time. In this way, by combining the bag-packing posture, the behavior of holding an item, and the preset posture corresponding to the target bag information, it is possible to accurately determine whether the human object has a bag-packing behavior, reduce the possibility of misjudgment, and improve the recognition efficiency of the bag-packing behavior.
[0122] Optionally, in order to improve the system execution speed, when no bag is detected in step 402, the bag-packing posture detection and the behavior detection of holding an item in step 403 may not be executed, saving the calculation amount and improving the resource utilization rate.
[0123] Taking the supermarket scenario as an example, the bag-packing behavior decision gives a conclusion on whether to pack the bag according to the detection results of "bag", the results of holding the commodity, and the detection results of the bag-packing posture. The decision conditions may include the following:
[0124] 1. The bag-packing posture detection module gives the bag-packing posture, the type of bag involved, and the time when the bag-packing posture occurs.
[0125] 2. The bag detection module detects the target bag information, and the type of the target bag is consistent with the type of bag involved in the bag-packing posture.
[0126] 3. Before the time when the bag-packing posture occurs, the item-holding module detects a commodity, and the distance between the commodity and the hand key points of the human object is less than a preset threshold (such as the length of the long side of the circumscribed rectangle of the commodity).
[0127] When the above three decision conditions are met at the same time, it can be determined that a customer has a behavior of packing the commodity in the supermarket. At this time, the bag-packing behavior decision module can send a prompt message, and the prompt message may include the human trajectory of the customer who has the bag-packing posture, and give a video segment of the bag-packing behavior according to the occurrence time of the bag-packing posture, which is convenient for relevant personnel to view and handle.
[0128] Optionally, the system can immediately push the risks existing in the bagging behavior to the loss prevention staff of the supermarket in various ways, so that the loss prevention staff can judge whether the decision result of the system is accurate by viewing the video, and then track the trajectory of the customer in the supermarket, judge whether the checkout behavior is completed, and take further measures to reduce the risk of theft and damage of goods in the supermarket.
[0129] The above method for identifying behavior information structures the monitoring video frames obtained by the camera to obtain the position of the human object in each frame of the video. The position of the bag in the human image is obtained through the "bag" detection algorithm, and the type of the bag is distinguished. Through the commodity detection algorithm, it is identified whether there is a commodity in the hand of the human object. For different types of bags, different bagging postures are configured, and then the bagging behavior decision is made by combining the bag type, the detection of the behavior of holding goods, and the detection of the bagging posture, so as to realize accurate identification of the bagging behavior.
[0130] As Figure 6 shown, it is a schematic flowchart of the method for identifying behavior information according to an embodiment of the present application. This method can be executed by the Figure 1 shown electronic device 1 and can be applied to the Figures 2 - 3 supermarket monitoring application scenario shown, so as to accurately judge whether a human object has a bagging behavior by combining the bagging posture, the behavior of holding an item, and the preset posture corresponding to the target bag information, reducing the possibility of misjudgment and improving the recognition efficiency of the bagging behavior. Taking the terminal 220 as the execution end in this embodiment as an example, this method includes the following steps:
[0131] Step 601: Obtain the monitoring video frames in the specified area.
[0132] Step 602: Identify the human trajectory information in the monitoring video frames.
[0133] Step 603: Intercept the human image from the monitoring video frames according to the human trajectory information. Then enter step 604.
[0134] Step 604: Input the human image into the preset bag detection model, and judge whether the preset bag detection model outputs the target bag information existing in the human image. The target bag information includes the bag position and / or the bag type. If so, enter step 605 and step 606 respectively. Otherwise, it means that no bag is detected, and the subsequent steps can be skipped and it can return to step 601.
[0135] Step 605: Input the human body image into the preset pose detection model, and determine whether the preset pose detection model outputs the packing pose existing in the human body image, the type of the bag for which the packing pose occurs, and the occurrence time T1 of the packing pose. If so, proceed to step 609; otherwise, it indicates that no packing pose is detected, and step 601 can be returned.
[0136] Step 606: Input the human body image into the preset behavior detection model, and determine whether the preset behavior detection model outputs the human body object's hand and the target item held by the hand in the human body image. If so, proceed to step 607; otherwise, step 601 can be returned.
[0137] Step 607: Determine whether the distance between the position of the target item and the position of the human body object's hand is less than the preset threshold. If so, proceed to step 608; otherwise, it indicates that the behavior detection model makes a misjudgment, and step 601 can be returned.
[0138] Step 608: Determine that the human body object has the behavior of holding an item, and determine the occurrence time T2 of the behavior of holding an item. Then proceed to step 610.
[0139] Step 609: Determine whether the packing pose is the preset pose corresponding to the target bag information. If so, proceed to step 610; otherwise, step 601 can be returned for re-detection.
[0140] Step 610: Determine whether the occurrence time T2 of the behavior of holding an item is before the occurrence time T1 of the packing pose. If so, proceed to step 611; otherwise, step 601 can be returned for re-detection.
[0141] Step 611: Determine that the human body object has the packing behavior in the specified area, and send a prompt message.
[0142] It should be noted that the execution order of the above steps 609 and 610 is only for example. In the actual scenario, step 610 can also be executed first, and when the judgment result of step 610 is yes, step 609 is executed; or step 609 and step 610 can also be executed simultaneously, and when the judgment results of both are yes, step 611 is entered. This embodiment does not make any limitations in this regard.
[0143] For the details of each step of the above method, reference can be made to the relevant descriptions of the above embodiments, which will not be elaborated here.
[0144] Please refer to Figure 7 which is the flowchart of the supermarket anti-theft detection method according to an embodiment of the present application. This method can be executed by Figure 1 the electronic device 1 shown, and can be applied to Figures 2 - 3In the supermarket monitoring application scenario shown, by combining the bagging posture, the behavior of holding items, and the preset postures corresponding to the target bag information, it is possible to accurately determine whether a human object has performed a bagging behavior, reducing the possibility of misjudgment and improving the recognition efficiency of the bagging behavior. In this embodiment, taking the behavior detection scenario of bagging winning products in a supermarket as an example, the method includes the following steps:
[0145] Step 701: Obtain the monitoring video frames in the supermarket. Here, the supermarket can be a large shopping mall, a retail store, or other forms of offline physical stores.
[0146] Step 702: Identify the target bag information existing in the human object area in the monitoring video frames.
[0147] Step 703: Detect the bagging posture and the behavior of holding goods of the human object.
[0148] Step 704: If the bagging posture is the preset posture corresponding to the target bag information, and the human object has the behavior of holding goods before the bagging posture occurs, determine that the human object has performed the behavior of bagging goods in the supermarket, and send out an anti-theft prompt message.
[0149] The solution of the embodiment of the present application realizes the detection of bagging behavior and discovers suspicious bagging behavior by attaching computing power through the existing monitoring cameras in the supermarket. The modification to the site is small and the implementation is simple. For the field of anti-loss in the whole supermarket, it can autonomously discover the bagging risk. By combining bag detection, bag type recognition, bagging posture recognition, and holding goods recognition, a comprehensive decision on the bagging behavior is made, which has the characteristic of high recognition accuracy. By corresponding different bag types to different bagging postures, a multi-category bagging posture recognition model is trained, and the alarm accuracy is improved by matching the bag type and the bagging posture.
[0150] Although in the related art, in an unmanned store, by identifying the goods taken by customers and combining the goods that have been checked out, bagging detection can be realized. However, such solutions rely on a large number of hardware devices, including densely arranged cameras, deep reading sensors, weighing shelves, etc., and at the same time require extremely high computing power, and the deployment cost is high, so they cannot be widely applied. In contrast, the purpose of the solution of the embodiment of the present application is to discover commodity losses through bagging behavior, and subsequent human confirmation and measures can be combined to further improve the anti-loss efficiency. It can be realized only by using the existing monitoring cameras, and the cost is relatively lower.
[0151] Compared with the method based on picking and placing goods recognition, the risk link of the bagging behavior recognition in the present application is short, and it can rely only on the behavior recognition problem of the human body under a single camera, and the implementation is relatively simple.
[0152] Compared with the acousto-magnetic anti-theft door, in the actual scenario, only a small number of goods are attached with magnetic tags, and the tags can be removed. The solution of the embodiment of the present application can detect the bagging behavior of all goods through bagging behavior recognition, with higher efficiency.
[0153] For the detailed steps of the above method, reference can be made to the relevant descriptions of the above embodiments, and details are not described herein again.
[0154] Please refer to Figure 8 , which is an identification device 800 for behavior information according to an embodiment of the present application. The device can be applied to Figure 1 the electronic device 1 shown in Figures 2 - 3 and can be applied to the supermarket monitoring application scenario shown in
[0155] The obtaining module 801 is configured to obtain a monitoring video frame in a specified area.
[0156] The recognition module 802 is configured to recognize target bag information existing in the human object area in the monitoring video frame.
[0157] The detection module 803 is configured to detect the bagging posture and the behavior of holding an item of the human object.
[0158] The determination module 804 is configured to determine that the human object has a bagging behavior in the specified area and send a prompt message if the bagging posture is a preset posture corresponding to the target bag information and the human object has a behavior of holding an item before the bagging posture occurs.
[0159] In one embodiment, the recognition module 802 is configured to recognize the human trajectory information in the monitoring video frame. Intercept a human body image from the monitoring video frame according to the human trajectory information. Input the human body image into a preset bag detection model so that the preset bag detection model outputs the target bag information existing in the human body image, and the target bag information includes the bag position and / or the bag type.
[0160] In one embodiment, it further includes: a first training module, configured to obtain a bag sample image set before inputting the human body image into the preset bag detection model, and the bag position and type of different bag objects are marked in the bag sample image. Train the preset detector with the bag sample image set to obtain the preset bag detection model.
[0161] In one embodiment, the detection module 803 is configured to identify the human trajectory information in the monitored video frame. Intercept a human body image from the monitored video frame according to the human trajectory information. Input the human body image into a preset pose detection model, so that the preset pose detection model outputs the packing pose existing in the human body image, the type of the bag for the packing pose, and the occurrence time of the packing pose.
[0162] In one embodiment, it further includes: a second training module, configured to, before inputting the human body image into the preset pose detection model, obtain a set of packing sample images, in which various packing poses and non-packing poses corresponding to different target bag information are marked. Train a preset detector with the set of packing sample images to obtain a preset pose detection model.
[0163] In one embodiment, the detection module 803 is configured to identify the human trajectory information in the monitored video frame. Intercept a human body image from the monitored video frame according to the human trajectory information. Input the human body image into a preset behavior detection model, so that the preset behavior detection model outputs the hand of the human object in the human body image and the target item held by the hand. If the distance between the position of the target item and the position of the hand of the human object is less than a preset threshold, it is determined that the human object has performed the behavior of holding an item, and the occurrence time of the behavior of holding an item is determined.
[0164] In one embodiment, it further includes: a third training module, configured to, before inputting the human body image into the preset behavior detection model, obtain a set of sample images of holding an item, in which items with different external features are marked. Train a preset detector with the set of sample images of holding an item to obtain a preset behavior detection model.
[0165] In one embodiment, the bag type includes one or more of a backpack carried on the human shoulders, a single-shoulder bag with the bag above the human waist, a single-shoulder bag with the bag below the human waist, a bag carried on the human forearm, and a bag carried by the human hand.
[0166] For the detailed description of the recognition device 800 of the above behavior information, please refer to the description of the relevant method steps in the above embodiments. The implementation principles and technical effects are similar, and will not be elaborated here in this embodiment.
[0167] Figure 9 The following is a schematic structural diagram of a cloud device 90 provided by an exemplary embodiment of the present application. The cloud device 90 can be used to run the method provided in any of the above embodiments. As Figure 9 shown, the cloud device 90 may include: a memory 904 and at least one processor 905, Figure 9 Here, one processor is taken as an example.
[0168] A memory 904 is used to store computer programs and can be configured to store various other data to support operations on the cloud device 90. The memory 904 can be an Object Storage Service (OSS).
[0169] The memory 904 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0170] A processor 905 is coupled to the memory 904 and is used to execute the computer program in the memory 904 to implement the solutions provided by any of the above method embodiments. The specific functions and achievable technical effects are not described herein again.
[0171] Furthermore, as Figure 9 shown, the cloud device further includes other components such as a firewall 901, a load balancer 902, a communication component 906, and a power supply component 903. Figure 9 Only some components are schematically shown in Figure 9 and it does not mean that the cloud device only includes
[0172] In one embodiment, the above Figure 9 communication component 906 is configured to facilitate communication between the device where the communication component 906 is located and other devices in a wired or wireless manner. The device where the communication component 906 is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G, LTE (Long Term Evolution), 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component 906 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 906 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0173] In one embodiment, the aboveFigure 9 The power supply component 903 provides power for various components of the device where the power supply component 903 is located. The power supply component 903 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.
[0174] An embodiment of the present application also provides a computer-readable storage medium storing computer-executable instructions, and when a processor executes the computer-executable instructions, the method of any of the foregoing embodiments is implemented.
[0175] An embodiment of the present application also provides a computer program product including a computer program, and when the computer program is executed by a processor, the method of any of the foregoing embodiments is implemented.
[0176] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0177] The integrated modules implemented in the form of software function modules as described above can be stored in a computer-readable storage medium. The above software function modules are stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods of various embodiments of the present application.
[0178] It should be understood that the above processor may be a central processing unit (CPU for short), and may also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly embodied as being executed by a hardware processor, or implemented by a combination of hardware and software modules in the processor. The memory may include high-speed RAM (Random Access Memory) memory, and may also include non-volatile storage NVM (Nonvolatile memory for short), such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.
[0179] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0180] An exemplary storage medium is coupled to the processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a master device.
[0181] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, clothing or device including a series of elements includes not only those elements but also other elements not explicitly listed, or also includes elements inherent to such process, method, clothing or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, clothing or device including the element.
[0182] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0183] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of the present application.
[0184] In the technical solution of the present application, the processing of collection, storage, use, processing, transmission, provision, and disclosure of user data and other information complies with the provisions of relevant laws and regulations and does not violate public order and good customs.
[0185] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A method for identifying behavior information, characterized in that: The method comprises: Get surveillance video frames in the specified area; Identify target bag information existing in a human object area in the surveillance video frame, wherein the target bag information includes a type of the target bag; Detecting the bagging posture and the holding behavior of the human subject, wherein the bagging posture refers to the posture of the human hand or the object in contact with the bag; If the packing posture is a preset posture corresponding to the target bag information, and the human subject has a behavior of holding an item before the packing posture occurs, it is determined that the human subject has a packing behavior in the designated area, and a prompt message is issued; The method further comprises: Determining the type of bag in which the packing posture occurs; If the packing posture is a preset posture corresponding to the target bag information, and the human subject has a behavior of holding an item before the packing posture occurs, determining that the human subject has a packing behavior in the designated area also includes: If the type of the target bag is consistent with the type of the bag in which the packing gesture occurs, and the human subject has held a commodity before the packing gesture occurs, it is determined that the human subject has performed a packing action in the designated area.
2. The method according to claim 1, characterized in that The identifying target bag information existing in the human body object area in the surveillance video frame includes: Identifying human body trajectory information in the surveillance video frame; Intercepting a human body image from the monitoring video frame according to the human body trajectory information; The human body image is input into a preset bag detection model so that the preset bag detection model outputs target bag information existing in the human body image, wherein the target bag information includes a bag position and / or a bag type.
3. The method according to claim 2, characterized in that Before inputting the human body image into the preset bag detection model, the method further includes: Acquire a bag sample image set, wherein the bag sample images are marked with positions and types of different bag objects; The preset detector is trained using the bag sample image set to obtain the preset bag detection model.
4. The method according to claim 1, characterized in that: The detecting of the packing posture of the human object comprises: Identifying human body trajectory information in the surveillance video frame; Intercepting a human body image from the monitoring video frame according to the human body trajectory information; The human body image is input into a preset posture detection model, so that the preset posture detection model outputs the packing posture existing in the human body image, the type of bag in which the packing posture occurs, and the occurrence time of the packing posture.
5. The method according to claim 4, characterized in that Before inputting the human body image into a preset posture detection model, the method further includes: Acquire a set of packaging sample images, wherein the packaging sample images are marked with a plurality of packaging postures and non-packaging postures corresponding to different target bag information; The packaged sample image set is used to train a preset detector to obtain the preset posture detection model.
6. The method according to claim 1, characterized in that Detecting the human subject's behavior of holding an object, including: Identifying human body trajectory information in the surveillance video frame; Intercepting a human body image from the monitoring video frame according to the human body trajectory information; Inputting the human body image into a preset behavior detection model so that the preset behavior detection model outputs the hands of the human object in the human body image and the target object held by the hands; If the distance between the position of the target object and the hand position of the human subject is less than a preset threshold, it is determined that the human subject has performed the hand-holding object behavior, and the occurrence time of the hand-holding object behavior is determined.
7. The method according to claim 6, characterized in that Before inputting the human body image into a preset behavior detection model, the method further includes: Acquire a set of sample images of handheld objects, wherein objects with different appearance features are marked in the sample images of handheld objects; The preset detector is trained using the handheld object sample image set to obtain the preset behavior detection model.
8. The method according to any one of claims 2 to 5, characterized in that: The bag types include one or more of a backpack, a shoulder bag with the bag located above the waist, a shoulder bag located below the waist, a forearm bag and a handbag.
9. A supermarket anti-theft detection method, characterized in that: include: Get surveillance video frames in the supermarket; Identify target bag information existing in a human object area in the surveillance video frame, wherein the target bag information includes a type of the target bag; Detecting the bagging posture and the commodity holding behavior of the human object, wherein the bagging posture refers to the posture of the human hand or the commodity in contact with the bag; If the packing posture is a preset posture corresponding to the target bag information, and the human subject has a behavior of holding goods before the packing posture occurs, it is determined that the human subject has a behavior of packing goods in the supermarket, and an anti-theft prompt message is issued; The method further comprises: Determining the type of bag in which the packing posture occurs; If the packing posture is a preset posture corresponding to the target bag information, and the human subject has a behavior of holding a commodity before the packing posture occurs, determining that the human subject has a behavior of packing the commodity in the supermarket also includes: If the type of the target bag is consistent with the type of the bag in which the packing gesture occurs, and the human subject has held the commodity before the packing gesture occurs, it is determined that the human subject has packed the commodity in the designated area.
10. A system for identifying behavioral information, characterized in that: include: A bag detection module, used to obtain a surveillance video frame in a specified area, and identify target bag information existing in a human object area in the surveillance video frame, wherein the target bag information includes a type of the target bag; A hand-held object detection module, used to detect hand-held object behavior of the human subject; A packing posture detection module, used to detect the packing posture of the human subject and determine the type of bag in which the packing posture occurs, wherein the packing posture refers to the posture of contact between a human hand or a commodity and a bag; a packing behavior decision module, connected to the bag detection module, the handheld item detection module and the packing posture detection module respectively, for determining that the human subject has a packing behavior in the specified area and issuing a prompt message if the packing posture is a preset posture corresponding to the target bag information and the human subject has a handheld item behavior before the packing posture occurs; The packing behavior decision module is further used to determine that the human body object has performed a packing behavior in the designated area if the type of the target bag is consistent with the type of the bag in which the packing gesture occurs, and the human body object has held the commodity before the packing gesture occurs.
11. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the electronic device to perform the method according to any one of claims 1 to 9.
12. A cloud device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the cloud device to execute the method described in any one of claims 1-9.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the method according to any one of claims 1 to 9 is implemented.
14. A computer program product, characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 9 when the computer program is executed by a processor.
Citation Information
Patent Citations
Intelligent supermarket with intelligent monitoring function
CN113781730A