Method, device and storage medium for recognizing behavioral information
By combining image and depth information to obtain the movement trajectories of items and user limbs in self-checkout devices, overlapping items can be identified and operational behaviors can be judged, thus solving the problem of low recognition accuracy of self-checkout devices and achieving more efficient theft and damage recognition and improved user experience.
Patent Information
- Application Number
- CN202311625388.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2043-11-29
AI Technical Summary
Existing self-checkout equipment has low accuracy in identifying theft and damage, which cannot meet the needs of supermarkets and stores, resulting in the need for additional manpower to supervise and waste resources.
By combining image and depth information, the movement trajectories of objects and user limbs within the target area are obtained, overlapping objects are identified and their operational behaviors are determined, and depth information is obtained using binocular or depth cameras. Combined with image processing and trajectory tracking technologies, the accuracy of behavior recognition is improved.
It improves the accuracy of theft and damage identification, reduces reliance on manpower, and enhances user experience and supermarket efficiency.
Smart Images

Figure CN117437264B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of information processing, and in particular to a behavior information identification method, device and storage medium. BACKGROUND
[0002] With the continuous breakthrough and maturity of information technology and artificial intelligence technology, new retail business is also developing towards convenience and intelligence. Self-service cash register devices can allow customers to scan goods and check out in a self-service manner, free from the queuing process and the one-to-one bondage between cashiers and cash registers. Currently, self-service cash registers have been successfully deployed in various supermarkets. Self-service cash registers have brought efficiency improvement and good customer shopping experience to merchants, and the checkout mode based on self-service cash registers has gradually replaced the mode based on manual cash registers.
[0003] However, the deployment of self-service cash register devices has brought more loss problems to supermarkets. In order to prevent the occurrence of loss events, supermarkets need to arrange additional manpower to supervise customer self-service operations in real time, resulting in waste of resources.
[0004] Currently, various hardware manufacturers and algorithm software manufacturers of self-service cash register systems are actively developing automatic loss prevention functions in order to automatically identify the occurrence of loss events by machines and timely handle and recover losses. However, the recognition accuracy of various schemes in actual execution is low, which cannot meet the use requirements of supermarket customers. SUMMARY
[0005] The main purpose of the embodiments of the present application is to provide a behavior information identification method, device and storage medium, which realizes the combination of image depth information to obtain a more accurate tracking trajectory of a tracking object, thereby improving the accuracy of risk behavior identification and improving user experience.
[0006] In a first aspect, the embodiments of the present application provide a behavior information identification method, comprising: acquiring an original image in a target region and depth information corresponding to the original image; determining a first motion trajectory of at least one to-be-processed object and a second motion trajectory of a key limb part of a user in the target region according to the original image and the depth information; determining a current object existing in the to-be-processed object and overlapping with the key limb part according to the first motion trajectory and the second motion trajectory; and in response to current order information, identifying whether the current object has an abnormal behavior according to the current order information and a target motion trajectory of the current object.
[0007] In an embodiment, the obtaining the original image in the target region and the depth information corresponding to the original image comprises: obtaining at least two original images of the target region under different angles of view through a preset binocular camera, wherein the target region is within the shooting range of the binocular camera; performing binocular ranging processing according to the at least two original images to determine the depth information of the original image; and / or obtaining the original image of the target region through a preset depth camera, wherein the original image comprises depth information.
[0008] In an embodiment, the determining the first motion trajectory of at least one to-be-processed object in the target region and the second motion trajectory of a key limb part of a user according to the original image and the depth information comprises: identifying at least one to-be-processed object and / or a key limb part of a user contained in the target region according to the original image and the depth information; performing trajectory tracking on each to-be-processed object according to the original image and the depth information to generate the first motion trajectory of each to-be-processed object; and / or performing trajectory tracking on the key limb part according to the original image and the depth information to obtain the second motion trajectory of the key limb part.
[0009] In an embodiment, the identifying at least one to-be-processed object contained in the target region according to the original image and the depth information comprises: performing classification identification on the original image to determine the position of a first object in the original image; performing position matching on the position of the first object and the depth information to determine a target pixel region in the depth information that does not fall within the position range of the first object; determining the position of a second object corresponding to the target pixel region according to the depth information, wherein the to-be-processed object comprises the first object and the second object.
[0010] In an embodiment, the determining the position of the second object corresponding to the target pixel region according to the depth information comprises: performing position aggregation on the pixel points in the target pixel region, grouping pixel points with adjacent positions and the same depth value into the same aggregation region to obtain at least one aggregation region corresponding to the target pixel region; calculating the maximum circumscribed rectangle of each aggregation region, determining the maximum circumscribed rectangle as the position of a new object identified, and determining the position of the second object corresponding to the target pixel region according to the position of the new object.
[0011] In an embodiment, the determining the position of the second object corresponding to the target pixel region according to the new object position comprises: screening the new object position to remove the position of a redundant object, and determining the position of the second object corresponding to the target pixel region according to the screened new object position, wherein the redundant object comprises one or more of a background object, an object with a position range greater than a first threshold, and / or an object with a position unit smaller than a second threshold, and the first threshold is greater than the second threshold.
[0012] In an embodiment, the original image comprises an image sequence of the target region; and the trajectory tracking of each of the to-be-processed objects according to the original image and the depth information to generate the first motion trajectory of each of the to-be-processed objects comprises: for a current frame image in the image sequence, performing position matching on a current position of the to-be-processed object in the current frame image and corresponding depth information to determine a depth value set of all pixel points in the current position range; determining a depth value with the largest proportion in the depth value set as current depth information of the to-be-processed object in the current frame image; generating trajectory position information of the to-be-processed object in the current frame image according to the current position and the current depth information; and statistically generating the first motion trajectory corresponding to the to-be-processed object according to a set of trajectory position information of the to-be-processed object in the image sequence.
[0013] In an embodiment, the determining the current object from the to-be-processed objects according to the first motion trajectory and the second motion trajectory, which overlaps with the key limb part, comprises: determining a candidate object from the to-be-processed objects according to the first motion trajectory and the second motion trajectory, which overlaps with the key limb part; if the candidate object is one, determining the candidate object as the current object; and if the candidate object is multiple, sorting the multiple candidate objects according to corresponding depth values from small to large, and determining a candidate object with a preset ranking as the current object.
[0014] In an embodiment, the determining, according to the first motion track and the second motion track, candidate items that exist in overlapping relationship with the key body part from the to-be-processed items, comprises: calculating an overlap degree between each of the to-be-processed items and the key body part respectively according to a position of the to-be-processed item in the first motion track and a position of the key body part in the second motion track; determining, as a candidate item that exists in overlapping relationship with the key body part, a to-be-processed item corresponding to an overlap degree greater than a preset threshold; and / or determining, as a candidate item that exists in overlapping relationship with the key body part, a to-be-processed item corresponding to a position range containing the center point position of the key body part according to the second motion track.
[0015] In an embodiment, the identifying, according to the current order information and the target motion track of the current item, whether the current item has an abnormal behavior in response to the current order information, comprises: if it is determined that the current item has implemented a code scanning behavior according to the target motion track of the current item and the second motion track, determining whether code scanning information of the current item is received according to the current order information, and if the code scanning information of the current item is not received, determining that the current item has an abnormal behavior.
[0016] In an embodiment, the method further comprises: detecting a bag position existing in the target region according to the original image; and the identifying, according to the current order information and the target motion track of the current item, whether the current item has an abnormal behavior in response to the current order information, comprises: if it is determined that the current item has implemented a bagging behavior according to the target motion track of the current item and the bag position, determining whether code entry information about the current item is received according to the current order information, and if the code entry information about the current item is not received, determining that the current item has an abnormal behavior.
[0017] In an embodiment, the identifying, according to the current order information and the target motion track of the current item, whether the current item has an abnormal behavior in response to the current order information, further comprises: if it is determined that a moving distance of the current item exceeds a preset distance threshold according to the target motion track of the current item, determining whether code entry information about the current item is received according to the current order information, and if the code entry information about the current item is not received, determining that the current item has an abnormal behavior.
[0018] In a second aspect, the embodiments of the present application provide a behavior information identification method, applied to a self-service checkout system, the system comprising a checkout area and a depth image collector, the checkout area being within the shooting range of the depth image collector, the method comprising: acquiring an original image of the checkout area and depth information corresponding to the original image through the depth image collector; determining a first motion trajectory of at least one to-be-processed commodity in the checkout area and a second motion trajectory of a key limb part of a user according to the original image and the depth information; determining a current commodity existing in the to-be-processed commodity and overlapping with the key limb part according to the first motion trajectory and the second motion trajectory; and in response to current order information, identifying whether the current commodity has an abnormal behavior according to the current order information and a target motion trajectory of the current commodity.
[0019] In a third aspect, the embodiments of the present application provide a behavior information identification device, comprising:
[0020] An acquisition module is configured to acquire an original image in a target area and depth information corresponding to the original image.
[0021] A trajectory tracking module is configured to determine a first motion trajectory of at least one to-be-processed commodity in the target area and a second motion trajectory of a key limb part of a user according to the original image and the depth information.
[0022] A determination module is configured to determine a current commodity existing in the to-be-processed commodity and overlapping with the key limb part according to the first motion trajectory and the second motion trajectory.
[0023] An identification module is configured to identify whether the current commodity has an abnormal behavior according to current order information and a target motion trajectory of the current commodity in response to the current order information.
[0024] In an embodiment, the acquisition module is configured to acquire at least two original images of the target area at different angles through a preset binocular camera, wherein the target area is within the shooting range of the binocular camera; perform binocular ranging processing according to the at least two original images to determine the depth information of the original image; and / or acquire the original image of the target area through a preset depth camera, the original image comprising depth information.
[0025] In an embodiment, the trajectory tracking module is configured to identify at least one to-be-processed object and / or a key body part of a user included in the target region according to the original image and the depth information; perform trajectory tracking on each to-be-processed object according to the original image and the depth information to generate the first motion trajectory of each to-be-processed object; and / or perform trajectory tracking on the key body part according to the original image and the depth information to obtain the second motion trajectory of the key body part.
[0026] In an embodiment, the trajectory tracking module is specifically configured to perform classification identification on the original image to determine the position of a first object in the original image; perform position matching on the position of the first object and the depth information to determine a target pixel region in the depth information that does not fall within the position range of the first object; determine the position of a second object corresponding to the target pixel region according to the depth information, wherein the to-be-processed object includes the first object and the second object.
[0027] In an embodiment, the trajectory tracking module is specifically configured to perform position aggregation on pixel points in the target pixel region, and group pixel points at adjacent positions and with the same depth value into the same aggregated region to obtain at least one aggregated region corresponding to the target pixel region; for each aggregated region, calculate a maximum circumscribed rectangle of the aggregated region, determine the maximum circumscribed rectangle as a position of a new object identified, and determine the position of the second object corresponding to the target pixel region according to the position of the new object.
[0028] In an embodiment, the trajectory tracking module is specifically configured to perform screening on the position of the new object to remove the position of a redundant object, and determine the position of the second object corresponding to the target pixel region as the position of the new object after screening, wherein the redundant object includes one or more of a background object, an object with a position range greater than a first threshold value, and / or an object with a position unit less than a second threshold value, and the first threshold value is greater than the second threshold value.
[0029] In an embodiment, the original image comprises an image sequence of the target region; the trajectory tracking module is specifically configured to, for a current frame image in the image sequence, perform position matching between a current position of the to-be-processed object in the current frame image and corresponding depth information, determine a set of depth values of all pixel points in the current position range, determine a depth value with the largest proportion in the set of depth values as current depth information of the to-be-processed object in the current frame image, and generate trajectory position information of the to-be-processed object in the current frame image according to the current position and the current depth information; and statistically generate a set of trajectory position information of the to-be-processed object in the image sequence, and generate the first motion trajectory corresponding to the to-be-processed object according to the set of trajectory position information.
[0030] In an embodiment, the determination module is configured to determine, according to the first motion trajectory and the second motion trajectory, a candidate object that overlaps with the key body part from the to-be-processed objects; if the candidate object is one, determine the candidate object as the current object; and if the candidate object is multiple, sort the multiple candidate objects according to corresponding depth values from small to large, and determine a candidate object with a preset ranking as the current object.
[0031] In an embodiment, the determination module is specifically configured to calculate an overlap degree between each to-be-processed object and the key body part according to the position of the to-be-processed object in the first motion trajectory and the position of the key body part in the second motion trajectory; and determine, as a candidate object that overlaps with the key body part, a to-be-processed object corresponding to an overlap degree greater than a preset threshold; and / or, the determination module is specifically configured to determine a center point position of the key body part according to the second motion trajectory, and determine, as a candidate object that overlaps with the key body part, a to-be-processed object corresponding to a position range containing the center point position.
[0032] In an embodiment, the identification module is configured to, if it is determined that the current object has performed a code scanning behavior according to a target motion trajectory of the current object and the second motion trajectory, determine whether code scanning information of the current object is received according to the current order information; and if the code scanning information of the current object is not received, determine that the current object has an abnormal behavior.
[0033] In an embodiment, the device further comprises a detection module configured to detect a bag position in the target region according to the original image; and the identification module is specifically further configured to determine, if the current item is determined to have a bagging behavior according to the target motion trajectory of the current item and the bag position, whether the current item has an abnormal behavior according to whether the code entry information of the current item is received according to the current order information.
[0034] In an embodiment, the identification module is specifically further configured to determine, if the moving distance of the current item exceeds a preset distance threshold according to the target motion trajectory of the current item, whether the current item has an abnormal behavior according to whether the code entry information of the current item is received according to the current order information.
[0035] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising:
[0036] at least one processor; and
[0037] a memory connected with the at least one processor in communication;
[0038] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the electronic device to perform the method in any of the above aspects.
[0039] In a fifth aspect, an embodiment of the present application provides a cloud device, comprising:
[0040] at least one processor; and
[0041] a memory connected with the at least one processor in communication;
[0042] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the cloud device to perform the method in any of the above aspects.
[0043] In a sixth aspect, an embodiment of the present application provides a computer readable storage medium, which stores computer execution instructions, and when a processor executes the computer execution instructions, the method in any of the above aspects is implemented.
[0044] In a seventh aspect, an embodiment of the present application provides a computer program product, which comprises a computer program, and when the computer program is executed by a processor, the method in any of the above aspects is implemented.
[0045] The behavior information identification method, device and storage medium provided by the embodiment of the application, by combining the original image in the target region and the depth information corresponding to the original image, performing trajectory detection and tracking on the user and the to-be-processed article in the target region, obtaining the motion trajectory of the to-be-processed article and the user respectively, then selecting a current article from at least one to-be-processed article, the current article is an article being operated by the user, and the user is likely to perform an unsafe operation behavior on the current article, so that in response to the current order information, the operation behavior is identified in combination with the target motion trajectory of the current article, and it is determined whether the user performs an unsafe operation behavior on the current article. In this way, the tracking trajectory of the tracked object can be obtained more accurately in combination with the image depth information, so that the accuracy of risk behavior identification is improved, and the user experience is improved. BRIEF DESCRIPTION OF DRAWINGS
[0046] The accompanying drawings, which are incorporated herein and constitute part of the specification, illustrate embodiments consistent with the application and, together with the description, serve to explain the principles of the application. It is clear that the drawings described below are some embodiments of the application, and those skilled in the art can obtain other drawings from these drawings without creative labor.
[0047] Figure 1 A structural schematic diagram of an electronic device provided by the embodiment of the application;
[0048] Figure 2 An application scenario schematic diagram of a behavior information identification system provided by the embodiment of the application;
[0049] Figure 3 A scene top view schematic diagram of a self-service cash register system provided by the embodiment of the application;
[0050] Figure 4 A work flow architecture schematic diagram of a self-service cash register system provided by the embodiment of the application;
[0051] Figure 5 A flow schematic diagram of a behavior information identification method provided by the embodiment of the application;
[0052] Figure 6 A flow schematic diagram of a behavior information identification method provided by the embodiment of the application;
[0053] Figure 7 A method flow schematic diagram of article motion trajectory tracking provided by the embodiment of the application;
[0054] Figure 8A self-service cash register abnormal action judgment process schematic diagram provided by an embodiment of the present application;
[0055] Figure 9 A behavior information recognition method flow schematic diagram provided by an embodiment of the present application;
[0056] Figure 10 A behavior information recognition device structure schematic diagram provided by an embodiment of the present application;
[0057] Figure 11 A cloud device structure schematic diagram provided by an embodiment of the present application.
[0058] The above figures have shown the specific embodiments of the present application, which will be described in more detail hereinafter. These figures and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0059] The exemplary embodiments will be described in detail herein with reference to the accompanying drawings. Unless otherwise indicated, the same numbers on different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application.
[0060] The term "and / or" is used herein to describe the association relationship of associated objects, which specifically represents that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone.
[0061] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.
[0062] In order to clearly describe the technical solutions of the embodiments of the present application, first, the terms involved in the present application are explained:
[0063] POS: Point of Sales, generally refers to a cash register terminal device, in the embodiments of the present application, it can refer to a self-service code scanning cash register all-in-one machine device used by a supermarket for customers.
[0064] RGB: is a color standard, is through the change of red (R), green (G), blue (B) three color channels and their mutual superposition to get all kinds of colors, RGB is to represent the red, green, blue three channel color.
[0065] IOU: Intersection over Union, overlap.
[0066] TOF: Time of flight, time of flight method.
[0067] As Figure 1 shown, the embodiment provides an electronic device 1, comprising: at least one processor 11 and a memory 12, Figure 1 The processor 11 and the memory 12 are connected through the bus 10. The memory 12 stores instructions executable by the processor 11, and the instructions are executed by the processor 11 to enable the electronic device 1 to execute all or part of the processes of the method in the following embodiments, so as to realize the tracking trajectory with better accuracy by combining image depth information to track the object, thereby improving the accuracy of risk behavior identification and improving user experience.
[0068] In an embodiment, the electronic device 1 can be a POS machine, a mobile phone, a tablet computer, a notebook computer, a desktop computer, or a large-scale computing system composed of multiple computers.
[0069] Figure 2 A schematic diagram of an application scenario 200 of a behavior information identification system provided by the embodiment is shown in the figure. Figure 2 As shown, the system includes: a server 210 and a terminal 220, wherein:
[0070] The server 210 can be a data platform providing behavior information identification services, such as a data service platform of a supermarket. In actual scenarios, a data service platform can have multiple servers 210, Figure 2 In an embodiment, one server 210 is taken as an example.
[0071] The terminal 220 can be a POS machine, a computer, a mobile phone, a tablet, etc. used by a user when logging in to a data service platform, and the terminal 220 can also have multiple, Figure 2 In an embodiment, two terminals 220 are taken as an example for illustration.
[0072] The terminal 220 and the server 210 can transmit information through the Internet, so that the terminal 220 can access the data on the server 210. The terminal 220 and / or the server 210 can be realized by the electronic device 1.
[0073] The behavior information identification scheme of the embodiments of the present application can be deployed on the server 210, or on the terminal 220, or partially on the server 210 and partially on the terminal 220. The actual scene can be selected based on actual needs, and the embodiments are not limited.
[0074] When the behavior information identification scheme is deployed on the server 210 in whole or in part, the terminal 220 can be opened to call an interface to provide algorithm support for the terminal 220.
[0075] The method provided by the embodiments of the present application can be implemented by corresponding software code executed by an electronic device 1, and is realized by data interaction with a server. The electronic device 1 can be a local terminal device. When the method is run on the server, the method can be realized and executed based on a cloud interaction system, wherein the cloud interaction system includes a server and a client device.
[0076] In one possible implementation, the method provided by the embodiments of the present application provides a graphical user interface through a terminal device, wherein the terminal device can be the aforementioned local terminal device, or the aforementioned client device in the cloud interaction system.
[0077] The behavior information identification scheme of the embodiments of the present application can be applied to any field that needs risk behavior identification.
[0078] Taking the self-service cash register scene of retail business as an example, with the continuous breakthrough and maturity of information technology and artificial intelligence technology, new retail business is also developing towards convenience and intelligence. Self-service cash register devices can allow customers to scan goods and pay in a self-service manner, free from the queuing process, and free from the one-to-one bondage relationship between cashiers and cash registers. Currently, self-service cash register devices have been successfully deployed in various supermarkets. Self-service cash register devices bring efficiency improvement and good customer shopping experience to merchants, and the checkout mode mainly based on self-service cash register gradually replaces the mode mainly based on manual cash register.
[0079] However, the deployment of self-service cash register devices also brings more theft and loss problems to supermarkets. In order to prevent theft and loss events, supermarkets need to arrange additional manpower to supervise customer self-service operations in real time, causing waste of resources. Currently, POS hardware manufacturers and algorithm software manufacturers are actively developing automatic loss prevention functions, in order to automatically identify the occurrence of theft and loss events through machines, and timely handle and recover losses.
[0080] In the related art, obtaining two-dimensional image information by erecting a common monitoring camera above a POS machine and using a deep learning algorithm for behavior action analysis has become the mainstream of the automatic cash register loss prevention scheme. However, this scheme is limited by the imperfect performance of the deep learning algorithm, such as detection algorithm missed detection, tracking algorithm split serial number, and complex scene problems such as product stacking and product occlusion in the self-service cash register scene. Often, the product cannot be effectively detected or the product trajectory tracking fails, the complete and accurate product motion features cannot be obtained, and thus the risk report is inaccurate or not timely.
[0081] To solve the above problems, the embodiment of the present application provides a behavior information recognition scheme. The original image in the target area and the depth information corresponding to the original image are combined to detect and track the trajectories of the user and the to-be-processed articles in the target area, and the motion trajectories of the to-be-processed articles and the user are obtained respectively. Then, according to the first motion trajectory of the to-be-processed articles and the second motion trajectory of the key limb parts of the user, a current article that overlaps with the key limb parts of the user is selected from at least one to-be-processed article. The current article is the article being operated by the user. The user is likely to perform an unsafe operation behavior on the current article. Therefore, in response to the current order information, the operation behavior is identified in combination with the target motion trajectory of the current article, and it is determined whether the user has performed an unsafe operation behavior on the current article. In this way, the tracking trajectory of the tracking object can be obtained more accurately in combination with the image depth information, thereby improving the accuracy of risk behavior identification and improving the user experience.
[0082] As shown in Figure 3 , it is a scene top view schematic diagram of a self-service cash register system of an embodiment of the present application. The self-service cash register system includes a cash register area and a depth image collector. The depth image collector takes a binocular camera as an example. The self-service cash register takes a POS machine as an example. The binocular camera and the computing unit can be installed on the POS machine, or the binocular camera and the computing unit can be directly integrated in the POS machine, and the cash register area is within the shooting range of the depth image collector. The cash register area can include an operation table, which can be used to temporarily place the to-be-processed articles of the user. Here, the to-be-processed articles can be to-be-processed articles taken by the user from the supermarket shelf. The user can perform order settlement by scanning the articles on the POS machine. The binocular camera can be erected directly above the operation table top of the POS machine, and a vertical shooting mode is adopted to collect the image and the depth information of the image of the cash register area, so as to identify the operation behavior of the user on the articles according to the image depth information.
[0083] As shown in Figure 4The diagram shown is a schematic of the workflow architecture of a self-service checkout system according to an embodiment of this application. The system uses a binocular camera to collect video image information of the checkout area, which is then input into an image processing module to identify basic information such as people, goods, and behaviors. The system also uses binocular ranging technology to calculate the depth information of pixels in the image. Combined with customer transaction log information output by the POS machine, the system uses a risk action recognition module to identify the user's risky actions related to the items and promptly report the risks so that loss prevention personnel can handle them immediately.
[0084] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Where there is no conflict between the embodiments, the following embodiments and features can be combined with each other. Furthermore, the timing of the steps in the following method embodiments is merely an example and not a strict limitation.
[0085] Please refer to Figure 5 This is a method for recognizing behavioral information according to an embodiment of this application. The method can be derived from... Figure 1 The electronic device 1 shown is used to perform this action and can be applied to... Figures 2-4 In the application scenario of the self-service checkout system shown, the goal is to combine image depth information to obtain a more accurate tracking trajectory of the object being tracked, thereby improving the accuracy of risk behavior identification and enhancing user experience. This embodiment uses terminal 220 as the execution end as an example, and the method includes the following steps:
[0086] Step 501: Obtain the original image and the corresponding depth information of the target area.
[0087] In this step, the target area refers to the area where user interaction with the item needs to be identified and detected. For example, it could be... Figures 3-4 The image shows the checkout area in a self-checkout scenario. The original image refers to the image information captured from the scene of the target area; the original image can be video information or an image sequence. In 3D computer graphics and computer vision, a depth map is an image or image channel that contains information related to the distance from the surface of scene objects to the viewpoint. The depth information of the original image can be represented using a depth map of the original image.
[0088] In one embodiment, step 501 may specifically include: acquiring at least two original images of the target area from different viewpoints using a preset binocular camera, wherein the target area is within the shooting range of the binocular camera. Binocular ranging processing is then performed based on the at least two original images to determine the depth information of the original images.
[0089] In the embodiment, the original image of the target region and the depth information of the original image can be acquired by a depth image collector. The depth image collector includes but is not limited to a binocular camera, a depth camera and the like. For example, at least two original images of the target region under different visual angles can be acquired by the binocular camera, so as to perform binocular ranging based on the at least two original images under different visual angles, and the depth map of the original image is calculated. Figure 3 For example, in the self-service cash register scene shown, the binocular camera can be erected directly above the POS machine operating table surface, and can be used in a vertical shooting manner to acquire the original image sequence of the cash register region, and to perform binocular ranging on the original image sequence to obtain the depth map of the original image sequence. For example, the binocular camera acquires the images of the left and right cameras of the cash register region, and according to the pre-calibration information, the depth image information of the cash register region in the image is obtained by a binocular ranging matching algorithm.
[0090] The binocular camera is a camera designed to simulate the human eye, which acquires two visual angle images of a scene through two cameras at different positions, and then calculates the depth information of the scene. The binocular camera mainly consists of two cameras, a computer and a binocular vision algorithm. Through the computer algorithm, the distance and depth information of the scene can be accurately calculated. Binocular camera ranging is a computer vision technology for estimating the distance between an object and a lens (i.e. depth information). It can use two cameras to simultaneously shoot the same scene, and calculate the distance of the object by measuring the pixel difference of the same object in the image field of view of the two cameras. Binocular ranging generally mainly includes the following four steps:
[0091] 1) Camera calibration: the internal and external parameters, position change matrix of the two cameras are obtained through calibration, so as to determine the geometric relationship between the two cameras.
[0092] 2) Binocular correction: according to the results of camera calibration, the original images shot by the cameras are corrected, so that the original images under two different visual angles are located on the same plane and are parallel to each other.
[0093] 3) Binocular matching: assuming that the original image includes left and right views of the target scene, the corresponding pixel points on the left and right views of the same scene can be matched to obtain the disparity map of the left and right views (the shift of the same pixel point in the two images is called disparity).
[0094] 4) Calculate depth map: according to the disparity map and the principle of similar triangles, the depth of each pixel in the original image is calculated to obtain the depth map of the original image.
[0095] In an optional embodiment, the depth image collector can also be a depth camera, and the original image of the target region is acquired by the preset depth camera, and the original image directly includes the depth information. This scheme is more efficient. For example, a depth camera based on TOF (Time of Flight) can be selected to provide the depth information of the objects in the picture.
[0096] Step 502: determining a first motion trajectory of at least one to-be-processed object in the target region and a second motion trajectory of a key limb part of the user according to the original image and the depth information.
[0097] In this embodiment, the original image can be subjected to image recognition based on the depth information, the to-be-processed object and the key limb part of the user in the target region are determined, and the position of the to-be-processed object and the key limb part is detected to determine the corresponding motion trajectories. Taking the self-service checkout scenario as an example, the to-be-processed object can be an unaccounted-for commodity brought to the operation table by the user, and the key limb part of the user can be the limb part of the user operating the commodity in relation to the checkout operation in front of the operation table, such as the palm and / or arm of the user. The motion trajectories of the commodity and the palm of the user are determined based on the original image and the depth information, so that the depth information of the commodity and the palm of the user is included in the motion trajectories, and the accuracy of the motion trajectories is improved.
[0098] In an embodiment, the trajectory determination process of the to-be-processed object in step 502 can specifically include: identifying at least one to-be-processed object and a key limb part of the user contained in the target region according to the original image and the depth information. The trajectory of each to-be-processed object is tracked according to the original image and the depth information to generate a first motion trajectory of each to-be-processed object, and the trajectory of the key limb part is tracked according to the original image and the depth information to obtain a second motion trajectory of the key limb part.
[0099] In this embodiment, the to-be-processed object and the key limb part of the user contained in the target region are first identified, and then the motion trajectories of each to-be-processed object and the key limb part are tracked according to the original image and the depth image. Taking the self-service checkout scenario as an example, any one camera of a binocular camera can be selected as a main view lens, and the RGB image and the depth information of the checkout region collected based on the main view lens are used to identify the positions of all the commodities brought to the operation table and the feature information thereof, and the position of the commodity can be represented by a commodity frame. Then, the trajectory of each to-be-checked commodity is tracked based on the commodity frame to obtain a first motion trajectory of each commodity. At the same time, the position information of the user currently operating in front of the self-service checkout machine and the position of the palm of the user, such as the position of the joint of the palm, are detected. Then, the trajectory of the palm of the user is tracked based on the position of the joint of the palm to obtain a second motion trajectory of the palm of the user. The motion trajectories obtained in combination with the depth information are more accurate.
[0100] In an embodiment, the at least one to-be-processed object contained in the target region is identified according to the original image and the depth information in step 502, which can specifically include: performing classification identification on the original image to determine the position of the first object in the original image; performing position matching on the position of the first object and the depth information to determine a target pixel region in the depth information that does not fall within the position range of the first object; determining the position of the second object corresponding to the target pixel region according to the depth information, and the to-be-processed object including the first object and the second object.
[0101] In the embodiment, the first object position is the position of the to-be-processed object identified in the original image. Since the identification of the original image can have missed detection, the position of the second object can be determined based on the depth information to supplement the detection, and the second object refers to the to-be-processed object detected in combination with the depth information and the original image. In this way, the to-be-processed object is supplemented by the depth information to reduce the missed detection and improve the recall rate of object detection.
[0102] Taking a self-service checkout scenario as an example, any one camera of a binocular camera can be selected as a main view lens, and the original RGB image of the checkout area collected by the main view lens is input into a deep learning algorithm module to detect the position and feature information of the current goods brought to the operation table from the original RGB image. The position of the goods can be represented by a goods frame, and the goods frame identified based on the original RGB image is the position of the first goods. In order to solve the problem that the original RGB image can have missed detection, the depth map of the original RGB image can be position-matched with the RGB image of the main view lens to confirm the depth information of each first goods frame from the lens. The pixels in the depth map that are not matched with the first goods frame (i.e., not within the range of the first goods frame) are classified as a target pixel region, and the target pixel region can have missed detection of the goods object. Therefore, the target pixel region is further detected and identified in combination with the depth information to determine the second goods frame that can exist in the target pixel region, i.e., the position of the second goods. Meanwhile, the feature information of the second goods can also be extracted in combination with the original image. In this way, the first goods and the second goods both belong to the to-be-processed goods in the checkout area that need to be tracked.
[0103] In an embodiment, the position of the second object corresponding to the target pixel region is determined according to the depth information in step 502, which includes: performing position aggregation on the pixel points in the target pixel region, and classifying the pixel points with adjacent positions and the same depth value into the same aggregation region to obtain at least one aggregation region corresponding to the target pixel region. For each aggregation region, a maximum circumscribed rectangle frame of the aggregation region is calculated, the maximum circumscribed rectangle frame is determined as the position of a new object identified, and the position of the second object corresponding to the target pixel region is determined according to the position of the new object.
[0104] In the embodiment, for the pixel points of the target pixel region not in the first article position range in the depth map, the pixel points of adjacent positions and the same depth value are aggregated, that is, the pixel points of adjacent positions and the same depth value are classified into the same aggregation region, to obtain one or more aggregation regions corresponding to the target pixel region, the aggregation region is generally an irregular region, and then a maximum circumscribed rectangle of the aggregation region is calculated. The maximum circumscribed rectangle is determined as a new object position recognized, and the new object position is likely to be a to-be-processed article missed in the original image. Therefore, the position of the second article corresponding to the target pixel region can be determined according to the new object position. Taking a self-service checkout scene as an example, for the pixel points not in the first commodity frame, the pixel points of adjacent positions and the same depth value are classified into the same aggregation region in combination with the depth map, and a maximum circumscribed rectangle of the aggregation region is determined as a possible new commodity frame. In this way, the position and size information of the aggregation circumscribed rectangle of adjacent positions and the same depth value are used to supplement the commodity tracking trajectory result, so that the recall rate of commodity trajectory tracking is improved.
[0105] In an embodiment, the position of the second article corresponding to the target pixel region is determined according to the new object position, including: screening the new object position, and removing the positions of redundant objects, and determining the positions of the second article corresponding to the target pixel region as the screened new object positions, wherein the redundant objects include one or more of a background object, an object with a position range greater than a first threshold value, and / or an object with a position unit less than a second threshold value, and the first threshold value is greater than the second threshold value.
[0106] In the embodiment, the redundant object refers to an object not belonging to the to-be-processed article. The first threshold value can be a maximum position range threshold value set based on the actual possible size of the to-be-processed article, and the second threshold value can be a minimum position range threshold value set based on the actual possible size of the to-be-processed article. Taking a self-service checkout scene as an example, the maximum circumscribed rectangle obtained after aggregation can also be a position frame of an interfering object, such as a table, a floor, a shelf, and the like, which are background objects collected by the binocular camera. These background objects do not belong to the to-be-checked articles and are redundant objects, so it is unnecessary to perform trajectory tracking. In addition, some of the rectangular frames are very large and have exceeded the maximum size of the commodity (that is, the first threshold value), or some of the rectangular frames are very small and are smaller than the minimum size of the commodity (that is, the second threshold value). The rectangular frames that are too large or too small do not belong to the to-be-processed articles and are redundant objects that will interfere. Therefore, the maximum circumscribed rectangle of each aggregation region obtained after aggregation can be filtered, the background frames such as the floor frame and the tabletop frame are removed according to the initial depth information determined when the binocular camera is initialized, and the rectangular frames that are too large or too small and do not conform to the actual size of the commodity are removed. The remaining rectangular frames are confirmed as real new commodity frames, that is, the second commodity frames, and the depth information can be added to the trajectory of the second commodity frame, so that the accuracy of commodity recognition is further improved.
[0107] In an embodiment, the original image includes an image sequence of the target region. In step 502, the trajectory tracking of each to-be-processed article is performed according to the original image and the depth information, to generate a first motion trajectory of each to-be-processed article, including: for a current frame image in the image sequence, performing position matching on a current position of the to-be-processed article in the current frame image and corresponding depth information, to determine a set of depth values of all pixel points in the current position range. The depth value with the largest proportion in the set of depth values is determined as the current depth information of the to-be-processed article in the current frame image. The trajectory position information of the to-be-processed article in the current frame image is generated according to the current position and the current depth information. A set of trajectory position information of the to-be-processed article in the image sequence is counted, and the first motion trajectory corresponding to the to-be-processed article is generated according to the set of trajectory position information.
[0108] In this embodiment, the tracking process of the article trajectory combines the depth information of the image, which can improve the accuracy of trajectory tracking. In the related art, the RGB image is simply relied on for commodity trajectory tracking. In this way, in the scenarios of fast commodity movement, commodity occlusion, and commodity deformation, etc., the problems of commodity missed detection and commodity trajectory breakage are likely to occur, and when the commodities are stacked, the judgment result of the current commodity held by the user is also likely to be wrong, which further leads to false positives or false negatives when the risk behavior is judged according to the commodity motion trajectory.
[0109] In this embodiment, in order to solve the above problems, during the user checkout process through the self-service checkout system, an image sequence of the checkout area is collected, such as a main view lens collecting video frames of the checkout area. For a current frame image, first, the current commodity frame position of a to-be-processed commodity A in the current frame image is obtained according to the image recognition algorithm based on deep learning. The current commodity frame is positionally matched with a depth map of the current frame image, and the depth values of each pixel point falling into the current commodity frame are counted to obtain a set of depth values, and a depth value histogram of the pixels contained in the current commodity frame is calculated. The depth value with the largest proportion is taken as the depth value of the to-be-processed commodity A in the current frame image. In this way, the depth value of each to-be-processed commodity in the current frame image is determined, and the actual trajectory position information of the commodity A in the current frame image is updated according to the current commodity frame position and the depth value. The trajectory position of the commodity A is counted for each frame image in the image sequence, and a first motion trajectory of the commodity A is generated according to the set of trajectory position information of the commodity A in the image sequence. Similarly, the first motion trajectory of each to-be-processed commodity can be generated. In this way, the depth information is combined to track the commodity trajectory, which improves the accuracy of trajectory tracking.
[0110] In an embodiment, for tracking the trajectory of the key body part of the user, the same method as the above-mentioned trajectory tracking of the to-be-processed article can be used to track the user's body in combination with the depth information, thereby improving the accuracy of the user behavior trajectory tracking.
[0111] Step 503: determining, from the to-be-processed articles, a current article that overlaps with the key body part according to the first motion trajectory and the second motion trajectory.
[0112] In this step, the current article refers to the article that the user is operating. For example, in a self-service checkout scenario, the current article is the current commodity that the user is holding. Since there can be multiple to-be-processed articles before the user reaches the POS machine, and the current commodity held by the user is more likely to have abnormal behavior, such as possibly missing the commodity code when checking out, resulting in a billing error. Therefore, the current commodity that the user is operating can be selected from the multiple to-be-processed articles, and the risk behavior identification is performed for the current commodity, thereby saving the computing resources and improving the identification efficiency.
[0113] In an embodiment, step 503 can specifically include: determining, from the to-be-processed articles, a candidate article that overlaps with the key body part according to the first motion trajectory and the second motion trajectory. If the candidate article is one, the candidate article is determined as the current article. If the candidate article is multiple, the multiple candidate articles are sorted in ascending order according to the corresponding depth values, and the candidate article ranked in the front of the preset name is determined as the current article.
[0114] In this embodiment, there can be multiple to-be-processed items that overlap with the key body part of the user, such as multiple items stacked and each overlapping with the key body part, in which case the stacked candidate items can be screened according to the depth information of each item in the image, and one or more items closest to the lens are selected as the current item being operated by the user, to improve the accuracy of the current item identification result. Taking a self-checkout scenario as an example, assuming that the key body part of the user is the palm, after identifying the trajectory of the to-be-processed goods in the checkout area and the trajectory of the user's palm, the trajectory of each good can be compared with the trajectory of the user's palm in real time, and one or more candidate goods overlapping with the user's palm are selected from the to-be-processed items brought by the user to the checkout area. If there is only one candidate good overlapping with the user's palm, the candidate good can be directly determined as the current good being operated by the user. If multiple candidate goods are stacked and the multiple candidate goods overlap with the same palm of the user, the depth values of the multiple candidate goods can be sorted in ascending order, and the candidate goods ranked in the top pre-set positions are determined as the current goods, that is, the first N (N is a positive integer) candidate goods closest to the camera are determined as the current goods being operated. For example, N = 1, indicating that the candidate good closest to the camera is determined as the current good being operated by the user.
[0115] In an embodiment, the step 503 of determining the candidate goods overlapping with the key body part from the to-be-processed goods according to the first motion trajectory and the second motion trajectory comprises: calculating the overlap degree between each to-be-processed good and the key body part according to the position of the to-be-processed good in the first motion trajectory and the position of the key body part in the second motion trajectory. The to-be-processed goods corresponding to the overlap degree greater than the pre-set threshold are determined as the candidate goods overlapping with the key body part.
[0116] In this embodiment, the selection of the candidate goods having the interaction behavior with the user can be judged in the manner of overlap degree. Taking a self-checkout scenario as an example, assuming that the key body part of the user is the palm, after identifying the behavior trajectory of the user's palm and the motion trajectory of each good, the overlap degree IOU value between the position frame of the user's palm and the position frame of each good can be calculated according to the position information of the user's palm and the position information of each to-be-processed good. If the IOU between the position frame of a good A and the position frame of the user's palm is greater than a pre-set threshold, the good A is determined as the candidate good overlapping with the user's palm. The pre-set threshold can be set based on actual needs.
[0117] In an embodiment, the determining, from the to-be-processed item, the candidate item that overlaps with the key body part according to the first motion track and the second motion track in step 503 can further include: determining a center point position of the key body part according to the second motion track, and determining the to-be-processed item that contains the center point position in a corresponding position range as the candidate item that overlaps with the key body part.
[0118] In the embodiment, the selection of the candidate item that has the interaction behavior with the user can track and determine the center point position of the key body part. Taking the self-checkout scenario as an example, assuming that the key body part of the user is the palm, the center point position of the palm can be determined according to the motion track of the palm of the user, and whether the palm of the user overlaps with the goods can be determined by judging whether the center point position of the palm of the user is in the position frame of the goods. If the center point position of the palm of the user is in the position frame of the goods A, it is determined that the goods A is the candidate goods that overlaps with the palm of the user. Otherwise, it is determined that the goods A is not the candidate goods that overlaps with the palm of the user. The overlapping judgment mode of the plurality of goods and the palm can improve the flexibility of calculation.
[0119] Step 504: in response to the current order information, determining whether the current item has an abnormal behavior according to the current order information and the target motion track of the current item.
[0120] In this step, when the user performs the order operation, the risk behavior is identified according to the order information and the target motion track of the current item that the user is operating, corresponding to the current order information of the user. Taking the self-checkout scenario as an example, assuming that the key body part of the user is the palm, and the current goods are the goods currently held by the user, when the user holds the current goods on the operation table to perform the checkout order operation, the current order information is generated in the transaction log of the POS machine. If the user scans the code or inputs the information of the current goods on the POS machine, the order information will contain the billing information of the current goods, so whether the current item has an abnormal behavior can be identified according to the current order information and the target motion track of the current item. The abnormal behavior here refers to the behavior that the current item fails to perform the normal order settlement process.
[0121] In an embodiment, step 504 can specifically include: if it is determined that the current item has a code scanning behavior according to the target motion track of the current item and the second motion track, determining whether the code scanning information of the current item is received according to the current order information, and if the code scanning information of the current item is not received, determining that the current item has an abnormal behavior.
[0122] In the embodiment, if it is determined that the current item is implemented with the code scanning behavior, but the code scanning information of the current item does not exist in the current order information, it indicates that the current item fails in the code scanning, or the user implements the false code scanning behavior, for example, the user holds the goods to perform the code scanning action in the code scanning area of the POS machine, but the POS machine does not receive the code scanning signal, which generally means that the user uses the hand to shield the bar code of the current goods or intentionally uses the non-bar code position close to the code scanning gun to perform the code scanning action. In this case, the current goods are implemented with the abnormal behavior, and an alarm can be issued to facilitate the timely processing.
[0123] In an embodiment, the method can further include detecting the bag position existing in the target region according to the original image. Step 504 can further include detecting the bag position existing in the target region according to the original image. If it is determined that the current item is implemented with the bagging behavior according to the target motion trajectory of the current item and the bag position, it is determined whether the code entry information about the current item is received according to the current order information, and if the code entry information of the current item is not received, it is determined that the current item has the abnormal behavior.
[0124] In the embodiment, the bag includes the shopping bag provided by the supermarket, the handbag brought by the user, the backpack, and the like. Taking the self-service checkout scenario as an example, the bag in the checkout region can be detected according to the original image, and the position of the bag is tracked. The target motion trajectory of the current goods is compared with the bag trajectory to determine whether the current goods are implemented with the bagging behavior. If it is determined that the current goods are implemented with the bagging behavior, it is further determined whether the code scanning information of the current goods or the code entry information manually input by the user is received according to the current order information. The code entry information can be the bar code information of the current goods input through the interactive interface of the mobile phone or the POS, and the like. If the code scanning information or the code entry information is not received, it is determined that the current goods are directly bagged without normal checkout, for example, the user directly puts the current goods into the bag without performing the code scanning behavior, manually inputting the bar code, or manually adding the quantity of the same goods, and the like. In this case, it is determined that the current item has the abnormal behavior, and an alarm can be issued to facilitate the timely processing.
[0125] In an embodiment, step 504 can further include the following steps. If it is determined that the moving distance of the current item exceeds the preset distance threshold according to the target motion trajectory of the current item, it is determined whether the code entry information about the current item is received according to the current order information, and if the code entry information of the current item is not received, it is determined that the current item has the abnormal behavior.
[0126] In the embodiment, the abnormal behavior can also include that the current item is directly moved to a relatively far place, such as a place outside the cash register area or a place in the scanned code area. Taking the self-service cash register scenario as an example, if the moving distance of the current item exceeds a preset distance threshold, such as exceeding the range of the cash register area or exceeding the range of the unscanned code area, and there is no scanning behavior, it is determined whether the code entry information of the current item has been received. If not, it is determined that the current item has a long-distance moving risk, such as the user directly placing the current item from the unscanned code area to the scanned code area, or directly placing the current item from the left side of the operation table to the right side of the operation table, without scanning the code, manually inputting the barcode, or manually adding the quantity of the same item. At this time, it is determined that the current item has an abnormal behavior, and an alarm can be issued to facilitate timely processing.
[0127] The above-mentioned behavior information recognition method uses a binocular camera to collect the product operation information in front of the self-service cash register, and acquires the depth information of the product based on the original RGB information, so that the product detection and tracking are more accurate, thereby making the risk early warning of the present solution more accurate and the customer experience better. Based on the more accurate product trajectory, more risk action types can be defined, and the risk types are more abundant.
[0128] As shown in Figure 6 , it is a flowchart of a behavior information recognition method according to an embodiment of the present application. The method can be executed by the electronic device 1 as shown in Figure 1 , and can be applied to the application scenario of the self-service cash register system as shown in Figures 2-4 , to realize the combination of image depth information to obtain a more accurate tracking trajectory of the product, thereby improving the accuracy of risk behavior recognition and improving the user experience. The present embodiment takes the self-service cash register scenario as an example, and the method includes the following steps:
[0129] Step 601: image acquisition. The binocular camera installed on the POS machine can be used to collect images of the cash register area to obtain RGB images of different perspectives of the cash register area.
[0130] Step 602: binocular distance measurement. Binocular distance measurement is performed according to the RGB images of different perspectives to obtain a depth map of the cash register area. For example, the left and right camera images obtained by the binocular camera are used to obtain the depth image information of the cash register area by a binocular matching algorithm according to the pre-calibration information.
[0131] Step 603: human body detection. Any one of the binocular cameras can be selected as the main lens, and the RGB image collected by the main lens is input into a deep learning algorithm module to detect the position information of the person operating in front of the self-service cash register and the position information of each human body joint.
[0132] Step 604: commodity detection, any one of the binocular camera can be selected as the main lens, and the RGB image collected is input into the deep learning algorithm module to detect the positions of all commodities currently brought to the operating table and to perform feature extraction on the commodities.
[0133] Step 605: commodity and human body tracking, the behavior trajectory of the human body and the motion trajectory of each commodity are identified through the above information.
[0134] Step 606: confirmation of commodity and human body interaction information, whether the commodity is in the hand-held state is determined through the position of the human hand and the position of the commodity, such as calculating the IOU value of the human hand position box and the commodity position box, and the commodity box with an IOU greater than a preset threshold is determined as a candidate commodity interacting with the human body. Or through the palm center position in the commodity position box, it is determined that the commodity interacts with the human hand.
[0135] Step 607: commodity depth information and position information matching, the depth map obtained in step 602 is matched with the RGB image of the main lens, and the depth information of each commodity in the RGB image is confirmed according to the commodity position box output by the RGB image tracking algorithm.
[0136] Step 608: adding or updating commodity motion trajectory and attribute information, the depth information of each commodity is determined according to step 607, and the depth information of the trajectory of each commodity (i.e., the first commodity) identified in steps 604-605 and the interaction information with the human hand are updated. For the target pixel area in the depth map that does not fall into the commodity position box identified in step 604, the pixel points with the same depth value are aggregated in position, and the tracking trajectory result of the newly added commodity (i.e., the second commodity) is supplemented according to the position and size information of the maximum circumscribed rectangle of the aggregated area. For details, please refer to the embodiment shown in Figure 7 Step 611.
[0137] Step 609: bag detection, the position of the shopping bag or the customer's handbag or backpack in the current transaction state can be detected through the preset depth algorithm module using the RGB image.
[0138] Step 610: POS machine transaction log, current order information is obtained.
[0139] Step 611: risk action recognition, the commodity trajectory result information updated in step 608 and the POS machine transaction log are simultaneously input into the risk action analysis module for risk recognition. For details, please refer to the detailed description in the following Figure 8
[0140] In an actual scenario, the human body detection module in step 603, the commodity detection module in step 604, and the bag detection module in step 609 can be combined into one detection module, and one algorithm module can be used to output detection results of different categories, or different implementation modules can be separately used, and the embodiments of the present application do not limit this.
[0141] As shown in FIG. 8, an implementation method flow diagram of the commodity motion trajectory tracking in step 608 provided by the embodiments of the present application is shown, and depth information of the commodity is added to improve the commodity trajectory and the corresponding motion attribute, including the following steps. Figure 7
[0142] Step 701: Initialize the depth information. The depth information in the initial state can be determined by image acquisition when there is no commodity and person in the cash register area after binocular camera calibration, and the depth information of the operation table, the ground, and the background object is initialized. For example, the depth information is calculated in the initial stage of transaction, and the ground and the position and depth information of the POS machine desktop area are recorded. Go to step 708.
[0143] Step 702: Obtain the current image depth information and the trajectory of the first commodity recognized in the current image. For example, in the transaction process, the trajectory of the first commodity in the current image is obtained according to the depth learning algorithm, including the position of the first commodity frame, and the depth map of the current image is obtained by binocular distance measurement.
[0144] Step 703: Count the depth information of each pixel point in each first commodity frame. For each first commodity frame recognized in step 702, the depth value of each pixel point in the corresponding commodity frame in the current depth map is counted.
[0145] Step 704: Calculate the depth value histogram of each first commodity frame.
[0146] Step 705: Confirm the depth of the first commodity. The depth value with the largest proportion in the corresponding histogram is taken as the current depth value of the corresponding first commodity. Then go to step 710.
[0147] Step 706: Count the depth information of the non-commodity region pixel point. According to the current image depth information and the recognized first commodity trajectory, the target pixel region in the current depth map that does not fall into the first commodity frame is determined.
[0148] Step 707: Aggregate the same depth value to obtain the maximum bounding rectangle. The target pixel point in the depth map that does not match the position of the first commodity trajectory is aggregated according to the condition of position adjacency and same depth value, the maximum bounding rectangle of the aggregated irregular aggregation region is obtained after aggregation, and a new rectangular frame is obtained.
[0149] Step 708: Remove background / large / small rectangular frame, remove the background frame such as ground and table from the newly added rectangular frame according to the initialization depth information in step 701, and remove the rectangular frame which is too large or too small and does not conform to the actual size of the commodity.
[0150] Step 709: Add new commodity position, the remaining rectangular frame after step 708 is determined as a newly added second commodity frame, which is added to the trajectory set of the recognized commodity to be processed, and depth information is added for each recognized commodity.
[0151] Step 710: Confirm the interaction information between the commodity and the hand, and match the commodity trajectory and the hand position obtained in steps 705 and 709 again. If the commodity is stacked and multiple commodities overlap with the same hand, the first N commodities closest to the camera are determined as the current commodities being operated.
[0152] Step 711: Update the trajectory information of the commodity to be processed.
[0153] The scheme of the embodiment adds depth information to the RGB image acquisition camera, which can make the acquired commodity motion information more accurate, and thus make the risk behavior report more accurate.
[0154] As shown in Figure 8 FIG. 6 is a flowchart of an abnormal action determination process in step 611 provided by the embodiment of the present application. Common self-checkout theft and damage behaviors are divided into three categories, which specifically include the following steps:
[0155] Step 801: Input the commodity motion trajectory, which refers to the target motion trajectory of the current commodity overlapping with the hand.
[0156] Step 802: Determine whether the current commodity has a code scanning behavior. If yes, go to step 803, otherwise go to step 804.
[0157] Step 803: Determine whether the code scanning signal of the current commodity is received. If yes, it means that the current commodity is safe, otherwise it means that there is an abnormal behavior, and a risk prompt is reported.
[0158] Step 804: Determine whether the current commodity has a bagging behavior. If yes, go to step 805, otherwise go to step 806.
[0159] Step 805: Determine whether a manual addition signal about the current commodity is received, such as a manual addition of a barcode. If yes, it means that the current commodity is safe, otherwise it means that there is an abnormal behavior, and a risk prompt is reported.
[0160] Step 806: Determine whether the current commodity has a long-distance movement behavior, if yes, go to step 807, otherwise, it means that the current commodity is safe. It can return to step 802 for further detection.
[0161] Step 807: Determine whether a manual addition signal about the current commodity is received, such as manually adding barcode information. If received, it means that the current commodity is safe, and it can return to step 802 for further detection. Otherwise, it means that there is an abnormal behavior, and a risk prompt is reported.
[0162] The above behavior recognition method divides the common self-checkout theft behavior into three categories. The first category is customer fake scanning risk, that is, the customer holds the commodity in the scanning area to perform the scanning action, but the POS machine does not receive the scanning signal. Generally, the barcode is shielded by hand or the scanning action is performed near the scanning gun using a non-barcode position. The second category is direct packing risk, that is, the customer directly packs the commodity into the packing bag without scanning behavior or manually inputting the barcode or manually adding the quantity of the same commodity. The third category is long-distance movement risk, that is, the customer directly places the commodity from the unscanned area to the scanned area, or directly from the left to the right without scanning behavior or manually inputting the barcode or manually adding the quantity of the same commodity. The typical self-checkout risk action is defined and effectively recognized, and the action recognition category is more abundant compared with the traditional scheme.
[0163] Please refer to Figure 9 , which is a flowchart of a behavior information recognition method of an embodiment of the application. The method can be executed by the electronic device 1 shown in Figure 1 , and can be applied to the application scenario of the self-checkout system shown in Figures 2-4 to achieve better tracking trajectory of the tracking object by combining image depth information, thereby improving the accuracy of risk behavior recognition and improving user experience. In this embodiment, the self-checkout scenario of a supermarket is taken as an example. The method includes the following steps:
[0164] Step 901: Obtain the original image of the checkout area and the depth information corresponding to the original image through the depth image collector.
[0165] Step 902: Determine the first motion trajectory of at least one commodity to be processed and the second motion trajectory of the key limb part of the user in the checkout area according to the original image and the depth information.
[0166] Step 903: Determine the current commodity which has an overlap with the key limb part from the commodity to be processed according to the first motion trajectory and the second motion trajectory.
[0167] Step 904: In response to the current order information, determine whether the current commodity has an abnormal behavior according to the current order information and the target motion trajectory of the current commodity.
[0168] The above-mentioned steps of the method can refer to the related descriptions of the above-mentioned embodiments for details, which will not be repeated here.
[0169] Please refer to Figure 10 , which is a behavior information identification device 1000 of an embodiment of the present application. The device can be applied to Figure 1 The electronic device 1 shown in the figure, and can be applied to Figures 2-4 The application scenario of the self-service checkout system shown in the figure, so as to obtain a more accurate tracking trajectory of the tracked object by combining image depth information, thereby improving the accuracy of risk behavior identification and improving user experience. The device comprises an acquisition module 1001, a trajectory tracking module 1002, a determination module 1003 and an identification module 1004, and the functions and principles of each module are as follows:
[0170] The acquisition module 1001 is configured to acquire an original image in a target region and depth information corresponding to the original image.
[0171] The trajectory tracking module 1002 is configured to determine, according to the original image and the depth information, a first motion trajectory of at least one to-be-processed item in the target region and a second motion trajectory of a key limb part of a user.
[0172] The determination module 1003 is configured to determine, according to the first motion trajectory and the second motion trajectory, a current item that overlaps with the key limb part from the to-be-processed items.
[0173] The identification module 1004 is configured to, in response to current order information, identify whether the current item has an abnormal behavior according to the current order information and a target motion trajectory of the current item.
[0174] In an embodiment, the acquisition module 1001 is configured to acquire at least two original images of the target region at different angles of view through a preset binocular camera, wherein the target region is within the shooting range of the binocular camera. The depth information of the original image is determined by performing binocular ranging processing according to the at least two original images. And / or, the original image of the target region is acquired through a preset depth camera, and the original image comprises depth information.
[0175] In an embodiment, the trajectory tracking module 1002 is configured to identify at least one to-be-processed item and / or a key limb part of a user contained in the target region according to the original image and the depth information. The first motion trajectory of each to-be-processed item is generated by performing trajectory tracking on each to-be-processed item according to the original image and the depth information. And / or, the second motion trajectory of the key limb part is obtained by performing trajectory tracking on the key limb part according to the original image and the depth information.
[0176] In an embodiment, the trajectory tracking module 1002 is specifically configured to classify and identify the original image, and determine a position of the first object in the original image. The position of the first object is positionally matched with the depth information, and a target pixel region in the depth information that does not fall within the position range of the first object is determined. The position of the second object corresponding to the target pixel region is determined according to the depth information, and the object to be processed includes the first object and the second object.
[0177] In an embodiment, the trajectory tracking module 1002 is specifically configured to positionally aggregate the pixel points in the target pixel region, and group the pixel points at adjacent positions and with the same depth value into the same aggregated region, to obtain at least one aggregated region corresponding to the target pixel region. For each aggregated region, a maximum circumscribed rectangle of the aggregated region is calculated, and the maximum circumscribed rectangle is determined as a position of a new object identified. The position of the new object is used to determine the position of the second object corresponding to the target pixel region.
[0178] In an embodiment, the trajectory tracking module 1002 is specifically configured to screen the position of the new object, and remove the position of a redundant object. The position of the new object after screening is determined as the position of the second object corresponding to the target pixel region, wherein the redundant object includes one or more of a background object, an object with a position range greater than a first threshold value, and / or an object with a position unit less than a second threshold value, and the first threshold value is greater than the second threshold value.
[0179] In an embodiment, the original image includes an image sequence of a target region. The trajectory tracking module 1002 is specifically configured to, for a current frame image in the image sequence, positionally match a current position of the object to be processed in the current frame image with corresponding depth information, and determine a set of depth values of all pixel points within the current position range. A depth value with the largest proportion in the set of depth values is determined as current depth information of the object to be processed in the current frame image. Trajectory position information of the object to be processed in the current frame image is generated according to the current position and the current depth information. A set of trajectory position information of the object to be processed in the image sequence is counted, and a first motion trajectory corresponding to the object to be processed is generated according to the set of trajectory position information.
[0180] In an embodiment, the determination module 1003 is configured to determine a candidate object that overlaps with the key limb part from the object to be processed according to the first motion trajectory and the second motion trajectory. If the candidate object is one, the candidate object is determined as the current object. If the candidate object is multiple, the multiple candidate objects are sorted in ascending order according to the corresponding depth values, and a candidate object with a preset ranking in the front is determined as the current object.
[0181] In one embodiment, the determining module 1003 is specifically used to calculate the overlap between each item to be processed and the key limb part based on the position of the item to be processed in the first motion trajectory and the position of the key limb part in the second motion trajectory. Items to be processed with an overlap greater than a preset threshold are determined as candidate items that overlap with the key limb part. And / or, the determining module 1003 is specifically used to determine the center point position of the key limb part based on the second motion trajectory, and determine items to be processed within the corresponding position range that include the center point position as candidate items that overlap with the key limb part.
[0182] In one embodiment, the identification module 1004 is used to determine whether the current item has been scanned based on the target movement trajectory and the second movement trajectory of the current item, and to determine whether the scanning information of the current item has been received based on the current order information. If the scanning information of the current item has not been received, the module determines that the current item has abnormal behavior.
[0183] In one embodiment, the device further includes: a detection module, used to detect the location of a bag in the target area based on the original image. The identification module 1004 is further configured to, if it is determined that the current item has been packaged based on the target movement trajectory and bag location of the current item, determine whether coding information for the current item has been received based on the current order information; if no coding information for the current item has been received, determine that the current item exhibits abnormal behavior.
[0184] In one embodiment, the identification module 1004 is further configured to determine whether the current item has received code entry information based on the target movement trajectory of the current item if the moving distance of the current item exceeds a preset distance threshold, and if no code entry information for the current item is received, determine that the current item has abnormal behavior.
[0185] For a detailed description of the behavior information recognition device 1000, please refer to the description of the relevant method steps in the above embodiments. The implementation principle and technical effect are similar, and will not be repeated here.
[0186] Figure 11 This is a schematic diagram of the structure of a cloud device 110 provided for an exemplary embodiment of this application. The cloud device 110 can be used to run the methods provided in any of the above embodiments. Figure 11 As shown, the cloud device 110 may include: a memory 1104 and at least one processor 1105. Figure 11 Let's take a processor as an example.
[0187] The memory 1104 is configured to store computer programs and can be configured to store other various data to support operations on the cloud device 1100. The memory 1104 can be an Object Storage Service (OSS).
[0188] The memory 1104 can be implemented by any type of volatile or nonvolatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0189] The processor 1105 is coupled to the memory 1104 and is configured to execute computer programs in the memory 1104 to implement the solutions provided by any of the method embodiments described above. The specific functions and technical effects that can be achieved are not described here.
[0190] Further, as Figure 11 The cloud device further includes a firewall 1101, a load balancer 1102, a communication component 1106, a power supply component 1103, and other components. Figure 11 Some components are only schematically shown in the figure, and it does not mean that the cloud device only includes Figure 11 the components shown.
[0191] In an embodiment, the communication component 1106 in the above Figure 11 The communication component 1106 is configured to facilitate wired or wireless communication between the device where the communication component 1106 is located and other devices. The device where the communication component 1106 is located can access a wireless network based on a communication standard, such as a WiFi, 2G, 3G, 4G, LTE (Long Term Evolution, LTE for short), 5G, or other mobile communication network, or a combination thereof. In an example embodiment, the communication component 1106 receives a broadcast signal or broadcast related information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component 1106 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0192] In one embodiment, the above Figure 11 The power supply component 1103 provides power to various components of the device in which it resides. The power supply component 1103 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component resides.
[0193] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method of any of the foregoing embodiments.
[0194] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the method of any of the foregoing embodiments.
[0195] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed.
[0196] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.
[0197] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as execution by a hardware processor, or execution by a combination of hardware and software modules within the processor. The memory may include high-speed RAM (Random Access Memory), and may also include non-volatile memory (NVM), such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk, or optical disc, etc.
[0198] The storage medium can be implemented by any type of volatile or non-volatile storage devices or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic storage, flash memory, magnetic or optical disk. The storage medium can be any available medium that can be accessed by a general or special purpose computer.
[0199] An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. Of course, the storage medium can be a component of the processor. Suitable processors include, by way of example, both general and special purpose microprocessors. Suitable processors also include both general and special purpose microprocessors. The processor can be a single processor or a plurality of processors. It is further noted that the processor can be a component of a distributed computer system, cloud computing system, or any other component or combination of components of a distributed computer system.
[0200] It should be noted that, in the present document, the terms "comprising", "containing" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not only include those elements, but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, an element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0201] The above-mentioned sequence number of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments.
[0202] Through the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned example method can be realized by means of software and a necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product in essence or in the form of a contribution to the prior art. The computer software product is stored in a storage medium (such as a ROM / RAM, a magnetic disk, an optical disk), and includes a plurality of instructions for causing a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the method of each embodiment of the present application.
[0203] In the technical solutions of the present application, the collection, storage, use, processing, transmission, provision and disclosure of user data and other information involved in the processing comply with the relevant legal regulations and do not violate public order and good customs.
[0204] The above is only a preferred embodiment of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent process transformation using the content of the specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method of recognizing behavioral information, characterized by, The method comprises the following steps: acquiring an original image and depth information corresponding to the original image in a target region through a depth image collector; determining a first motion trajectory of at least one to-be-processed article and a second motion trajectory of a key limb part of a user in the target region according to the original image and the depth information; determining a current article existing in overlap with the key limb part from the to-be-processed articles according to the first motion trajectory and the second motion trajectory; in response to current order information, determining whether the current article has an abnormal behavior according to the current order information and a target motion trajectory of the current article, the abnormal behavior including a code scanning behavior, a bag packing behavior and a long distance moving behavior; the step of determining a current article existing in overlap with the key limb part from the to-be-processed articles according to the first motion trajectory and the second motion trajectory comprises the following steps: calculating an overlap degree between each to-be-processed article and the key limb part according to the position of the to-be-processed article in the first motion trajectory and the position of the key limb part in the second motion trajectory, respectively; determining a candidate article existing in overlap with the key limb part as the to-be-processed article corresponding to the overlap degree greater than a preset threshold value; determining a candidate article existing in overlap with the key limb part as the to-be-processed article corresponding to the position range containing the center point position of the key limb part according to the center point position of the key limb part determined according to the second motion trajectory; if the candidate article is one, determining the candidate article as the current article; if the candidate article is multiple, sorting the multiple candidate articles according to the corresponding depth value from small to large, and determining the candidate article ranked in the front preset position as the current article; the step of determining a first motion trajectory of at least one to-be-processed article in the target region according to the original image and the depth information comprises the following steps: performing classification recognition on the original image to determine the position of a first article in the original image, the to-be-processed article including the first article and a second article; performing position matching on the position of the first article and the depth information to determine a target pixel region in the depth information not falling within the position range of the first article; performing position aggregation on the pixel points in the target pixel region, and grouping the pixel points with adjacent positions and the same depth value into the same aggregation region to obtain at least one aggregation region corresponding to the target pixel region; for each aggregation region, calculating a maximum circumscribed rectangle of the aggregation region, determining the maximum circumscribed rectangle as a recognized new object position, performing screening on the new object position to remove the position of a redundant object, and determining the new object position after screening as the position of the second article corresponding to the target pixel region, wherein the redundant object includes one or more of a background object, an object with a position range greater than a first threshold value and / or an object with a position unit smaller than a second threshold value, the first threshold value being greater than the second threshold value. According to the original image and the depth information, a first motion trajectory of each of the to-be-processed articles is generated by performing trajectory tracking on each of the to-be-processed articles.
2. The method of claim 1, wherein, The original image and the depth information of the target region are acquired by the depth image collector, and the original image includes the depth information. At least two original images of the target region under different visual angles are acquired by a preset binocular camera, wherein the target region is within the shooting range of the binocular camera. The depth information of the original image is determined by performing binocular distance measurement processing according to the at least two original images. And / or, the original image of the target region is acquired by a preset depth camera, and the original image includes depth information.
3. The method of claim 1, wherein, According to the original image and the depth information, a second motion trajectory of a key limb part of the user is determined, including: According to the original image and the depth information, a key limb part of the user in the target region is identified; According to the original image and the depth information, a second motion trajectory of a key limb part of the user is determined, including:
4. The method of claim 1, wherein, The original image includes an image sequence of the target region; and according to the original image and the depth information, a first motion trajectory of each of the to-be-processed articles is generated by performing trajectory tracking on each of the to-be-processed articles, including: For a current frame image in the image sequence, the current position of the to-be-processed article in the current frame image is matched with the corresponding depth information to determine a set of depth values of all pixel points within the current position range; The depth value with the largest proportion in the set of depth values is determined as the current depth information of the to-be-processed article in the current frame image; According to the current position and the current depth information, trajectory position information of the to-be-processed article in the current frame image is generated; A set of trajectory position information of the to-be-processed article in the image sequence is counted, and the first motion trajectory corresponding to the to-be-processed article is generated according to the set of trajectory position information.
5. The method of claim 1, wherein, In response to the current order information, whether the current article has an abnormal behavior is identified according to the current order information and the target motion trajectory of the current article, including: If it is determined according to the target motion trajectory of the current article and the second motion trajectory that the current article has implemented a code scanning behavior, whether the code scanning information of the current article is received is determined according to the current order information, and if the code scanning information of the current article is not received, it is determined that the current article has an abnormal behavior; And / or, according to the original image, the position of a bag in the target region is detected; If it is determined according to the target motion trajectory of the current article and the position of the bag that the current article has implemented a bagging behavior, whether the encoding entry information about the current article is received is determined according to the current order information, and if the encoding entry information of the current article is not received, it is determined that the current article has an abnormal behavior; And / or, if it is determined according to the target motion trajectory of the current item that the moving distance of the current item exceeds a preset distance threshold, it is determined according to the current order information whether the code entry information about the current item is received, and if the code entry information of the current item is not received, it is determined that the current item has abnormal behavior.
6. A method of recognizing behavioral information, characterized by, The method is applied to a self-service checkout system, the system comprises a checkout area and a depth image collector, the checkout area is within the shooting range of the depth image collector, and the method comprises: Obtaining an original image of the checkout area and depth information corresponding to the original image through the depth image collector; According to the original image and the depth information, determining a first motion trajectory of at least one to-be-processed item in the checkout area and a second motion trajectory of a key limb part of a user; According to the first motion trajectory and the second motion trajectory, determining a current item that overlaps with the key limb part from the to-be-processed items; In response to current order information, according to the current order information and a target motion trajectory of the current item, identifying whether the current item has abnormal behavior, the abnormal behavior including code scanning behavior, bagging behavior, and long-distance moving behavior; The method according to the first motion trajectory and the second motion trajectory, determining a current item that overlaps with the key limb part from the to-be-processed items, comprises: According to the position of the to-be-processed item in the first motion trajectory and the position of the key limb part in the second motion trajectory, respectively calculating the overlap degree between each to-be-processed item and the key limb part; Determining the to-be-processed item corresponding to the overlap degree greater than a preset threshold as a candidate item that overlaps with the key limb part; According to the second motion trajectory, determining the center point position of the key limb part, and determining the to-be-processed item containing the center point position in the corresponding position range as a candidate item that overlaps with the key limb part; If the candidate item is one, determining the candidate item as the current item; If the candidate item is multiple, sorting multiple candidate items according to the corresponding depth value from small to large, and determining the candidate item ranked in the front preset position as the current item; According to the original image and the depth information, determining a first motion trajectory of at least one to-be-processed item in the checkout area, comprises: Performing classification identification on the original image to determine the position of a first item in the original image, the to-be-processed items including the first item and a second item; Positionally matching the position of the first item with the depth information to determine a target pixel region in the depth information that does not fall within the position range of the first item; Positionally aggregating the pixel points in the target pixel region, grouping the pixel points with adjacent positions and the same depth value into the same aggregated region, to obtain at least one aggregated region corresponding to the target pixel region; For each of the aggregation areas, a maximum circumscribed rectangular frame of the aggregation area is calculated, and the maximum circumscribed rectangular frame is determined as a recognized new object position; the new object positions are screened, and positions of redundant objects are removed, and the screened new object positions are determined as positions of the second articles corresponding to the target pixel area, wherein the redundant objects include one or more of a background object, an object with a position range greater than a first threshold, and / or an object with a position unit less than a second threshold, the first threshold being greater than the second threshold; According to the original image and the depth information, trajectory tracking is performed on each of the to-be-processed articles to generate the first motion trajectory of each of the to-be-processed articles.
7. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to cause the electronic device to perform the method of any one of claims 1-6.
8. A cloud device, characterized by The cloud device comprises: at least one processor; and a memory connected in communication with the at least one processor; 9. A computer-readable storage medium, characterized in that, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to cause the cloud device to perform the method of any one of claims 1-6.
10. A computer program product, characterised in that, The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the method of any one of claims 1-6 is implemented. The computer program is executed by the processor to implement the method of any one of claims 1-6.
Citation Information
Patent Citations
Risk identification method, device and system for self-service cashier
CN114529850A
Track detection method and device based on binocular vision
CN115311323A
Article delivery confirmation method, device and system
CN115601686A