Method for identifying items in a display case, display case, and storage medium

CN117011625BActive Publication Date: 2026-09-22BEIJING GENKI FOREST BEVERAGE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210449454.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-26
Publication Date
2026-09-22
Estimated Expiration
2042-04-26

AI Technical Summary

Technical Problem

在主摄像头的安装角度无法对陈列柜的陈列区域进行全面检测时,现有陈列柜将无法对陈列柜内部进行全面采集,从而导致现有的检测方式可靠性较差

Benefits of technology

[0009]本公开实施例提供了一种陈列柜中物品识别方法、陈列柜以及存储介质。在本公开实施例中,在获取到陈列柜中各摄像头采集到的视频流后,就可以对每个视频流的各视频帧进行检测,从而在各视频帧中识别包含目标动作的目标视频帧,并基于该目标视频帧确定拿取物品的第一物品信息。由于本申请中陈列柜中所安装的每个摄像头对应不同的安装角度,因此,可以得到对应不同角度的视频流。通过上述对视频流进行处理的方式,可以实现利用从不同角度采集到的视频流来对目标对象所拿取的物品进行识别,从而提高了陈列柜中物品识别的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117011625B_ABST
    Figure CN117011625B_ABST
Patent Text Reader

Abstract

The present disclosure provides a display cabinet, a display cabinet article identification method and a storage medium, wherein the method comprises: obtaining video streams collected by each camera in the display cabinet to obtain a plurality of video streams; the plurality of video streams are videos collected by each camera from a target period when the cabinet door of the display cabinet is opened to closed; detecting each video frame of each video stream to obtain a detection result; the detection result includes a detection result of a limb part of a target object taking an article, and / or a detection result of an article included in each video frame; identifying a target video frame containing a target action in each video frame based on the detection result; the target action is an action of the limb part taking the article; and determining first article information of the article taken by the limb part based on the target video frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image processing technology, and more specifically, to a method for identifying items in a display case, the display case, and a storage medium. Background Technology

[0002] In recent years, display cases have become widely used in convenience stores and supermarkets, such as unmanned self-service display cases and refrigerated display cases for displaying refrigerated goods. In existing display cases, a single camera is typically installed as the main camera to identify the items inside. When the main camera is damaged, cannot detect a target, or is obstructed, a backup camera is usually used for item identification. If the main camera's installation angle does not allow for a comprehensive view of the display area, existing display cases cannot capture complete data inside the case, resulting in poor reliability of current detection methods. Summary of the Invention

[0003] This disclosure provides at least one method for identifying items in a display case, a display case, and a storage medium. This disclosure can utilize video streams captured from different angles to identify items taken by a target object, thereby improving the accuracy of item identification in a display case.

[0004] In a first aspect, embodiments of this disclosure provide a method for identifying items in a display case, comprising: acquiring video streams captured by cameras in the display case to obtain multiple video streams; the multiple video streams being videos captured by each camera during a target time period from the opening to the closing of the display case door; detecting each video frame of each video stream to obtain detection results; the detection results including detection results of a limb part in a target object taking an item, and / or, detection results of items contained in each video frame; identifying a target video frame containing a target action in each video frame based on the detection results; the target action being an action of a limb part taking an item; and determining first item information of the item taken by the limb part based on the target video frame.

[0005] Secondly, embodiments of this disclosure provide an item recognition device in a display case, comprising: an acquisition unit, configured to acquire video streams captured by cameras in the display case to obtain multiple video streams; the multiple video streams are videos captured by each camera during a target time period from the opening to the closing of the display case door; a detection unit, configured to detect each video frame of each video stream to obtain a detection result; the detection result includes the detection result of a limb part in the target object taking an item, and / or the detection result of the item contained in each video frame; an identification unit, configured to identify a target video frame containing a target action in each video frame based on the detection result; the target action is the action of a limb part taking an item; and a determination unit, configured to determine first item information of the item taken by the limb part based on the target video frame.

[0006] Thirdly, this disclosure provides a display case, including: a display case body, multiple cameras, and a processor; wherein the multiple cameras are mounted on the display case body at different positions, and each camera is communicatively connected to the processor; each camera is configured to capture a video stream of the interior of the display case from the time the cabinet door is opened to the time the cabinet door is closed; the processor is configured to acquire the video streams captured by each camera in the display case, resulting in multiple video streams; to detect each video frame of each video stream, resulting in a detection result; the detection result includes the detection result of a limb part of a target object taking an item, and / or the detection result of an item contained in each video frame; based on the detection result, a target video frame containing a target action is identified in each video frame; the target action is the action of a limb part taking an item; and first item information of the item taken by the limb part is determined based on the target video frame.

[0007] Fourthly, embodiments of this disclosure also provide a display case, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the display case is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps of the first aspect above, or any possible implementation of the first aspect, are performed.

[0008] Fifthly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the first aspect or any possible implementation thereof.

[0009] This disclosure provides a method for identifying items in a display case, the display case itself, and a storage medium. In this embodiment, after acquiring video streams from each camera in the display case, each video frame of each video stream can be detected to identify a target video frame containing the target action. Based on this target video frame, the information of the first item to be picked up is determined. Since each camera installed in the display case corresponds to a different installation angle, video streams corresponding to different angles can be obtained. Through the above-described video stream processing, the identification of items picked up by a target object can be achieved using video streams acquired from different angles, thereby improving the accuracy of item identification in the display case.

[0010] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.

[0012] Figure 1 A flowchart of a method for identifying items in a display case provided by an embodiment of this disclosure is shown;

[0013] Figure 2 A flowchart illustrating the detection of each video frame in each video stream provided in an embodiment of this disclosure is shown;

[0014] Figure 3 This illustration shows a flowchart of identifying target video frames containing target actions in each video frame based on detection results, according to an embodiment of the present disclosure.

[0015] Figure 4 This invention discloses a flowchart illustrating the process of determining first item information based on a target video frame to identify a target object's limb parts taking up a target item, according to an embodiment of this disclosure.

[0016] Figure 5 A schematic diagram of an item identification device in a display case provided by an embodiment of this disclosure is shown;

[0017] Figure 6 A schematic diagram of the structure of a display cabinet provided according to an embodiment of this disclosure is shown;

[0018] Figure 7 A schematic diagram of another display cabinet provided in an embodiment of this disclosure is shown. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0020] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0021] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0022] Research has found that display cases have become widely used in convenience stores and supermarkets in recent years. Existing display cases typically employ a single main camera to identify items within the case. When the main camera malfunctions, fails to detect a target, or is obstructed, a backup camera is usually used for item identification. However, if the main camera's installation angle prevents a comprehensive view of the display area, existing display cases cannot capture complete internal data, leading to unreliable detection methods.

[0023] Based on the above research, in this embodiment, after acquiring the video streams captured by each camera in the display case, each video frame of each video stream can be detected to identify the target video frame containing the target action, and the first item information for the item being picked up can be determined based on the target video frame. Since each camera installed in the display case in this application corresponds to a different installation angle, video streams corresponding to different angles can be obtained. Through the above-described method of processing the video streams, it is possible to identify the item picked up by the target object using video streams acquired from different angles, thereby improving the accuracy of item identification in the display case.

[0024] In the embodiments of this disclosure, the display cabinet can be any type of cabinet with doors that can be opened and closed. For example, the display cabinet can be a display cabinet that supports unmanned self-service functions, or a refrigerated cabinet used in supermarkets to store refrigerated items. In addition, the display cabinet can be a warm cabinet, a freezer used to store frozen items, or any type of cabinet with doors that can be opened and closed. This disclosure does not specifically limit the type of display cabinet.

[0025] To facilitate understanding of this embodiment, a method for identifying items in a display case disclosed in this disclosure will first be described in detail. The execution subject of the method for identifying items in a display case provided in this disclosure is generally a display case with certain computing power, wherein the display case is a device installed in the display case.

[0026] See Figure 1 The diagram shows a flowchart of a method for identifying items in a display case according to an embodiment of this disclosure. The method includes steps S101 to S107, wherein:

[0027] S101: Obtain the video streams captured by each camera in the display case to obtain multiple video streams; the multiple video streams are the videos captured by each camera from the opening to the closing of the display case door during the target time period.

[0028] In this embodiment, multiple cameras are installed in the display case. Each camera can be installed at a different location within the display case, meaning each camera has a different viewing angle. This processing method allows for the capture of video footage of a user taking items from the display case from different angles using these multiple cameras, resulting in corresponding video streams. For example, two cameras can be installed in the display case, one on the left and one on the right of the cabinet door, with their lenses facing inwards. These two cameras can capture video footage of the user taking items from the left and right sides of the display case, respectively, resulting in two video streams. It should be understood that the number of cameras installed in the display case is not limited in this embodiment; the user can set the number according to their actual needs. Furthermore, the capture areas of any two cameras can partially overlap or not overlap at all.

[0029] In this embodiment, a sensor is pre-installed in the display case. This sensor is used to monitor the opening and closing of the display case door. For example, the sensor can be a magnetic door switch or other device capable of opening and closing the door. This disclosure does not specifically limit the application to this, but only to what is feasible. The display case includes a processor, which controls the multiple cameras. The processor and the sensor are communicatively connected.

[0030] In practice, when the sensor detects a user opening the cabinet door, it generates an opening signal and sends it to the processor. Upon receiving the opening signal, the processor controls the camera to turn on. Similarly, when the sensor detects a user closing the cabinet door, it generates a closing signal and sends it to the processor. Upon receiving the closing signal, the processor controls the camera to turn off.

[0031] S103: Detect each video frame of each video stream to obtain detection results; the detection results include the detection results of the limb parts of the target object that are picking up the item, and / or the detection results of the items contained in each video frame.

[0032] In this embodiment of the disclosure, several video frames can be extracted from each video stream first, and then the extracted video frames can be detected to obtain the detection result for each video frame. Here, several video frames can be extracted in the following ways:

[0033] Method 1: Determine all video frames in each video stream as the extracted video frames.

[0034] Method 2: Extract the corresponding video frames from each video stream at preset intervals to obtain a number of video frames.

[0035] Method 3: Extract key video frames from each video stream and determine the extracted key video frames as a number of video frames. The key video frames are the video frames where the key actions of the target object are moving or changing.

[0036] After extracting several video frames as described above, each of these video frames can be detected. Specifically, the limb part of the object being held can be detected in each video frame to obtain a corresponding detection result. For example, the hand of the object being held can be detected in each video frame to obtain a corresponding detection result, which includes the type information of the limb part and its position information in the video frame (e.g., the bounding box of the hand), such as the left hand of the object. Alternatively, the object being held can be detected in each video frame to obtain a corresponding detection result, which includes the position information of the object being held in the video frame (e.g., the bounding box of the object).

[0037] S105: Based on the detection results, identify the target video frame containing the target action in each video frame; the target action is the action of a limb part picking up an object.

[0038] In this embodiment of the disclosure, after detecting each video frame in the manner described above and obtaining the detection results, target video frames containing target actions can be identified in each video frame based on the detection results. Specifically, target video frames containing target actions can be identified in each video frame based on the position information of limb parts and the position information of objects in the video frame from the detection results.

[0039] For example, a preset position requirement can be determined in advance based on the action type of the target action. This preset position requirement indicates the positional relationship between the position of the limb part of the target object that is taking an item in the video frame and the position of the item in the video frame when the target object is performing the target action. For instance, if the preset position requirement is met between the position information of the limb part in the video frame and the position information of the item in the video frame, a video frame containing the target action is determined to be detected.

[0040] Here, the positional information of the limb part in the video frame and the positional information of the object in the video frame can be determined to meet the preset positional requirement through the following methods, specifically including:

[0041] Method 1:

[0042] If the intersection-union ratio between the bounding box of the limb and the bounding box of the object meets the preset intersection-union ratio requirement (i.e., the preset position requirement), then the preset position requirement is satisfied.

[0043] Method 2:

[0044] If the distance between the bounding box of the limb part and the bounding box of the object meets the preset distance requirement (i.e., preset position requirement), then the preset position requirement is satisfied.

[0045] In the embodiments of this disclosure, the action of the above-mentioned limb component to take an item (i.e., the target action) can be divided into a variety of types, and the type of limb component can be divided into a variety of types. For example, the limb component can be a hand, and the type of hand can be divided into left hand and right hand. Each type of limb component can correspond to at least one action of taking an item. For example, when the limb component is a hand, the corresponding action of taking an item can be grasping an item.

[0046] In the process of identifying target video frames containing target actions in each video frame based on the detection results, if any limb part is detected to take an object in a video frame, it can be determined that a target action has been detected, and the video frame is identified as the target video frame.

[0047] S107: Determine the first item information of the item picked up by the limb part based on the target video frame.

[0048] In this embodiment of the disclosure, firstly, the bounding box of the object being held by a limb in the target video frame can be determined, and based on the bounding box, the image of the object being held by the target object is extracted from the target video frame. Then, the image of the object can be recognized to determine the object information of the object held by the target object, i.e., the first object information. Here, the first object information may include at least one of the following: object category information, object price information.

[0049] As described above, in this embodiment, after acquiring the video streams captured by each camera in the display case, each video frame of each video stream can be detected to identify the target video frame containing the target action, and the first item information for the item being picked up can be determined based on the target video frame. Since each camera installed in the display case in this application corresponds to a different installation angle, video streams corresponding to different angles can be obtained. Through the above-described method of processing the video streams, it is possible to identify the item picked up by the target object using video streams acquired from different angles, thereby improving the accuracy of item identification in the display case.

[0050] In an optional implementation, if the above detection results include a first detection result and a second detection result, such as Figure 2 As shown, step S103 above, which involves detecting each video frame of each video stream to obtain detection results, specifically includes the following process:

[0051] S1031: The video frames are identified using a limb recognition model to obtain the first detection result, wherein the first detection result is used to indicate the location information and / or type information of the identified limb parts.

[0052] In this embodiment of the disclosure, limb parts in each video frame can first be identified using a limb recognition model to obtain a first detection result. Specifically, the limb recognition model can be used as a neural network model to identify preset limb types.

[0053] Before identifying body parts in each video frame using the body recognition model, the model needs to be trained. The specific training process is described below:

[0054] First, sample images are acquired, which are images captured by cameras in the display case while the user is taking items from the display case. Then, the limb parts of the user taking the items are labeled in the acquired sample images; for example, the bounding box of the limb part and the type information of the limb part (e.g., right hand) can be labeled. Next, the labeled sample images are used to train the limb recognition model to be trained, thereby obtaining the trained limb recognition model. Finally, the trained limb recognition model can be used to identify limb parts in each video frame, thereby obtaining a first detection result, wherein the first detection result includes the position information of the limb part and / or the type information of the limb part (e.g., right hand).

[0055] After training the limb recognition model as described above, the model can be used to detect limb parts of the target object in each video frame, obtaining a first detection result. Based on this first detection result, the type and / or location information of the limb parts in the video frame can be determined. For example, the first detection result could be: left hand [(x1,y1) / (x2,y2)], where (x1,y1) / (x2,y2) are the coordinates of the bounding box of the left hand, such as the coordinates of the upper left and lower right corners of the bounding box.

[0056] In this embodiment of the disclosure, the limb recognition model can choose to recognize all limb parts in each video frame. For example, it can recognize each limb part in each video frame to obtain type information and / or location information of each limb part (i.e., the first detection result mentioned above). Then, based on the first detection result, the limb part that picks up the item can be determined from the recognized limb parts, for example, the left hand.

[0057] In addition, the limb recognition model can also selectively identify specific limb parts in each video frame. For example, the model can choose to identify the target object's hands (left and right hands). If the target object's hands are not identified in a video frame, the model outputs a first detection result indicating recognition failure.

[0058] In this embodiment of the present disclosure, before performing limb recognition on each video frame, a region of interest can be determined in each video frame. Then, a limb recognition model is used to perform limb recognition on the image within the region of interest, thereby obtaining a first detection result.

[0059] S1032: The video frames are identified using an object recognition model to obtain the second detection result, wherein the second detection result is used to indicate the location information and / or item information of the identified object.

[0060] In this embodiment of the disclosure, an object recognition model can be used to identify objects in each video frame to obtain a second detection result. The object recognition model is a neural network model used to identify objects contained in each video frame. Here, the limb recognition model and the object recognition model are different neural network models.

[0061] Before identifying objects in each video frame using the object recognition model, the model needs to be trained. The specific training process is described below:

[0062] First, sample images are acquired, specifically images captured by cameras within the display case while the user is retrieving items. Then, the items retrieved are labeled within these sample images, for example, by adding bounding boxes to the items. Next, the labeled sample images are used to train the item recognition model, resulting in a trained model. Finally, the trained model is used to identify items in each video frame, yielding a second detection result that includes the item's location information.

[0063] Here, the position information of an object in a video frame can be the bounding box of the object, that is, the coordinate information of the object in the video frame. If the bounding box is rectangular, each vertex corresponds to a sub-coordinate information. The coordinate information of the bounding box of the limb can be determined based on the determined sub-coordinate information, and the position information of the object in the video frame can be determined based on the determined coordinate information.

[0064] In this embodiment of the disclosure, the item recognition model can choose to recognize all items in each video frame. For example, it can recognize each item in each video frame to obtain the location information of each item (i.e., the second detection result mentioned above). Then, based on the second detection result, the item taken by the target object can be determined from the recognized items.

[0065] In addition, the object recognition model can also selectively identify specific objects in each video frame. For example, in this embodiment of the present disclosure, before performing object recognition on each video frame, a region of interest can be determined in each video frame. Then, the object recognition model performs limb recognition on the objects within the region of interest to obtain a second detection result.

[0066] As described above, by identifying each video frame of the video stream using both a limb recognition model and an object recognition model, the first detection result of the limb parts contained in each video frame and the second detection result of the identified objects can be obtained. This enables the detection of relevant information about the limb parts and objects containing the target object in each video frame, thus providing a corresponding technical basis for identifying the target video frame in each video frame.

[0067] In an optional implementation, the above method further includes the following process:

[0068] (1) Determine the second video frame in each video stream, and determine the first adjacent video frame of the second video frame in the video stream, wherein the second video frame is the video frame in the corresponding video stream that has not been identified and obtained the first detection result;

[0069] (2) Using the position information of the limbs in the first adjacent video frame, predict the first position prediction information of the limbs in the second video frame;

[0070] (3) Use the first position prediction information to determine the type information of limb parts in the second video frame.

[0071] In this embodiment of the disclosure, a second video frame in each video stream for which the first detection result was not identified can first be determined; then, a first adjacent video frame of the second video frame is determined in the video stream. For example, the first adjacent video frame can be determined from video frames located before and / or after the second video frame in the video stream. For instance, N frames before and / or M frames after the second video frame in the video stream can be determined as the first adjacent video frame, where N and M can be the same or different.

[0072] After determining the first adjacent video frame, the positional information of the limb parts in the first adjacent video frame is determined. If the positional information of the limb parts in the first adjacent video frame is determined, the positional information of the limb parts in the second video frame can be predicted based on this positional information, thus obtaining the first positional prediction information. Here, the first positional prediction information of the limb parts in the second video frame can be predicted using the Kalman filter algorithm.

[0073] After determining the first position prediction information, the bounding box of the limb component in the second video frame can be determined based on the first position prediction information, and the limb component can be identified based on the bounding box, thereby determining the type information of the limb component in the second video frame.

[0074] In this embodiment of the disclosure, if each video stream contains multiple second video frames, the first position prediction information of the limb parts in the second video frames can be predicted according to the acquisition order of each second video frame in the video stream.

[0075] For example, video stream A contains a second video frame AN, a second video frame AN+1, and a second video frame AY. The second video frame AN ​​and the second video frame AN+1 are adjacent video frames in the video stream, and the second video frame AN+1 is a video frame captured after the second video frame AN. In this case, the first adjacent video frames of the second video frame AN ​​can be determined in the video stream. For example, the first adjacent video frames AN-2 and AN-1 in video stream A before the second video frame AN ​​can be determined, and the first adjacent video frames AN+1 and AN+2 in video stream A after the second video frame AN ​​can also be determined. Since the first adjacent video frame AN+1 is the second video frame, the first position prediction information of the limb parts in the second video frame N can be predicted based on the first adjacent video frames AN-2 and AN-1. For the second video frame AN+1, the first position prediction information of the limb parts in the second video frame N+1 can be predicted using the first adjacent video frames AN-1, AN, AN+2, and AN+3.

[0076] After predicting the first position prediction information of limb parts in the second video frame N as described above, the type information of limb parts in the second video frame can be determined using the first position prediction information. For example, the type information of the corresponding limb part can be obtained by predicting the image located within the first position prediction information using a limb recognition model.

[0077] As described above, by determining the position information of limb parts in the first adjacent video frame in the second video frame, the first position prediction information of the limb parts in the second video frame can be predicted. Thus, in the event of a missed detection, the position of the limb parts in the missed video frame can be predicted.

[0078] In an optional implementation, the above method further includes the following process:

[0079] (1) Determine the third video frame in each video stream, and determine the second adjacent video frame of the third video frame in the video stream, wherein the third video frame is the video frame in the corresponding video stream that has not been identified with the second detection result;

[0080] (2) Based on the position information of the items in the second adjacent video frame, predict the second position prediction information of the items in the third video frame;

[0081] (3) Use the second position prediction information to determine the item information of the item in the third video frame.

[0082] In this embodiment of the disclosure, a third video frame in the video stream that has not been identified as having a second detection result can first be determined. The third video frame can be a video frame in the video stream in which a limb part takes an object.

[0083] Next, it is necessary to determine the second adjacent video frame of the third video frame. For example, the second adjacent video frame can be determined from the video frames in the video stream that are located before and / or after the third video frame. For instance, N frames before and / or M frames after the third video frame can be determined as the second adjacent video frames, where N and M can be the same or different. Then, based on the position information of the item in the second adjacent video frame, the second position prediction information of the item in the third video frame can be predicted.

[0084] In practice, the position information of the object in the third video frame can be predicted first by using the Kalman filter algorithm, thereby obtaining the second position prediction information, which can be the bounding box of the object in the third video frame. Then, the object can be identified based on the bounding box, thereby determining the object information of the object in the third video frame.

[0085] In this embodiment of the disclosure, if each video stream contains multiple third video frames, the second position prediction information of the items in the third video frames can be predicted according to the acquisition order of each third video frame in the video stream.

[0086] For example, video stream B contains third video frame AP, third video frame AP+1, and third video frame AQ. Third video frame AP and third video frame AP+1 are adjacent video frames in the video stream, and third video frame AP+1 is the video frame captured after third video frame AP. In this case, the second adjacent video frame of third video frame AP can be determined in the video stream. For example, determine the second adjacent video frame AP-2 and the second adjacent video frame AP-1 in video stream B that precedes third video frame AP, and determine the second adjacent video frame AP+1 and the second adjacent video frame AP+2 in video stream A that follows third video frame AP. Since the second adjacent video frame AP+1 is the third video frame, the second position prediction information of the limb parts in third video frame P can be predicted based on the second adjacent video frames AP-2 and AP-1. For third video frame AP+1, the second position prediction information of the limb parts in third video frame P+1 can be predicted using the second adjacent video frames AP-1, AP, AP+2, and AP+3.

[0087] After predicting the second position prediction information of the item in the third video frame P as described above, the position information of the item in the third video frame can be determined using the second position prediction information. For example, an item recognition model can be used to predict the image located within the second position prediction information to obtain the item information of the corresponding item.

[0088] As described above, by determining the position information of the item in the second adjacent video frame in the third video frame, the position information of the item in the third video frame can be predicted. Thus, in the event of a missed detection, the position of the item in the missed video frame can be predicted, and the item information can be determined based on the prediction result.

[0089] In an optional implementation, step S105 above, which identifies target video frames containing target actions in each video frame based on the detection results, specifically includes the following process:

[0090] (1) Based on the detection results, determine the position information of the limb parts that pick up the items in each video frame to obtain the first position information, and determine the position information of the items contained in each video frame to obtain the second position information;

[0091] (2) Determine the video frame in each video frame whose first position information and second position information satisfy the preset position requirements, and determine the video frame that satisfies the preset position requirements as the target video frame.

[0092] In this embodiment of the disclosure, the first position information of the limb part that picks up the item in each video frame and the second position information of the item contained in each video frame can be determined respectively, and a preset position requirement can be obtained, wherein the preset position requirement can be a preset distance requirement.

[0093] In practice, based on the first and second location information, the distance between the limbs that take the items and the items in each video frame can be calculated. Then, items with a distance less than the preset distance requirement can be identified as the items taken by the target object, and the video frame of that item can be identified as the target video frame.

[0094] In another alternative implementation, the preset position requirement can be a preset Intersection over Union (IoU) requirement, where the preset IoU can be used to characterize the overlap rate between the limb bounding box of a limb component and the item bounding box of an item.

[0095] In specific implementation, the bounding box of the limb component in the video frame can be determined based on the first location information mentioned above, and the bounding box of the item in the video frame can be determined based on the second location information mentioned above. It should be understood that when the limb component picks up the item, at least a portion of the bounding box of the limb and the bounding box of the item can overlap. Therefore, the intersection-over-union ratio (IoU) between the limb and item bounding boxes can be calculated based on the overlapping portion, and the video frame whose IoU meets the aforementioned preset IoU requirement is determined as the target video frame, thereby ensuring that the target video frame is the video frame of the limb component picking up the item.

[0096] As described above, by filtering video frames according to preset position requirements, target video frames that meet the preset position requirements are determined, thereby ensuring that the obtained target video frames are video frames of limb parts taking objects.

[0097] In one optional implementation, the aforementioned limb component includes multiple sub-parts, and the aforementioned target action comprises multiple sub-actions, each sub-action corresponding to a sub-part. Therefore, as follows: Figure 3 As shown, step S105 above, which identifies target video frames containing target actions in each video frame based on the detection results, specifically includes the following process:

[0098] S1051: Using the detection results, determine the video frames in each video stream that contain each of the sub-parts, and obtain multiple first video frame groups, wherein the video frames in each first video frame group contain the same sub-parts.

[0099] In this embodiment of the disclosure, the video frames in the video stream can first be classified based on the detection results. Specifically, the video frames can be classified based on the limb types identified by the limb recognition model to obtain at least one first video frame group.

[0100] For example, the sub-parts corresponding to the limb components in the above video stream can be divided into the left hand and the right hand. Therefore, the first video frame group can be the left hand video frame group and the right hand video frame group respectively. The sub-parts contained in the video frames of the left hand video frame group are the left hand, and the sub-parts contained in the video frames of the right hand video frame group are the right hand.

[0101] S1052: Based on the detection results, detect video frames containing corresponding sub-actions in each of the first video frame groups to obtain a second video frame group.

[0102] In this embodiment of the disclosure, the video frames in each of the first video frame groups can be traversed to determine the video frames in each first video frame group that contain the sub-actions corresponding to the sub-parts, i.e., the second video frame groups. It should be understood that the sub-actions for different sub-parts to pick up objects can be different; for example, the sub-actions corresponding to picking up objects with the left hand and the right hand are different.

[0103] S1053: Determine the video frame in the second video frame group as the target video frame.

[0104] In this embodiment of the disclosure, after detection is performed based on each first video frame group, a video frame containing the sub-action corresponding to the corresponding sub-part is detected in each first video frame group, thereby determining a second video frame group based on the detected video frame, and determining the second video frame group as the target video frame.

[0105] It should be understood that for a first video frame group in which no corresponding sub-action is detected, the first video frame group does not contain a corresponding second video frame group.

[0106] As described above, the video frames are divided into multiple first video frame groups based on different sub-parts, and the sub-actions corresponding to the sub-parts in each first video frame group are identified. This allows for the fusion of video frames from different video streams captured by multiple cameras, and the fused video frames are then classified. This fully utilizes the video streams captured by each camera and improves the accuracy of obtaining the target video frames.

[0107] In one alternative implementation, such as Figure 4 As shown, step S107 above, which determines the first item information of the target object's limb parts taking the target item based on the target video frame, specifically includes the following process:

[0108] S1071: Determine the first video frame that meets the recognition requirements in each of the second video frame groups.

[0109] In this embodiment of the disclosure, the above-mentioned identification requirements can be requirements set for items in the video, wherein the identification requirements may include item detection requirements and / or quantity requirements.

[0110] When the above identification requirements include an item detection degree requirement, the item detection degree requirement can be divided into a first integrity requirement and / or a clarity requirement. Specifically, video frames in the second video frame group whose item integrity meets the first integrity requirement can be identified firstly, and / or, video frames whose clarity meets the clarity requirement can be identified from the identified video frames, and the identified video frames can be identified as the first video frames.

[0111] In practice, when determining a video frame whose completeness meets the first completeness requirement, the first completeness requirement can first be obtained, for example, it can be 80%. When determining the completeness of an item in the video frame, the proportion of the occluded part of the item in the corresponding item image can be determined. For example, when the proportion of the occluded part of the item in the corresponding item image is 10%, the completeness of the item can be determined to be 90%, and the video frame corresponding to the item meets the aforementioned first completeness requirement.

[0112] Furthermore, the aforementioned clarity requirement can be used to meet the recognition conditions for item information. Specifically, if the portion of the image corresponding to the item that can be used to identify the item information can be recognized, then the video frame corresponding to the item can be determined to meet the clarity requirement. For example, the portion used to identify the item information can be the item's name.

[0113] In this embodiment of the disclosure, the requirement for the quantity to be taken can be understood as taking the maximum quantity. In specific implementation, the video frame containing the largest number of items taken can be determined in the second video frame group, and the first video frame can be determined based on that video frame.

[0114] Based on the aforementioned requirements for completeness, clarity, and quantity, the priority of these three requirements can be determined, and the first video frame that meets the recognition requirements can be found in the second video frame group according to this priority.

[0115] S1072: Perform item identification on each of the first video frames to obtain at least one item sub-information.

[0116] In this embodiment of the disclosure, since there may be a situation where the target object holds multiple items, for example, the target object holds one item in its left hand and another item in its right hand, at least one item sub-information can be obtained for each of the aforementioned first video frames.

[0117] For example, the first video frame mentioned above is a video frame in the left-hand video frame group, in which the user picks up two items with his left hand. At this time, the video frames containing these two items (i.e., the first video frame) can be analyzed separately to obtain the corresponding item sub-information.

[0118] S1073: Determine the first item information based on the at least one item sub-information.

[0119] In this embodiment of the disclosure, during the process of determining the first item information based on the above-mentioned at least one item sub-information, the at least one item sub-information can first be deduplicated, that is, duplicate items in the obtained item sub-information can be identified to obtain multiple duplicate item sub-information. Next, the multiple duplicate item sub-information can be deduplicated to output an item sub-information and the number of times the item corresponding to the item sub-information has been taken, and the output item sub-information is determined as the first item information.

[0120] For example, if a user takes item A twice, the above item sub-information may contain two duplicate item sub-information entries for item A. After removing duplicate entries from the duplicate item sub-information, the item sub-information for item A and the number of times item A was taken can be obtained.

[0121] As described above, by identifying the first video frame in the second video frame group and performing item identification on the first video frame to obtain at least one item sub-information, and determining the first item information based on at least one item sub-information, this solution can be applied to multiple situations of taking items, thus increasing the applicability of this solution.

[0122] In an optional implementation, the above method further includes the following process:

[0123] (1) Determine the quantity change information of the items displayed in the display cabinet during the target time period;

[0124] (2) When it is determined that the quantity change information meets the preset quantity requirement, the step of determining the video frames of each sub-part contained in each video stream based on the detection result is executed to obtain multiple first video frame groups.

[0125] In this embodiment of the disclosure, the first step is to detect changes in the quantity of items displayed in the display case. For example, if a few items have been removed from the display case, this can be achieved through the following methods:

[0126] Method 1: Detect by measuring changes in the weight of the displayed items.

[0127] This method is applicable when the displayed items are of equal weight. In practice, the total weight reduction of the displayed items in the display case during the target time period can be obtained by the weight sensor in the display case. The total weight is then divided by the weight of each displayed item to determine the reduction in the number of displayed items during the target time period. Based on this reduction, the quantity change information of the displayed items can be determined.

[0128] Method 2: Detect the distance between the items in each display aisle and the cabinet door.

[0129] In this embodiment, the display cabinet includes multiple longitudinal display aisles, each containing several display items. After a user takes an item from an aisle, the distance between the item and the cabinet door changes. Since the length of each item is fixed, by dividing the distance between the item and the cabinet door, the decrease in the number of displayed items within a target time period can be determined, and the quantity change information of the displayed items can be determined based on this decrease.

[0130] Method 3: The quantity of items is determined by the input panel on the display case selected by the user.

[0131] In this embodiment of the disclosure, before opening the cabinet door to take out items, the user can first enter the quantity of items to be taken in the input panel of the display cabinet, and then take out items from the display cabinet based on the quantity of items.

[0132] After determining the quantity change information of the items displayed in the display case, a preset quantity requirement can be obtained. This preset quantity requirement can be determined according to the user's target quantity. When the quantity change information meets the preset quantity requirement, the video frames of each sub-part in each video stream can be determined based on the detection results, resulting in multiple first video frame groups. The specific process of obtaining multiple first video frame groups is as described in step S1051 above, and will not be repeated here.

[0133] As described above, by determining whether the quantity change information of the items displayed in the display cabinet within the target time period meets the preset quantity requirements, the process of determining whether to execute the process of determining the video frames of each sub-part in each video stream based on the detection results is carried out, thereby ensuring that the quantity change information meets the preset quantity requirements.

[0134] In an optional implementation, step S107 above, which determines the first item information of the item picked up by the limb part based on the target video frame, specifically includes the following process:

[0135] (1) When there are multiple target video frames, obtain the detection confidence of the bounding box of the item contained in each target video frame to obtain multiple detection confidences; wherein, the detection confidence is used to characterize the completeness of the target item contained in the bounding box;

[0136] (2) Perform item identification on the target video frames that meet the confidence requirements among the multiple detection confidence levels to obtain the first item information.

[0137] In this embodiment of the disclosure, when there are multiple target video frames, the bounding box of each target video frame can be obtained separately, and the image of the included item can be determined based on the bounding box. The image of the included item is then detected to determine the completeness of the included target item. It should be understood that the completeness corresponding to the confidence level requirement should be higher than the value corresponding to the first completeness requirement mentioned above. For example, if the value corresponding to the first completeness requirement is 80%, then the completeness corresponding to this detection confidence level requirement can be 80%.

[0138] Specifically, the method for determining the integrity of the target item is as described in step S1071 above, and will not be repeated here.

[0139] As described above, by detecting the confidence level, target video frames are filtered to obtain target video frames that meet the confidence level requirements. Item recognition is then performed on the determined target video frames, ensuring the accuracy of the recognition results.

[0140] In an optional implementation, the above method further includes the following process:

[0141] (1) Obtain the target item selected by the user, determine the item information of the target item, and obtain the second item information;

[0142] (2) Based on the first item information and the second item information, determine whether the item taken by the target object is of the same type as the target item;

[0143] (3) If the same condition is determined, obtain the quantity of items taken by the target object, and perform item settlement operation based on the quantity of items.

[0144] In this embodiment of the disclosure, before opening the display case door to take an item, the user can first select the second item they want to take through the input panel on the display case. Next, the second item information can be obtained, and based on the first item information identified in the target video frame, it can be determined whether the item type represented by the first item information is the same as the item type represented by the second item information.

[0145] If the same conditions are determined, the quantity of items taken by the user can be determined based on the above-mentioned quantity change information of the items displayed in the display cabinet within the determined target time period. The amount that the user needs to pay can then be calculated based on the quantity of the items, so that the user can perform the above-mentioned item settlement operation.

[0146] As described above, when a user selects a target item, the system can determine whether the type of the item the user takes is the same as the target item based on the second item information of that target item, thus ensuring the accuracy of the item settlement operation.

[0147] The following section will provide a detailed introduction to the item recognition method in the above display cabinet, using specific scenarios as examples. Here, we assume that the display cabinet is a self-service display cabinet.

[0148] First, the system can detect the user's selection of items in the display case. The specific detection process can be described as follows:

[0149] For example, a user's selection of items in a display case can be detected through an application, where the application is the program corresponding to that display case. In practice, the user can open the application through a terminal device and select the item they want to purchase within the application.

[0150] For example, the display case can be used to detect a user's selection of items within it. The outer cabinet of the display case can be equipped with a touchscreen, allowing users to select the items they wish to purchase.

[0151] Secondly, it can detect the user's opening and closing of the display cabinet door. Here, a sensor is pre-installed in the display cabinet to listen to the opening / closing action of the display cabinet door. For example, the sensor can be a door magnetic switch.

[0152] Subsequently, after detecting a user's opening of the display case door, multiple cameras within the display case begin capturing corresponding video, resulting in multiple video streams. These multiple video streams are then sent to the display case's controller for processing. After acquiring the video streams captured by each camera in the display case, the controller executes steps S103 to S107, specifically including: detecting each video frame of each video stream to obtain detection results; the detection results include detection results of limb parts in the target object taking an item, and / or, detection results of items contained in each video frame; and based on the detection results, identifying target video frames containing target actions in each video frame; the target action is the action of a limb part taking an item; finally, determining the first item information of the item taken by the limb part based on the target video frame.

[0153] Here, the process of the controller in the display case executing steps S103 to S107 is as described in the above embodiment of the item recognition method in the display case, and will not be described in detail here.

[0154] If a user selects an item they wish to purchase through the application, the display case controller can then send that item information to the server corresponding to the application via the network. Upon receiving this item information, the application's server can determine whether the item selected by the user matches the item listed in the first item information.

[0155] If the items are identical, the quantity of items taken by the user can be determined based on the multiple video streams mentioned above. The specific determination process is described in the embodiment corresponding to the item recognition method in the display case above, and will not be described in detail here. Then, item settlement information can be generated based on the quantity of items and pushed to the application so that the user can perform the item settlement operation based on the settlement information, which may include the amount the user needs to pay. If the items are different, an instruction message can be sent to the user to instruct them to take items from the display case again.

[0156] If a user selects an item they want to buy through the display case, the display case controller can directly determine whether the item selected by the user is the same as the item corresponding to the first item information.

[0157] If the items are identical, the quantity of items taken by the user can be determined based on the multiple video streams mentioned above. The specific determination process is described in the embodiment corresponding to the item recognition method in the display case above, and will not be described in detail here. Then, item settlement information can be generated based on the quantity of items and pushed to the application so that the user can perform the item settlement operation based on the settlement information, which may include the amount the user needs to pay. If the items are different, an instruction message can be sent to the user to instruct them to take items from the display case again.

[0158] In summary, in this embodiment, after acquiring the video streams captured by each camera in the display case, each video frame of each video stream can be detected to identify the target video frame containing the target action, and the first item information for the item being picked up can be determined based on the target video frame. Since each camera installed in the display case in this application corresponds to a different installation angle, video streams corresponding to different angles can be obtained. Through the above-described method of processing the video streams, it is possible to identify the item picked up by the target object using video streams acquired from different angles, thereby improving the accuracy of item identification in the display case.

[0159] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0160] Based on the same inventive concept, this disclosure also provides an item identification device in a display cabinet corresponding to the item identification method in the display cabinet. Since the principle of the device in this disclosure is similar to the item identification method in the display cabinet described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0161] Reference Figure 5 The illustration shows a schematic diagram of an item identification device in a display case provided in an embodiment of this disclosure. The device includes: an acquisition unit 51, a detection unit 52, an identification unit 53, and a determination unit 54; wherein,

[0162] The acquisition unit 51 is used to acquire the video streams captured by each camera in the display case to obtain multiple video streams; the multiple video streams are the videos captured by each camera from the opening to the closing of the display case door during the target time period.

[0163] The detection unit 52 is used to detect each video frame of each video stream and obtain detection results; the detection results include the detection results of the limb parts of the target object that are picking up the item, and / or the detection results of the items contained in each video frame;

[0164] The identification unit 53 is used to identify target video frames containing target actions in each video frame based on the detection results; the target action is the action of a limb part picking up an object;

[0165] The determining unit 54 is used to determine the first item information of the item picked up by the limb part based on the target video frame.

[0166] In this embodiment, after acquiring the video streams captured by each camera in the display case, each video frame of each video stream can be detected to identify the target video frame containing the target action, and the first item information for the item being picked up can be determined based on the target video frame. Since each camera installed in the display case corresponds to a different installation angle, video streams corresponding to different angles can be obtained. Through the above-described video stream processing method, the item picked up by the target object can be identified using video streams acquired from different angles, thereby improving the accuracy of item identification in the display case.

[0167] In one possible implementation, the limb component includes multiple sub-parts, the target action includes multiple sub-actions, each sub-action corresponds to a sub-part, and the recognition unit 53 is further configured to:

[0168] Using the detection results, video frames containing each of the sub-parts are determined in each video stream, resulting in multiple first video frame groups, wherein the video frames in each first video frame group contain the same sub-parts;

[0169] Based on the detection results, video frames containing the corresponding sub-actions are detected in each of the first video frame groups to obtain the second video frame group;

[0170] The video frame in the second video frame group is determined as the target video frame.

[0171] In one possible implementation, the determining unit 54 is further configured to:

[0172] In each of the second video frame groups, a first video frame that meets the recognition requirements is determined;

[0173] Item identification is performed on each of the first video frames to obtain at least one item sub-information;

[0174] The first item information is determined based on the at least one item sub-information.

[0175] In one possible implementation, the identification unit 53 is further configured to:

[0176] Determine the change in the quantity of items displayed in the display case during the target time period;

[0177] If the quantity change information is determined to meet the preset quantity requirement, the step of determining the video frames containing each of the sub-parts in each video stream based on the detection results is executed to obtain multiple first video frame groups.

[0178] In one possible implementation, the determining unit 54 is further configured to:

[0179] When there are multiple target video frames, the detection confidence of the bounding box of the item contained in each target video frame is obtained, resulting in multiple detection confidences; wherein, the detection confidence is used to characterize the completeness of the target item contained in the bounding box;

[0180] Item identification is performed on the target video frames that meet the confidence requirements among the multiple detection confidence levels to obtain the first item information.

[0181] In one possible implementation, the detection result includes a first detection result and a second detection result, and the detection unit 52 is further configured to:

[0182] The first detection result is obtained by identifying each video frame using a limb recognition model, wherein the first detection result is used to indicate the location information and / or type information of the identified limb parts;

[0183] The second detection result is obtained by identifying each video frame using an object recognition model. The second detection result is used to indicate the location information and / or item information of the identified object.

[0184] In one possible implementation, the detection unit 52 is further configured to:

[0185] A second video frame is determined in each video stream, and a first adjacent video frame of the second video frame is determined in the video stream, wherein the second video frame is a video frame in the corresponding video stream that has not been identified and has not received a first detection result;

[0186] The positional information of limb parts in the first adjacent video frame is used to predict the first positional prediction information of limb parts in the second video frame;

[0187] The type information of limb parts in the second video frame is determined using the first position prediction information.

[0188] In one possible implementation, the detection unit 52 is further configured to:

[0189] A third video frame is determined in each video stream, and a second adjacent video frame of the third video frame is determined in the video stream, wherein the third video frame is a video frame in the corresponding video stream that has not been identified with a second detection result;

[0190] Based on the position information of the items in the second adjacent video frame, predict the second position prediction information of the items in the third video frame;

[0191] The item information of the item in the third video frame is determined using the second location prediction information.

[0192] In one possible implementation, the identification unit 53 is further configured to:

[0193] Based on the detection results, the position information of the limb parts that pick up the item in each video frame is determined to obtain the first position information, and the position information of the item contained in each video frame is determined to obtain the second position information.

[0194] Among the video frames, determine the video frame in which the first position information and the second position information satisfy the preset position requirements, and determine the video frame in which the preset position requirements are satisfied as the target video frame.

[0195] In one possible implementation, the device is also used for:

[0196] Obtain the target item selected by the user, determine the item information of the target item, and obtain the second item information;

[0197] Based on the first item information and the second item information, determine whether the item taken by the target object is of the same type as the target item;

[0198] If the same conditions are determined, obtain the quantity of items taken by the target object, and perform an item settlement operation based on the quantity of items.

[0199] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.

[0200] Reference Figure 6 The diagram shown is a structural schematic of a display cabinet according to an embodiment of this disclosure. The display cabinet includes a cabinet body 111, multiple cameras 112, and a processor 113. The multiple cameras are mounted on the cabinet body at different locations, and each camera is communicatively connected to the processor.

[0201] Each of the cameras is configured to capture a video stream of the interior of the display case from the time the case door is opened until it is closed.

[0202] The processor is configured to acquire video streams captured by cameras in the display case, resulting in multiple video streams; to detect each video frame in each video stream, obtaining detection results; the detection results include detection results of limb parts in the target object picking up items, and / or detection results of items contained in each video frame; to identify target video frames containing target actions in each video frame based on the detection results; the target action is the action of limb parts picking up items; and to determine first item information of the item picked up by the limb parts based on the target video frames.

[0203] In this embodiment, after acquiring the video streams captured by each camera in the display case, each video frame of each video stream can be detected to identify the target video frame containing the target action, and the first item information for the item being picked up can be determined based on the target video frame. Since each camera installed in the display case corresponds to a different installation angle, video streams corresponding to different angles can be obtained. Through the above-described video stream processing method, the item picked up by the target object can be identified using video streams acquired from different angles, thereby improving the accuracy of item identification in the display case.

[0204] Corresponding to Figure 1 This disclosure also provides a display cabinet 700 for identifying items in a display case. Figure 7 The diagram shown is a structural schematic of the display cabinet 700 provided in an embodiment of this disclosure, including:

[0205] The system includes a processor 71, a memory 72, and a bus 73. The memory 72 stores execution instructions and includes a main memory 721 and an external memory 722. The main memory 721, also known as internal memory, temporarily stores the computational data in the processor 71, as well as data exchanged with external storage devices such as hard disks 722. The processor 71 exchanges data with the external storage devices 722 through the main memory 721. When the display case 700 is running, the processor 71 communicates with the memory 72 through the bus 73, causing the processor 71 to execute the following instructions:

[0206] The video streams captured by each camera in the display case are obtained to obtain multiple video streams; the multiple video streams are the videos captured by each camera from the time the display case door is opened to the time it is closed during the target period.

[0207] Each video frame of each video stream is detected to obtain detection results; the detection results include the detection results of the limb parts of the target object that are picking up the item, and / or the detection results of the items contained in each video frame;

[0208] Based on the detection results, target video frames containing target actions are identified in each video frame; the target action is the action of a limb part picking up an object.

[0209] The first item information of the item picked up by the limb part is determined based on the target video frame.

[0210] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the display case identification method described in the above-described method embodiments. The storage medium can be either volatile or non-volatile computer-readable storage.

[0211] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the display case identification method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.

[0212] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0213] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0214] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0215] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0216] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a display case (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0217] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.

Claims

1. A method for identifying items in a display case, characterized in that, include: The video streams captured by each camera in the display case are obtained, resulting in multiple video streams; The multiple video streams are videos captured by each camera during the target time period from when the display case door opens to when it closes; Each video frame of each video stream is detected to obtain detection results; the detection results include the detection results of the limb parts of the target object that are picking up the item, and / or the detection results of the items contained in each video frame; Based on the detection results, target video frames containing target actions are identified in each video frame; the target action is the action of a limb part picking up an object. The first item information of the item picked up by the limb component is determined based on the target video frame; the limb component includes multiple sub-parts, the target action includes multiple sub-actions, and each sub-action corresponds to a sub-part; The step of identifying target video frames containing target actions in each video frame based on the detection results includes: Using the detection results, video frames containing each of the sub-parts are determined in each video stream, resulting in multiple first video frame groups, wherein the video frames in each first video frame group contain the same sub-parts; Based on the detection results, video frames containing the corresponding sub-actions are detected in each of the first video frame groups to obtain the second video frame group; The video frame in the second video frame group is determined as the target video frame; The first item information for determining the limb parts of the target object to take the target item based on the target video frame includes: In each of the second video frame groups, a first video frame that meets the recognition requirements is determined; wherein, the recognition requirements include an item detection rate requirement and / or a quantity requirement; the item detection rate requirement includes: a first integrity requirement and / or a clarity requirement; For each of the first video frames, perform item identification to obtain at least one item sub-information; determine the first item information based on the at least one item sub-information. The method further includes: Determine the quantity change information of the items displayed in the display cabinet during the target time period; wherein, the quantity change information is detected by the following methods: detecting by the weight change of the displayed items, and detecting by the distance between the items in each display aisle and the cabinet door; If the quantity change information is determined to meet the preset quantity requirement, the step of determining the video frames containing each of the sub-parts in each video stream based on the detection results is executed to obtain multiple first video frame groups.

2. The method according to claim 1, characterized in that, The first item information for determining the item held by the limb part based on the target video frame includes: When there are multiple target video frames, the detection confidence of the bounding box of the item contained in each target video frame is obtained, resulting in multiple detection confidences; wherein, the detection confidence is used to characterize the completeness of the target item contained in the bounding box; Item identification is performed on the target video frames that meet the confidence requirements among the multiple detection confidence levels to obtain the first item information.

3. The method according to claim 1, characterized in that, The test results include a first test result and a second test result; The step of detecting each video frame of each video stream to obtain detection results includes: The first detection result is obtained by identifying each video frame using a limb recognition model, wherein the first detection result is used to indicate the location information and / or type information of the identified limb parts; The second detection result is obtained by identifying each video frame using an object recognition model. The second detection result is used to indicate the location information and / or item information of the identified object.

4. The method according to claim 3, characterized in that, The method further includes: A second video frame is determined in each video stream, and a first adjacent video frame of the second video frame is determined in the video stream, wherein the second video frame is a video frame in the corresponding video stream that has not been identified and has not obtained a first detection result; The positional information of limb parts in the first adjacent video frame is used to predict the first positional prediction information of limb parts in the second video frame; The type information of limb parts in the second video frame is determined using the first position prediction information.

5. The method according to claim 3, characterized in that, The method further includes: A third video frame is determined in each video stream, and a second adjacent video frame of the third video frame is determined in the video stream, wherein the third video frame is a video frame in the corresponding video stream that has not been identified with a second detection result; Based on the position information of the items in the second adjacent video frame, predict the second position prediction information of the items in the third video frame; The item information of the item in the third video frame is determined using the second location prediction information.

6. The method according to claim 1, characterized in that, The step of identifying target video frames containing target actions in each video frame based on the detection results includes: Based on the detection results, the position information of the limb parts that pick up the item in each video frame is determined to obtain the first position information, and the position information of the item contained in each video frame is determined to obtain the second position information. Among the video frames, determine the video frame in which the first position information and the second position information satisfy the preset position requirements, and determine the video frame in which the preset position requirements are satisfied as the target video frame.

7. The method according to claim 1, characterized in that, The method further includes: Obtain the target item selected by the user, determine the item information of the target item, and obtain the second item information; Based on the first item information and the second item information, determine whether the item taken by the target object is of the same type as the target item; If the same conditions are determined, obtain the quantity of items taken by the target object, and perform an item settlement operation based on the quantity of items.

8. A display case, characterized in that, include: The display case includes a cabinet body, multiple cameras, and a processor; wherein the multiple cameras are mounted on the display case body at different locations, and each camera is communicatively connected to the processor. Each of the cameras is configured to capture a video stream of the interior of the display case from the time the case door is opened until it is closed. The processor is configured to acquire video streams captured by cameras in the display case, obtaining multiple video streams; to detect each video frame in each video stream, obtaining detection results; the detection results include detection results of limb parts in the target object that pick up items, and / or detection results of items contained in each video frame; to identify target video frames containing target actions in each video frame based on the detection results; the target action is the action of a limb part picking up an item; to determine first item information of the item picked up by the limb part based on the target video frame; the limb part includes multiple sub-parts, the target action includes multiple sub-actions, and each sub-action corresponds to a sub-part; the step of identifying target video frames containing target actions in each video frame based on the detection results includes: using the detection results to determine video frames in each video stream that contain each sub-part, obtaining multiple first video frame groups, wherein video frames in each first video frame group contain the same sub-part; and detecting video frames containing corresponding sub-actions in each first video frame group based on the detection results, obtaining a first... Two video frame groups; determining the video frames in the second video frame group as the target video frames; determining the first item information of the target object's limb parts taking the target item based on the target video frames includes: determining a first video frame that meets the recognition requirements in each of the second video frame groups; wherein, the recognition requirements include item detection degree requirements and / or quantity requirements; the item detection degree requirements include: a first integrity requirement and / or a clarity requirement; performing item recognition on each of the first video frames to obtain at least one item sub-information; determining the first item information based on the at least one item sub-information; determining the quantity change information of the items displayed in the display cabinet during the target time period; wherein, the quantity change information is detected by the following methods: detecting by the weight change of the displayed items, detecting by the distance between the items and the cabinet door in each display aisle in the display cabinet; when it is determined that the quantity change information meets the preset quantity requirements, performing the step of determining video frames containing each of the sub-parts in each video stream based on the detection results to obtain multiple first video frame groups.

9. A display case, characterized in that, include: The system includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, the processor communicates with the memory via the bus, and the machine-readable instructions, when executed by the processor, perform the steps of the display case identification method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the method for identifying items in a display case as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Commodity detection method and device and readable storage medium

    CN111626201A

  • Commodity pick-and-place process discrimination method and device, intelligent container and readable storage medium

    CN111723777A