Intelligent order generation method combining wide-area and local-area collection and intelligent vending machine

By installing a main camera and a sub-camera on the smart vending machine, and combining video data processing technology, the problem of generating abnormal orders in the smart vending machine was solved, achieving higher detection accuracy and a better user experience.

CN114022244BActive Publication Date: 2025-12-19YOPOINT SMART RETAIL TECH LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111304690.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-03
Publication Date
2025-12-19
Estimated Expiration
2041-11-03

AI Technical Summary

Technical Problem

The existing smart vending machines generate a large number of abnormal orders, which negatively impacts the user experience. This is mainly due to the low image clarity of the products caused by the fully open door design, making it impossible to accurately detect the user's shopping behavior.

Method used

A combination of wide-area and local acquisition methods is adopted, using a main camera and a sub-camera to collect video data. The sub-video is used to replace low-resolution image areas in the main video, and order information is generated by combining the object detection model to ensure the clarity of product images.

Benefits of technology

It improved the accuracy of order detection and user experience, reduced the occurrence of abnormal orders, and enhanced the operational efficiency and user satisfaction of smart vending machines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114022244B_ABST
    Figure CN114022244B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of image processing, and solves the technical problem that the imaging clarity of the goods in the shopping video of the prior art intelligent vending machine is constantly changing, which affects the final detection result, thereby forming an abnormal order and resulting in poor user experience, and provides an intelligent order generation method combining wide-area and local-area collection and an intelligent vending machine. The method comprises: acquiring a main video collected by a main camera for all goods regions and a sub-video collected by a sub-camera for local goods regions, replacing image data of an image region of the main video with image data of an image region of the sub-video to obtain a target video, and generating goods order information according to the goods information and the corresponding motion track of each goods detected in the target video. The application uses the sub-video to replace the image region with low clarity in the main video, ensures the clarity of the goods in each frame of the target video, and improves the detection accuracy and user experience effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image analysis, and particularly relates to an intelligent order generation method combining wide-area and local-area collection and an intelligent vending machine. BACKGROUND

[0002] With the continuous development of artificial intelligence technology, the selling method of the retail industry has also undergone tremendous changes. Intelligent vending machines have spread in various places in the city, including stations, shopping malls, tourist attractions or department stores, where various types of intelligent vending machines can be found. Intelligent vending machines greatly meet the shopping needs of users in special scenarios, as they do not require a dedicated attendant, users can automatically place orders, and shopping can be completed.

[0003] The existing full-door intelligent vending machine greatly meets the autonomous selection of users in the shopping process because it allows users to select and purchase multiple items at a time and replace items multiple times during a single shopping trip. However, because the full-door intelligent vending machine allows customers to take items from the shelves multiple times or replace items, the imaging clarity of the items in the shopping video changes constantly, which affects the final detection results due to low clarity, resulting in abnormal orders and affecting the user experience. SUMMARY

[0004] Therefore, the embodiments of the present application provide an intelligent order generation method combining wide-area and local-area collection and an intelligent vending machine to solve the technical problem of poor user experience caused by the large number of abnormal orders and the complicated operation of the existing intelligent vending machine automatically generated orders.

[0005] The technical solution adopted by the present application is:

[0006] The present application provides an intelligent order generation method combining wide-area and local-area collection, which comprises:

[0007] S1: obtaining a main video collected by a main camera for all item areas and a sub-video collected by at least one sub-camera for a local item area, wherein the frame rate of the main camera and each sub-camera is the same, and the main video and each sub-video collected by the main camera and each sub-camera have the same number of frames;

[0008] S2: covering the image area of the main video with the item area of the sub-video to obtain a target video, wherein the image area of the main video is the same as the sub-video;

[0009] S3: inputting each frame image of the target video into a preset target detection model to obtain item information of a plurality of target items and motion trajectories of each target item;

[0010] S4: generating item order information according to the item information and the motion trajectories of each target item;

[0011] The main camera and each of the sub-cameras are arranged along the direction of arranging shelves of the intelligent vending machine, and each of the sub-cameras is located below the main camera.

[0012] Preferably, the S1 comprises:

[0013] S11: dividing the product placement area of the intelligent vending machine into a plurality of virtual product areas along the direction of arranging shelves of the intelligent vending machine;

[0014] S12: controlling the camera corresponding to each product area to collect video data of the corresponding product area, to obtain the main video collected by the main camera and the sub-video collected by at least one of the sub-cameras.

[0015] Preferably, the S2 comprises:

[0016] S21: obtaining image frames that coincide when a product is taken out or put back in the main video collected by the main camera of the product placement area;

[0017] S22: dividing the main video into a plurality of video segments according to each of the coinciding image frames;

[0018] S23: deleting each of the sub-videos according to each of the video segments, to obtain optimized sub-videos;

[0019] S24: replacing the local image of each frame image of each of the optimized sub-videos with the local image of the corresponding local area image of each frame image of the corresponding each of the optimized sub-videos in each of the video segments, and then synthesizing a video to obtain the target video.

[0020] Preferably, the S22 comprises:

[0021] S221: obtaining a boundary line for defining whether a product belongs to shelving or unshelving;

[0022] S222: dividing the main video into each of the video segments corresponding to shelving and unshelving of products according to different state areas of the boundary line where the products are located in adjacent image frames, in combination with the corresponding coinciding image frames.

[0023] Preferably, the S3 comprises:

[0024] S31: inputting each frame image of the target video into a preset target detection model to obtain the product information of each product in each frame image;

[0025] S32: determining the motion trajectory of each product belonging to shelving or unshelving according to the video segment to which each frame image belongs.

[0026] Preferably, the S4 comprises:

[0027] S41: obtaining a target confidence threshold of a commodity corresponding to a valid order;

[0028] S42: determining, according to the motion trajectory, a detection result of each target commodity that is removed from each frame image;

[0029] S43: comparing the confidence of the detection result corresponding to the target commodity with the target confidence threshold, if the requirement is met, outputting the commodity order information, if the requirement is not met, outputting abnormal order information and the main video and / or each sub-video and / or the target video.

[0030] Preferably, before the S1, further comprising:

[0031] S01: acquiring a video of a current state of the intelligent vending machine collected by a third camera in real time;

[0032] S02: analyzing each frame image of the video of the current state of the intelligent vending machine, and determining whether the cabinet door of the intelligent vending machine is in an open state or a closed state;

[0033] S03: when it is detected that the cabinet door of the intelligent vending machine is in the open state, controlling the main camera and each sub-camera for collecting video information corresponding to the commodity area to be opened;

[0034] S04: when it is detected that the cabinet door of the intelligent vending machine is in the closed state, controlling the main camera and each sub-camera for collecting video information corresponding to the commodity area to be closed.

[0035] The application further provides an intelligent order generation device combining wide-area and local-area collection, comprising:

[0036] a video collection module: used for acquiring a main video collected by a main camera for all commodity areas and a sub-video collected by at least one sub-camera for a local commodity area;

[0037] a video synthesis module: used for covering the image area of the main video that is the same as the sub-video with the commodity area of the sub-video to obtain a target video;

[0038] a target detection module: used for inputting each frame image of the target video into a preset target detection model to obtain commodity information of several target commodities and a motion trajectory of each target commodity;

[0039] an order module: used for generating commodity order information according to the commodity information and the motion trajectory of each target commodity;

[0040] The main camera and the sub cameras are arranged along the arrangement direction of the shelves of the intelligent vending machine, and the sub cameras are arranged below the main camera.

[0041] The application further provides an intelligent vending machine, comprising at least one processor, at least one memory, and computer program instructions stored in the memory, which, when executed by the processor, implement the method of any one of the preceding method embodiments.

[0042] The application further provides a medium having computer program instructions stored thereon, which, when executed by a processor, implement the method of any one of the preceding method embodiments.

[0043] In summary, the application has the following beneficial effects:

[0044] The application provides an intelligent order generation method combining wide-area and local-area acquisition and an intelligent vending machine, the intelligent vending machine is provided with a main camera for acquiring all commodity areas and at least one sub camera for supplementing video data of the main camera, a main video acquired by the main camera for all commodity areas and a sub video acquired by the sub camera for a local commodity area are obtained, image data of an image area of the sub video is used to replace image data of a corresponding image area of the main video, a target video is obtained, commodity order information is generated according to commodity information and corresponding motion tracks of each commodity detected in the target video. The application uses the sub video to replace an image area with low definition in the main video, ensures the definition of commodities in each frame of the target video, and improves detection accuracy and user experience effect. BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings needed in the embodiments of the application will be briefly introduced as follows, and for those skilled in the art, other drawings can also be obtained on the premise of not creating labor, and these are within the protection scope of the application.

[0046] Figure 1 A flowchart of the intelligent order generation method combining wide-area and local-area acquisition in embodiment 1 is shown in the figure.

[0047] Figure 2 A structural diagram of the intelligent vending machine with multiple cameras with different angles in embodiment 1 is shown in the figure.

[0048] Figure 3 A flowchart of acquiring the main video and the sub video in embodiment 1 is shown in the figure.

[0049] Figure 4 A flowchart of synthesizing the target video in embodiment 1 is shown in the figure.

[0050] Figure 5 A flowchart for acquiring the motion trajectory of each commodity in Example 1;

[0051] Figure 6 A flowchart for acquiring the effective commodity order in Example 1;

[0052] Figure 7 A structural diagram of the intelligent order generation device combining wide-area and local-area collection in Example 2;

[0053] Figure 8 A structural diagram of the automatic settlement system of the intelligent vending machine in Example 3;

[0054] Figure 9 A structural diagram of the intelligent vending machine in Example 4;

[0055] Figures 1 to 9 Reference signs:

[0056] 1, cabinet; 11, shelf; 12, commodity area; 2, cabinet door; 3, camera. DETAILED DESCRIPTION

[0057] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. It should be noted that, in this document, relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. In the description of the present application, it should be understood that the orientations or positional relationships indicated by terms such as center, upper, lower, front, rear, left, right, vertical, horizontal, top, bottom, inner, and outer are based on the orientations or positional relationships shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and therefore cannot be understood as indicating or implying that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. Moreover, the terms “include”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the elements defined by the statement “include” do not exclude the presence of other identical elements in the process, method, article or device that includes the elements. If there is no conflict, the features of the present application and the embodiments can be combined with each other, and are all within the protection scope of the present application.

[0058] Embodiment 1

[0059] The existing full-open door intelligent vending machine has the advantages of convenient one-time purchase of multiple goods, multiple replacement of goods in one-time shopping, quick generation of an order during a complex shopping process of a user, and quick settlement through self-settlement, compared with the existing one-time purchase of one piece of goods through one-time code scanning and the inability to reselect after purchase. However, because the full-open door intelligent vending machine allows a user to purchase multiple goods in one-time shopping and multiple goods to be shelved and unshelved, a large number of goods may be blocked during shelving or unshelving due to different blocking positions, resulting in different detection results of the same goods before and after shelving or unshelving, and causing abnormal orders to affect the user experience and the credibility of a merchant.

[0060] The present application is based on the feasibility study of obtaining shopping videos of a user shopping from an intelligent vending machine from multiple angles, setting cameras for real-time monitoring of a goods area from different directions in the goods area of the intelligent vending machine, combining the shopping videos captured by the multiple cameras, and obtaining order information of the user shopping through picture splicing and comparative analysis, and then automatically settling through a server to improve the shopping experience of the user and reduce the manual settlement process.

[0061] Specifically, please refer to Figure 2 , Figure 2 is a structural schematic view of a full-open door intelligent vending machine, which includes a cabinet 1 and a cabinet door 2. The cabinet 1 and the cabinet door 2 are rotationally connected, and when the cabinet door 2 is in a closed state relative to the cabinet 1, the cabinet door 2 covers all the goods areas of the cabinet 1 for placing goods, i.e., no goods in the cabinet can be taken out. When the cabinet door 2 is opened, all the goods in the cabinet 1 are exposed to the user, and the user can select any goods or multiple goods in one-time shopping, take out the selected goods, or put back the selected goods that need to be put back. The cabinet 1 is provided with shelves 11, which can divide the cabinet 1 into multiple goods areas 12. Each goods area in the cabinet 1 is provided with a camera, so that shopping videos of a user shopping from the intelligent vending machine can be obtained from multiple angles, avoiding the problem of low reliability of the shopping videos collected from a single angle due to blocking. Figure 2 As shown in the figure, the intelligent vending machine is provided with multiple cameras on the left and right inner side walls of the vending machine, so that the shopping videos of the same goods area can be collected from opposite visual angle directions, improving the reliability of the video data.

[0062] Please refer to Figure 1 , Figure 1A flowchart of an intelligent order generation method combining wide-area and local-area collection in embodiment 1 of the present application, the method comprising:

[0063] S1: obtaining a main video collected by a main camera for all commodity areas and a sub-video collected by at least one sub-camera for a local commodity area, wherein the frame rate of the main camera and each sub-camera is the same and the number of frames of the main video and each sub-video collected is the same;

[0064] Specifically, the intelligent vending machine has multiple shelves, so the intelligent vending machine is divided into multiple commodity areas, each commodity area includes at least one shelf, and a main camera or a sub-camera is arranged opposite to each other in each commodity area; for ease of understanding, the camera arranged on the left side plate or the right side plate or the inner side of the top plate of the topmost commodity area is referred to as a main camera, and the camera arranged on the left side or the right side of a commodity area other than the topmost commodity area is referred to as a sub-camera; for example, the commodity areas of the intelligent vending machine are divided into upper, middle and lower commodity areas, a main camera and a second main camera are arranged opposite to each other on the top, the main camera has a viewing direction from left to right, and a sub-camera is arranged on the left side of the middle commodity area and the lower commodity area; when a user starts to purchase commodities through the intelligent vending machine, the video data of the commodity area is collected by controlling each camera to obtain a main video and each sub-video; it should be noted that the video frame rate of the main video and each sub-video is the same, and the number of image frames of the main video and the sub-video is the same in the same shopping event.

[0065] In an embodiment, referring to Figure 3 , the S1 comprises:

[0066] S11: dividing the commodity area of the intelligent vending machine into multiple virtual commodity areas along the arrangement direction of the shelves of the intelligent vending machine;

[0067] Specifically, the intelligent vending machine has multiple shelves, and the commodity area of the intelligent vending machine is divided into multiple commodity areas, each commodity area includes at least one shelf, and the viewing angle of the camera is arranged along the arrangement direction of the shelves; for example, the shelves of the intelligent vending machine include multiple shelves from top to bottom, each camera is arranged on the left side wall or the right side wall of the intelligent vending machine, and the viewing angle of each camera is from the top left to the bottom right, from the top right to the bottom left, or from top to bottom. Among them, the cameras arranged on different sides of the same commodity area have the same installation height.

[0068] S12: controlling the camera corresponding to each commodity area to collect video data of the corresponding commodity area to obtain the main video collected by the main camera and the sub-video collected by at least one sub-camera.

[0069] Specifically, the main camera capable of collecting the image information of the first view angle of all the commodity areas is arranged in the commodity area, and at least one sub-camera corresponding to the main camera is arranged to obtain clearer images of the commodity areas. The sub-camera mainly obtains the image information of a certain commodity area, and the view angle of the sub-camera is the second view angle. The first view angle and the second view angle have an angle difference γ, and 0° < γ < 90°. For example, the commodity area of the intelligent vending machine is divided into three commodity areas, i.e., upper, middle and lower commodity areas. The main camera is arranged at the top, and the view angle direction of the main camera is from left to right. The sub-camera is arranged at the left side of the middle and lower commodity areas, and the view angle direction of the sub-camera is from left to right. The cameras collect the image information of the corresponding commodity areas to obtain the main video and the sub-video.

[0070] In an embodiment, the S1 further includes:

[0071] Step 1: obtaining the frame rate and the number of cameras for collecting the video data of the commodity area;

[0072] Step 2: determining the interval time for each camera to start collecting the video data of the corresponding commodity area according to the frame rate and the number of cameras;

[0073] Step 3: obtaining the starting time for the main camera and each sub-camera to start collecting the video data according to the interval time;

[0074] Step 4: controlling the corresponding camera to collect the video data of the commodity area according to the starting time to obtain the main video and each sub-video.

[0075] Specifically, the frame rates of the cameras for collecting the video data are the same, such as 20 frames per second. The starting time for each camera to start collecting the video data is determined according to the number of cameras and the frame rate. The starting time for each camera or each group of cameras to start collecting the video data has a time interval, and the interval time is preferably an integer multiple of the time difference corresponding to two adjacent frames of images. For example, four cameras are included, and the starting collection time of each camera is separated by 1 / 4 of the time corresponding to the frame rate. Alternatively, the four cameras are divided into two groups, and the interval time for each group of cameras to start collecting the video data is 1 / 2 of the time corresponding to the frame rate. In this way, the image frame rate is increased, and more image information of the commodity area at different time points is collected to improve the detection accuracy.

[0076] S2: covering the image area of the main video with the commodity area of the sub-video to obtain a target video;

[0077] Specifically, because the main camera is arranged at the topmost commodity area, the picture of the topmost commodity area can be collected by the main camera without obstruction in the non-topmost commodity area. Therefore, each frame image of the main video contains the image area of each frame image of each sub-video. The image area of the non-topmost commodity area corresponding to each frame image of the main video is replaced by the corresponding frame image of each sub-video, so that the target video is obtained. In the picture corresponding to each commodity area in each frame image of the target video, the imaging of the commodity is clearer, and the detection accuracy is improved.

[0078] In an embodiment, referring to Figure 4 , the S2 comprises:

[0079] S21: acquiring an image frame in which there is overlap when the commodity is taken out or put back in the main video collected by the main camera in the commodity area;

[0080] S22: dividing the main video into multiple video segments according to each of the overlapping image frames;

[0081] Specifically, the detection is performed on each frame image of the main video, and the relative position relationship between the user's commodity taking part and the intelligent vending machine is mainly detected, such as detecting whether the user's hand enters the commodity area. The image frame in which the hand enters the commodity area is recorded as an overlapping image frame. Then, the main video is divided into multiple video segments by taking the image frame as a boundary, which is used to represent multiple put-in and taking operations of the user in selecting commodities.

[0082] In an embodiment, the S22 comprises:

[0083] S221: acquiring a boundary line used to define that the commodity belongs to shelving and unshelving;

[0084] Specifically, a virtual boundary line is arranged in the image area corresponding to the camera. The virtual boundary line can be closed or not closed. Taking the closed virtual boundary line as an example, when the commodity moves from inside the boundary line to outside the boundary line, it is recorded as a first state, and when the commodity moves from outside the boundary line to inside the boundary line, it is recorded as a second state. The first state is recorded as unshelving of the commodity, and the second state is recorded as shelving of the commodity.

[0085] S222: according to the different state areas of the commodity located in the boundary line in adjacent image frames, combining the corresponding overlapping image frames, dividing the main video into each of the video segments corresponding to the shelving and unshelving of the commodity.

[0086] Specifically, according to the different position information of the same commodity about the boundary line in two adjacent frames of images, it is determined that the commodity is in the on-shelf or off-shelf state. For example, if the boundary line is a rectangular frame, the commodity area is located in the rectangular frame, the commodity A in the previous frame of image belongs to the boundary line, and the commodity A in the next frame of image belongs to the boundary line, it is considered that the operation on the commodity A this time is to take the commodity, which is recorded as the commodity off-shelf, so as to obtain a video segment of taking the commodity. Similarly, if the commodity A in the previous frame of image belongs to the boundary line, and the commodity A in the next frame of image belongs to the boundary line, it is considered that the operation on the commodity A this time is to put back the commodity, which is recorded as the commodity on-shelf, so as to obtain a video segment of putting back the commodity.

[0087] S23: According to each video segment, each sub-video is pruned to obtain each optimized sub-video;

[0088] S24: The local area of each frame of image of each optimized sub-video is replaced with the local area image of the corresponding frame of image of each optimized sub-video in the corresponding video segment, and then a video is synthesized to obtain the target video.

[0089] Specifically, according to the image area involved in each video segment of the main video, each sub-video is pruned, and the image data irrelevant to the corresponding video segment is pruned. Each sub-video after pruning is recorded as an optimized sub-video. The video data of each optimized sub-video is used to replace the corresponding video data of each video segment of the main video, so as to obtain a target video. In the target video, the replaced video data is all in the area where the target commodity exists. Further, the image area where the target imaging is detected is merged, so as to avoid reducing the image quality by scaling the image and improve the detection accuracy.

[0090] In an embodiment, the S2 comprises:

[0091] First step: obtaining a target commodity area corresponding to a commodity change;

[0092] Second step: determining the sub-video containing the target commodity area according to the position information of the target commodity area;

[0093] Specifically, when a user takes or puts back a commodity from the intelligent vending machine, the position where the commodity is taken out or put in is determined, specifically, the commodity is put into or taken out of the commodity area, and then the nearest sub-camera and the corresponding sub-video of the area are determined. For example, the commodities of the intelligent vending machine are divided into upper, middle and lower commodity areas. When a user takes or puts back a commodity from the lower commodity area, the sub-video of the target commodity area is only the sub-video corresponding to the lower commodity area.

[0094] In an embodiment, the determination of the sub-video containing the target commodity area according to the position information of the target commodity area comprises:

[0095] First, analyze each frame image of the main video to determine each image frame corresponding to the occurrence of commodity change;

[0096] Second, divide the main video into multiple video segments according to each image frame corresponding to the occurrence of commodity change;

[0097] Finally, determine the sub-video corresponding to each video segment according to the target commodity region corresponding to the occurrence of commodity change in each video segment.

[0098] Third step: merge the sub-video containing the target commodity with the main video to obtain the target video.

[0099] Specifically, by analyzing each frame image of the main video, the position of each time the user picks up or puts down the commodity is determined, and the corresponding sub-video is matched for each time the commodity is picked up or put down. The redundant sub-video can be deleted, thereby reducing the data processing amount.

[0100] In an embodiment, the sub-video containing the target commodity region is recorded as a target sub-video, and the merging of the sub-video containing the target commodity with the main video to obtain the target video comprises:

[0101] First, according to the position information of the target sub-video, other sub-videos located between the target sub-video and the main video are determined as intermediate sub-videos;

[0102] Second, delete the region in the main video corresponding to the target sub-video and the intermediate sub-video to obtain an optimized first video;

[0103] Finally, merge the sub-video in the target sub-video and the intermediate sub-video with the first video to obtain the target video.

[0104] Specifically, the image region between the main video and the target sub-video is replaced with the image region of the corresponding sub-video, thereby ensuring the integrity of the image of the synthesized target video, so as to facilitate manual review by the back-end when an abnormal order occurs.

[0105] S3: input each frame image of the target video into a preset target detection model to obtain commodity information of each commodity in each frame image and motion trajectories of each commodity;

[0106] In an embodiment, please refer to Figure 5 , S3 comprises:

[0107] S31: input each frame image of the target video into a preset target detection model to obtain the commodity information of each commodity in each frame image;

[0108] S32: determining the motion trajectory of each commodity as being on-shelf or off-shelf according to the video segment to which each frame image belongs.

[0109] Specifically, each frame image is sent into a target detection model for detection to obtain the commodity category of each commodity and the price and other information of each commodity. According to the imaging position and imaging size of the commodity in different image frames, the motion direction of the commodity, i.e., whether the commodity is taken out or put into the intelligent vending machine, can be determined.

[0110] S4: generating commodity order information according to the commodity information and the motion trajectory of each commodity.

[0111] In an embodiment, referring to Figure 6 , the S4 comprises:

[0112] S41: obtaining a target confidence threshold of a commodity corresponding to a valid order;

[0113] Specifically, when it is preliminarily determined that the commodity belongs to the commodity purchased by the user this time, the confidence of the detected commodity needs to be judged to finally determine whether the order is an abnormal order.

[0114] S42: determining, according to the motion trajectory, the detection result of each target commodity detected from each frame image as being off-shelf;

[0115] Specifically, according to the motion trajectory of each frame image as being put in or taken out, it is determined which commodities in each frame image are finally put back into the intelligent vending machine and which commodities are finally selected by the user. Then, the detection result of each commodity selected by the user is determined, and the detection result includes the category and confidence of each time the commodity is detected.

[0116] S43: comparing the confidence of the detection result corresponding to the target commodity with the target confidence threshold. If the requirement is met, the commodity order information is output. If the requirement is not met, abnormal order information and the main video and / or each sub-video and / or the target video are output.

[0117] Specifically, when it is considered that the target commodity corresponding to the current detection result is abnormal, manual review is needed. At this time, manual review can be performed on the main video and / or each sub-video and / or the target video to ensure the accuracy of the order and improve the user experience.

[0118] In an embodiment, before the S10, the method further comprises:

[0119] S01: acquiring, in real time, a video of the current state of the intelligent vending machine collected by a third camera;

[0120] Specifically, the third camera is further arranged on the intelligent vending machine, and the third camera is used to detect whether the intelligent vending machine is opened or closed.

[0121] S02: analyzing each frame image of the video of the current state of the intelligent vending machine to determine whether the cabinet door of the intelligent vending machine is in an open state or a closed state;

[0122] S03: when it is detected that the cabinet door of the intelligent vending machine is in an open state, the main camera and each sub-main camera for collecting video information corresponding to the commodity area are controlled to be opened;

[0123] S04: when it is detected that the cabinet door of the intelligent vending machine is in a closed state, the main camera and each sub-camera for collecting video information corresponding to the commodity area are controlled to be closed.

[0124] Specifically, when the user performs automatic shopping, each frame image of the vending machine state video is analyzed to determine the state of the vending machine cabinet door, when it is detected that the vending machine cabinet door is opened, the main camera and each sub-camera are opened to obtain the video data of the commodity area, and the main video and each sub-video are obtained; when it is detected that the vending machine cabinet door is closed, the main camera and each sub-camera are closed, the invalid video for producing commodity orders is reduced, and the data processing efficiency is improved.

[0125] The wide-area and local-area acquisition combined intelligent order generation method of the embodiment is provided with a main camera for collecting all commodity areas and at least one sub-camera for supplementing video data of the main camera on the intelligent vending machine, the main video collected by the main camera for all commodity areas and the sub-video collected by the sub-camera for local commodity areas are obtained, the image data of the image area of the sub-video is replaced with the image data of the corresponding image area of the main video to obtain a target video, and commodity order information is generated according to the commodity information and the corresponding motion track of each commodity detected in the target video. The present application uses the sub-video to replace the image area with low definition in the main video, ensures the definition of commodities in each frame image of the target video, and improves the detection accuracy and user experience effect.

[0126] Embodiment 2

[0127] The embodiment 2 of the present application further provides a wide-area and local-area acquisition combined intelligent order generation device based on the method of the embodiment 1, please refer to Figure 7 , comprising:

[0128] The video acquisition module is configured to acquire a main video collected by a main camera for all commodity areas and a sub-video collected by at least one sub-camera for a local commodity area, wherein the main camera and each sub-camera have the same frame rate and collect the same number of frames of the main video and each sub-video;

[0129] The video synthesis module is configured to cover the same image area of the main video with the commodity area of the sub-video to obtain a target video;

[0130] The target detection module is configured to input each frame image of the target video into a preset target detection model to obtain commodity information of a plurality of target commodities and a motion trajectory of each target commodity;

[0131] The order module is configured to generate commodity order information according to the commodity information and the motion trajectory of each target commodity;

[0132] The main camera and each sub-camera are arranged along the arrangement direction of each layer of shelves of the intelligent vending machine, and each sub-camera is located below the main camera.

[0133] The wide-area and local-area acquisition combined intelligent order generation device of the embodiment is provided with a main camera for collecting all commodity areas and at least one sub-camera for supplementing video data of the main camera on the intelligent vending machine, a main video collected by the main camera for all commodity areas and a sub-video collected by the sub-camera for a local commodity area are acquired, image data of an image area of the sub-video is replaced with image data of a corresponding image area of the main video to obtain a target video, commodity order information is generated according to commodity information and a corresponding motion trajectory of each commodity detected in the target video. The present application replaces the image area with low definition in the main video with the sub-video, ensures the definition of commodities in each frame image of the target video, and improves the detection accuracy and user experience effect.

[0134] Embodiment 3

[0135] The present application provides an automatic settlement system of an intelligent vending machine, please refer to Figure 8The automatic settlement system includes the intelligent vending machine, the mobile terminal and the server, and can adopt the automatic shopping method in the above embodiments. The user identifies the identification code on the intelligent vending machine through the mobile terminal, the server establishes the shopping event of the user, the camera of different perspectives starts to collect the shopping video or starts to collect the shopping video after the cabinet door of the intelligent vending machine is opened or the camera starts to collect the shopping video when the user enters the preset range, and the camera stops collecting the shopping video and transmits the shopping video to the server when the user leaves the preset shopping range or the cabinet door of the intelligent vending machine is closed. The server generates the order information of the user according to the shopping video and sends the order information to the mobile terminal, and the user independently settles or sets the automatic settlement through the order information of the mobile terminal. The automatic settlement system has better selectivity of user independent shopping, high order accuracy, and can improve the shopping experience of the user.

[0136] Embodiment 4

[0137] The present application provides an intelligent vending machine device and a storage medium, as shown in the accompanying drawings, comprising at least one processor, at least one memory and computer program instructions stored in the memory. Figure 9

[0138] Specifically, the above-mentioned processor can include a central processing unit (CPU), or a specific integrated circuit (Application Specific Integrated Circuit, ASIC), or can be configured to implement one or more integrated circuits of the embodiments of the present application.

[0139] The memory can include a mass storage for data or instructions. By way of example and not limitation, the memory can include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive or a combination of two or more of these. Where appropriate, the memory can include removable or non-removable (or fixed) media. Where appropriate, the memory can be internal or external to the data processing device. In certain embodiments, the memory is a non-volatile solid-state memory. In certain embodiments, the memory includes read-only memory (ROM). Where appropriate, the ROM can be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these.

[0140] ​The processor realizes the intelligent order generation method combining wide-area and local-area collection in any one of the above embodiment modes by reading and executing computer program instructions stored in the memory.

[0141] In one example, the electronic device can further include a communication interface and a bus. The processor, the memory, and the communication interface are connected through the bus and complete communication with each other.

[0142] The communication interface is mainly used to realize the communication between the modules, devices, units and / or equipment in the embodiments of the application.

[0143] The bus includes hardware, software or both to couple the components of the electronic device to each other. By way of example, and not limitation, the bus can include an accelerated graphics port (AGP) or other graphics bus, an enhanced industry standard architecture (EISA) bus, a front-side bus (FSB), a HyperTransport (HT) interconnect, an industry standard architecture (ISA) bus, an InfiniBand (IB) interconnect, a low pin count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a peripheral component interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a serial advanced technology attachment (SATA) bus, a video electronics standards association local (VLB) bus, or another suitable bus or combination of two or more of these. Where appropriate, the bus can include one or more buses. Although the present embodiments describe and show a particular bus, the present application contemplates any suitable bus or interconnect.

[0144] In summary, the embodiments of the present application provide a wide-area and local-area collection combined intelligent order generation method, device, intelligent vending machine and storage medium.

[0145] It should be understood that the present application is not limited to the particular configurations and processes described and illustrated herein. Detailed descriptions of known methods are omitted for the sake of brevity. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order of the steps, after understanding the spirit of the present application.

[0146] The functional blocks shown in the structural block diagrams described above can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, functional cards, and the like. When implemented in software, the elements of the present application are program or code segments that are used to perform the required tasks. The program or code segments can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave over a transmission medium or communication link. A "machine-readable medium" includes any medium that can store or transfer information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, and the like. The code segments can be downloaded via computer networks such as the Internet, intranets, and the like.

[0147] Finally, it should be noted that the above-described embodiments are merely intended to illustrate the technical solutions of the present application, and are not intended to limit the present application; even though the present application has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the above-described embodiments, or make equivalent replacements for some or all of the technical features; and such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for intelligent order generation combining wide-area and local-area collection, characterized in that, The method comprises: S1: acquiring a main video collected by a main camera for a whole commodity area and a sub-video collected by at least one sub-camera for a local commodity area, wherein the main camera and each sub-camera have the same frame rate and collect the main video and each sub-video with the same number of frames; S2: covering the commodity area of the sub-video with the same image area of the main video to obtain a target video; S3: inputting each frame image of the target video into a preset target detection model to obtain commodity information of several target commodities and motion trajectories of each target commodity; S4: generating commodity order information according to the commodity information and the motion trajectories of each target commodity; Wherein, the main camera and each sub-camera are arranged along the arrangement direction of each layer of shelves of the intelligent vending machine, and each sub-camera is located below the main camera; The S2 comprises: S21: acquiring image frames that coincide when a commodity is taken out or put back in the main video collected by the main camera for a commodity area; S22: dividing the main video into multiple video segments according to each of the coinciding image frames; S23: deleting each sub-video according to each video segment to obtain each optimized sub-video; S24: replacing the local area image of each frame image of each optimized sub-video with the local area image of the corresponding frame image of each optimized sub-video in the corresponding video segment, and then synthesizing the video to obtain the target video; The S1 further comprises: Acquiring the frame rate and the number of cameras for collecting commodity area video data; According to the frame rate and the number of cameras, determining the interval time for each camera to start collecting video data of the corresponding commodity area; According to the interval time, respectively obtaining the starting time for the main camera and each sub-camera to start collecting video data; According to each starting time, controlling the corresponding camera to collect video data of the commodity area to obtain the main video and each sub-video; The S2 comprises: Acquiring a target commodity area where a commodity changes; According to the position information of the target commodity area, determining the sub-video containing the target commodity area; According to the position information of the target commodity area, determining the sub-video containing the target commodity area comprises: Analyzing each frame image of the main video to determine each image frame where a commodity changes; According to each image frame where a commodity changes, dividing the main video into multiple video segments; According to the target commodity area corresponding to each video segment where a commodity changes, determining the sub-video corresponding to each video segment; Merging the sub-video containing the target commodity with the main video to obtain the target video; The sub-video containing the target commodity area is recorded as a target sub-video, and the sub-video containing the target commodity is merged with the main video to obtain the target video, which comprises: According to the position information of the target sub-video, determining other sub-videos located between the target sub-video and the main video as intermediate sub-videos; Delete the region in the main video corresponding to the target sub-video and the intermediate sub-video, respectively, to obtain an optimized first video; Merge the sub-video in the target sub-video and the intermediate sub-video with the first video to obtain the target video.

2. The method of claim 1, wherein, The S1 comprises: S11: Divide the product placement area of the intelligent vending machine into a plurality of virtual product areas along the arrangement direction of the shelves of the intelligent vending machine; S12: Control the camera corresponding to each product area to collect video data of the corresponding product area to obtain the main video collected by the main camera and the sub-video collected by at least one sub-camera. 3.The method of claim 1, wherein, The S22 comprises: S221: Obtain a boundary line for defining whether a product is on a shelf or off a shelf; S222: According to the different state areas of the product located on the boundary line in adjacent image frames, combine the corresponding overlapping image frames, and divide the main video into each video segment corresponding to the product on the shelf and the product off the shelf.

4. The method of claim 3, wherein the method further comprises: The S3 comprises: S31: Input each frame image of the target video into a preset target detection model to obtain the product information of each product in each frame image; S32: According to the video segment to which each frame image belongs, determine the motion trajectory of each product belonging to the on-shelf or off-shelf. 5.The smart order generation method of claim 1, wherein, The S4 comprises: S41: Obtain a target confidence threshold of a product corresponding to a valid order; S42: According to the motion trajectory, determine the detection result of each target product belonging to the off-shelf detected from each frame image; S43: Compare the confidence of the detection result corresponding to the target product with the target confidence threshold. If it meets the requirements, output the product order information. If it does not meet the requirements, output abnormal order information and the main video and / or each sub-video and / or the target video. 6.The method of claim 1 to 5, wherein, Before the S1, it further comprises: S01: Real-time acquire a video of the current state of the intelligent vending machine collected by a third camera; S02: Analyze each frame image of the video of the current state of the intelligent vending machine to determine whether the cabinet door of the intelligent vending machine is in an open or closed state; S03: When it is detected that the cabinet door of the intelligent vending machine is in an open state, control the main camera and the sub-camera for collecting video information corresponding to the product area to be opened; S04: When it is detected that the cabinet door of the intelligent vending machine is in a closed state, control the main camera and each sub-camera for collecting video information corresponding to the product area to be closed.

7. A wide-area and local-area collection combined intelligent order generation device, characterized in that, The device comprises: A video acquisition module: configured to acquire a main video collected by a main camera for all product areas and a sub-video collected by at least one sub-camera for a local product area, wherein the frame rate of the main camera and each sub-camera is the same, and the main video and each sub-video collected by the main camera and each sub-camera have the same number of frames; A video synthesis module: configured to cover the image area of the main video with the product area of the sub-video to obtain a target video; A target detection module: configured to input each frame image of the target video into a preset target detection model to obtain the product information of a plurality of target products and the motion trajectory of each target product; An order module is configured to generate order information of the target goods according to the goods information and the motion track of each target good; The target video is obtained by covering the image region of the main video with the goods region of the sub-video in the same image region of the main video as the sub-video. An image frame that overlaps when the goods are taken out or put back in the main video of the main camera capturing the goods region is obtained. The main video is divided into a plurality of video segments according to each of the overlapping image frames. Each of the sub-videos is pruned according to each of the video segments to obtain an optimized sub-video. The target video is obtained by replacing the local image of each frame image of each of the optimized sub-videos with the local image of the corresponding local region of each frame image of the corresponding video segment of each of the optimized sub-videos, and then synthesizing the video. The main video captured by the main camera for all goods regions and the sub-video captured by at least one sub-camera for a local goods region further include: The frame rate and the number of cameras of the camera used to capture the goods region video data are obtained. According to the frame rate and the number of cameras, the interval time at which each camera starts to capture video data of the corresponding goods region is determined. According to the interval time, the starting time at which the main camera and each of the sub-cameras starts to capture video data is obtained. According to each of the starting times, the video data of the corresponding camera is controlled to capture the goods region, and the main video and each of the sub-videos are obtained. The target video is obtained by covering the image region of the main video with the goods region of the sub-video in the same image region of the main video as the sub-video. The target goods region where the goods change occurs is obtained. According to the position information of the target goods region, the sub-video containing the target goods region is determined. According to the position information of the target goods region, the sub-video containing the target goods region is determined. Each image frame where the goods change occurs is determined by analyzing each frame image of the main video. The main video is divided into a plurality of video segments according to each image frame where the goods change occurs. According to the target goods region in each of the video segments where the goods change occurs, the sub-video corresponding to each of the video segments is determined. The target video is obtained by merging the sub-video containing the target goods with the main video. According to the position information of the target sub-video, other sub-videos located between the target sub-video and the main video are determined as intermediate sub-videos. The region corresponding to the target sub-video and the intermediate sub-video in the main video is deleted to obtain an optimized first video. The target sub-video and the intermediate sub-video are merged with the first video to obtain the target video. The target video is obtained by covering the image region of the main video with the goods region of the sub-video in the same image region of the main video as the sub-video.

8. A machine according to claim 7, wherein the machine is a vending machine. ​ at least one processor, at least one memory, and computer program instructions stored in the memory that, when executed by the processor, implement the method of any of claims 1-6.

9. A medium having stored thereon computer program instructions, characterized in that, The computer program instructions, when executed by the processor, implement the method of any of claims 1-6.

Citation Information

Patent Citations

  • Automatic commodity identification and settlement vending cabinet, method and system

    CN109979130A

  • Shelf commodity detection method and system

    CN110472515A

  • Three-dimensional printing process monitoring method, device and equipment and storage medium

    CN112596981A