Order generation method, device and intelligent vending machine based on commodity trajectory segmentation

By adopting an order generation method based on product trajectory segmentation in smart vending machines, the problem of low detection accuracy due to product occlusion in smart vending machines is solved, and higher detection accuracy and user experience are achieved.

CN114627422BActive Publication Date: 2025-06-06YOPOINT SMART RETAIL TECH LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210298768.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-09
Publication Date
2025-06-06
Estimated Expiration
2041-11-09

AI Technical Summary

Technical Problem

Because existing smart vending machines rely on shopping videos for order settlement, they are prone to low detection accuracy due to product obstruction and abnormal orders.

Method used

The order generation method based on product track segmentation is adopted. By obtaining the target video of the user's shopping, the images of each frame are divided into first frame images and non-first frame images according to the acquisition time, the basic confidence is added, and the target detection network is used to detect product information to generate an order.

Benefits of technology

Improve detection accuracy, reduce the generation of abnormal orders, and improve user experience and merchant credibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627422B_ABST
    Figure CN114627422B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of image processing technology, solves the technical problem of low detection accuracy of intelligent vending machines in the prior art, which leads to abnormal orders, and provides an order generation method, device and intelligent vending machine based on commodity trajectory segmentation, obtains a target video of a user shopping through a smart vending machine, and divides each frame image of the target video into a first frame image and a non-first frame image according to commodity change information; then adds a basic confidence to each target in each frame image, and uses a target detection model to detect each frame image of the target video, and obtains the order information corresponding to each target commodity according to the detection result of each target combined with the corresponding basic confidence; uses the imaging distance difference between the first frame image and the non-first frame image to set the corresponding basic confidence, which can improve the image with close distance and many features to guide the target detection result, thereby improving the detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application filed on November 9, 2021, with the invention name "Multi-perspective identification of goods and intelligent order generation method, device and intelligent vending machine" and application number 202111318651.1. Technical Field

[0002] The present invention relates to the field of image analysis technology, and in particular to an order generation method, device and intelligent vending machine based on commodity trajectory segmentation. Background Art

[0003] With the continuous development of artificial intelligence technology, the sales methods of the retail industry have also undergone tremendous changes. Smart vending machines have been spread across various occasions in the city, including stations, shopping malls, tourist attractions or department stores. Various types of smart vending machines can be found. Smart vending machines do not require special supervision, users can automatically place orders and check out, which greatly meets the shopping needs of users in special scenarios.

[0004] However, existing smart vending machines include fully-opening smart vending machines. When the cabinet door of the fully-opening smart vending machine is open, users can perform multiple pick-up and drop operations in one shopping trip, and can select multiple items at a time and then settle the bill. This smart vending machine greatly facilitates the shopping needs of users. However, since this type of smart vending machine mainly relies on shopping videos to settle product orders, when users are picking up and dropping products, some features of the products will be blocked, which can easily lead to false detection and abnormal order generation. In order to eliminate the problem of false detection of products due to occlusion, shopping videos are usually collected from multiple angles, and then each frame of each video is detected, and the final product order is determined by comparing the detection results of each video. Because multiple videos need to be detected, the processor needs to have multi-threaded data processing capabilities, which requires greater computing power and cost, or single-threaded processing requires queuing and waiting, and the processing efficiency is low, which affects the user experience. Summary of the invention

[0005] In view of this, the embodiments of the present invention provide an order generation method, device and intelligent vending machine based on commodity trajectory segmentation, so as to solve the technical problem that the existing intelligent vending machines have low detection accuracy and cause abnormal orders.

[0006] The technical solution adopted by the present invention is:

[0007] The present invention provides an order generation method based on commodity trajectory segmentation, the method comprising:

[0008] Obtain target videos of users shopping through smart vending machines;

[0009] According to the commodity change information in each frame image of the target video, each frame image of the target video is divided into a first frame image and non-first frame images other than the first frame image according to the acquisition time;

[0010] According to the first frame image and the non-first frame images, adding a basic confidence to each target of each frame image of the target video;

[0011] According to the basic confidence of each commodity combined with the detection results of each frame image of the target video input into the target detection network, order information corresponding to each target commodity is obtained.

[0012] Preferably, obtaining a target video of a user shopping through a smart vending machine includes:

[0013] Get the status video of the current status of the smart vending machine;

[0014] Analyze each frame image of the status video to obtain status information of the smart vending machine cabinet door, wherein the status information includes an open state and a closed state;

[0015] The camera is controlled to collect video data of the commodity area according to the on state, and the camera is controlled to stop collecting video data of the commodity area according to the off state to obtain the target video.

[0016] Preferably, obtaining a target video of a user shopping through a smart vending machine includes:

[0017] Obtaining the direction of arrangement of the shelves of the smart vending machine to divide the area where the smart vending machine places goods into multiple virtual goods areas;

[0018] Obtain basic videos within the viewing angle range collected by cameras arranged opposite to each other in each commodity area;

[0019] Physically splicing each frame image of each basic video according to the corresponding frame image of the acquisition time sequence one by one to obtain the target video;

[0020] The physical stitching means that the stitched image is the sum of the sizes of all images involved in the stitching.

[0021] Preferably, the step of dividing each frame image of the target video into a first frame image and non-first frame images other than the first frame image according to the acquisition time according to the commodity change information in each frame image of the target video comprises:

[0022] Performing target quantity detection on each frame image of the target video to determine each frame image corresponding to a commodity change;

[0023] According to each image frame where the commodity changes, the target video is divided into a plurality of target sub-videos;

[0024] Each frame image of each target sub-video is divided into the first frame image and the non-first frame image according to the acquisition time.

[0025] Preferably, adding basic confidence to each target of each frame image of the target video according to the first frame image and the non-first frame image comprises:

[0026] Determine, based on the image information of the first frame image, the location information of the commodity area to which each commodity in the first frame image belongs;

[0027] According to the positioning information, a basic confidence is added to each target in each frame image of the target video.

[0028] Preferably, the step of obtaining order information corresponding to each target product by combining the basic confidence of each product with the detection result of each frame image of the target video input into the target detection network includes:

[0029] Using the target detection network to perform target detection on each frame image of the target video to obtain basic product information of each product;

[0030] According to the confidence of the basic product information of each product, combined with the basic confidence of each product, the product information including the target confidence of each product is obtained;

[0031] Deduplication of each of the products is performed according to the product location information of each of the product information to obtain each target product;

[0032] According to the product information of each of the target products, order information corresponding to each of the target products is output.

[0033] Preferably, the deduplication of each of the commodities according to the commodity location information of each of the commodity information to obtain each target commodity comprises:

[0034] Partitioning each frame image of the target video into image regions of each viewing angle to obtain image sub-regions corresponding to the image of each viewing angle;

[0035] According to the image sub-regions and the imaging size information corresponding to the commodities, a relative relationship of the imaging size information of the commodities belonging to the same image sub-region at different viewing angles is obtained;

[0036] The products are deduplicated according to the relative relationship of the imaging size information of each product to obtain the target products.

[0037] The present invention also provides an order generation device based on commodity trajectory segmentation, the device comprising:

[0038] Video acquisition module: used to obtain target videos of users shopping through smart vending machines;

[0039] Target detection module: used for dividing each frame image of the target video into a first frame image and non-first frame images other than the first frame image according to the commodity change information in each frame image of the target video according to the acquisition time;

[0040] A target processing module: used for adding a basic confidence level to each target of each frame image of the target video according to the first frame image and the non-first frame image;

[0041] Order generation module: used to obtain order information corresponding to each target product according to the basic confidence of each product combined with the detection results of each frame image of the target video input into the target detection network.

[0042] The present invention also provides an intelligent vending machine, comprising: at least one processor, at least one memory, and computer program instructions stored in the memory, and when the computer program instructions are executed by the processor, any of the above-mentioned methods is implemented.

[0043] The present invention also provides a medium on which computer program instructions are stored, and when the computer program instructions are executed by a processor, any of the above-mentioned methods is implemented.

[0044] In summary, the beneficial effects of the present invention are as follows:

[0045] The present invention provides an order generation method, device and intelligent vending machine based on commodity trajectory segmentation, which obtains a target video of a user shopping through a smart vending machine, and divides each frame image of the target video into a first frame image and a non-first frame image according to commodity change information; then adds a basic confidence to each target in each frame image, and uses a target detection model to detect each frame image of the target video, and obtains order information corresponding to each target commodity based on the detection result of each target combined with the corresponding basic confidence; uses the difference in imaging distance between the first frame image and the non-first frame image to set the corresponding basic confidence, which can improve the image with a short distance and many features to guide the target detection result, thereby improving the detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solution of the embodiment of the present invention, the drawings required for use in the embodiment of the present invention will be briefly introduced below. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work, and these are all within the protection scope of the present invention.

[0047] Figure 1 This is a flow chart of the method for intelligently generating orders for commodities through multi-view recognition in Example 1;

[0048] Figure 2 This is a schematic diagram of the structure of the smart vending machine with multiple cameras of different viewing angles in Example 1;

[0049] Figure 3 A schematic diagram of a process for obtaining a physically spliced ​​target video in Example 1;

[0050] Figure 4 This is a schematic diagram of the process of obtaining product information in Example 1;

[0051] Figure 5 This is a schematic diagram of the process of deduplication of commodities in Example 1;

[0052] Figure 6 It is a flow chart of the intelligent order generation method for processing video segments based on the weight change of commodity areas in Example 2;

[0053] Figure 7 This is a schematic diagram of the process of stitching a target video with a base video in Example 2;

[0054] Figure 8 A schematic diagram of a process for obtaining a target sub-video in Example 2;

[0055] Fig. 9 A schematic diagram of the process of generating order information in Example 2;

[0056] Fig.10 This is a schematic diagram of the process of the device for intelligently generating orders for commodities by multi-view recognition in Example 3;

[0057] Fig.11 Schematic diagram of the process of the intelligent order generation device for processing video segments based on the weight change of the commodity area in Example 4;

[0058] Fig.12 It is a structural schematic diagram of the automatic settlement system including the intelligent vending machine in Example 5;

[0059] Fig.13 This is a schematic diagram of the structure of the intelligent vending machine in Example 6;

[0060] Figures 1 to 13 Reference numerals:

[0061] 1. Cabinet body; 11. Shelves; 12. Merchandise area; 2. Cabinet doors; 3. Camera. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solution and advantages of the embodiment of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly and completely described in conjunction with the drawings in the embodiment of the present invention. It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. In the description of the present invention, it should be understood that the orientation or position relationship indicated by the terms "center", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc. is based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. Moreover, the term "include", "comprise" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article or device. In the absence of further restrictions, the elements defined by the phrase "comprising..." do not exclude the existence of other identical elements in the process, method, article or device comprising the elements. If there is no conflict, the various features of the present invention and the embodiments can be combined with each other, all within the protection scope of the present invention.

[0063] Example 1

[0064] Existing fully-opened smart vending machines have the advantages of being convenient for purchasing multiple items at one time and changing items multiple times during one shopping trip, as well as being able to quickly generate orders when users complete a complex shopping process and quickly settle accounts through autonomous settlement. Compared with the existing methods of only being able to purchase one item at a time by scanning a code and not being able to reselect after purchase, fully-opened smart vending machines have the advantages of being simple to operate and allowing users to have greater shopping autonomy. However, because fully-opened smart vending machines allow users to purchase multiple items during one shopping trip and can put items on and off the shelves multiple times, occlusion problems may result in a large number of items being put on or off the shelves. Due to different occlusion positions, different results may be detected for the same item before and after, resulting in abnormal orders that affect user experience and merchant reputation.

[0065] The present invention is based on a feasibility study of obtaining shopping videos of users shopping from smart vending machines from multiple angles. Cameras are set in the product area of ​​the smart vending machine to monitor the product area from different directions in real time. The shopping videos shot by multiple cameras are combined, and then the user's shopping order information is obtained through image stitching, comparative analysis, etc., and then automatic settlement is performed through the server, thereby improving the user's shopping experience and reducing the manual settlement process.

[0066] For details, see Figure 2 , Figure 2 The schematic diagram of the structure of a fully-open smart vending machine is shown in FIG. 1 , which includes a cabinet body 1 and a cabinet door 2. The cabinet body 1 and the cabinet door 2 are rotatably connected. When the cabinet door 2 is in a closed state relative to the cabinet body 1, the cabinet door 2 covers all commodity areas of the cabinet body 1 where commodities are placed, i.e., the commodities in the cabinet body cannot be taken out. When the cabinet door 2 is opened, all commodities in the cabinet body 1 are displayed to the user. The user can select any commodity in a shopping mall and can also select multiple commodities. The user can take out the selected commodity or put back the commodity that needs to be put back after selection. A shelf 11 is provided in the cabinet body 1. The shelf 11 can be a shelf that divides the cabinet body 1 into multiple commodity areas 12. A camera is provided in each commodity area inside the cabinet body 1. In this way, a shopping video of a user shopping from the smart vending machine can be obtained from multiple angles, thereby avoiding the problem that the shopping video collected from a single angle has low reliability due to occlusion. Figure 2 The smart vending machine shown is provided with multiple cameras on the left and right internal walls of the vending machine so that shopping videos can be collected from the same commodity area in relative viewing directions, thereby improving the reliability of video data.

[0067] See also Figure 1 , Figure 1 A flowchart of a method for intelligently generating orders for products through multi-view recognition is provided, wherein the method comprises:

[0068] S10: Obtain a target video of the commodity area, wherein the target video is composed of a video stream obtained by physically splicing each frame image corresponding to a plurality of basic videos, each of the basic videos is composed of images of the same event occurring in the commodity area captured from a perspective, and each basic video has a different perspective corresponding to the same event;

[0069] Specifically, cameras are set at different positions of the smart vending machine to obtain video data of the product area of ​​the smart vending machine from different perspectives. When the user starts shopping from the smart vending machine, each camera obtains basic video of the user taking or putting back the product from different angles, and each frame image of the basic video obtained by different cameras is physically spliced ​​according to the acquisition sequence to obtain a target video composed of spliced ​​images; wherein the same event is the entire process of a user's shopping.

[0070] It should be noted that physical stitching is to stitch two images into one image, and the stitched image is the sum of the sizes of the images involved in the stitching; at the same time, the physical stitching of each frame image of different videos is: the first frame image of the first video, the first frame image of the second video...the first frame image of the Nth video are stitched, the second frame image of the first video, the second frame image of the second video...the second frame image of the Nth video are stitched, and so on, the nth frame image of the first video, the nth frame image of the second video...the nth frame image of the Nth video are stitched to obtain the target video.

[0071] In one embodiment, see Figure 3 , the S10 comprises:

[0072] S101: Dividing the commodity placement area of ​​the smart vending machine into multiple virtual commodity areas along the shelf arrangement direction of the smart vending machine;

[0073] Specifically, the smart vending machine has multiple layers of shelves, and the area where the smart vending machine places goods is divided into multiple product areas, each product area includes at least one layer of shelves, and the camera's viewing angle is set along the arrangement direction of the shelves. For example, the shelves of the smart vending machine include multiple layers of shelves from top to bottom, and each camera is respectively set on the left and right side walls of the smart vending machine, and the viewing angle of each camera is from upper left to lower right or from upper right to lower left or from top to bottom. Among them, the cameras set on different sides of the same product area have the same installation height.

[0074] In one embodiment, the S101 includes:

[0075] S1011: Divide the commodity area of ​​the smart vending machine into an upper commodity area and a lower commodity area along the viewing direction of the camera from top to bottom;

[0076] S1012: A camera is provided on the left and right sides of the upper commodity area and on the left and right sides of the lower commodity area respectively;

[0077] Among them, the viewing angle direction of the camera on the left is from the upper left corner to the lower right corner, and the viewing angle direction of the camera on the right is from the upper right corner to the lower left corner.

[0078] Specifically, in a preferred embodiment, the shelves of the smart vending machine are divided into an upper merchandise area and a lower merchandise area, and a camera is provided on the left and right side walls of the upper and lower merchandise areas. The left and right cameras of the upper merchandise area can collect video data of the entire area, and the left and right cameras of the lower merchandise area can only collect video data within the range of the lower merchandise area.

[0079] It should be noted that: dividing the shelf into two upper and lower commodity areas, and setting a pair of cameras opposite to each other in each commodity area, can ensure that in a shopping event, shopping video data is obtained from four directions of up, down, left and right, thereby improving data reliability; this method can not only save costs, but also control the size of each frame image of the target video and reduce the amount of data processing.

[0080] S102: Obtain basic videos within the viewing angle range collected by cameras disposed opposite to each other in each commodity area; specifically, each camera obtains a video stream of a corresponding area to obtain a video of the user taking or putting back the commodity.

[0081] S103: physically splicing the frame images of the basic videos one by one according to the corresponding frame images of the acquisition time sequence to obtain the target video;

[0082] The physical stitching means that the stitched image is the sum of the sizes of all images involved in the stitching.

[0083] Specifically, each frame image of different basic videos is physically spliced. The specific splicing method refers to the above method to obtain the final target video.

[0084] In one embodiment, the S10 includes:

[0085] S105: Obtaining the frame rate and number of cameras used to collect video data;

[0086] S106: Determine, according to the frame rate and the number of cameras, the interval at which each camera starts to collect video data of the corresponding commodity area;

[0087] S107: Controlling each of the cameras to obtain corresponding basic videos according to each of the interval times;

[0088] Specifically, the frame rates of the cameras used to collect video data are the same, such as 20 frames per second. The time when each camera starts collecting video data is determined according to the number and frame rate of the cameras. It is assumed that there is a time interval between the start times of collecting video data between each camera or each group of cameras, wherein the preferred interval time is an integer multiple of the time difference corresponding to two adjacent frames of images. For example, if four cameras are included, the start time of collecting video data of each camera is spaced apart by 1 / 4 of the time corresponding to the frame rate, or the four cameras are divided into two groups, and the interval time when each group of cameras starts collecting video data is 1 / 2 of the time corresponding to the frame rate. This improves the image frame rate in disguise, ensuring that image information of more moments of the commodity area is collected, so as to improve the accuracy of detection.

[0089] S108: Physically splicing the frame images of each of the basic videos corresponding to the frame images in the acquisition time sequence one by one to obtain the target video.

[0090] S11: Input each frame image of the target video into a target detection network to perform target detection, and obtain commodity information of each commodity;

[0091] Specifically, the spliced ​​frames of images are sent to a target detection network for target detection to obtain product information of each detected target, wherein the product information includes at least one of the following: position information of the product in the image, area information of the detection frame, product category information, and confidence.

[0092] In one embodiment, see Figure 4 , the S11 comprises:

[0093] S111: dividing each frame image of the target video into a first frame image and non-first frame images other than the first frame image according to the acquisition time;

[0094] In one embodiment, the S111 includes:

[0095] S1111: Detect the target quantity of each frame image of the target video to determine each frame image corresponding to a commodity change;

[0096] S1112: dividing the target video into a plurality of target sub-videos according to the image frames where the commodity changes occur;

[0097] S1113: Divide each frame image of each target sub-video into the first frame image and the non-first frame image according to the acquisition time.

[0098] S112: Determine, based on the image information of the first frame image, the location information of the commodity area to which each commodity in the first frame image belongs;

[0099] Specifically, each frame image of the target video is divided into a first frame image and non-first frame images other than the first frame image, wherein the first frame image includes but is not limited to the first frame image of the complete target video, and may also be the target video divided into multiple video segments, and the first frame image is the first frame image of each video segment; the first frame image is subjected to preliminary target detection to determine which commodity area the target comes from, and the video data shot by the camera in the commodity area where the commodity comes from is recorded as the positioning video of the commodity.

[0100] S113: adding a basic confidence level to each target in each frame image of the target video according to the positioning information;

[0101] Specifically, after determining the positioning information of the product, a basic confidence level is added to each target in each frame image of the positioning video corresponding to the positioning information. For example, the smart vending machine is divided into multiple product areas, and each product area is equipped with a corresponding camera. It should be noted that when taking or putting products, not only the camera that takes out the corresponding product area or puts the corresponding product area will collect the corresponding video data, but the cameras in other product areas may also be able to collect the corresponding video data. For example, the product area of ​​the smart vending machine is divided into an upper product area and a lower product area. When taking products from the lower product area, the products exist in both the upper product area and the lower product area. The video also exists in the video of the camera at the lower commodity area. Because the commodity belongs to the lower commodity area, the imaging size of the commodity is larger in the video data of the camera at the lower commodity area, which is conducive to improving the detection accuracy. Therefore, a basic confidence can be added to the target in each frame image of the video taken by the camera in the lower commodity area, recorded as the first basic confidence, and a basic confidence can be added to the target in each frame image of the video corresponding to the camera in the upper commodity area, recorded as the second basic confidence, wherein the second basic confidence is less than the first basic confidence; similarly, when the commodity belongs to the upper commodity area, the second basic confidence is greater than the first basic confidence.

[0102] In a preferred embodiment, the S113 includes:

[0103] S1131: Obtaining a boundary line of commodity movement corresponding to commodity confidence enhancement;

[0104] S1132: segmenting the target video according to different state areas of the commodity located on the boundary line in adjacent image frames to obtain a first video segment with enhanced confidence and a second video segment with normal confidence;

[0105] S1133: In combination with the positioning video, add a basic confidence level to each target in each frame image of the first video segment corresponding to the product area to which the target product belongs.

[0106] Specifically, after the product is taken out from the product area, as the distance from the camera increases, the imaging size of the product in the image decreases, so the detection accuracy will be reduced. Therefore, each frame image close to the camera area is taken as the key detection object. Therefore, the target video is divided into a first video segment and a second video segment using a boundary line, and a basic confidence is increased for each target in the image area of ​​the positioning video in each frame image of the first video segment, or a first basic confidence is increased for each target in the image area of ​​the positioning video in each frame image of the first video segment, and a second basic confidence is increased for each target in the image area of ​​the positioning video in each frame image of the first video segment, wherein the first basic confidence is greater than the second basic confidence.

[0107] S114: performing target detection on each frame image of the target video using the target detection network to obtain basic product information of each product;

[0108] S115: The basic confidence of each product is superimposed on the confidence of the basic product information of each product to obtain the product information including the target confidence of each product.

[0109] Specifically, each frame image of the target video is sent to the target detection network for detection to obtain basic product information of each target in each frame image, wherein the basic product information includes at least one of the following: product category, confidence, and product location information of the detection frame representing the detected product. The confidence of each target detected this time is recorded as the actual confidence, and then the actual confidence of each target belonging to the first video segment is added to the basic confidence to obtain the target confidence of each target in the first video segment; the actual confidence of each target in the second video segment is used as the final target confidence, thereby obtaining the product information of each target composed of the target confidence.

[0110] S12: De-duplicate the commodities according to the commodity location information of the commodity information to obtain target commodities;

[0111] Specifically, according to the position information of the detection box used to represent the detected target in the product information, it is determined which ones in the videos shot by different cameras are the same product, so as to achieve product deduplication; the specific method of product deduplication includes but is not limited to using a classifier to detect the same product, and using the prior relationship between the imaging position and imaging size of the product in the same frame image in different videos, so as to distinguish the same product and achieve product deduplication.

[0112] In one embodiment, see Figure 5 , the S12 comprises:

[0113] S121: Acquire multiple positive samples and multiple negative samples, wherein each of the positive samples is a target belonging to the same product that appears at a different position in the image after images from different perspectives are stitched together, and each of the negative samples is a target belonging to different products that appears at a different position in the image after images from different perspectives are stitched together;

[0114] Specifically, each camera on the smart vending machine is controlled to obtain training videos of multiple shopping training times, and each target in each frame of the training video is manually labeled. The targets with different position information belonging to the same product in the image area captured by different cameras in each frame are recorded as positive samples, and the targets with other position information are taken as negative samples. That is, two products are taken at a time, recorded as product A and product B. Assuming that there are four cameras with different viewing angles, and in any frame all four cameras capture images of product A and product B, then the corresponding target frame images will have four positive samples consisting of product A, four positive samples consisting of product B, and a negative sample between product A and product B.

[0115] S122: Inputting the samples including the positive samples and the negative samples into a support vector machine for training, so as to obtain a product deduplication classifier that can distinguish whether the products under different viewing angles are the same product by using product location information;

[0116] Specifically, a classifier is obtained by training with a manually labeled sample set, which can distinguish whether the products are the same product based on the product location information. The classifier is then used to deduplicate each frame of the target video, so that the targets detected in each frame are independent products.

[0117] S123: De-duplication is performed using the product de-duplication classifier according to the product location information of each product information to obtain each target product;

[0118] Among them, each of the samples is an image obtained by physically stitching together each frame of images captured by each camera on the smart vending machine.

[0119] In one embodiment, the S12 includes:

[0120] Step 1: partition each frame image of the target video into image regions of each viewing angle to obtain image sub-regions corresponding to the image of each viewing angle;

[0121] Specifically, each frame image of the target video is obtained by physically stitching each frame image area taken by each camera. Therefore, each image area of ​​each frame image of the target video is divided into multiple image sub-areas. For example, the target video is stitched together by each frame image of 4 videos. Therefore, each frame image of the target video includes 4 image areas, which are recorded as the upper left image area, the upper right image area, the lower left image area and the lower right image area, and then each image area is divided into multiple image sub-areas.

[0122] Step 2: According to the image sub-regions and the imaging size information corresponding to the commodities, a relative relationship of the imaging size information of the commodities belonging to the same image sub-region at different viewing angles is obtained;

[0123] Specifically, the image sub-areas and imaging sizes corresponding to the imaging positions of the goods detected in the images of different viewing angles are compared to determine the relative relationship of the imaging sizes of the targets belonging to the same image sub-areas in the images of different viewing angles, such as: targets are detected in the lower right corners of the four image areas, wherein the image overlap of the targets in the upper left image area and the upper right image area is very high and greater than the overlap threshold, the target imaging range of the lower left image area belongs to the target imaging range of the upper left image area, and the target imaging range of the lower right image area belongs to the target imaging range of the upper right image area, and it can be determined that the targets detected in the four image areas are the same product; if the image overlap of the targets in the upper left image area and the upper right image area is very low and less than the overlap threshold, the target imaging range of the lower left image area belongs to the target imaging range of the upper left image area, and the target imaging range of the lower right image area belongs to the target imaging range of the upper right image area, it can be determined that the targets detected in the upper left and lower left image areas are the same product, and the targets detected in the upper right and lower right image areas are the same product; including but not limited to the above situations, which are not listed here one by one.

[0124] Step 3: De-duplicate products according to the relative relationship of the imaging size information of each product to obtain the target products.

[0125] Specifically, after determining multiple detection targets corresponding to the same product in each frame image, duplicate removal is performed, and targets with high confidence levels can be selected and retained as detection results of the same product.

[0126] S13: outputting order information corresponding to each of the target commodities according to the commodity information of each of the target commodities;

[0127] The product information includes at least one of the following: product category, confidence level, and product location information of a detection frame representing a detected product.

[0128] In one embodiment, before S10, the method further includes:

[0129] S01: Acquire the video of the current state of the smart vending machine collected by the third camera in real time;

[0130] Specifically, the smart vending machine is also provided with a third camera, which is used to detect whether the smart vending machine is turned on or off. The third camera can be turned on in real time or after the user makes a shopping request.

[0131] S02: Analyze each frame of the video of the current state of the smart vending machine to determine whether the cabinet door of the smart vending machine is in an open or closed state;

[0132] S03: When it is detected that the cabinet door of the smart vending machine is in an open state, the first main camera, the first sub-camera, the second main camera and the second sub-camera for collecting video information corresponding to the commodity area are controlled to be turned on;

[0133] S04: When it is detected that the cabinet door of the smart vending machine is in a closed state, the first main camera, the first sub-camera, the second main camera and the second sub-camera for collecting video information corresponding to the commodity area are controlled to be turned off.

[0134] Specifically, when the user performs automatic shopping, each frame image of the vending machine status video is analyzed to determine the status of the vending machine door. When it is detected that the vending machine door is open, the first main camera, the first sub-camera, the second main camera and the second sub-camera are turned on to obtain video data of the commodity area to obtain each basic video; when it is detected that the vending machine door is closed, the first main camera, the first sub-camera, the second main camera and the second sub-camera are turned off.

[0135] The multi-view commodity identification and intelligent order generation method of this embodiment is adopted to obtain the target video by physically stitching the videos of the commodity area from different viewpoints, and then perform target detection on each frame image of the target video and deduplicate the same commodity in the same frame image to obtain the target commodity for producing order information. By obtaining videos of shopping events from different viewpoints, this method can prevent order anomalies caused by obstruction of commodities and improve detection accuracy and user experience.

[0136] Example 2

[0137] In Example 1, the video data of the product area of ​​the smart vending machine is obtained from different perspectives, and then the frames of the different videos are spliced ​​to obtain the target video, and the frames of the target video are analyzed to obtain the product order information; however, in a shopping event, the user may have complex operations such as repeated selection or repeated switching, resulting in multiple events of picking and placing products, which often leads to false detection or mixed detection of products with high similarity, affecting the detection accuracy. Therefore, Example 2 of the present invention further improves the method for intelligently generating orders for multi-perspective product recognition based on Example 1; please refer to Figure 6 , the method comprising:

[0138] S20: Obtaining a target video of the commodity area and weight change information of the commodity area;

[0139] Specifically, the target video of the commodity area is the image data collected by the camera of the smart vending machine during a shopping process. The target video can be the video data corresponding to the commodity area collected by one camera, or the video data collected by multiple cameras, or the video data collected by multiple cameras. The weight change information of the commodity area includes weight increase and decrease information and time information. Only the weight increase and decrease information can be used to quickly determine the action of putting the commodity on the shelf or taking the commodity off the shelf, thereby providing guidance for the neural network to perform target detection, and even a small amount of image analysis can be combined to directly determine whether the user finally takes the commodity. For example, if the weight decrease is detected at the first moment, it means that the user has taken the commodity. Through image analysis, it is found that the user has only taken one commodity. At the second moment, the weight increase is detected. If no other weight change information is detected during this period, it can be determined that the user puts the commodity taken out for the first time directly back on the shelf, and then the video data from the first moment to the second moment can be deleted to reduce the subsequent data processing volume.

[0140] In one embodiment, see Figure 7 , the S20 includes:

[0141] S201: Acquire basic videos of the product area captured from different viewing angles;

[0142] Specifically, cameras are set at different positions of the smart vending machine to obtain basic videos of the product area of ​​the smart vending machine from different perspectives, which can provide more reliable image data for target detection during shopping and improve the accuracy of target detection.

[0143] In one embodiment, the S201 includes:

[0144] S2011: obtaining a method of dividing the commodity placement area of ​​the intelligent vending machine into a plurality of virtual commodity areas along the shelf arrangement direction of the intelligent vending machine;

[0145] S2012: Obtain basic videos within the viewing angle range collected by cameras disposed opposite to each other in each commodity area.

[0146] Specifically, the smart vending machine has multiple layers of shelves, and the area where the smart vending machine places goods is divided into multiple product areas, each product area includes at least one layer of shelves, and the camera's view is set along the arrangement direction of the shelves. For example, the shelves of the smart vending machine include multiple layers of shelves from top to bottom, and each camera is respectively set on the left and right side walls of the smart vending machine, and the view of each camera is from the upper left to the lower right or from the upper right to the lower left or from top to bottom. Among them, the cameras set on different sides of the same product area have the same installation height; when the user is shopping, each camera is turned on to obtain the video of this shopping, so as to obtain each basic video of different perspectives of this shopping.

[0147] In one embodiment, the S201 includes:

[0148] S2014: Obtaining the frame rate and number of cameras used to collect video data;

[0149] S2015: Determine, according to the frame rate and the number of cameras, the interval at which each camera starts to collect video data of the corresponding commodity area;

[0150] S2016: According to each of the interval times, control each of the cameras to obtain corresponding basic videos.

[0151] Specifically, the frame rates of the cameras used to collect video data are the same, such as 20 frames per second; the time when each camera starts collecting video data is determined according to the number and frame rate of the cameras, and it is assumed that there is a time interval between the start times of collecting video data between each camera or each group of cameras, wherein the preferred interval time is an integer multiple of the time difference corresponding to two adjacent frames of images, such as: including 4 cameras, the start time of collecting video data of each camera is 1 / 4 of the time corresponding to the frame rate, or the 4 cameras are divided into two groups, and the interval time for each group of cameras to start collecting video data is 1 / 2 of the time corresponding to the frame rate; thereby increasing the image frame rate in disguised form, ensuring that image information of more moments in the commodity area is collected, so as to improve the accuracy of detection; after detecting the start of shopping, each camera obtains the video data of the corresponding area according to its own start time to obtain each basic video, wherein the signal for starting shopping is the user's verification through an identification code, such as: a QR code, a bar code, etc., or after the cabinet door of the smart vending machine is opened, and the specific signal for starting shopping is not limited here.

[0152] S202: Physically splicing the frame images of the basic videos one by one according to the frame images corresponding to the acquisition time sequence to obtain the target video;

[0153] The physical stitching means that the stitched image is the sum of the sizes of all images involved in the stitching.

[0154] Specifically, when a user is shopping, each camera obtains a basic video of the user taking or putting back the product from different angles, and physically stitches the frames of the basic video obtained by different cameras to obtain a target video composed of stitched images; wherein the same event is the entire process of a user's shopping.

[0155] It should be noted that physical stitching is to stitch two images into one image, and the stitched image is the sum of the sizes of the images involved in the stitching; at the same time, the physical stitching of each frame image of different basic videos is: the first frame image of the first video, the first frame image of the second video...the first frame image of the Nth video are stitched, the second frame image of the first video, the second frame image of the second video...the second frame image of the Nth video are stitched, and so on, the nth frame image of the first video, the nth frame image of the second video...the nth frame image of the Nth video are stitched to obtain the target video.

[0156] S21: segmenting the target video according to the weight change information to obtain a plurality of target sub-videos;

[0157] Specifically, the weight change information of the commodity area is detected in real time. When the weight change is detected, it is considered that there is an action of picking up and putting in the commodity. The target video between the current weight change and the previous weight change is used as a target sub-video, thereby dividing the target video into multiple target sub-videos; target detection can be performed on each target sub-video to obtain multiple commodity information corresponding to each target sub-video, thereby improving the accuracy of commodity orders.

[0158] In one embodiment, see Figure 8 , the S21 comprises:

[0159] S211: segmenting the target video into a plurality of first videos according to each time information of the weight change information;

[0160] S212: According to the increase and decrease information of the weight change information, each of the first videos is divided into the target sub-videos corresponding to the product listing and the product delisting.

[0161] Specifically, when a weight change is detected in the commodity area, the current image frame is determined based on the time information of the weight change to obtain a target sub-video, and then the target sub-video is determined to be a commodity placing video or a commodity taking video based on the increase or decrease information of the weight change information.

[0162] In one embodiment, the S212 includes:

[0163] S2121: Obtain the boundary line used to define whether a product belongs to the shelf or not;

[0164] S2122: According to different state areas of the goods located at the boundary line in adjacent image frames, combined with the corresponding increase and decrease information of the weight change information, each of the first videos is divided into the target sub-videos corresponding to the goods being put on the shelves and the goods being taken off the shelves.

[0165] Specifically, a virtual boundary line is set in the door area of ​​the smart vending machine. The boundary line is used to determine the taking or putting of goods in combination with the weight change information. This is because when selecting goods, users may take and put goods in the same area multiple times in a very short time. The picture at this time belongs to the inside of the smart vending machine. These goods are severely obstructed. If the video is segmented only by weight change, a large number of extremely short videos will be generated. The analysis of these extremely short videos alone not only increases the amount of calculation, but also has little meaning and may even increase the detection error rate. The accuracy of detection can be improved by dividing the video segments based on the goods leaving or entering the boundary line.

[0166] S22: Inputting each frame image of each target sub-video into a target detection network to perform target detection, and obtaining commodity information corresponding to each target sub-video;

[0167] S23: Output order information according to the product information corresponding to each target sub-video.

[0168] In one embodiment, see Fig. 9 , the S23 comprises:

[0169] S231: Deduplication of the same product in each frame image is performed according to the product position information of the product information to obtain target products after deduplication;

[0170] For details, please refer to the deduplication method in the embodiment, which will not be described here.

[0171] In one embodiment, the S231 includes:

[0172] S2311: Acquire multiple positive samples and multiple negative samples, wherein each of the positive samples is a target belonging to the same product that appears at a different position in the image after images from different perspectives are stitched together, and each of the negative samples is a target belonging to different products that appears at a different position in the image after images from different perspectives are stitched together;

[0173] S2312: Input the samples including the positive samples and the negative samples into a support vector machine for training, so as to obtain a product deduplication classifier that can distinguish whether the products under different viewing angles are the same product based on the product location information;

[0174] S2313: De-duplicate the product using the product de-duplicate classifier according to the product location information of each product information to obtain each target product;

[0175] Among them, each of the samples is an image obtained by physically stitching together each frame of images captured by each camera on the smart vending machine.

[0176] Specifically, for product deduplication, refer to the method in Example 1, which will not be described in detail here.

[0177] In one embodiment, the S231 includes:

[0178] Step 1: partition each frame image of the target video into image regions of each viewing angle to obtain image sub-regions corresponding to the image of each viewing angle;

[0179] Step 2: According to the image sub-regions and the imaging size information corresponding to the commodities, a relative relationship of the imaging size information of the commodities belonging to the same image sub-region at different viewing angles is obtained;

[0180] Step 3: De-duplicate products according to the relative relationship of the imaging size information of each product to obtain the target products.

[0181] Specifically, after determining multiple detection targets corresponding to the same product in each frame image, duplicate removal is performed, and targets with high confidence levels can be selected and retained as detection results of the same product.

[0182] Specifically, for product deduplication, refer to the method in Example 1, which will not be described in detail here.

[0183] S232: Outputting the order information corresponding to each of the target products according to the product information of each of the target products.

[0184] The intelligent order generation method based on the video segmentation processing of the weight change of the commodity area of ​​the present embodiment is adopted. By real-time collecting the weight change information of the commodity area during the user's shopping process, the shopping target video is segmented according to the weight change information to obtain multiple target sub-videos, and then each frame image of each target sub-video is used for target detection by using the target detection network, and finally the commodity information purchased by the user is obtained, thereby generating shopping order information; this method divides the complete target video into multiple target sub-videos for target detection according to the weight change, which can avoid the mutual influence of taking out or putting in events, and improve the detection accuracy and user experience effect.

[0185] Example 3

[0186] Embodiment 3 of the present invention further provides a device for intelligently generating orders for goods by identifying goods from multiple perspectives based on the methods of Embodiments 1 to 2. Fig.10 ,include:

[0187] Video acquisition module: used to obtain a target video of a commodity area, wherein the target video is composed of a video stream obtained by physically splicing each frame image corresponding to a number of basic videos, each of which is composed of images of the same event occurring in the commodity area captured from a certain perspective, and each basic video has a different perspective corresponding to the same event;

[0188] Target detection module: used to input each frame image of the target video into the target detection network for identification, and obtain the commodity information of each commodity;

[0189] Target processing module: used for removing duplicates of each commodity according to the commodity location information of each commodity information to obtain each target commodity;

[0190] An order generation module: used for outputting order information corresponding to each target commodity according to the commodity information of each target commodity;

[0191] The product information includes at least one of the following: product category, confidence level corresponding to each product category, and product location information of a detection frame representing a detected product.

[0192] The order generation device based on multi-view image analysis of this embodiment is adopted to obtain the target video by physically stitching the videos of the commodity area from different viewpoints, and then perform target detection on each frame image of the target video and deduplicate the same commodity in the same frame image to obtain the target commodity for producing order information; this method can prevent order anomalies caused by obstruction of commodities by obtaining videos of shopping events from different viewpoints, thereby improving detection accuracy and user experience.

[0193] It should be noted that the device also includes the remaining technical solutions described in Examples 1 and 2, which will not be repeated here.

[0194] Example 4

[0195] Embodiment 4 of the present invention further provides an intelligent order generation device for processing video segments based on the weight change of commodity areas based on the methods of Embodiments 1 to 2, see Fig.11 ,include:

[0196] Video acquisition module: used to obtain the target video of the commodity area and the weight change information of the commodity area;

[0197] Video segmentation module: used for segmenting the target video according to the weight change information to obtain multiple target sub-videos;

[0198] Data processing module: used for inputting each frame image of each target sub-video into the target detection network for target detection, and obtaining the commodity information corresponding to each target sub-video;

[0199] Order generation module: used to output order information according to the product information corresponding to each target sub-video.

[0200] The intelligent order generation device based on the weight change of the commodity area in the embodiment of the present invention is used to collect the weight change information of the commodity area in the user's shopping process in real time, and the shopping target video is segmented according to the weight change information to obtain multiple target sub-videos, and then each frame image of each target sub-video is used for target detection by using the target detection network, and finally the commodity information purchased by the user is obtained, thereby generating shopping order information; this method divides the complete target video into multiple target sub-videos for target detection according to the weight change, which can avoid the mutual influence of the taking out or putting in events, and improve the detection accuracy and user experience effect.

[0201] It should be noted that the device also includes the remaining technical solutions recorded in Example 4, which will not be repeated here.

[0202] Example 5

[0203] The present invention provides an automatic settlement system for intelligent vending machines, see Fig.12 The automatic settlement system includes an intelligent vending machine, a mobile terminal and a server. The automatic settlement system can adopt the automatic shopping method described in the above embodiment. The user identifies the identification code on the intelligent vending machine through the mobile terminal, the server establishes the user's shopping event, the camera with different viewing angles starts to collect shopping videos, or the camera starts to collect shopping videos after the door of the intelligent vending machine is opened, or the camera starts to collect shopping videos when the user enters the preset range. When the user leaves the preset shopping range or the door of the intelligent vending machine is closed, the camera stops collecting shopping videos and transmits the shopping videos to the server. The server generates the user's order information based on the shopping video and sends it to the mobile terminal. The user performs autonomous settlement or sets automatic settlement through the order information of the mobile terminal; the automatic settlement system has better selectivity for autonomous shopping by users, high order accuracy, and can improve the user's shopping experience.

[0204] Example 6

[0205] The present invention provides an intelligent vending machine device and a storage medium, such as Fig.13 As shown, the system comprises at least one processor, at least one memory and computer program instructions stored in the memory.

[0206] Specifically, the processor may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present invention. The smart vending machine is provided with cabinet doors that can cover the entire merchandise area at the merchandise area, and the cabinet doors are movable cabinet doors that can be opened and closed. At the same time, the smart vending machine also includes identification devices such as cameras, QR codes, bar codes, etc. that are convenient for shopping.

[0207] The memory may include a large capacity memory for data or instructions. By way of example and not limitation, the memory may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. Where appropriate, the memory may include a removable or non-removable (or fixed) medium. Where appropriate, the memory may be inside or outside a data processing device. In a particular embodiment, the memory is a non-volatile solid-state memory. In a particular embodiment, the memory includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM) or a flash memory or a combination of two or more of these.

[0208] The processor reads and executes computer program instructions stored in the memory to implement any one of the multi-perspective commodity recognition and intelligent order generation methods in the first embodiment and the intelligent order generation method based on video segmentation processing based on commodity area weight changes.

[0209] In one example, the electronic device may further include a communication interface and a bus, wherein the processor, the memory, and the communication interface are connected via the bus and communicate with each other.

[0210] The communication interface is mainly used to implement communication between the modules, devices, units and / or equipment in the embodiments of the present invention.

[0211] Bus includes hardware, software or both, and the parts of electronic equipment are coupled to each other.For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industrial standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In suitable cases, bus may include one or more buses. Although the embodiment of the present invention describes and shows a specific bus, the present invention considers any suitable bus or interconnection.

[0212] In summary, the embodiments of the present invention provide a method for intelligently generating orders for commodities through multi-perspective recognition, a method for intelligently generating orders through video segmentation processing based on weight changes in commodity areas, a device, an intelligent vending machine, and a storage medium.

[0213] It should be clear that the present invention is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present invention.

[0214] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0215] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An order generation method based on commodity trajectory segmentation, It is characterized in that The method comprises: Obtain target videos of users shopping through smart vending machines; According to the commodity change information in each frame image of the target video, each frame image of the target video is divided into a first frame image and non-first frame images other than the first frame image according to the acquisition time; According to the first frame image and the non-first frame images, adding a basic confidence to each target of each frame image of the target video; According to the basic confidence of each commodity and the detection result of each frame image of the target video input into the target detection network, order information corresponding to each target commodity is obtained; the target video of the user shopping through the smart vending machine is obtained, including: Obtaining the direction of arrangement of the shelves of the smart vending machine to divide the area where the smart vending machine places goods into multiple virtual goods areas; Obtain basic videos within the viewing angle range collected by cameras arranged opposite to each other in each commodity area; Physically splicing each frame image of each basic video according to the corresponding frame image of the acquisition time sequence one by one to obtain the target video; Wherein, the physical stitching means that the stitched image is the sum of the sizes of all images involved in the stitching; The step of obtaining order information corresponding to each target product by combining the basic confidence of each product with the detection result of each frame image of the target video input into the target detection network includes: Using the target detection network to perform target detection on each frame image of the target video to obtain basic product information of each product; According to the confidence of the basic product information of each product, combined with the basic confidence of each product, the product information including the target confidence of each product is obtained; Deduplication of each of the products is performed according to the product location information of each of the product information to obtain each target product; Outputting order information corresponding to each of the target commodities according to the commodity information of each of the target commodities; Deduplication of each of the commodities is performed according to the commodity location information of each of the commodity information to obtain each target commodity, including: Partitioning each frame image of the target video into image regions of each viewing angle to obtain image sub-regions corresponding to the image of each viewing angle; According to the image sub-regions and imaging size information corresponding to the commodities, a relative relationship of the imaging size information of the commodities belonging to the same image sub-region at different viewing angles is obtained; The products are deduplicated according to the relative relationship of the imaging size information of each product to obtain the target products.

2. The order generation method based on commodity trajectory segmentation according to claim 1, It is characterized in that The obtaining of the target video of the user shopping through the smart vending machine includes: Get the status video of the current status of the smart vending machine; Analyze each frame image of the status video to obtain status information of the smart vending machine cabinet door, wherein the status information includes an open state and a closed state; The camera is controlled to collect video data of the commodity area according to the on state, and the camera is controlled to stop collecting video data of the commodity area according to the off state to obtain the target video.

3. The order generation method based on commodity trajectory segmentation according to claim 1, It is characterized in that The step of dividing each frame image of the target video into a first frame image and non-first frame images other than the first frame image according to the commodity change information in each frame image of the target video according to the acquisition time includes: Performing target quantity detection on each frame image of the target video to determine each frame image corresponding to a commodity change; According to each image frame where the commodity changes, the target video is divided into a plurality of target sub-videos; Each frame image of each target sub-video is divided into the first frame image and the non-first frame image according to the acquisition time.

4. The order generation method based on commodity trajectory segmentation according to claim 1 or 3, It is characterized in that The adding basic confidence to each target of each frame image of the target video according to the first frame image and the non-first frame image includes: Determine, based on the image information of the first frame image, the location information of the commodity area to which each commodity in the first frame image belongs; According to the positioning information, a basic confidence is added to each target in each frame image of the target video.

5. An order generation device based on commodity trajectory segmentation, It is characterized in that The device comprises: Video acquisition module: used to obtain target videos of users shopping through smart vending machines; Target detection module: used for dividing each frame image of the target video into a first frame image and non-first frame images other than the first frame image according to the commodity change information in each frame image of the target video according to the acquisition time; A target processing module: used for adding a basic confidence level to each target of each frame image of the target video according to the first frame image and the non-first frame image; An order generation module is used to obtain order information corresponding to each target product according to the basic confidence of each product and the detection result of each frame image of the target video input into the target detection network; The obtaining of the target video of the user shopping through the smart vending machine includes: Obtaining the direction of arrangement of the shelves of the smart vending machine to divide the area where the smart vending machine places goods into multiple virtual goods areas; Obtain basic videos within the viewing angle range collected by cameras arranged opposite to each other in each commodity area; Physically splicing each frame image of each basic video according to the corresponding frame image of the acquisition time sequence one by one to obtain the target video; Wherein, the physical stitching means that the stitched image is the sum of the sizes of all images involved in the stitching; The step of obtaining order information corresponding to each target product by combining the basic confidence of each product with the detection result of each frame image of the target video input into the target detection network includes: Using the target detection network to perform target detection on each frame image of the target video to obtain basic product information of each product; According to the confidence of the basic product information of each product, combined with the basic confidence of each product, the product information including the target confidence of each product is obtained; Deduplication of each of the products is performed according to the product location information of each of the product information to obtain each target product; Outputting order information corresponding to each of the target commodities according to the commodity information of each of the target commodities; Deduplication of each of the commodities is performed according to the commodity location information of each of the commodity information to obtain each target commodity, including: Partitioning each frame image of the target video into image regions of each viewing angle to obtain image sub-regions corresponding to the image of each viewing angle; According to the image sub-regions and imaging size information corresponding to the commodities, a relative relationship of the imaging size information of the commodities belonging to the same image sub-region at different viewing angles is obtained; The products are deduplicated according to the relative relationship of the imaging size information of each product to obtain the target products.

6. A smart vending machine, It is characterized in that include: At least one processor, at least one memory and computer program instructions stored in the memory, when the computer program instructions are executed by the processor, implement the method according to any one of claims 1 to 4.

7. A medium having computer program instructions stored thereon, It is characterized in that When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Multi-angle image object fusion method and system

    CN103914821A

  • Commodity identification method, automatic vending machine and computer readable storage medium

    CN109003390A

  • Commodity detection method and device and readable storage medium

    CN111626201A