Order generation method based on multi-view image contrast detection and intelligent vending machine

By setting up multi-view cameras on smart vending machines to collect and merge video data, and combining it with the target detection model to generate orders, the problem of abnormal orders caused by product occlusion in smart vending machines is solved, and the detection accuracy and user experience are improved.

CN114648715BActive Publication Date: 2025-09-09YOPOINT SMART RETAIL TECH LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210276704.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-03
Publication Date
2025-09-09
Estimated Expiration
2041-11-03

AI Technical Summary

Technical Problem

Existing smart vending machines are prone to misdetection due to product obstruction during the shopping process, generating abnormal orders, or are cumbersome to operate, resulting in a poor user experience.

Method used

A multi-view image comparison detection method is adopted. By setting the first main camera and the second main camera on the smart vending machine to collect video data of the same product area from different perspectives, the data are merged for target detection and order information is generated.

Benefits of technology

It improves detection accuracy, reduces the occurrence of abnormal orders, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114648715B_ABST
    Figure CN114648715B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of image processing technology, and solves the technical problem of poor user experience caused by cumbersome operation or abnormal orders in existing smart vending machines, and provides an order generation device method and smart vending machine with multi-view image comparison detection. The method includes: obtaining a first main video of the product area and a second main video with a different perspective from the first main video; merging each frame image of the first main video and the second main video to obtain a target video; performing image analysis on each frame image of the target video to obtain each target product for generating product order information, thereby obtaining the final product order information. The present invention also includes a device and a smart vending machine for executing the above method. The present invention can avoid the impact of product occlusion on detection accuracy and the generation of abnormal orders by obtaining shopping videos from different perspectives.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application filed on November 3, 2021, with the invention name “Intelligent order generation method and intelligent vending machine after fusion of multi-perspective image acquisition” and application number 202111291693.0. Technical Field

[0002] The present invention relates to the field of image analysis technology, and in particular to an order generation method for multi-view image comparison detection and an intelligent vending machine. Background Art

[0003] With the continuous development of artificial intelligence technology, the sales methods of the retail industry have also undergone tremendous changes. Smart vending machines have been widely used in various places in the city, including stations, shopping malls, tourist attractions or department stores. Various types of smart vending machines can be found. Smart vending machines do not require special supervision, users can automatically place orders and check out, which greatly meets the shopping needs of users in special scenarios.

[0004] However, existing smart vending machines include fully-open smart vending machines and smart vending machines that place orders by buttons. Among them, users of fully-open smart vending machines can perform multiple pick-up and put-down operations in one shopping trip when the smart vending machine door is open, and can select multiple items at a time and then settle the bill in a unified manner. This smart vending machine greatly facilitates the shopping needs of users. However, because this type of smart vending machine mainly relies on shopping videos to settle product orders, when users are picking up and putting down products, some features of the products will be blocked, which easily leads to false detection and abnormal order generation. For the other type of smart vending machines that place orders by buttons, when users need to shop, they need to determine all the products they need to buy in advance, and then purchase one item and settle it once before purchasing and settling the second item. All shopping can be completed after multiple orders and settlements, and products cannot be returned, so there is a problem of cumbersome operation. Summary of the Invention

[0005] In view of this, an embodiment of the present invention provides an order generation method and an intelligent vending machine based on multi-view image contrast detection, which are used to solve the technical problems of poor user experience caused by cumbersome operations or abnormal orders in existing intelligent vending machines.

[0006] The technical solution adopted in the present invention is:

[0007] The present invention provides a method for generating an order for multi-view image contrast detection, the method comprising:

[0008] Obtaining a first main video captured by a first main camera from a first perspective and a second main video of the same shopping event captured by a second main camera from a second perspective different from the first perspective;

[0009] Physically splicing the first main video and the second main video frames in sequence according to the acquisition timing of the frames in the first main video and the second main video to obtain a target video;

[0010] Detecting each frame image of the target video using a target detection model, and outputting a first detection result corresponding to each frame image of the first main video and a second detection result corresponding to each frame image of the second main video in the target video;

[0011] According to the first detection result and the second detection result, the product order information of each target product corresponding to the current shopping event is output.

[0012] Preferably, obtaining a first main video captured by a first main camera from a first perspective and a second main video of the same shopping event captured by a second main camera from a second perspective different from the first perspective includes:

[0013] Obtain the current status video of the smart vending machine captured by the third camera;

[0014] Analyze each frame of the vending machine status video and output cabinet door status information of the intelligent vending machine, wherein the cabinet door status information includes open state information and closed state information;

[0015] Control the first main camera and the second main camera to collect video data of the commodity area of ​​the smart vending machine according to the open status information, and control the first main camera and the second main camera to stop collecting video data of the commodity area of ​​the smart vending machine according to the closed status information to obtain the first main video and the second main video.

[0016] Preferably, obtaining a first main video captured by a first main camera from a first perspective and a second main video of the same shopping event captured by a second main camera from a second perspective different from the first perspective includes:

[0017] Obtaining the interval between each frame of image captured by the first main camera and the second main camera in the same time sequence;

[0018] According to the interval time, respectively controlling the first main camera and the second main camera to capture video images of the same shopping event to obtain the first main video and the second main video;

[0019] The interval time is smaller than the time interval between two frames of images corresponding to the frame rate.

[0020] Preferably, physically splicing the first main video and the second main video frame images in sequence according to the acquisition timing of each frame image in the first main video and the second main video to obtain the target video includes:

[0021] Obtain a first starting time corresponding to the first main video and a second starting time corresponding to the second main video;

[0022] Merging, one by one, each frame image of the first main video with each frame image corresponding to a time sequence in the second main video according to the first starting time, the second starting time, and the interval time to generate the target video;

[0023] The size of each frame image of the target video is the sum of the size of each frame image of the first main video and the size of each frame image corresponding to the second main video.

[0024] Preferably, the detecting each frame image of the target video using the target detection model and outputting a first detection result corresponding to each frame image of the first main video and a second detection result corresponding to each frame image of the second main video in the target video includes:

[0025] Dividing each frame image of the target video into a first image region belonging to the first main video and a second image region belonging to the second main video;

[0026] Detecting each frame of the target video using the target detection model to obtain a first detection result corresponding to the first image area and a second detection result corresponding to the second image area;

[0027] The first detection result and the second detection result both include: the category of each commodity, the confidence corresponding to each commodity, and the average confidence corresponding to each category of commodities detected in all frame images.

[0028] Preferably, the detecting each frame image of the target video using the target detection model to obtain a first detection result corresponding to the first image area and a second detection result corresponding to the second image area includes:

[0029] Detecting each frame of the target video using the target detection model to obtain a category and confidence score of each commodity corresponding to the first image region in each frame, and a category and confidence score of each commodity corresponding to the second image region in each frame;

[0030] Obtaining, based on the confidence levels of the commodities detected in the first image region and the second image region, an average value of first confidence levels corresponding to each commodity category among all commodities detected in the first image region across all image frames, and an average value of second confidence levels corresponding to each commodity category among all commodities detected in the second image region;

[0031] Obtaining, based on the categories of the commodities detected in the first image area and the second image area, first commodity information corresponding to all commodities detected in the first image area for each frame of image, and second commodity information corresponding to all commodities detected in the second image area for each frame of image;

[0032] The first detection result is obtained based on each piece of first product information and each first confidence average value corresponding to each piece of first product information, and the second detection result is obtained based on each piece of second product information and each second confidence average value corresponding to each piece of second product information.

[0033] Preferably, outputting the product order information of each target product corresponding to the current shopping event according to the first detection result and the second detection result includes:

[0034] Get the target confidence threshold for items used to generate valid orders;

[0035] Compare the product categories and quantities included in the first detection result with the product categories and quantities included in the second detection result. If the product categories of the first detection result and the second detection result are the same, output the product with a higher confidence level in each category of the first detection result and the second detection result as the target product; otherwise, generate an abnormal order;

[0036] Compare the confidence level of the target product with the target confidence level threshold. If the confidence level meets the requirements, produce the product order information corresponding to each target product. Otherwise, generate an abnormal order.

[0037] Among them, if it is an abnormal order, the abnormal order information is output and the first main video and / or the second main video and / or the target video are output at the same time.

[0038] The present invention also provides an order generation device for multi-view image contrast detection, the device comprising:

[0039] Video acquisition module: used to acquire a first main video captured by a first main camera from a first perspective and a second main video of the same shopping event captured by a second main camera from a second perspective different from the first perspective;

[0040] A video synthesis module is configured to physically splice the first main video and the second main video frames in sequence according to the acquisition timing of the frames in the first main video and the second main video to obtain a target video;

[0041] An image analysis module is configured to detect each frame image of the target video using a target detection model, and output a first detection result corresponding to each frame image of the first main video and a second detection result corresponding to each frame image of the second main video in the target video;

[0042] Order generation module: used to output product order information of each target product corresponding to the current shopping event based on the first detection result and the second detection result.

[0043] The present invention also provides an intelligent vending machine, comprising: at least one processor, at least one memory, and computer program instructions stored in the memory, wherein when the computer program instructions are executed by the processor, any of the above-mentioned methods is implemented.

[0044] The present invention further provides a medium having computer program instructions stored thereon, which implement any of the above methods when the computer program instructions are executed by a processor.

[0045] In summary, the beneficial effects of the present invention are as follows:

[0046] The present invention provides an order generation device method and an intelligent vending machine for multi-view image contrast detection. A first main camera and a second main camera are set on the intelligent vending machine to obtain video data of the same product area from different perspectives. Then, a first main video obtained by the first main camera and a second main video obtained by the second main camera are merged, and target detection is performed on the merged target video to obtain a first detection result corresponding to the first main video and a second detection result corresponding to the second main video. The product order information is determined by combining the first detection result and the second detection result, which can prevent order anomalies caused by obstruction of the product and improve detection accuracy and user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work, and these are all within the scope of protection of the present invention.

[0048] Figure 1 Schematic diagram of the process of intelligent order generation method after fusion of multi-view image acquisition in Example 1;

[0049] Figure 2 This is a schematic diagram of the structure of the smart vending machine with multiple cameras of different viewing angles in Example 1;

[0050] Figure 3 Schematic diagram of the process of intelligently generating orders by merging images captured from multiple perspectives in Example 2;

[0051] Figure 4 This is a flow chart of the intelligent order generation device for fusing images acquired from multiple perspectives in Example 3;

[0052] Figure 5 This is a schematic diagram of the process of the intelligent order generation device for merging images after multi-view capture in Example 4;

[0053] Figure 6 This is a schematic diagram of the structure of the automatic settlement system including the intelligent vending machine in Example 5;

[0054] Figure 7 This is a schematic diagram of the structure of the intelligent vending machine in Example 6;

[0055] Figures 1 to 7 Reference numerals:

[0056] 1. Cabinet body; 11. Partition; 12. Merchandise area; 2. Cabinet door; 3. Camera. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. In the description of the present invention, it should be understood that the orientation or position relationship indicated by the terms "center", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like is based on the orientation or position relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the present invention. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further limitations, elements defined by the phrase "comprising..." do not preclude the presence of other identical elements in the process, method, article, or apparatus comprising the elements. The various features of the present invention and its embodiments may be combined with each other if there is no conflict, and all are within the scope of protection of the present invention.

[0058] Example 1

[0059] Existing fully-opening smart vending machines have the advantages of being convenient for users to purchase multiple items at one time and change items multiple times during one shopping trip. At the same time, they can quickly generate orders when users complete a complex shopping process and quickly settle accounts through autonomous settlement. Compared with the existing methods that can only purchase one item by scanning a code once and cannot be re-selected after purchase, fully-opening smart vending machines have the advantages of simple operation and greater user autonomy in shopping. However, because fully-opening smart vending machines allow users to purchase multiple items in one shopping trip and can put and take items off the shelves multiple times, the occlusion problem will cause a large number of items to be detected differently before and after the same item is put on or taken off the shelves due to different occlusion positions, resulting in abnormal orders that affect user experience and merchant reputation.

[0060] The present invention is based on a feasibility study of obtaining shopping videos of users shopping from smart vending machines from multiple angles. By setting cameras in the product area of ​​the smart vending machine to monitor the product area from different directions in real time, combining the shopping videos shot by multiple cameras, and then obtaining the user's shopping order information through image splicing, comparative analysis, etc., and then automatically settling the bill through the server, the user's shopping experience is improved, while reducing the manual settlement process.

[0061] For details, see Figure 2 , Figure 2 The schematic diagram of the structure of a fully-open smart vending machine is as follows: the smart vending machine includes a cabinet 1 and a cabinet door 2. The cabinet 1 and the cabinet door 2 are rotatably connected. When the cabinet door 2 is in a closed state relative to the cabinet 1, the cabinet door 2 covers all the commodity areas of the cabinet 1 where commodities are placed, that is, the commodities in the cabinet cannot be taken out. When the cabinet door 2 is opened, all commodities in the cabinet 1 are displayed in front of the user. The user can select any commodity in a shopping mall, and can also select multiple commodities. The user can take out the selected commodity or put back the commodity that needs to be put back after selection. A partition 11 is provided in the cabinet 1. The partition 11 can be a shelf that divides the cabinet 1 into multiple commodity areas 12. Among them, a camera is provided in each commodity area inside the cabinet 1, so that the shopping video of the user shopping from the smart vending machine can be obtained from multiple angles, avoiding the problem that the shopping video collected from a single angle is not reliable due to occlusion. Figure 2 The smart vending machine shown is equipped with multiple cameras on the left and right internal walls of the vending machine so that shopping videos can be collected from the same product area in relative viewing directions, thereby improving the reliability of the video data.

[0062] See Figure 1 , Figure 1 This is a flow chart of a method for intelligently generating orders by fusing images acquired from multiple perspectives in Example 1 of the present invention. The method includes:

[0063] S10: Obtain a first main video captured by a first main camera at a first viewing angle of a product area, and a second main video captured by a second main camera at a second viewing angle different from the first viewing angle of the product area, wherein the first main camera and the second main camera have the same frame rate and the same number of frames for capturing the first and second main videos;

[0064] Specifically, a plurality of cameras are provided on the smart vending machine, and the plurality of cameras are divided into two groups of cameras, which are recorded as a first group of cameras and a second group of cameras. The first group of cameras includes a first main camera, and the first group of cameras may also include other first sub-cameras. The second group of cameras includes a second main camera, and the second group of cameras may also include other second sub-cameras. When only the first main camera and the second main camera are included, the first main camera and the second main camera are installed in the space corresponding to the top partition of all partitions where goods are placed, specifically, on the left and right panels inside the cabinet of the smart vending machine, and located above the top partition, so that the first main camera and the second main camera can both capture videos of goods being taken out from all areas of the smart vending machine; the viewing angle ranges of the first main camera and the second main camera can be the same or different. The viewing angles of the first and second main cameras are different, but can cover all product areas; therefore, in a preferred embodiment, the first main camera is set on the left wall inside the cabinet, and the second main camera is set on the right wall inside the cabinet, so that the first main camera shoots a video of the product area from the left, and the second main camera shoots a video of the product area from the right, thereby obtaining a first main video of the product area collected from the left by the first main camera at a first viewing angle and a second main video of the product area collected from the right by the second main camera at a second viewing angle. By obtaining videos of the product area from different viewing angles, when some features of the product are blocked by the user's product, more features of the product can be obtained from other viewing angles, thereby improving the accuracy of target detection; it should be noted that: the video frame rates of the first and second main videos are the same.

[0065] S11: Merging each frame image of the first main video with each frame image corresponding to the acquisition time sequence of the second main video one by one to generate a target video, wherein the size of each frame image of the target video is the sum of the size of each frame image of the first main video and the size of each frame image corresponding to the second main video;

[0066] Specifically, each frame image of the first main video and each frame image of the second main video are merged one by one according to the acquisition timing, so as to obtain a target video, wherein the merging method of each frame image of the first main video and the second main video includes but is not limited to: merging the image frames of the first main video corresponding to the same shooting moment with the image frames corresponding to the second main video, and merging the image frames of the first main video corresponding to the shooting moment staggered with the image frames of the same frame number corresponding to the second main video, specifically merging the first frame image of the first main video with the first frame image of the second main video, specifically merging the second frame image of the first main video with the second frame image of the second main video, and so on, specifically merging the Nth frame image of the first main video with the Nth frame image of the second main video, to obtain a target video composed of merged images; such as: the starting time of starting shooting of the first main video is earlier than the starting time of starting shooting of the second main video, it should be noted that: the starting time of starting shooting of the first main video and the starting time of starting shooting of the second main video are earlier than the starting time of starting shooting of the second main video. The time interval between the starting moments of shooting is less than the interval between two adjacent frames of images, such as the interval is 1 / 2, 1 / 3, 1 / 4, etc., of the interval between two adjacent frames of images. Then, the image frames of the first main video and the image frames of the second main video are merged in the order of shooting. More images at the corresponding shooting time interval under a given frame rate can be obtained, thereby enriching the number of samples and improving the accuracy of detection. It should be noted that because the frame rate of the video is related to the response of the human eye, if the frame rate is too low, the picture may be discontinuous. If the frame rate is too high, the eyes may not react quickly, causing eye discomfort. Therefore, a camera will set a frame rate that meets the requirements when capturing video, generally 20 frames / second to 30 frames / second. Taking 20 frames / second as an example, there is 0.05s between two adjacent frames of images. That is, there are only two last frames of images in 0.05 seconds. By using staggered shooting, more image frames at the 0.05s time interval can be obtained, thereby increasing the number of effective image frames without causing eye discomfort.

[0067] In one embodiment, the S11 includes:

[0068] S111: Obtain a first starting time corresponding to when the first main camera starts capturing the first main video and a second starting time corresponding to when the second main camera starts capturing the second main video, which is different from the first starting time;

[0069] S112: Merging, based on the first starting time and the second starting time, each frame image of the first main video acquired within a preset time with each frame image of the second main video corresponding to the time sequence, one by one, to generate the target video;

[0070] The time interval between the first starting moment and the second starting moment is smaller than the time interval between two adjacent frames of images corresponding to the camera frame rate.

[0071] Specifically, different starting times are set for the first main camera and the second main camera to start collecting video data of the product area. The starting time of the first main camera is set as the first starting time, and the starting time of the second main camera to start shooting is set as the second starting time. The time difference between the first starting time and the second starting time is less than the interval time between two adjacent frames corresponding to the video data collected by the camera; then, each frame image of the first main video and the second main video is merged one by one according to the acquisition sequence to obtain the target video. The specific merging method is not repeated here.

[0072] S12: Inputting each frame image of the target video into a preset target detection model to obtain product information of a number of target products;

[0073] Specifically, a preset target detection model is used to detect each frame image of the synthesized target video, so as to determine the product information of the target product taken out of the smart vending machine and the target product taken out of the smart vending machine and put back. The target detection model is composed of a large number of frames of videos corresponding to products being taken at different angles, held in different ways, and taken at different speeds, and each frame image is manually marked to obtain a sample set, and then the sample set is used to train the model to obtain the target detection model.

[0074] In one embodiment, the S12 includes:

[0075] S121: Divide each frame image of the target video into a first image region belonging to the first main video and a second image region belonging to the second main video;

[0076] Specifically, each frame image in the target video is partitioned, the image area belonging to the first main video is recorded as the first image area, and the image area belonging to the second main video is recorded as the second image area, so as to facilitate the statistics of the detection results of each frame image and avoid confusion of the detection results.

[0077] S122: Detecting each frame of the target video using the target detection model to obtain a first detection result corresponding to the first image area and a second detection result corresponding to the second image area;

[0078] Specifically, the detection results of the first image area of ​​each frame image are analyzed to obtain the first detection result belonging to the first main video and the second detection result belonging to the second main video in the target video. The detection results include the category of each commodity and the confidence of each commodity. The confidence can be the confidence in the detection result corresponding to each detected target, or it can be the average confidence corresponding to each category of commodities calculated from the confidence of each target in each category of commodities after all detected targets are classified according to the category information of the detection results; if the detection results of all frame images of the target video include commodity A detected 3 times, commodity B detected 5 times, and commodity C detected 4 times, then the average confidence corresponding to commodity A is the average of the confidence of commodity A detected each time. Similarly, the average confidence of commodities B and C is obtained.

[0079] In one embodiment, the S122 includes:

[0080] S1221: Detecting each frame of the target video using the target detection model to obtain a category and confidence score of each commodity corresponding to the first image region in each frame, and a category and confidence score of each commodity corresponding to the second image region in each frame;

[0081] S1222: Based on the confidence levels of the commodities detected in the first image area and the second image area, obtain an average value of first confidence levels corresponding to each commodity type among all commodities detected in the first image area for all frames of images, and an average value of second confidence levels corresponding to each commodity type among all commodities detected in the second image area.

[0082] Specifically, each image frame has a first image region and a second image region. The confidence values ​​of all objects corresponding to each type of commodity detected in the first image region and the second image region are respectively calculated to obtain an average confidence value for the corresponding commodity category. The average confidence value of each type of commodity in the first image region is recorded as the first average confidence value, and the average confidence value of each type of commodity in the second image region is recorded as the second average confidence value. For example, if the target video includes N frames of images, three objects are identified in the first image region, namely, commodity A, commodity B, and commodity C. Among them, commodity A is detected a total of 5 times, recorded as A1, A2, A3, A4, and A5, with confidence values ​​of 0.6, 0.7, 0.7, 0.75, and 0.75, respectively. Then, the first average confidence value of commodity A in the first image region is 0.7. Similarly, the average confidence value of each type of commodity for commodities B and C is calculated based on their respective detection times and corresponding confidence values. The same method is used to obtain the second average confidence value of each type of commodity in the second image region.

[0083] In one embodiment, the S1222 includes:

[0084] S12221: Obtain the detection count threshold corresponding to the number of valid detections of the product;

[0085] Specifically, each frame of the video is detected, and the target product of this order is determined based on the detection results of each frame. A corresponding detection number threshold is set for the same product in different image frames. If it is greater than or equal to the threshold, the product is considered to be true. If it is less than the threshold, the product is considered to be abnormal and requires manual review by the background.

[0086] S12222: Compare the number of detections corresponding to each category of goods with the detection number threshold, calculate the average of the confidence levels of the goods that meet the requirements, and obtain the average confidence level of each first product and the average confidence level of each second product.

[0087] Specifically, when the detection count threshold is set to 80% of the total number of image frames contained in the target video, taking the target video containing 20 frames of images as an example, the detection count threshold is 16 times. When product A is detected in 10 frames of images, the average confidence of product A is not calculated. If product B is detected in 18 frames of images, the average confidence of product B is calculated based on the confidence of each detection. This can avoid the misdetection of products due to angle problems and improve the accuracy of the detection results.

[0088] S1223: Based on the categories of the commodities detected in the first image area and the second image area, obtain first commodity information corresponding to all commodities detected in the first image area for each frame of image, and second commodity information corresponding to all commodities detected in the second image area for each frame of image;

[0089] Specifically, the categories of commodities detected in each frame image are counted separately according to the first image area and the second image area, thereby obtaining first commodity information corresponding to the first image area and second commodity information corresponding to the second image area of ​​the target video.

[0090] S1224: Obtain the first detection result based on each piece of first product information and each first confidence average value corresponding to each piece of first product information; obtain the second detection result based on each piece of second product information and each second confidence average value corresponding to each piece of second product information.

[0091] Specifically, the average confidence level of each first product and the average confidence level of each second product are compared with the confidence threshold respectively, and classified according to the first image area and the second image area. The products corresponding to the average confidence level of each first product that meets the requirements are used as the first detection results, and the products corresponding to the average confidence level of each second product that meets the requirements are used as the second detection results.

[0092] S123: Compare the first detection result and the second detection result to obtain the product information of each target product used to generate product order information;

[0093] The first detection result and the second detection result both include: the category of each commodity, the confidence corresponding to each commodity, and the average confidence corresponding to each category of commodities detected in all frame images.

[0094] Specifically, the first detection result and the second detection result are statistics of the detection results of multiple frames of images. If the commodity categories in the first detection result and the second detection result are the same, and the confidence of the target commodity in the detection result meets the requirements, the target commodity for generating order information is obtained. If the first detection result and the second detection result are different, the commodities in the detection result with more quantity and category are used as target commodities (because of angle problems, there will be occlusion, resulting in the detection process. Some commodities do not meet the requirements at specific angles and are not counted, so the detection results with more quantity and meeting the requirements are used). If the confidence does not meet the requirements, it is considered to be an abnormal order, and the target video is manually reviewed to obtain the target commodity for generating order information. Because it is a continuous frame detection, if the commodity category and / or quantity of the first detection result and the second detection result are inconsistent, it is considered that the commodity that does not appear in one detection result does belong to the blind spot of the camera, and the detected detection result is directly used as the basis for generating order information, effectively avoiding missed detection caused by the blind spot of a single camera and improving the accuracy of the order.

[0095] S13: Generate product order information based on the product information of each target product.

[0096] Specifically, after obtaining the category of each target product, the product database is traversed to determine the product price, thereby generating product order information. The user can pay at the terminal, which includes a traditional checkout counter or a third-party App software.

[0097] In one embodiment, the S13 includes:

[0098] S131: Obtaining a target confidence threshold for a product used to generate a valid order;

[0099] S132: Compare the commodity categories and quantities included in the first detection result with the commodity categories and quantities included in the second detection result. If the commodity categories of the first detection result and the second detection result are the same, output the commodity with a higher confidence level in each category of the first detection result and the second detection result as the target commodity; otherwise, generate an abnormal order;

[0100] S133: Compare the confidence level corresponding to the target product with the target confidence level threshold. If the confidence level meets the requirements, produce the product order information corresponding to each target product. Otherwise, generate an abnormal order.

[0101] Among them, if it is an abnormal order, the abnormal order information is output and the first main video and / or the second main video and / or the target video are output at the same time.

[0102] Specifically, when it is believed that the target product corresponding to the current detection result is abnormal, manual review is required. At this time, the first main video and / or the second main video and / or the target video can be reviewed to ensure the accuracy of the order and improve the user experience.

[0103] In one embodiment, the first detection area corresponding to the first main camera includes the second detection area corresponding to the second main camera.

[0104] In one embodiment, the first detection area and the second detection area both cover the entire commodity area, and the viewing angle direction of the first main camera is arranged to face the viewing angle direction of the second main camera.

[0105] Specifically, the detection areas of the first main camera and the second main camera are inclusive, that is, the first detection area and the second detection area are the same, or one detection area belongs to the detection range of the other detection area. At the same time, the viewing angles of the first main camera and the second main camera are different, such as the first main camera is from left to right, and the second main camera is from right to left. It should be noted that the first main camera and the second main camera are set in the same commodity area. For example, there are three commodity areas of upper, middle and lower in the smart vending machine. If the first main camera is installed in the upper commodity area, the second main camera should also be installed in the upper commodity area. Similarly, if the first main camera is installed in the middle commodity area, the second main camera should also be installed in the middle commodity area. If the first main camera is installed in the lower commodity area, the second main camera should also be installed in the lower commodity area to avoid commodity obstruction caused by partition partitions at different stages and improve detection accuracy.

[0106] In one embodiment, before S10, the method further includes:

[0107] S01: Obtain the video of the current state of the smart vending machine captured by the third camera in real time;

[0108] Specifically, the smart vending machine is also provided with a third camera, which is used to detect whether the smart vending machine is turned on or off. The third camera can be turned on in real time or after the user makes a shopping request.

[0109] S02: Analyze each frame of the video of the current state of the smart vending machine to determine whether the door of the smart vending machine is in an open or closed state;

[0110] S03: When it is detected that the door of the smart vending machine is in an open state, controlling the first main camera and the second main camera for collecting video information corresponding to the commodity area to be turned on;

[0111] S04: When it is detected that the cabinet door of the smart vending machine is in a closed state, the first main camera and the second main camera for collecting video information corresponding to the commodity area are controlled to be turned off.

[0112] Specifically, when a user performs automatic shopping, each frame of the vending machine status video is analyzed to determine the status of the vending machine door. When it is detected that the vending machine door is open, the first main camera and the second main camera are turned on to obtain video data of the commodity area, thereby obtaining the first main video and the second main video; when it is detected that the vending machine door is closed, the first main camera and the second main camera are turned off.

[0113] The intelligent order generation method of multi-perspective image acquisition and fusion in this embodiment is adopted. By setting a first main camera and a second main camera on the intelligent vending machine to obtain video data of the same product area from different perspectives, and then merging the first main video obtained by the first main camera and the second main video obtained by the second main camera, target detection is performed on the merged target video to obtain a first detection result corresponding to the first main video and a second detection result corresponding to the second main video. The product order information is determined by combining the first detection result and the second detection result, which can prevent order anomalies caused by obstruction of the product and improve detection accuracy and user experience.

[0114] Example 2

[0115] In Example 1, the first and second main cameras with different viewing angles are set for the product areas of the smart vending machine. However, the imaging positions of products in different product areas are different in each frame image, which often leads to false detection or mixed detection of products with high similarity, affecting the detection accuracy. Therefore, Example 2 of the present invention further improves the method for automatically generating order information for smart vending machines based on Example 1; please refer to Figure 3 , the method comprising:

[0116] S20: Acquire a first main video obtained by capturing the product area with a first main camera, a second main video obtained by capturing the product area with a second main camera, a first sub-video obtained by capturing the product area with at least one first sub-camera, and a second sub-video obtained by capturing the product area with at least one second sub-camera;

[0117] In one embodiment, the S20 includes:

[0118] S201: Dividing the merchandise area of ​​the smart vending machine into multiple virtual merchandise areas along the arrangement direction of the shelves of the smart vending machine;

[0119] S202: Control the camera corresponding to each commodity area to collect video data of the corresponding commodity area, and obtain the first main video collected by the first main camera and the second main video collected by the second main camera, as well as the first sub-video collected by at least one of the first sub-camera and the second sub-video collected by at least one of the second sub-camera.

[0120] Specifically, the smart vending machine has multiple layers of shelves, so the smart vending machine is divided into multiple commodity areas, each commodity area includes at least one layer of shelves, and a first main camera and a second main camera are arranged opposite to each other in each commodity area, or a first sub-camera and a second sub-camera are arranged opposite to each other; for ease of understanding, the camera arranged on the left side of the topmost commodity area is recorded as the first main camera and the camera arranged on the right side is recorded as the second main camera, the camera arranged on the left side of the non-top commodity area is recorded as the first sub-camera and the camera arranged on the right side is recorded as the second sub-camera. For example, the commodity area of ​​the smart vending machine is divided into three commodity areas, upper, middle and lower. The first main camera and the second main camera are arranged opposite to each other at the top, the viewing direction of the first main camera is from left to right, and the viewing direction of the second main camera is from right to left. The first sub-camera is arranged on the left side of the middle commodity area and the lower commodity area, and the second sub-camera is arranged on the right side; when the user starts to purchase commodities through the smart vending machine, each camera is controlled to collect video data of the commodity area to obtain the first main video, the second main video, each first sub-video and each second sub-video.

[0121] In one embodiment, the S20 includes:

[0122] S203: Obtaining the frame rate and number of cameras used to collect video data of the commodity area;

[0123] S204: Determine, based on the frame rate and the number of cameras, the interval at which each camera starts collecting video data of the corresponding product area;

[0124] S205: Obtaining, based on the interval time, starting times for the first main camera, the second main camera, each of the first sub-cameras, and each of the second sub-cameras to start collecting video data;

[0125] S206: According to each of the starting moments, control the corresponding camera to collect video data of the product area to obtain the first main video, the second main video, each of the first sub-videos, and each of the second sub-videos.

[0126] Specifically, the frame rates of the cameras used to collect video data are the same, such as 20 frames per second; the starting time for each camera to start collecting video data is determined based on the number and frame rate of the cameras, and a time interval exists between the starting time of collecting video data between each camera or each group of cameras, wherein the preferred interval time is an integer multiple of the time difference corresponding to two adjacent frames of images, such as: including 4 cameras, the start time of collection of each camera is spaced apart by 1 / 4 of the time corresponding to the frame rate, or the 4 cameras are divided into two groups, and the interval time for each group of cameras to start collecting video data is 1 / 2 of the time corresponding to the frame rate; thereby indirectly increasing the image frame rate, ensuring that image information of the commodity area at more moments is collected, so as to improve the accuracy of detection. S21: Cover the image area in the first main video that is the same as the first sub-video with the commodity area of ​​the first sub-video to obtain the first target video, and cover the image area in the second main video that is the same as the second sub-video with the commodity area corresponding to the second sub-video to obtain the second target video;

[0127] Specifically, because the first main camera and the second main camera are set in the top-level commodity area, if there is no obstruction in the non-top-level commodity area, the picture of the commodity area can also be captured by the first main camera and the second main camera. Therefore, each frame image in the first main video contains the image area of ​​each frame image in each first sub-video, and the image area of ​​the non-top-level commodity area corresponding to each frame image in the first main video is replaced with the corresponding frame image in each first sub-video, thereby obtaining the first target video. The second target video is obtained in the same way. In the pictures corresponding to each commodity area in each frame image of the first target video and the second target video, the imaging of the commodity is clearer, thereby improving the accuracy of detection.

[0128] In one embodiment, the S21 includes:

[0129] S211: Obtain the target product area corresponding to the product change;

[0130] S212: Determine the first sub-video and the second sub-video including the target product area according to the location information of the target product area;

[0131] Specifically, when a user takes out or puts goods from the smart vending machine, the location where the goods are taken out or put in is determined, specifically the product area where the goods are put in or taken out, and then the nearest sub-camera in the area and the corresponding first sub-video and / or second sub-video are determined. For example, the goods of the smart vending machine are divided into three product areas: upper, middle and lower. When the user takes out or puts in goods from the lower product area, the sub-video of the target product area is the first sub-video corresponding to the first sub-camera of the lower product area and the second sub-video corresponding to the second sub-camera at the beginning; when the goods are continuously taken out and enter the area corresponding to the middle product area, the sub-video becomes the first sub-video and the second sub-video corresponding to the first sub-camera and the second sub-camera of the middle product area.

[0132] In one embodiment, the S212 includes:

[0133] S2121: Analyze each frame of the first main video and / or the second main video to determine each image frame corresponding to a product change;

[0134] S2122: Divide the first main video and the second main video into multiple video segments according to each image frame in which a product change occurs;

[0135] S2123: Determine the first sub-video and the second sub-video corresponding to each of the video segments according to the target product area corresponding to the product change in each of the video segments.

[0136] Specifically, each frame of the first main video and / or the second main video is analyzed to determine the adjacent image frames in which the product has changed in each frame, and the first main video and / or the second main video are segmented according to the adjacent image frames in which the product has changed. After performing position analysis on each video segment, the product area closest to the product in each video segment at each moment is determined, and the sub-video corresponding to the product at each moment is filtered out for merging, avoiding merging all sub-videos, thereby reducing the amount of data processing.

[0137] S213: Merge the first sub-video containing the target product with the first main video to obtain the first target video, and merge the second sub-video containing the target product with the second main video to obtain the second target video.

[0138] In one embodiment, the first sub-video and the second sub-video including the target product area are recorded as target sub-videos, and S213 includes:

[0139] S2131: Determine, based on the position information of the target sub-video, other first sub-videos or second sub-videos located between the target sub-video and the first main video or the second main video as intermediate sub-videos;

[0140] S2132: Deleting areas of the first main video and the second main video that do not correspond to the target sub-video and the middle sub-video, to obtain optimized first and second videos, respectively;

[0141] Specifically, when merging each frame image of the sub-video into the image area corresponding to each frame image of the main video, not all frame images of the sub-video are merged. It is necessary to delete each frame image without the target product and only merge each frame image with the target product. Furthermore, the image area where the target is detected is merged to avoid reducing the image quality due to scaling the image and improve the detection accuracy.

[0142] S2133: Merge the target sub-video and each intermediate sub-video belonging to the first sub-video with the first video to obtain the first target video, and merge the target sub-video and each intermediate sub-video belonging to the second sub-video with the second video to obtain the second target video.

[0143] Specifically, the image area between the main video and the target sub-video is replaced with the image area of ​​the corresponding sub-video, thereby ensuring the integrity of the synthesized first target video and / or second target video image, so as to facilitate back-end manual review when abnormal orders occur.

[0144] S22: Merging the first target video and each frame image of the second target video to obtain a target video;

[0145] Specifically, each frame image of the first target video and each frame image of the second target video are merged one by one according to the acquisition timing, so as to obtain a target video, wherein the merging method of each frame image of the first target video and the second target video includes but is not limited to: merging the image frames of the first target video corresponding to the same shooting moment with the image frames corresponding to the second target video, and merging the image frames of the first target video corresponding to the shooting moment staggered with the image frames of the same frame number corresponding to the second target video, specifically merging the first frame image of the first target video with the first frame image of the second target video, merging the second frame image of the first target video with the second frame image of the second target video, and so on, merging the Nth frame image of the first target video with the Nth frame image of the second target video to obtain a target video composed of merged images; such as: the starting moment of starting shooting of the first target video is earlier than the starting moment of starting shooting of the second target video, it should be noted that: the starting moment of starting shooting of the first target video is earlier than the starting moment of starting shooting of the second target video. The time interval between the start times of the target video shooting is less than the interval between two adjacent frames, such as 1 / 2, 1 / 3, 1 / 4 of the interval between two adjacent frames, and then the image frames of the first target video are merged with the image frames of the second target video in the order of shooting. In this way, more images at the corresponding shooting time interval can be obtained at a given frame rate, thereby enriching the number of samples and improving the accuracy of detection. It should be noted that because the frame rate of the video is related to the human eye response, if the frame rate is too low, the picture may be discontinuous, and if the frame rate is too high, the eyes may not react quickly, causing eye discomfort. Therefore, a camera will set a frame rate that meets the requirements when capturing video, generally 20 frames / second to 30 frames / second. Taking 20 frames / second as an example, the interval between two adjacent frames is 0.05s, that is, there are only two final frames in 0.05 seconds. By using staggered shooting, more image frames at the 0.05s time interval can be obtained, thereby increasing the number of effective image frames without causing eye discomfort.

[0146] In one embodiment, the S22 includes:

[0147] S221: Acquire a first starting time for starting to capture the first target video and a second starting time for starting to capture the second target video;

[0148] S222: Merge each frame image of the first target video with each frame image corresponding to the acquisition timing of the second target video one by one according to the first starting time and the second starting time to obtain the target video.

[0149] Specifically, the starting time interval of the first main camera and its first sub-camera and the second main camera and its second sub-camera is set to be an integer multiple of the interval time between two adjacent frames of images, and then the first frame image of the first target video is merged with the first frame image of the second target video, and the second frame image of the first target video is merged with the second frame image of the second target video, and so on, the Nth frame image of the first target video is merged with the Nth frame image of the second target video to obtain a target video composed of merged images; thereby ensuring that the product obtains images of more moments within the interval time between two adjacent frames corresponding to the fixed frame rate, reducing the loss of product information caused by occlusion, and thus improving the accuracy of detection.

[0150] S23: Inputting each frame image of the target video into a preset target detection model to obtain product information of each target product;

[0151] Specifically, for a method of identifying each frame image of a target video according to a target detection model to obtain product information of a target product, please refer to the embodiment for details, which will not be described in detail here.

[0152] S24: Generate product order information based on the product information of each target product;

[0153] Among them, the range of the commodity area captured by the first main camera and the second main camera is the same, the range of the commodity area captured by each of the first sub-cameras and the second sub-cameras belongs to the partial area corresponding to the commodity area captured by the first main camera or the second main camera, the first main camera and each of the first sub-cameras are arranged along the arrangement direction of each layer of shelves of the smart vending machine, and each of the first sub-cameras is located below the first main camera, the second main camera and each of the second sub-cameras are arranged along the arrangement direction of each layer of shelves of the smart vending machine, and each of the second sub-cameras is located below the second main camera.

[0154] Specifically, the method for generating product order information is described in the embodiment and will not be repeated here.

[0155] The intelligent order generation method of merging images after multi-perspective capture of the present embodiment is adopted. A first main camera and a second main camera are set on the intelligent vending machine to obtain video data of the same product area with different perspectives, and a first sub-camera and a second sub-camera are also provided for the first camera and the second camera to obtain video data of the product area. The first sub-video obtained by the first sub-camera replaces the data of the corresponding area in the first main video obtained by the first main camera to obtain a first target video. The same method is used to obtain the second target video, and then the first target video and the second target video are merged to obtain the target video, ensuring the imaging clarity of the target object in each frame image in the target video. Target detection is performed on the merged target video to obtain a first detection result corresponding to the first target video and a second detection result corresponding to the second target video. The product order information is determined in combination with the first detection result and the second detection result, which can prevent order anomalies caused by obstruction of the product and improve detection accuracy and user experience.

[0156] Example 3

[0157] Example 3 of the present invention further provides an intelligent order generation device based on the methods of Example 1 to Example 2, which can be seen in the following example: Figure 4 ,include:

[0158] A video acquisition module is configured to acquire a first main video obtained by capturing a product area from a first viewing angle by a first main camera, and a second main video obtained by capturing the product area from a second viewing angle different from the first viewing angle by a second main camera, wherein the first main camera and the second main camera have the same frame rate and capture the same number of frames for the first and second main videos;

[0159] A video synthesis module is configured to merge each frame image of the first main video with each frame image corresponding to the acquisition time sequence of the second main video one by one to generate a target video, wherein the size of each frame image of the target video is the sum of the size of each frame image of the first main video and the size of each frame image corresponding to the second main video;

[0160] Image analysis module: used to input each frame image of the target video into a preset target detection model to obtain product information of several target products;

[0161] Order generation module: used to generate product order information based on the product information of each target product;

[0162] The range of the commodity area captured by the first main camera is the same as the range of the commodity area captured by the second main camera.

[0163] The intelligent order generation device that fuses multi-perspective image acquisition in this embodiment is adopted. By setting a first main camera and a second main camera on the intelligent vending machine to obtain video data of the same product area from different perspectives, and then merging the first main video obtained by the first main camera and the second main video obtained by the second main camera, target detection is performed on the merged target video to obtain a first detection result corresponding to the first main video and a second detection result corresponding to the second main video. The product order information is determined by combining the first detection result and the second detection result, which can prevent order anomalies caused by obstruction of the product and improve detection accuracy and user experience.

[0164] It should be noted that the device also includes the remaining technical solutions described in Examples 1 and 2, which will not be repeated here.

[0165] Example 4

[0166] Example 4 of the present invention further provides an intelligent order generation device for merging images after multi-view acquisition based on the methods of Example 1 to Example 2, see Figure 5 ,include:

[0167] Video acquisition module: used to obtain a first main video obtained by a first main camera capturing a product area, a second main video obtained by a second main camera capturing a product area, as well as a first sub-video obtained by at least one first sub-camera capturing a product area, and a second sub-video obtained by at least one second sub-camera capturing a product area;

[0168] Video splicing module: used to overlay the same image area in the first main video as the first sub-video with the product area of ​​the first sub-video to obtain a first target video, and overlay the same image area in the second main video as the second sub-video with the product area of ​​the second sub-video to obtain a second target video;

[0169] A video merging module is configured to merge each frame image of the first target video with each frame image corresponding to the acquisition timing of the second target video to obtain a target video, wherein the size of each frame image of the target video is the sum of the size of each frame image of the first target video and the size of each frame image corresponding to the second target video;

[0170] Data analysis module: used to input each frame image of the target video into a preset target detection model to obtain product information of each target product;

[0171] Order information module: Generates product order information based on the product information of each target product;

[0172] Among them, the range of the commodity area captured by the first main camera and the second main camera is the same, the range of the commodity area captured by each of the first sub-cameras and the second sub-cameras belongs to the partial area corresponding to the commodity area captured by the first main camera or the second main camera, the first main camera and each of the first sub-cameras are arranged along the arrangement direction of each layer of shelves of the smart vending machine, and each of the first sub-cameras is located below the first main camera, the second main camera and each of the second sub-cameras are arranged along the arrangement direction of each layer of shelves of the smart vending machine, and each of the second sub-cameras is located below the second main camera.

[0173] The intelligent order generation device of the multi-view image acquisition and merging of the present embodiment is adopted, by setting a first main camera and a second main camera on the intelligent vending machine to obtain video data of the same product area with different perspectives, and further providing a first sub-camera and a second sub-camera for the first camera and the second camera to obtain video data of the product area, using the first sub-video obtained by the first sub-camera to replace the data of the corresponding area in the first main video obtained by the first main camera to obtain a first target video, using the same method to obtain the second target video, and then merging the first target video and the second target video to obtain the target video, ensuring the imaging clarity of the target object in each frame image in the target video, performing target detection on the merged target video, obtaining a first detection result corresponding to the first target video and a second detection result corresponding to the second target video, and determining the product order information in combination with the first detection result and the second detection result, thereby preventing order anomalies caused by obstruction of the product and improving detection accuracy and user experience.

[0174] It should be noted that the device also includes the remaining technical solutions described in Example 4, which will not be repeated here.

[0175] Example 5

[0176] The present invention provides an automatic settlement system for intelligent vending machines. Figure 6The automatic settlement system includes an intelligent vending machine, a mobile terminal, and a server. The automatic settlement system can adopt the automatic shopping method described in the above embodiment. The user recognizes the identification code on the intelligent vending machine through the mobile terminal, and the server establishes the user's shopping event. The cameras with different viewing angles begin to capture shopping videos. The cameras begin to capture shopping videos after the intelligent vending machine door is opened or after the user enters a preset range. When the user leaves the preset shopping range or the intelligent vending machine door is closed, the cameras stop capturing shopping videos and transmit the shopping videos to the server. The server generates the user's order information based on the shopping video and sends it to the mobile terminal. The user performs self-settlement or sets automatic settlement through the order information on the mobile terminal. The automatic settlement system provides users with better self-shopping selectivity and high order accuracy, which can improve the user's shopping experience.

[0177] Example 6

[0178] The present invention provides an intelligent vending machine device and a storage medium, such as Figure 7 As shown, the system includes at least one processor, at least one memory, and computer program instructions stored in the memory.

[0179] Specifically, the above-mentioned processor may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of an embodiment of the present invention, and the electronic device includes at least one of the following: a camera, a mobile device with a camera, and a wearable device with a camera.

[0180] The memory may include a large capacity memory for data or instructions. By way of example and not limitation, the memory may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory may include a removable or non-removable (or fixed) medium. Where appropriate, the memory may be inside or outside the data processing device. In a specific embodiment, the memory is a non-volatile solid-state memory. In a specific embodiment, the memory includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.

[0181] The processor reads and executes computer program instructions stored in the memory to implement any one of the intelligent order generation methods for fusing multi-perspective image acquisition and the intelligent order generation method for merging multi-perspective image acquisition in the above-mentioned embodiment mode 1.

[0182] In one example, the electronic device may further include a communication interface and a bus, wherein the processor, the memory, and the communication interface are connected via the bus and communicate with each other.

[0183] The communication interface is mainly used to implement communication between the modules, devices, units and / or equipment in the embodiments of the present invention.

[0184] Bus comprises hardware, software or both, couples the parts of electronic equipment to each other.For example, and not limitation, bus can comprise accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations.In suitable cases, bus can comprise one or more buses.Although the embodiment of the present invention describes and shows specific bus, the present invention considers any suitable bus or interconnection.

[0185] In summary, the embodiments of the present invention provide a method for intelligently generating orders by fusing images captured from multiple perspectives, a method, device, intelligent vending machine and storage medium for intelligently generating orders by merging images captured from multiple perspectives.

[0186] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.

[0187] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in unit, a function card or the like. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0188] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating orders for multi-view image contrast detection, characterized in that: The method comprises: Obtaining a first main video captured by a first main camera from a first perspective and a second main video of the same shopping event captured by a second main camera from a second perspective different from the first perspective; Physically splicing the first main video and the second main video frames in sequence according to the acquisition timing of the frames in the first main video and the second main video to obtain a target video; Detecting each frame image of the target video using a target detection model, and outputting a first detection result corresponding to each frame image of the first main video and a second detection result corresponding to each frame image of the second main video in the target video; Outputting product order information of each target product corresponding to the current shopping event according to the first detection result and the second detection result, specifically including: obtaining a target confidence threshold for products used to generate a valid order; comparing the product categories and product quantities contained in the first detection result with the product categories and product quantities contained in the second detection result; if the product categories of the first detection result and the second detection result are the same, outputting the product with a higher confidence in each type of product in the first detection result and the second detection result as the target product; otherwise, generating an abnormal order; comparing the confidence corresponding to the target product with the target confidence threshold; if it meets the requirements, generating the product order information corresponding to each target product; otherwise, generating an abnormal order; wherein, if it is an abnormal order, outputting the abnormal order information and simultaneously outputting the first main video and / or the second main video and / or the target video, wherein both the first detection result and the second detection result include: the category of each product, the confidence corresponding to each product one-to-one, and the average confidence corresponding to each type of product detected in all frame images.

2. The order generation method for multi-view image contrast detection according to claim 1, characterized in that: The acquiring of a first main video captured by a first main camera from a first perspective and a second main video of the same shopping event captured by a second main camera from a second perspective different from the first perspective includes: Obtain the current status video of the smart vending machine captured by the third camera; Analyze each frame of the vending machine status video and output cabinet door status information of the intelligent vending machine, wherein the cabinet door status information includes open state information and closed state information; Control the first main camera and the second main camera to collect video data of the commodity area of ​​the smart vending machine according to the open status information, and control the first main camera and the second main camera to stop collecting video data of the commodity area of ​​the smart vending machine according to the closed status information to obtain the first main video and the second main video.

3. The order generation method for multi-view image contrast detection according to claim 1 or 2, characterized in that: The acquiring of a first main video captured by a first main camera from a first perspective and a second main video of the same shopping event captured by a second main camera from a second perspective different from the first perspective includes: Obtaining the interval between each frame of image captured by the first main camera and the second main camera in the same time sequence; According to the interval time, respectively controlling the first main camera and the second main camera to capture video images of the same shopping event to obtain the first main video and the second main video; The interval time is smaller than the time interval between two frames of images corresponding to the frame rate.

4. The order generation method for multi-view image contrast detection according to claim 3, characterized in that: The step of physically splicing the first main video and the second main video frame images in sequence according to the acquisition timing of the frame images in the first main video and the second main video to obtain the target video includes: Obtain a first starting time corresponding to the first main video and a second starting time corresponding to the second main video; Merging, one by one, each frame image of the first main video with each frame image corresponding to a time sequence in the second main video according to the first starting time, the second starting time, and the interval time to generate the target video; The size of each frame image of the target video is the sum of the size of each frame image of the first main video and the size of each frame image corresponding to the second main video.

5. The order generation method for multi-view image contrast detection according to claim 3, characterized in that: The detecting each frame image of the target video using the target detection model and outputting a first detection result corresponding to each frame image of the first main video and a second detection result corresponding to each frame image of the second main video in the target video includes: Dividing each frame image of the target video into a first image region belonging to the first main video and a second image region belonging to the second main video; The target detection model is used to detect each frame image of the target video to obtain a first detection result corresponding to the first image area and a second detection result corresponding to the second image area.

6. The order generation method for multi-view image contrast detection according to claim 5, characterized in that: The detecting each frame image of the target video using the target detection model to obtain a first detection result corresponding to the first image area and a second detection result corresponding to the second image area includes: Detecting each frame of the target video using the target detection model to obtain a category and confidence score of each commodity corresponding to the first image region in each frame, and a category and confidence score of each commodity corresponding to the second image region in each frame; Obtaining, based on the confidence levels of the commodities detected in the first image region and the second image region, an average value of first confidence levels corresponding to each commodity category among all commodities detected in the first image region across all image frames, and an average value of second confidence levels corresponding to each commodity category among all commodities detected in the second image region; Obtaining, based on the categories of the commodities detected in the first image area and the second image area, first commodity information corresponding to all commodities detected in the first image area for each frame of image, and second commodity information corresponding to all commodities detected in the second image area for each frame of image; The first detection result is obtained based on each piece of first product information and each first confidence average value corresponding to each piece of first product information, and the second detection result is obtained based on each piece of second product information and each second confidence average value corresponding to each piece of second product information.

7. An order generation device for multi-view image contrast detection, characterized in that: The device comprises: Video acquisition module: used to acquire a first main video captured by a first main camera from a first perspective and a second main video of the same shopping event captured by a second main camera from a second perspective different from the first perspective; A video synthesis module is configured to physically splice the frames of the first main video and the second main video in sequence according to the acquisition timing of the frames of the first main video and the second main video to obtain a target video; An image analysis module is configured to detect each frame image of the target video using a target detection model, and output a first detection result corresponding to each frame image of the first main video and a second detection result corresponding to each frame image of the second main video in the target video; An order generation module is configured to output the product order information of each target product corresponding to the current shopping event based on the first detection result and the second detection result, specifically to: obtain a target confidence threshold for the product used to generate a valid order; compare the product categories and product quantities contained in the first detection result with the product categories and product quantities contained in the second detection result; if the product categories of the first detection result and the second detection result are the same, output the product with a higher confidence in each category of the first detection result and the second detection result as the target product; otherwise, generate an abnormal order; compare the confidence corresponding to the target product with the target confidence threshold; if it meets the requirements, generate the product order information corresponding to each target product; otherwise, generate an abnormal order; wherein, if it is an abnormal order, output the abnormal order information and simultaneously output the first main video and / or the second main video and / or the target video, wherein both the first detection result and the second detection result include: the category of each product, the confidence corresponding to each product, and the average confidence corresponding to each category of products detected in all frame images.

8. An intelligent vending machine, characterized in that: include: At least one processor, at least one memory, and computer program instructions stored in the memory, which implement the method according to any one of claims 1 to 6 when the computer program instructions are executed by the processor.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Article detection method, system and computer-readable storage medium

    CN109308460A

  • Commodity detection method and device and readable storage medium

    CN111626201A