Settlement Method for AI Recognition of Items in a Vending Cabinet

By obtaining multi-view images in the sales cabinet and constructing a local item model to compare with the reference image model, the problem of inaccurate identification caused by the obstruction of items is solved, and efficient item identification and settlement is achieved.

CN120014617BActive Publication Date: 2025-07-22ZHEJIANG HI CONVENIENCE NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510487135.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-22
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

In the prior art, the items are blocked when the items are identified by the sales cabinet, resulting in inaccurate identification, which affects settlement efficiency.

Method used

By acquiring multiple objects images from different angles, building local object models and comparing them with the reference image model, combining the dynamic contour comparison mechanism to accurately capture the changes in object characteristics, and realizing multi-view image acquisition and accurate recognition.

Benefits of technology

It significantly improves the identification and settlement efficiency of sales cabinet items, ensures the integrity and feature recognition of image acquisition, and improves the accuracy and efficiency of settlement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014617B_ABST
    Figure CN120014617B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of image recognition, specifically a settlement method for AI recognition of items in a vending cabinet. The method includes the following steps: obtaining a video stream of a user purchasing items in a vending cabinet, intercepting multiple images of items for sale from the video stream to obtain an item image set, and extracting local feature information of the items for sale based on the item image set; constructing a local item model corresponding to the items for sale based on the local feature information; collecting reference image information of multiple reference items in the vending cabinet, and constructing a reference image model corresponding to the reference items based on the reference image information; comparing the local item model with multiple reference image models, extracting the reference image models that meet the preset conditions, and using the reference items corresponding to the reference image models as the target items corresponding to the items for sale, which can accurately identify the items in the vending cabinet. Through an automated recognition and settlement process, the settlement efficiency of the vending cabinet can be significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and specifically to a settlement method for AI recognition of items in a vending cabinet. Background Art

[0002] In the field of intelligent vending cabinets, accurately identifying product information and constructing its three-dimensional model is of great significance for enhancing the user experience and optimizing inventory management. AI recognition of items in the vending cabinet mainly relies on a camera to collect continuous images of the user's purchase to form a video stream, and identifies the items for sale from the video stream for settlement. However, in the prior art, when identifying items in the vending cabinet, when the user takes an item, the item is easily blocked, so only a partial image of the item for sale can be recognized, which affects the accuracy of item recognition and further affects the settlement efficiency of the item for sale.

[0003] For this reason, we propose a settlement method for AI recognition of items in a vending cabinet to solve the above problems. Summary of the Invention

[0004] The purpose of the present invention is to provide a settlement method for AI recognition of items in a vending cabinet to solve the problems raised in the above background art.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A settlement method for AI recognition of items in a vending cabinet, the method comprising the following steps:

[0006] Obtain a video stream of the user purchasing items in the vending cabinet, and intercept multiple images of the items for sale from the video stream to obtain an item image set, wherein the item image set includes at least two images of the item from different angles;

[0007] Extract local feature information of the item for sale based on the item image set; construct a local item model corresponding to the item for sale based on the local feature information;

[0008] Collect reference image information of multiple reference items in the vending cabinet, and construct a reference image model corresponding to the reference item based on the reference image information;

[0009] Compare the local item model with multiple reference image models, extract the reference image models that meet the preset conditions, and use the reference items corresponding to the reference image models as the target items corresponding to the items for sale;

[0010] Obtain the item information of the target item in the video stream and the price corresponding to the target item; settle the item for sale based on the price corresponding to the target item and the item information.

[0011] Preferably, the steps of obtaining a video stream of a user purchasing items from a vending cabinet, and intercepting multiple images of items for sale from the video stream to obtain an item image set, where the item image set includes images of the item from at least two different angles, are as follows:

[0012] Set image acquisition points corresponding to the vending cabinet, and based on the image acquisition points, shoot a full-process video stream of the user purchasing items from the vending cabinet to obtain a continuous image sequence containing the motion information of the items for sale;

[0013] Analyze the motion trajectory of the item for sale from the video stream, and determine the initial end position of the motion trajectory; acquire an image of the item for sale at the initial end position, denoted as the first item image;

[0014] Calibrate the feature points of the item for sale in the first item image to generate an initial feature point set;

[0015] Successively acquire the contour lines of the item for sale at multiple positions along the motion trajectory to form a dynamic contour line sequence;

[0016] Perform feature point matching and comparison between the currently acquired contour line and the contour line at the previous position, extract video frames that meet the preset conditions from the video stream, and intercept the corresponding images of the item for sale, denoted as the second item images;

[0017] Repeat the contour line acquisition and comparison steps until the image interception of the entire path of the motion trajectory is completed, and obtain second item images from multiple perspectives;

[0018] Associate and store the first item image and multiple second item images to form an item image set containing images of the item from at least two different angles.

[0019] Preferably, the steps of performing feature point matching and comparison between the currently acquired contour line and the contour line at the previous position, extracting video frames that meet the preset conditions from the video stream, and intercepting the corresponding images of the item for sale, denoted as the second item images, are as follows:

[0020] Extract the contour line of the item for sale from the first item image to obtain an initial contour line, denoted as the first contour line;

[0021] Set multiple refresh points based on the first contour line;

[0022] Successively acquire images at multiple positions along the motion trajectory of the item for sale to obtain a sequence of candidate item images;

[0023] Extract the contour lines of the item for sale corresponding to each candidate item image to form a sequence of contour lines to be compared; perform image interception based on the sequence of contour lines to be compared to obtain the second item images.

[0024] Preferably, the step of performing image capture based on the sequence of contour lines to be compared to obtain a second article image includes:

[0025] Obtain the current contour line to be compared, and compare the feature points of the current contour line to be compared with the contour line captured last time;

[0026] When a new feature point is detected and the disappearance of the new feature point does not occur, continue to collect the image at the next position;

[0027] When a new feature point is detected and the disappearance of the new feature point also exists, trigger the image capture condition;

[0028] Count the number of new feature points that appear in each candidate article image when the capture condition is met, and select the candidate article image with the highest frequency of new feature points appearing for capture to obtain a second article image.

[0029] Preferably, the step of constructing a local article model corresponding to the article for sale based on local feature information includes:

[0030] Extract the local feature points and the relative position relationship between the local feature points in each article image from the article image set;

[0031] Obtain the descriptions of the feature points in two continuously captured images, perform feature point matching between adjacent images, and identify and extract the overlapping feature points;

[0032] Predict the relative pose between adjacent images based on the overlapping feature points, recover the three-dimensional feature points based on the overlapping feature points and the relative pose, and align the three-dimensional feature points to a unified coordinate system to generate a local article three-dimensional model.

[0033] Preferably, the step of collecting the reference image information of multiple reference articles in the vending cabinet and constructing a reference image model corresponding to the reference articles based on the reference image information includes:

[0034] Obtain the reference image information of each reference article in the vending cabinet at different placement angles;

[0035] Input the reference image information of each reference article at different placement angles into the initial image model in sequence to gradually perform iterative training on the initial image model to obtain a reference image model corresponding to each reference article.

[0036] Preferably, the step of comparing the local article model with multiple reference image models, extracting the reference image models that meet the preset conditions, and using the reference articles corresponding to the reference image models as the target articles corresponding to the articles for sale includes:

[0037] Identify and mark the article feature points in the local article model, and identify and mark the reference feature points in the reference image model;

[0038] Coincide the item feature points with the reference feature points, and use the reference image model where the reference feature points corresponding to the coincidence degree meeting the preset conditions are located as the reference image model corresponding to the item feature points.

[0039] Preferably, the step of coinciding the item feature points with the reference feature points and using the reference image model where the reference feature points corresponding to the coincidence degree meeting the preset conditions are located as the reference image model corresponding to the item feature points includes:

[0040] Obtain the item feature points and the reference feature points, where the item feature points include the first item feature points and the second item feature points, and the reference feature points include the first reference feature points corresponding to the first item feature points and the second reference feature points corresponding to the second item feature points;

[0041] Coincide the first item feature points with the first reference feature points in multiple reference image models respectively, and calculate the first coincidence degree corresponding to each reference image model;

[0042] Extract the reference image models corresponding to the first coincidence degrees that meet the preset first coincidence degree threshold to obtain a set of preselected reference image models;

[0043] Judge whether the number of preselected reference image models in the set of preselected reference image models is 1;

[0044] If it is 1, use the preselected reference image model as the reference image model corresponding to the item feature points;

[0045] If it is not 1, compare the second item feature points with the second reference feature points of multiple preselected reference image models in the set of preselected reference image models respectively, calculate the second coincidence degree corresponding to each preselected reference image model, and use the preselected reference image model corresponding to the second coincidence degree that meets the preset second coincidence degree threshold as the reference image model corresponding to the item feature points.

[0046] Preferably, the step of obtaining the item information of the target item in the video stream and the price corresponding to the target item; and settling the items for sale based on the price corresponding to the target item and the item information includes:

[0047] Obtain the type information of the target items appearing in the video stream and identify the quantity of each target item;

[0048] Count the specific quantity of each type of target item, and use the quantity of the target item and the total number of target item types as the item information of the target item;

[0049] For each identified target item, obtain its corresponding price information; based on the obtained price information, calculate the total price of the same type of target items;

[0050] Count the types of all target items, and synthesize the total price of each type of target item to complete the settlement of the items for sale.

[0051] Compared with the prior art, the beneficial effects of the present invention are:

[0052] 1. Through the dynamic contour line comparison mechanism, dynamically intercept video frames that meet preset conditions to form multi-view item images, accurately capture the changes in the characteristics of moving items, and ensure the accuracy of the image interception timing through the dual determination conditions of the appearance and disappearance of feature points, effectively improving the integrity and feature recognition rate of item image acquisition;

[0053] 2. By intercepting multiple item images from different angles in the video stream, constructing a local item model, comparing the local item model with multiple reference image models, extracting the reference image models whose similarity meets the preset conditions, and using the reference items corresponding to the reference image models as the target items corresponding to the items for sale, through multi-model comparison, the items in the vending cabinet can be accurately identified, and through the automated identification and settlement process, the settlement efficiency of the vending cabinet can be significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0055] Figure 1 It is a schematic flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0057] Embodiment: Please refer to Figure 1 , the present invention provides a technical solution for the settlement method of AI recognition of items in a vending cabinet: including the following steps:

[0058] S1: Obtain the video stream of the user purchasing items from the vending cabinet, and intercept multiple images of the items for sale from the video stream to obtain an item image set, where the item image set includes images of the items from at least two different angles.

[0059] The steps of obtaining the video stream of the user purchasing items from the vending cabinet, and intercepting multiple images of the items for sale from the video stream to obtain an item image set, where the item image set includes images of the items from at least two different angles are as follows:

[0060] Set image acquisition points corresponding to the vending cabinet, and based on the image acquisition points, shoot the entire process video stream of the user purchasing items from the vending cabinet to obtain a continuous image sequence containing the motion information of the items for sale; analyze the motion trajectory of the items for sale from the video stream, and determine the initial end position of the motion trajectory; acquire an image of the item for sale at the initial end position, denoted as the first item image; calibrate the feature points of the item for sale in the first item image to generate an initial feature point set; sequentially acquire the contour lines of the item for sale at multiple positions along the motion trajectory to form a dynamic contour line sequence; perform feature point matching and comparison between the currently acquired contour line and the contour line at the previous position, extract video frames that meet the preset conditions from the video stream, and intercept the corresponding images of the items for sale, denoted as the second item images; repeat the contour line acquisition and comparison steps until the image interception of the entire path of the motion trajectory is completed to obtain second item images from multiple perspectives; associate and store the first item image and the multiple second item images to form an item image set containing images of the items from at least two different angles.

[0061] Specifically, for the preset image acquisition points of the vending cabinet, the acquisition points are distributed at preset positions around the vending cabinet to cover different visual angles of the items for sale. The whole process video stream of the user purchasing items from the vending cabinet is captured through the acquisition points, and a continuous image sequence containing the motion information of the items for sale is extracted. The motion information of the items for sale at least includes tracking the motion trajectory of the items for sale, providing a high-quality data basis for subsequent applications such as item recognition and behavior analysis. Analyze the motion trajectory of the items for sale from the video stream, and determine the initial segment position of the motion trajectory. Collect the first item image at the initial end position for subsequent judgment of the contour lines at other positions. Sequentially collect the contour lines of the items for sale at multiple positions along the motion trajectory to form a dynamic contour line sequence. Perform feature point matching and comparison between the currently collected contour line and the contour line at the previous position, and execute the following determination: When it is detected that new feature points appear on the contour line and there are disappearing feature points, it is determined that the preset condition is met. Record the video frame position corresponding to when the preset condition is met. According to the determination result, extract the video frames that meet the preset condition from the video stream, and intercept the corresponding item images of the items for sale to obtain the second item images. Associatively store the first item image and multiple second item images to form an image set containing at least two item images at different angles. The image set meets the following conditions: It contains at least one front view image at the initial position; it contains at least one side view / top view image in the motion trajectory; there is a verifiable feature point correspondence relationship between the images. Through the motion trajectory tracking and dynamic contour line analysis technology, the automatic acquisition of multi-view images of the vending cabinet items is realized, and the accuracy of the image interception timing is ensured through the feature point change determination mechanism, significantly improving the integrity and recognition reliability of the item image acquisition.

[0062] The steps of performing feature point matching and comparison between the currently collected contour line and the contour line at the previous position, extracting the video frames that meet the preset condition from the video stream, and intercepting the corresponding item images of the items for sale, which are recorded as the second item images, include: Extract the contour line of the item for sale from the first item image to obtain the initial contour line, which is recorded as the first contour line; Set multiple refresh points based on the first contour line; Along the motion trajectory of the item for sale, sequentially collect the images at multiple positions to obtain a sequence of candidate item images; Extract the contour line of the item for sale corresponding to each candidate item image to form a sequence of contour lines to be compared; Perform image interception based on the sequence of contour lines to be compared to obtain the second item images.

[0063] The steps of intercepting an image based on the sequence of contour lines to be compared to obtain a second object image include: obtaining the current contour line to be compared, and comparing the feature points of the current contour line to be compared with the contour line intercepted last time; when new feature points are detected and the disappearance of the new feature points does not occur, continue to collect images at the next position; when new feature points are detected and the disappearance of the new feature points also exists, trigger the image interception condition; count the number of new feature points that appear in each candidate object image when the interception condition is met, and select the candidate object image with the highest frequency of new feature points appearing for interception to obtain the second object image.

[0064] It should be noted that the image interception condition is: new feature points and the disappearance of feature points exist simultaneously in the current contour line;

[0065] Specifically, multiple refresh points are evenly selected on the first contour line, and their initial coordinates are recorded as the feature point benchmark. The refresh points are used as feature reference points during the subsequent contour line comparison process. When judging the contour line with the contour line at the previous position, the current contour line is compared with the contour line intercepted last time. Here, it is compared with the contour line intercepted last time. During the comparison, as the position of the item changes, the contour line of the item for sale will also change accordingly. When new feature points appear on the contour line, the refresh points on the contour line are refreshed. The new feature points here are the newly emerged feature points obtained by comparing with the image at the previous position, rather than the image intercepted last time. Continuously intercept the images of the item for sale at the positions corresponding to the newly emerged feature points as candidate item images, and continue to detect whether new feature points disappear when new feature points appear. For example, when the picture at the first position is the first item image, during the comparison between the picture at the first position and the picture at the second position, new feature points appear on the picture at the second position, but at the same time, it is possible that the feature points on the picture at the first position disappear. Since the feature points on the picture at the first position have been recorded by the first item image, their disappearance does not affect the subsequent use of the feature points. Therefore, there is no need to intercept the picture at the second position, and continue to compare the picture at the third position with the picture at the first position. If new feature points appear at this time, and the new feature points cover the new feature points at the second position, and at the same time, there is also a situation where new feature points disappear, then it is necessary to intercept the image corresponding to the contour line at this position, so as to ensure that the intercepted image can contain the most new feature points, so that each intercepted picture has common feature points while having more different feature points, which is used for subsequent 3D modeling of the item for sale, thus facilitating the subsequent identification of the item for sale and improving the accuracy of identification. Through the dynamic contour line comparison mechanism, the accurate capture of the feature changes of moving items is realized. Through the dual judgment conditions of the appearance and disappearance of feature points, the accuracy of the image interception timing is ensured, and the integrity and feature recognition rate of the item image acquisition are effectively improved;

[0066] For example, initially capture the upright image (the first item image) of a beverage on the shelf, extract the contour line, and set 4 refresh points. Capture 5 frames of images along the movement trajectory. It is found that new feature points appear in the 3rd frame (such as the label text becomes visible due to rotation), but there is no vanishing point, so continue to capture. In the 4th frame, both new feature points appear (new features appear, for example, the bottle cap was not visible before, but due to the change in the angle of the item brought about by the movement trajectory, the bottle cap can be seen, and the feature points on the bottle cap are used as new feature points) and feature points disappear (some feature points are blocked, which belongs to the disappearance of feature points), triggering the capture condition. Count the frequency of new feature points in each frame. The frequency in the 4th frame is the highest, and it is captured as the second item image. Finally, store 3 images including the upright, side view, and top view angles for use by the recognition model. Through the dual determination of the addition / disappearance of feature points, accurately capture the timing of the perspective change of the item, and the refresh point mechanism reduces the influence of light changes on the extraction of the contour line.

[0067] S2: Extract the local feature information of the item for sale based on the item image set; construct a local item model corresponding to the item for sale based on the local feature information.

[0068] The steps of constructing a local item model corresponding to the item for sale based on the local feature information include: extracting the local feature points in each item image and the relative position relationship between the local feature points from the item image set; obtaining the descriptions of the feature points in two continuously captured images, performing feature point matching between adjacent images, and identifying and extracting the overlapping feature points; predicting the relative pose between adjacent images based on the overlapping feature points, restoring the 3D feature points based on the overlapping feature points and the relative pose, and aligning the 3D feature points to a unified coordinate system to generate a local item 3D model.

[0069] It should be noted that the overlapping feature points refer to the feature points that appear in both of the two continuously captured pictures. The specific content of predicting the relative pose between adjacent images based on the overlapping feature points, restoring the 3D feature points based on the overlapping feature points and the relative pose, and aligning the 3D feature points to a unified coordinate system to generate a local item 3D model: extracting feature points from adjacent images, finding the matching feature point pairs, and extracting the feature points and descriptors, which can be implemented using feature detectors such as SIFT and ORB; performing feature matching using FLANN or BFMatcher, applying RANSAC to remove the wrong matches and retain the inliers; estimating the relative pose (rotation matrix and translation vector) between adjacent images based on the matching feature points. The pose can be obtained by decomposing the essential matrix, and the essential matrix is calculated. Decomposing the essential matrix gives four possible pose solutions. Through triangulation verification, select the correct solution; project the feature points into the normalized camera coordinate system, and use the linear triangulation method to solve for the 3D points. The coordinates of the 3D points can be restored from the overlapping feature points of adjacent images, thereby locally constructing the local 3D model of the item corresponding to the feature points in the image;

[0070] Specifically, when performing image capture, it is necessary to ensure that there are overlapping feature points in each image, that is, there are overlapping local regions. Control the appearance and disappearance times of the feature points to ensure that adjacent regions overlap. Extract the description of the feature points for each item image, match the feature points between adjacent images, estimate the relative pose between adjacent images based on the matching points, use the matching points and the camera pose, recover the 3D points through triangulation, and convert the 3D points to a unified coordinate system. Use triangulation to generate a local 3D model of the item, and through local feature modeling technology, achieve efficient and accurate identification of the items in the vending cabinet.

[0071] S3: Collect the reference image information of multiple reference items in the vending cabinet, and construct a reference image model corresponding to the reference item based on the reference image information.

[0072] The steps of collecting the reference image information of multiple reference items in the vending cabinet and constructing a reference image model corresponding to the reference item based on the reference image information include: obtaining the reference image information of each reference item in the vending cabinet at different placement angles; sequentially inputting the reference image information of each reference item at different placement angles into the initial image model to gradually perform iterative training on the initial image model to obtain the reference image model corresponding to each reference item.

[0073] Specifically, obtain the reference image information of each reference item in the vending cabinet at different placement angles. This can be achieved by setting multiple cameras in the vending cabinet or taking pictures of the goods in the cabinet when the vending cabinet is opened or closed. Perform preprocessing on the collected reference image information, including operations such as denoising and enhancing contrast, to improve the image quality; initialize an image model, which can be constructed based on a deep learning framework (such as convolutional neural network CNN) as the initial image model; sequentially input the reference image information of each reference item at different placement angles into the initial image model, and optimize the model through iterative training; during the training process, use a suitable loss function (such as mean square error MSE or Euclidean distance) to evaluate the difference between the model prediction result and the true value, and adjust the model parameters accordingly; after multiple iterative trainings, obtain the reference image model corresponding to each reference item, which can accurately identify and predict the image features of the reference item at different placement angles; based on the obtained reference image model, combine 3D modeling technologies (such as stereo vision, structured light, etc.) to construct the standard 3D model of each reference item in the vending cabinet. Perform optimization processing on the constructed 3D model, including operations such as smoothing the surface and removing noise, to improve the model quality; by collecting the standard reference information of the reference items in the vending cabinet and constructing a high-precision 3D model and performing feature point annotation based on this, effectively improve the recognition accuracy of the items in the intelligent vending cabinet.

[0074] S4: Compare the local item model with multiple reference image models, extract the reference image models that meet the preset conditions, and use the reference items corresponding to the reference image models as the target items corresponding to the items to be sold.

[0075] The steps of comparing the local item model with multiple reference image models, extracting the reference image models that meet the preset conditions, and using the reference items corresponding to the reference image models as the target items corresponding to the items to be sold include: identifying and marking the item feature points in the local item model; identifying and marking the reference feature points in the reference image model; overlapping the item feature points with the reference feature points, and using the reference image model where the reference feature points corresponding to the overlapping degree that meets the preset conditions are located as the reference image model corresponding to the item feature points.

[0076] The steps of overlapping the item feature points with the reference feature points and using the reference image model where the reference feature points corresponding to the overlapping degree that meets the preset conditions are located as the reference image model corresponding to the item feature points include: obtaining the item feature points and the reference feature points, where the item feature points include the first item feature points and the second item feature points, and the reference feature points include the first reference feature points corresponding to the first item feature points and the second reference feature points corresponding to the second item feature points; overlapping the first item feature points with the first reference feature points in multiple reference image models respectively, and calculating the first overlapping degree corresponding to each reference image model; extracting the reference image models corresponding to the first overlapping degrees that meet the preset first overlapping degree threshold to obtain a set of preselected reference image models; judging whether the number of preselected reference image models in the set of preselected reference image models is 1; if it is 1, using the preselected reference image model as the reference image model corresponding to the item feature points; if it is not 1, comparing the second item feature points with the second reference feature points of multiple preselected reference image models in the set of preselected reference image models respectively, calculating the second overlapping degree corresponding to each preselected reference image model, and using the preselected reference image model corresponding to the second overlapping degree that meets the preset second overlapping degree threshold as the reference image model corresponding to the item feature points;

[0077] It should be noted that the preset first coincidence degree threshold and the preset second coincidence degree threshold are preset according to the actual application scenario, the distribution characteristics of the item feature points, and the accuracy of the reference image model, and the preset second coincidence degree threshold can be higher than or equal to the preset first coincidence degree threshold to ensure the accuracy of the finally determined reference image model; coincidence includes, but is not limited to, calculating the Euclidean distance based on the feature point coordinates, calculating the feature vector similarity, or predicting the matching degree based on the machine learning model. When calculating the coincidence degree, it is necessary to ensure that the position information of each item feature point and the relative position information between the feature points correspond to the position information of the reference feature points and the relative position information between the reference feature points, and calculate the coincidence degree after correspondence;

[0078] The specific content of calculating the coincidence degree corresponding to each reference image model: When calculating the coincidence degree, it is necessary to normalize the feature vector or distance to ensure the consistency of comparison between different feature points or different images. Suppose there are two item feature points and , and their corresponding reference feature points and . Calculate the Euclidean distance between them and normalize it to the coincidence degree: , , where represents the coordinates of the first item feature point, where is the abscissa, is the ordinate, represents the coordinates of the second item feature point, is the abscissa, is the ordinate, represents the coordinates of the first reference feature point corresponding to the first item feature point , where is the abscissa, is the ordinate, represents the coordinates of the second reference feature point corresponding to the second item feature point , where is the abscissa, is the ordinate, represents the Euclidean distance between the first item feature point and the corresponding first reference feature point, represents the first item feature point and its corresponding reference feature point The coincidence degree between them. The coincidence degree is obtained by normalizing the Euclidean distance; Denote the maximum possible distance. Here, the "maximum possible distance" is a preset constant or a distance value determined according to the application scenario, which is used to convert the distance into a coincidence degree (a value between 0 and 1). The calculation method of the second coincidence degree is the same as that of the first coincidence degree, except that and coordinates are used;

[0079] Specifically, the first item feature points and the first reference feature points are respectively marked for the corner points and edge points in the local item model and multiple reference image models; the second item feature points and the second reference feature points are respectively set on the texture lines in the corresponding local item model and multiple reference image models to determine the texture orientation; by preferentially selecting the position of the first item feature points from multiple reference image models, the calculation amount of the overall similarity or coincidence rate can be reduced. When the number of local item models screened by the first coincidence degree is greater than 1, the texture comparison is performed; the corner points and texture contour points are used as the significant feature points on the item, which are easy to extract and match. Through two rounds of coincidence degree calculation and threshold judgment, the accuracy of the finally determined reference image model is ensured, thereby improving the recognition efficiency of the item for sale.

[0080] S5: Obtain the item information of the target item in the video stream and the price corresponding to the target item; settle the item for sale based on the price corresponding to the target item and the item information.

[0081] The steps of obtaining the item information of the target item in the video stream and the price corresponding to the target item and settling the item for sale based on the price corresponding to the target item and the item information include: obtaining the type information of the target item appearing in the video stream and identifying the quantity of each target item; counting the specific quantity of each type of target item and taking the quantity of the target item and the total number of target item types as the item information of the target item; for each identified target item, obtaining its corresponding price information; calculating the total price of the same type of target items according to the obtained price information; counting the types of all target items and summarizing the total price of each type of target item to complete the settlement of the item for sale.

[0082] Specifically, in the video stream, the user will purchase multiple different types of target items, obtain the unit price corresponding to each target item, ensure the continuous and accurate counting of the quantity of each type of target item in the video stream through a tracking algorithm, multiply the quantity by the unit price to obtain the total price of one type of target item, and synthesize the total prices of multiple types of target items to obtain the total price of the items for sale purchased by the user. Generate a settlement list containing the quantity, unit price, total price, and total settlement amount of each type of target item. According to the settlement list, complete the settlement process of the items for sale; obtain the unit price corresponding to each target item, query the pre-set price database to obtain the price corresponding to each target item; and / or obtain the price information of the target item in real time through an external API interface; ensure the accuracy and real-time nature of the price information to support accurate settlement, and generate a settlement list containing the quantity, unit price, total price, and total settlement amount of each type of target item; provide multiple payment methods for the user to complete the payment according to the settlement list; update the inventory information, reduce the quantity of the sold items, and record the settlement information for subsequent query and auditing.

[0083] In the present invention, at least two images with different angles are intercepted from the video, all different features in the images are extracted, and feature point fusion is performed according to the same feature points in adjacent pictures, including the relative positions between feature points, to construct a local commodity three-dimensional model. The local commodity model is compared with the reference model to find the reference item corresponding to the reference image model that is lower than the preset condition as the target item, and then the corresponding item price is determined for settlement; through a dynamic contour line comparison mechanism, video frames that meet the preset conditions are dynamically intercepted to form multi-perspective item images, accurately capturing the changes in the features of moving items. Through the dual determination conditions of the appearance and disappearance of feature points, the accuracy of the image interception timing is ensured, effectively improving the integrity and feature recognition rate of item image acquisition; by intercepting multiple item images with different angles from the video stream, constructing a local item model, comparing the local item model with multiple reference image models, extracting the reference image models whose similarity meets the preset conditions, and using the reference items corresponding to the reference image models as the target items corresponding to the items for sale. Through multi-model comparison, the items in the vending cabinet can be accurately identified, and through an automated identification and settlement process, the settlement efficiency of the vending cabinet can be significantly improved. Through AI technology, automatic identification and rapid settlement of commodities are realized, improving the shopping experience and operation efficiency.

[0084] In the description of this specification, the descriptions referring to the terms "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in a suitable manner in any one or more embodiments or examples.

[0085] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. Settlement method for AI recognition of items in a vending cabinet, characterized in that, Including the following steps: Obtain a video stream of a user purchasing items from a vending cabinet, and intercept multiple images of the items for sale from the video stream to obtain an item image set, where the item image set includes at least two images of the items from different angles; Set image acquisition points corresponding to the vending cabinet, and based on the image acquisition points, shoot a full-process video stream of the user purchasing items from the vending cabinet to obtain a continuous image sequence containing the motion information of the items for sale; Parse the motion trajectory of the items for sale from the video stream, and determine the initial end position of the motion trajectory; Acquire an image of the item for sale at the initial end position, denoted as the first item image; Calibrate the feature points of the item for sale in the first item image to generate an initial feature point set; Sequentially acquire the contour lines of the item for sale at multiple positions along the motion trajectory to form a dynamic contour line sequence; Perform feature point matching and comparison between the currently acquired contour line and the contour line at the previous position, and execute the following determination: Obtain the currently to-be-compared contour line, and perform feature point comparison between the currently to-be-compared contour line and the previously intercepted contour line; When new feature points are detected and the disappearance of the new feature points does not occur, continue to acquire the image at the next position; When new feature points are detected and the disappearance of the new feature points also exists, trigger the image interception condition; Count the number of new feature points that appear in each candidate item image when the interception condition is met, and select the candidate item image with the highest frequency of new feature points appearing for interception to obtain the second item image; Repeat the contour line acquisition and comparison steps until the image interception of the entire path of the motion trajectory is completed, and obtain second item images from multiple perspectives; Associate and store the first item image with the multiple second item images to form an image set containing at least two item images from different angles; Extract the local feature information of the item for sale based on the item image set; Construct a local item model corresponding to the item for sale based on the local feature information; Acquire the reference image information of multiple reference items in the vending cabinet, and construct a reference image model corresponding to the reference items based on the reference image information; Compare the local item model with multiple reference image models, extract the reference image models that meet the preset conditions, and use the reference items corresponding to the reference image models as the target items corresponding to the items for sale; Obtain the item information of the target item in the video stream and the price corresponding to the target item; Settle the items for sale based on the price corresponding to the target item and the item information.

2. The settlement method for AI recognition of vending cabinet items according to claim 1, wherein: Extract the contour line of the item for sale from the first item image to obtain an initial contour line, denoted as the first contour line; Set multiple refresh points based on the first contour line; Along the motion trajectory of the item for sale, sequentially acquire images at multiple positions to obtain a sequence of candidate item images; Extract the contour lines of the item for sale corresponding to each candidate item image to form a sequence of to-be-compared contour lines; Perform image interception based on the sequence of to-be-compared contour lines to obtain the second item image.

3. The settlement method for AI recognition of items in a vending cabinet according to claim 1, wherein: The step of constructing a local item model corresponding to the item for sale based on the local feature information includes: Extract local feature points in each item image and the relative position relationship between the local feature points from the item image set; Obtain the descriptions of the feature points in two continuously intercepted images, perform feature point matching between adjacent images, and identify and extract overlapping feature points; Predict the relative pose between adjacent images based on the overlapping feature points, restore the three-dimensional feature points based on the overlapping feature points and the relative pose, and align the three-dimensional feature points to a unified coordinate system to generate a local item three-dimensional model.

4. The settlement method for AI recognition of items in a vending cabinet according to claim 1, wherein: The step of collecting reference image information of multiple reference items in the vending cabinet and constructing a reference image model corresponding to the reference item based on the reference image information includes: Obtain the reference image information of each reference item in the vending cabinet at different placement angles; Input the reference image information of each reference item at different placement angles into the initial image model in sequence to gradually perform iterative training on the initial image model to obtain a reference image model corresponding to each reference item.

5. The settlement method for AI recognition of items in a vending cabinet according to claim 1, characterized in that: The step of comparing the local item model with multiple reference image models, extracting the reference image models that meet the preset conditions, and using the reference items corresponding to the reference image models as the target items corresponding to the items for sale includes: Identify and mark the item feature points in the local item model, and identify and mark the reference feature points in the reference image model; Coincide the item feature points with the reference feature points, and use the reference image model where the reference feature points corresponding to the coincidence degree that meets the preset conditions are located as the reference image model corresponding to the item feature points.

6. The settlement method for AI recognition of items in a vending cabinet according to claim 5, characterized in that: The step of coinciding the item feature points with the reference feature points, and using the reference image model where the reference feature points corresponding to the coincidence degree that meets the preset conditions are located as the reference image model corresponding to the item feature points includes: Obtain the item feature points and the reference feature points, where the item feature points include the first item feature point and the second item feature point, and the reference feature points include the first reference feature point corresponding to the first item feature point and the second reference feature point corresponding to the second item feature point; Coincide the first item feature point with the first reference feature points in multiple reference image models respectively, and calculate the first coincidence degree corresponding to each reference image model; Extract the reference image models corresponding to the first coincidence degrees that meet the preset first coincidence degree threshold to obtain a set of preselected reference image models; Judge whether the number of preselected reference image models in the set of preselected reference image models is 1; If it is 1, use the preselected reference image model as the reference image model corresponding to the item feature points; If it is not 1, compare the second item feature point with the second reference feature points of multiple preselected reference image models in the set of preselected reference image models respectively, calculate the second coincidence degree corresponding to each preselected reference image model, and use the preselected reference image model corresponding to the second coincidence degree that meets the preset second coincidence degree threshold as the reference image model corresponding to the item feature points.

7. The settlement method for AI recognition of items in a vending cabinet according to claim 1, wherein: Obtain the item information of the target item in the video stream and the price corresponding to the target item; The step of settling the items for sale based on the price corresponding to the target item and the item information includes: Obtain the type information of the target items appearing in the video stream and identify the quantity of each type of target item; Count the specific quantity of each type of target item, and use the quantity of the target items and the total number of target item types as the item information of the target items; For each identified target item, obtain its corresponding price information; calculate the total price of the same type of target items based on the obtained price information; Count the types of all target items, and synthesize the total price of each type of target item to complete the settlement of the items for sale.

Citation Information

Patent Citations

  • Intelligent identification method for commodity image of vending machine

    CN119360132A

  • Video processing method and apparatus, device, and storage medium

    US20230362416A1