Settlement method for AI identification of goods in vending cabinet
By obtaining multi-view item images in smart sales cabinets, building local item models and comparing them with reference image models, the problem of inaccurate identification caused by obstruction of items is solved, and efficient and accurate item identification and settlement are achieved.
Patent Information
- Application Number
- CN202510487135.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-18
AI Technical Summary
In smart sales cabinets, items are easily identified inaccurately due to items being blocked, which affects settlement efficiency.
By obtaining multiple item images from different angles in the video stream, local feature information is extracted, local item models are constructed, and compared with multiple reference image models, a reference image model that meets the preset conditions is extracted to accurately identify items for sale.
The accuracy of item recognition and settlement efficiency are improved. Through dynamic contour comparison mechanism and multi-model comparison, the accuracy of image interception timing and the integrity of item image acquisition are ensured.
Smart Images

Figure CN120014617A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a settlement method for AI recognition of items in a vending machine. Background Art
[0002] In the field of smart vending machines, accurately identifying product information and building its three-dimensional model is of great significance for improving user experience and optimizing inventory management. The AI recognition of items in vending machines mainly relies on the camera to collect continuous images of user purchases to form a video stream, and identify the items for sale from the video stream for settlement. However, in the prior art, when identifying items in the vending machine, when the user takes the item, it is easy for the item to be blocked, so only a partial image of the item for sale can be recognized, which affects the accuracy of item recognition and further affects the settlement efficiency of the items for sale.
[0003] To this end, we propose a settlement method based on AI recognition of container items to solve the above problems. Summary of the invention
[0004] The purpose of the present invention is to provide a settlement method for AI identification of vending machine items to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above object, the present invention provides the following technical solution: a settlement method for vending machine item AI recognition, the method comprising the following steps: Acquire a video stream of a user purchasing items from a vending machine, and capture multiple images of the items for sale from the video stream to obtain an item image set, wherein the item image set includes at least two images of the items taken from different angles; Extracting local feature information of the items for sale based on the item image set; constructing a local item model corresponding to the items for sale based on the local feature information; Collecting reference image information of multiple reference objects in the vending machine, and constructing reference image models corresponding to the reference objects based on the reference image information; Comparing the local object model with multiple reference image models, extracting the reference image model that meets the preset conditions, and using the reference object corresponding to the reference image model as the target object corresponding to the object for sale; Obtaining the item information of the target item in the video stream and the price corresponding to the target item; and performing settlement on the item for sale based on the price corresponding to the target item and the item information.
[0006] Preferably, the step of obtaining a video stream of a user purchasing items from a vending machine, and capturing multiple images of the items for sale from the video stream to obtain an item image set, wherein the item image set includes at least two images of the items at different angles, comprises: An image acquisition point is set corresponding to the vending cabinet, and a video stream of the entire process of a user purchasing items from the vending cabinet is captured based on the image acquisition point to obtain a continuous image sequence containing movement information of the items for sale; Analyze the motion trajectory of the item for sale from the video stream to determine the initial end position of the motion trajectory; capture an image of the item for sale at the initial end position, which is recorded as a first item image; Calibrate feature points of the item for sale in the first item image to generate an initial feature point set; sequentially collecting contour lines of the item for sale at multiple positions along the motion trajectory to form a dynamic contour line sequence; Perform feature point matching and comparison between the currently collected contour line and the contour line at the previous position, extract the video frame that meets the preset conditions from the video stream, and intercept the corresponding image of the item for sale, which is recorded as the second item image; Repeat the steps of contour line acquisition and comparison until the image capture of the entire motion trajectory is completed to obtain the second object image from multiple perspectives; The first object image is associated with the plurality of second object images and stored to form an object image set including at least two object images at different angles.
[0007] Preferably, the step of matching and comparing the currently collected contour line with the contour line at the previous position, extracting a video frame that meets a preset condition from the video stream, and capturing a corresponding image of the item for sale, which is recorded as the second item image, comprises: Extracting the contour of the item for sale from the first item image to obtain an initial contour, recorded as a first contour; Setting a plurality of refresh points based on the first contour line; Along the moving trajectory of the object to be sold, images of multiple positions are collected in sequence to obtain an image sequence of the object to be selected; The contour lines of the items for sale corresponding to each of the items to be selected are extracted to form a sequence of contour lines to be compared; and image capture is performed based on the sequence of contour lines to be compared to obtain a second item image.
[0008] Preferably, the step of performing image interception based on the sequence of contour lines to be compared to obtain the second object image comprises: Obtain the current contour line to be compared, and compare the feature points of the current contour line to be compared with the contour line intercepted last time; When a new feature point is detected and no new feature point disappears, continue to collect the next position image; When a new feature point is detected and at the same time a new feature point disappears, the image capture condition is triggered; The number of new feature points appearing in each candidate object image when the interception condition is met is counted, and the candidate object image with the highest frequency of new feature points is selected for interception to obtain a second object image.
[0009] Preferably, the step of constructing a local item model corresponding to the item for sale based on the local feature information includes: Extracting local feature points in each object image and the relative position relationship between the local feature points from the object image set; Obtain descriptions of feature points in two consecutively captured images, perform feature point matching between adjacent images, and identify and extract overlapping feature points; The relative pose between adjacent images is predicted based on the overlapping feature points, and the three-dimensional feature points are restored based on the overlapping feature points and the relative pose. The three-dimensional feature points are aligned to a unified coordinate system to generate a three-dimensional model of the local object.
[0010] Preferably, the step of collecting reference image information of a plurality of reference objects in the vending cabinet and constructing a reference image model corresponding to the reference objects based on the reference image information comprises: Obtain reference image information of each reference object in the vending cabinet at different placement angles; The reference image information of each reference object at different placement angles is input into the initial image model in turn, so as to gradually iteratively train the initial image model and obtain the reference image model corresponding to each reference object.
[0011] Preferably, the step of comparing the local object model with a plurality of reference image models, extracting a reference image model that meets a preset condition, and using the reference object corresponding to the reference image model as the target object corresponding to the object for sale comprises: Identify and mark the object feature points in the local object model, and identify and mark the reference feature points in the reference image model; The feature points of the object are overlapped with the reference feature points, and the reference image model where the reference feature points corresponding to the overlap degree that meets the preset conditions are located is used as the reference image model corresponding to the feature points of the object.
[0012] Preferably, the step of coinciding the feature point of the object with the reference feature point and using the reference image model where the reference feature point corresponding to the degree of coincidence that satisfies a preset condition is located as the reference image model corresponding to the feature point of the object comprises: Acquire item feature points and reference feature points, wherein the item feature points include a first item feature point and a second item feature point, and the reference feature points include a first reference feature point corresponding to the first item feature point and a second reference feature point corresponding to the second item feature point; Overlapping the first object feature point with the first reference feature points in the plurality of reference image models respectively, and calculating a first degree of overlap corresponding to each reference image model; Extracting reference image models corresponding to a first degree of overlap that satisfies a preset first degree of overlap threshold, and obtaining a set of preselected reference image models; Determining whether the number of pre-selected reference image models in the pre-selected reference image model set is 1; If it is 1, the pre-selected reference image model is used as the reference image model corresponding to the feature point of the object; If it is not 1, the second object feature point is compared with the second reference feature points of multiple pre-selected reference image models in the pre-selected reference image model set, the second degree of overlap corresponding to each pre-selected reference image model is calculated, and the pre-selected reference image model corresponding to the second degree of overlap that meets the preset second degree of overlap threshold is used as the reference image model corresponding to the object feature point.
[0013] Preferably, the step of obtaining the item information of the target item in the video stream and the price corresponding to the target item; and settling the item for sale based on the price corresponding to the target item and the item information comprises: Obtain information about the types of target objects appearing in the video stream and identify the quantity of each target object; Counting the specific number of each type of target items, and using the number of the target items and the total number of target item types as item information of the target items; For each identified target item, obtain its corresponding price information; based on the obtained price information, calculate the total price of similar target items; Count all the types of target items and combine the total price of each type of target items to complete the settlement of the items for sale.
[0014] Compared with the prior art, the present invention has the following beneficial effects: 1. Through the dynamic contour comparison mechanism, the video frames that meet the preset conditions are dynamically intercepted to form multi-view object images, accurately capture the feature changes of moving objects, and ensure the accuracy of image interception timing through the dual judgment conditions of the appearance and disappearance of feature points, effectively improving the integrity of object image acquisition and feature recognition; 2. By capturing multiple object images from different angles from the video stream, a local object model is constructed, and the local object model is compared with multiple reference image models to extract the reference image model whose similarity meets the preset conditions. The reference object corresponding to the reference image model is used as the target object corresponding to the object for sale. Through multi-model comparison, the objects in the vending machine can be accurately identified. Through automated identification and settlement processes, the settlement efficiency of the vending machine can be significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0016] Figure 1 It is a schematic diagram of the method flow of the present invention. DETAILED DESCRIPTION
[0017] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0018] Example: See Figure 1 The present invention provides a technical solution for a settlement method of AI identification of vending machine items, comprising the following steps: S1: Obtain a video stream of a user purchasing items from a vending machine, and capture multiple images of items for sale from the video stream to obtain an item image set, wherein the item image set includes at least two images of the items taken from different angles.
[0019] The steps of obtaining a video stream of a user purchasing items from a vending machine, and capturing multiple images of the items for sale from the video stream to obtain an item image set, wherein the item image set includes at least two images of the items at different angles, include: An image acquisition point is set corresponding to the vending machine, and a video stream of the entire process of a user purchasing items from the vending machine is captured based on the image acquisition point to obtain a continuous image sequence containing movement information of the items for sale; the movement trajectory of the items for sale is parsed from the video stream to determine the initial end position of the movement trajectory; an image of the items for sale is acquired at the initial end position and recorded as a first item image; feature points of the items for sale in the first item image are calibrated to generate an initial feature point set; contours of the items for sale at multiple positions are sequentially acquired along the movement trajectory to form a dynamic contour sequence; feature points of the currently acquired contours are matched and compared with the contours of the previous position, a video frame that meets a preset condition is extracted from the video stream, and a corresponding image of the items for sale is captured and recorded as a second item image; the contour acquisition and comparison steps are repeated until the image capture of the entire movement trajectory is completed to obtain second item images from multiple perspectives; the first item image is associated with the multiple second item images and stored to form an item image set containing at least two item images at different angles.
[0020] Specifically, image acquisition points are preset for the vending cabinet, and the acquisition points are distributed at preset positions around the vending cabinet to cover different visual angles of the items for sale. The video stream of the entire process of users purchasing items from the vending cabinet is captured through the acquisition points, and a continuous image sequence containing movement information of the items for sale is extracted. The movement information of the items for sale at least includes tracking the movement trajectory of the items for sale, so as to provide a high-quality data basis for subsequent applications such as item recognition and behavior analysis; the movement trajectory of the items for sale is parsed from the video stream, and the initial segment position of the movement trajectory is determined, and the first item image is acquired at the initial end position for subsequent judgment of the contour lines of other positions; the contour lines of the items for sale at multiple positions are sequentially acquired along the movement trajectory to form a dynamic contour line sequence; the contour line currently acquired is matched and compared with the contour line of the previous position for feature point matching, and the following judgment is performed: when a new contour line is detected, When feature points are added and some feature points disappear, it is determined that the preset conditions are met; the corresponding video frame position when the preset conditions are met is recorded; according to the determination result, the video frame that meets the preset conditions is extracted from the video stream, and the corresponding image of the item for sale is intercepted to obtain a second item image; the first item image is associated with multiple second item images and stored to form an image set containing at least two item images at different angles; the image set meets the following conditions: at least one front view angle image at the initial position; at least one side view / top view angle image in the motion trajectory; there is a verifiable feature point correspondence between the images; through the motion trajectory tracking and dynamic contour line analysis technology, the automatic collection of multi-view images of the items in the vending machine is realized, the accuracy of the image interception timing is ensured through the feature point change judgment mechanism, and the integrity of the item image collection and the recognition reliability are significantly improved.
[0021] The steps of matching and comparing the currently collected contour line with the contour line at the previous position with feature points, extracting video frames that meet preset conditions from the video stream, and intercepting the corresponding image of the item for sale, which is recorded as the second item image, include: extracting the contour line of the item for sale from the first item image to obtain an initial contour line, which is recorded as the first contour line; setting multiple refresh points based on the first contour line; collecting images at multiple positions in sequence along the movement trajectory of the item for sale to obtain a sequence of images of items to be selected; extracting the contour line of the item for sale corresponding to each image of the item to be selected to form a sequence of contour lines to be compared; and intercepting images based on the sequence of contour lines to be compared to obtain the second item image.
[0022] The steps of performing image interception based on the sequence of contour lines to be compared to obtain the second object image include: obtaining the current contour line to be compared, and performing feature point comparison between the current contour line to be compared and the contour line intercepted last time; when a new feature point is detected and has not disappeared, continuing to collect the next position image; when a new feature point is detected and has disappeared at the same time, triggering an image interception condition; counting the number of new feature points appearing in each image of the selected object when the interception condition is met, selecting the image of the selected object with the highest frequency of new feature points appearing for interception, and obtaining the second object image.
[0023] It should be noted that the image capture condition is: the current contour line has newly added feature points and feature points disappear at the same time; Specifically, multiple refresh points are evenly selected on the first contour line, and their initial coordinates are recorded as feature point benchmarks. The refresh points are used as feature reference points in the subsequent contour line comparison process; when judging the contour line with the contour line of the previous position, the current contour line is compared with the contour line intercepted last time, and here it is compared with the contour line intercepted last time. When comparing, as the position of the item changes, the contour line of the item for sale will also change accordingly. When a new feature point appears on the contour line, the refresh point on the contour line is refreshed. The new feature point here is a newly appeared feature point obtained by comparing with the image of the previous position, not the image intercepted last time; multiple images of the item for sale at positions corresponding to the newly appeared feature points are continuously intercepted as the item images to be selected, and whether the new feature points disappear when the new feature points appear is continuously detected. For example, when the picture at the first position is the first item image, in the process of comparing the picture at the first position with the picture at the second position, a new feature point appears on the picture at the second position, but at the same time, the first position may If the feature points on the picture disappear, since the feature points on the picture at the first position have been recorded by the first object image, their disappearance will not affect the use of subsequent feature points. Therefore, there is no need to intercept the picture at the second position and continue to compare the picture at the third position with the picture at the first position. If new feature points appear at this time, and the new feature points cover the new feature points at the second position, and the new feature points disappear, it is necessary to intercept the image corresponding to the contour line at this position, so as to ensure that the intercepted image can contain the most new feature points, so that each intercepted picture has common feature points and more different feature points, which are used for subsequent three-dimensional modeling of the item for sale, thereby facilitating the subsequent identification of the item for sale and improving the accuracy of identification. Through the dynamic contour line comparison mechanism, the accurate capture of the feature changes of the moving item is achieved. Through the dual judgment conditions of the appearance and disappearance of feature points, the accuracy of the image interception timing is ensured, and the integrity and feature recognition of the item image acquisition are effectively improved; For example, initially collect an upright image of a beverage on a shelf (the first item image), extract the contour line and set 4 refresh points. Collect 5 frames of images along the motion trajectory, and find that new feature points appear in the third frame (such as the label text is visible due to rotation), but there is no vanishing point, so continue to collect. In the fourth frame, new feature points appear at the same time (new features appear, such as the bottle cap was not visible before, but due to the change in the angle of the item caused by the change in the motion trajectory, the bottle cap can be seen, and the feature points on the bottle cap are used as new feature points) and feature point disappearance (some feature points are blocked and feature point disappearance), triggering the interception condition. Count the frequency of new feature points in each frame, and the fourth frame has the highest frequency, which is intercepted as the second item image. Finally, three images containing upright, side, and top view angles are stored for use by the recognition model. Through the dual judgment of feature point addition / disappearance, the timing of object perspective change is accurately captured, and the refresh point mechanism reduces the impact of lighting changes on contour line extraction.
[0024] S2: extracting local feature information of the items for sale based on the item image set; and constructing a local item model corresponding to the items for sale based on the local feature information.
[0025] The steps of constructing a local item model corresponding to the item for sale based on the local feature information include: extracting local feature points in each item image and the relative position relationship between the local feature points from the item image set; obtaining the description of the feature points in two consecutively captured images, matching the feature points between adjacent images, identifying and extracting overlapping feature points; predicting the relative posture between adjacent images based on the overlapping feature points, restoring the three-dimensional feature points based on the overlapping feature points and the relative posture, and aligning the three-dimensional feature points to a unified coordinate system to generate a local item three-dimensional model.
[0026] It should be noted that overlapping feature points refer to feature points that appear in two consecutively captured images. The relative pose between adjacent images is predicted based on the overlapping feature points. The three-dimensional feature points are restored based on the overlapping feature points and the relative pose. The three-dimensional feature points are aligned to a unified coordinate system to generate the specific content of the local object three-dimensional model: extract feature points from adjacent images, find matching feature point pairs, and extract feature points and descriptors. This can be achieved using feature detectors such as SIFT and ORB; use FLANN or BFMatcher for feature matching, apply RANSAC to remove mismatches, and retain internal points; estimate the relative pose (rotation matrix and translation vector) between adjacent images based on the matched feature points. The pose can be obtained by using the essential matrix decomposition and the essential matrix can be calculated. Decompose the essential matrix to obtain four possible pose solutions. Select the correct solution through triangulation verification; project the feature points to the normalized camera coordinate system, use the linear triangulation method to solve the three-dimensional points, and restore the coordinates of the three-dimensional points from the overlapping feature points of adjacent images, thereby locally constructing a local three-dimensional model of the object corresponding to the feature points in the image; Specifically, when capturing images, it is necessary to ensure that there are overlapping feature points in each image, that is, there are overlapping local areas, control the time of appearance and disappearance of feature points, thereby ensuring that adjacent areas overlap, extract feature point descriptions for each object image, match feature points between adjacent images, estimate the relative pose between adjacent images based on matching points, use matching points and camera pose to restore three-dimensional points through triangulation, convert the three-dimensional points to unified coordinates, use triangulation to generate a three-dimensional model of the local object, and achieve efficient and accurate recognition of vending machine items through local feature modeling technology.
[0027] S3: Collect reference image information of multiple reference objects in the vending machine, and construct reference image models corresponding to the reference objects based on the reference image information.
[0028] The steps of collecting reference image information of multiple reference items in a vending machine and constructing reference image models corresponding to the reference items based on the reference image information include: obtaining reference image information of each reference item in the vending machine at different placement angles; inputting the reference image information of each reference item at different placement angles into an initial image model in turn, so as to gradually iteratively train the initial image model and obtain reference image models corresponding to each reference item.
[0029] Specifically, obtain reference image information of each reference object in the vending cabinet at different placement angles. This can be achieved by setting up multiple cameras in the vending cabinet, or taking pictures of the goods in the cabinet when the vending cabinet is opened or closed. Preprocess the collected reference image information, including denoising, contrast enhancement and other operations to improve image quality; initialize an image model, which can be built based on a deep learning framework (such as a convolutional neural network CNN) as an initial image model; input the reference image information of each reference object at different placement angles into the initial image model in turn, and optimize the model through iterative training; during the training process, use a suitable loss function (such as mean square error MSE or Euclidean distance) to evaluate the difference between the model prediction result and the true value, and adjust the model parameters accordingly; after multiple iterative training, obtain the reference image model corresponding to each reference object, which can accurately identify and predict the image features of the reference object at different placement angles; based on the obtained reference image model, combined with three-dimensional modeling technology (such as stereoscopic vision, structured light, etc.), construct a standard three-dimensional model of each reference object in the vending cabinet. The constructed three-dimensional model is optimized, including surface smoothing, noise removal and other operations to improve the model quality; by collecting standard reference information of reference items in the vending machine, and based on this, a high-precision three-dimensional model is constructed and feature points are marked, which effectively improves the recognition accuracy of items in the smart vending machine.
[0030] S4: Compare the local object model with multiple reference image models, extract the reference image model that meets the preset conditions, and use the reference object corresponding to the reference image model as the target object corresponding to the object for sale.
[0031] The steps of comparing a local object model with a plurality of reference image models, extracting a reference image model that meets preset conditions, and using the reference object corresponding to the reference image model as the target object corresponding to the object for sale include: identifying and marking object feature points in the local object model; identifying and marking reference feature points in the reference image model; overlapping the object feature points with the reference feature points, and using the reference image model where the reference feature points corresponding to the overlap degree that meets preset conditions are located as the reference image model corresponding to the object feature points.
[0032] The steps of overlapping the feature points of the item with the reference feature points and taking the reference image model where the reference feature points corresponding to the overlap degree that satisfies the preset conditions are located as the reference image model corresponding to the feature points of the item include: obtaining the feature points of the item and the reference feature points, wherein the feature points of the item include the first feature points of the item and the second feature points of the item, and the reference feature points include the first reference feature points corresponding to the first feature points of the item and the second reference feature points corresponding to the second feature points of the item; overlapping the first feature points of the item with the first reference feature points in the plurality of reference image models respectively, and calculating the first overlap degree corresponding to each reference image model; extracting the reference image model that satisfies the preset first overlap degree; and obtaining the reference image model corresponding to the reference image model. a reference image model corresponding to a first overlap degree threshold, and obtain a preselected reference image model set; determine whether the number of preselected reference image models in the preselected reference image model set is 1; if it is 1, use the preselected reference image model as the reference image model corresponding to the feature point of the item; if it is not 1, compare the second item feature point with the second reference feature points of multiple preselected reference image models in the preselected reference image model set, calculate the second overlap degree corresponding to each preselected reference image model, and use the preselected reference image model corresponding to the second overlap degree that meets the preset second overlap degree threshold as the reference image model corresponding to the item feature point; It should be noted that the preset first coincidence threshold and the preset second coincidence threshold are pre-set according to the actual application scenario, the distribution characteristics of the feature points of the item and the accuracy of the reference image model, and the preset second coincidence threshold may be higher than or equal to the preset first coincidence threshold to ensure the accuracy of the reference image model finally determined; coincidence includes but is not limited to Euclidean distance calculation based on feature point coordinates, feature vector similarity calculation or matching degree prediction based on a machine learning model. When calculating the coincidence, it is necessary to ensure that the position information of each item feature point and the relative position information between the feature points correspond to the position information of the reference feature points and the relative position information between the reference feature points, and calculate the coincidence after correspondence; Calculate the specific content of the coincidence corresponding to each reference image model: When calculating the coincidence, the feature vector or distance needs to be normalized to ensure the consistency between different feature points or different images. Suppose there are two object feature points and , and their corresponding reference feature points and . Compute the Euclidean distance between them and normalize it to coincidence: , ,in, Represents the coordinates of the first item feature point, where is the horizontal axis, is the vertical coordinate, Represents the coordinates of the second item feature point, is the horizontal axis, is the vertical coordinate, Indicates the first item feature point The coordinates of the first reference feature point corresponding to is the horizontal axis, is the vertical coordinate, Indicates the feature point of the second item The coordinates of the corresponding second reference feature point, where is the horizontal axis, is the vertical coordinate, represents the Euclidean distance between the first item feature point and the corresponding first reference feature point, Indicates the first item feature point Corresponding reference feature points The overlap between them is obtained by normalizing the Euclidean distance; Indicates the maximum possible distance. Here, the "maximum possible distance" is a preset constant or a distance value determined according to the application scenario. It is used to convert the distance into a coincidence degree (a value between 0 and 1). The calculation method of the second coincidence degree is the same as the calculation method of the first coincidence degree, except that and The coordinates of Specifically, the first item feature point and the first reference feature point are used to mark the corner points and edge points in the local item model and multiple reference image models respectively; the second item feature point and the second reference feature point are respectively set on the texture lines in the corresponding local item model and multiple reference image models to determine the texture direction; by preferentially selecting the position of the first item feature point from multiple reference image models, the calculation amount of the overall similarity or overlap rate can be reduced, and when the number of local item models screened out by the first overlap is greater than one, the texture comparison is performed; the corner points and texture contour points are used as significant feature points on the item, which are easy to extract and match. Through two rounds of overlap calculation and threshold judgment, the accuracy of the reference image model finally determined is ensured, thereby improving the recognition efficiency of the items for sale.
[0033] S5: Obtaining the item information of the target item in the video stream and the price corresponding to the target item; and performing settlement for the item for sale based on the price corresponding to the target item and the item information.
[0034] The steps of obtaining the item information of the target item in the video stream and the price corresponding to the target item; and settling the items for sale based on the price and item information corresponding to the target item include: obtaining the type information of the target items appearing in the video stream, and identifying the quantity of each type of target item; counting the specific quantity of each type of target item, and taking the quantity of the target items and the total number of target item types as the item information of the target item; obtaining the corresponding price information for each identified target item; calculating the total price of the same type of target items based on the obtained price information; counting the types of all target items, and combining the total price of each type of target item to complete the settlement of the items for sale.
[0035] Specifically, in the video stream, the user acquires and purchases multiple different types of target items, obtains the unit price corresponding to each target item, ensures that the quantity of each type of target items in the video stream is continuously and accurately counted through a tracking algorithm, multiplies the quantity by the unit price to obtain the total price of one type of target items, and combines the total prices of multiple types of target items to obtain the total price of the items for sale purchased by the user, generates a settlement list containing the quantity, unit price, total price and total settlement amount of each type of target items, and completes the settlement process of the items for sale according to the settlement list; obtains the unit price corresponding to each target item, queries a pre-set price database to obtain the price corresponding to each target item; and / or obtains the price information of the target items in real time through an external API interface; ensures the accuracy and real-time nature of the price information to support accurate settlement, and generates a settlement list containing the quantity, unit price, total price and total settlement amount of each type of target items; provides multiple payment methods so that users can complete payment according to the settlement list; updates inventory information, reduces the number of items sold, and records settlement information for subsequent inquiries and audits.
[0036] The present invention captures at least two images of different angles from a video, extracts all different features in the images, fuses feature points according to the same feature points in adjacent images, including the relative positions between feature points, constructs a local commodity three-dimensional model, compares the local commodity model with a reference model, finds a reference item corresponding to a reference image model below a preset condition as a target item, and then determines the corresponding item price for settlement; through a dynamic contour line comparison mechanism, dynamically captures video frames that meet the preset conditions to form multi-view item images, accurately captures the feature changes of moving items, and uses the dual judgment conditions of the appearance and disappearance of feature points to determine the target item price. It ensures the accuracy of the image capture timing and effectively improves the integrity and feature recognition of object image acquisition; by capturing multiple object images from different angles from the video stream, building a local object model, comparing the local object model with multiple reference image models, extracting the reference image model whose similarity meets the preset conditions, and using the reference object corresponding to the reference image model as the target object corresponding to the object for sale. Through multi-model comparison, it can accurately identify the objects in the vending machine, and through automated identification and settlement processes, it can significantly improve the settlement efficiency of the vending machine. Through AI technology, it can realize automatic identification and rapid settlement of goods, thereby improving shopping experience and operational efficiency.
[0037] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0038] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A settlement method for vending machine items AI recognition, characterized in that: The following steps are involved: Acquire a video stream of a user purchasing items from a vending machine, and capture multiple images of the items for sale from the video stream to obtain an item image set, wherein the item image set includes at least two images of the items taken from different angles; Extracting local feature information of the items for sale based on the item image set; constructing a local item model corresponding to the items for sale based on the local feature information; Collecting reference image information of multiple reference objects in the vending machine, and constructing reference image models corresponding to the reference objects based on the reference image information; Comparing the local object model with multiple reference image models, extracting the reference image model that meets the preset conditions, and using the reference object corresponding to the reference image model as the target object corresponding to the object for sale; Obtaining the item information of the target item in the video stream and the price corresponding to the target item; and performing settlement on the item for sale based on the price corresponding to the target item and the item information.
2. The settlement method for vending machine item AI recognition according to claim 1, characterized in that: The step of obtaining a video stream of a user purchasing items from a vending machine and capturing multiple images of the items for sale from the video stream to obtain an item image set, wherein the item image set includes at least two images of the items taken at different angles, comprises: An image acquisition point is set corresponding to the vending cabinet, and a video stream of the entire process of a user purchasing items from the vending cabinet is captured based on the image acquisition point to obtain a continuous image sequence containing movement information of the items for sale; Analyze the motion trajectory of the item for sale from the video stream to determine the initial end position of the motion trajectory; capture an image of the item for sale at the initial end position, which is recorded as a first item image; Calibrate feature points of the item for sale in the first item image to generate an initial feature point set; sequentially collecting contour lines of the item for sale at multiple positions along the motion trajectory to form a dynamic contour line sequence; Perform feature point matching and comparison between the currently collected contour line and the contour line at the previous position, extract the video frame that meets the preset conditions from the video stream, and intercept the corresponding image of the item for sale, which is recorded as the second item image; Repeat the steps of contour line acquisition and comparison until the image capture of the entire motion trajectory is completed to obtain the second object image from multiple perspectives; The first object image is associated with the plurality of second object images and stored to form an object image set including at least two object images at different angles.
3. The settlement method for vending machine item AI recognition according to claim 2, characterized in that: The step of matching and comparing the currently collected contour line with the contour line at the previous position with feature points, extracting a video frame that meets the preset conditions from the video stream, and intercepting the corresponding image of the item for sale, which is recorded as the second item image, includes: Extracting the contour of the item for sale from the first item image to obtain an initial contour, recorded as a first contour; Setting a plurality of refresh points based on the first contour line; Along the moving trajectory of the object to be sold, images of multiple positions are collected in sequence to obtain an image sequence of the object to be selected; The contour lines of the items for sale corresponding to each of the items to be selected are extracted to form a sequence of contour lines to be compared; and image capture is performed based on the sequence of contour lines to be compared to obtain a second item image.
4. The settlement method for vending machine item AI recognition according to claim 3, characterized in that: The step of performing image interception based on the sequence of contour lines to be compared to obtain a second object image comprises: Obtain the current contour line to be compared, and compare the feature points of the current contour line to be compared with the contour line intercepted last time; When a new feature point is detected and no new feature point disappears, continue to collect the next position image; When a new feature point is detected and at the same time a new feature point disappears, the image capture condition is triggered; The number of new feature points appearing in each candidate object image when the interception condition is met is counted, and the candidate object image with the highest frequency of new feature points is selected for interception to obtain a second object image.
5. The settlement method for vending machine item AI recognition according to claim 1, characterized in that: The step of constructing a local item model corresponding to the item for sale based on the local feature information comprises: Extracting local feature points in each object image and the relative position relationship between the local feature points from the object image set; Obtain descriptions of feature points in two consecutively captured images, perform feature point matching between adjacent images, and identify and extract overlapping feature points; The relative pose between adjacent images is predicted based on the overlapping feature points, and the three-dimensional feature points are restored based on the overlapping feature points and the relative pose. The three-dimensional feature points are aligned to a unified coordinate system to generate a three-dimensional model of the local object.
6. The settlement method for vending machine item AI recognition according to claim 1, characterized in that: The step of collecting reference image information of a plurality of reference objects in the vending cabinet and constructing a reference image model corresponding to the reference objects based on the reference image information comprises: Obtain reference image information of each reference object in the vending cabinet at different placement angles; The reference image information of each reference object at different placement angles is input into the initial image model in turn, so as to gradually iteratively train the initial image model and obtain the reference image model corresponding to each reference object.
7. The settlement method for vending machine item AI recognition according to claim 1, characterized in that: The step of comparing the local object model with a plurality of reference image models, extracting a reference image model that meets a preset condition, and using the reference object corresponding to the reference image model as the target object corresponding to the object for sale comprises: Identify and mark the object feature points in the local object model, and identify and mark the reference feature points in the reference image model; The feature points of the object are overlapped with the reference feature points, and the reference image model where the reference feature points corresponding to the overlap degree that meets the preset conditions are located is used as the reference image model corresponding to the feature points of the object.
8. The settlement method for vending machine item AI recognition according to claim 7, characterized in that: The step of coinciding the feature point of the object with the reference feature point and taking the reference image model where the reference feature point corresponding to the degree of coincidence that satisfies the preset condition is located as the reference image model corresponding to the feature point of the object comprises: Acquire item feature points and reference feature points, wherein the item feature points include a first item feature point and a second item feature point, and the reference feature points include a first reference feature point corresponding to the first item feature point and a second reference feature point corresponding to the second item feature point; Overlapping the first object feature point with the first reference feature points in the plurality of reference image models respectively, and calculating a first degree of overlap corresponding to each reference image model; Extracting reference image models corresponding to a first degree of overlap that meets a preset first degree of overlap threshold, and obtaining a set of preselected reference image models; Determining whether the number of pre-selected reference image models in the pre-selected reference image model set is 1; If it is 1, the pre-selected reference image model is used as the reference image model corresponding to the feature point of the object; If it is not 1, the second object feature point is compared with the second reference feature points of multiple pre-selected reference image models in the pre-selected reference image model set, the second degree of overlap corresponding to each pre-selected reference image model is calculated, and the pre-selected reference image model corresponding to the second degree of overlap that meets the preset second degree of overlap threshold is used as the reference image model corresponding to the object feature point.
9. The settlement method for vending machine item AI recognition according to claim 1, characterized in that: The step of obtaining the object information of the target object in the video stream and the price corresponding to the target object; The steps of settling the items for sale based on the price and item information corresponding to the target item include: Obtain information about the types of target objects appearing in the video stream and identify the quantity of each target object; Counting the specific number of each type of target items, and using the number of the target items and the total number of target item types as item information of the target items; For each identified target item, obtain its corresponding price information; based on the obtained price information, calculate the total price of similar target items; Count all the types of target items and combine the total price of each type of target items to complete the settlement of the items for sale.
Citation Information
Patent Citations
Settlement method based on goods identification of unmanned vending cabinet
CN117253077A
Intelligent identification method for commodity image of vending machine
CN119360132A
Video processing method and apparatus, device, and storage medium
US20230362416A1
Cited By
Supplier purchase management method based on AI
CN121212581A