Image recognition-based vending machine inventory management method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-27
- Publication Date
- 2026-08-11
AI Technical Summary
[0006]本发明的目的是提供一种基于图像识别的售货机库存管理方法,以解决现有技术存在视觉盲区、硬件成本高、环境适应性差及实时性不足的问题
本发明通过多视角视觉传感器的协同布置与深度卷积神经网络的深度融合,有效攻克了传统库存管理技术在面对商品多层堆叠、部分遮挡以及货道深处光线不足时的识别难题;采用自适应局部直方图均衡化技术,使得系统在强反光或弱光环境下依然能够提取到清晰的商品视觉特征,使商品识别的综合准确率达到极高水平,显著降低了因视觉盲区导致的库存统计偏差。
Smart Images

Figure CN122551265A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer technology, specifically relating to a vending machine inventory management method based on image recognition. Background Technology
[0002] With the rapid transformation of the new retail industry, vending machines, as an important terminal of unmanned retail, are playing an increasingly significant role in improving consumer convenience and reducing operating costs. An efficient inventory management system is the core of stable vending machine operation, enabling real-time monitoring of product sales and providing fundamental data support for replenishment decisions and logistics scheduling. With the advancement of IoT and AI technologies, utilizing sensing devices to achieve digital monitoring of inventory status has become a key path to drive the intelligent and automated upgrading of the retail industry.
[0003] Among them, the vending machine inventory management solution based on image recognition deploys visual sensors inside the cabinet and uses image processing algorithms to extract and analyze the visual features of the goods on the shelves, thereby achieving automated verification of product types and inventory. This solution aims to replace traditional manual inventory counting or monitoring by a single physical sensor with non-contact visual perception, and uses computer vision technology to analyze the status of the product aisles, in order to build a more accurate and efficient digital retail closed loop.
[0004] Existing technologies typically use fixed-position cameras to capture images of the outermost edge of the vending machine's aisle for inventory statistics. However, due to the limited internal space of vending machines, this simplified acquisition method struggles to effectively handle blind spots caused by stacked or partially obscured merchandise, leading to significant discrepancies in inventory statistics. Furthermore, some solutions rely excessively on complex mechanical lifting structures or specific triggering mechanisms, increasing hardware costs and maintenance complexity. They also have limited image feature resolution capabilities in densely packed merchandise or low-light conditions, resulting in poor system versatility. In addition, intermittent image acquisition strategies cannot support the dynamic monitoring needs of high-frequency transactions, making it difficult to guarantee the real-time nature of inventory data and consequently impacting the timeliness of supply chain scheduling.
[0005] To address the above problems, this invention proposes a vending machine inventory management method based on image recognition. Summary of the Invention
[0006] The purpose of this invention is to provide an image recognition-based vending machine inventory management method to solve the problems of visual blind spots, high hardware costs, poor environmental adaptability, and insufficient real-time performance in existing technologies.
[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution: A vending machine inventory management method based on image recognition, comprising: Step 1: Collect raw image data of the vending machine's internal aisles: Through multiple vision sensors deployed on the top floor of the vending machine cabinet and between each shelf layer, multi-view raw images of all aisle areas are acquired synchronously at preset time intervals or when triggered by door lock status. The installation angle of the vision sensors is pre-calibrated to cover the entire space in the depth direction of the aisles. Step 2: Perform multi-dimensional preprocessing on the original image data: use the median filtering algorithm to remove impulse noise in the image, and use adaptive local histogram equalization to compensate for uneven ambient light. Then, use the perspective transformation operator to correct the non-frontal view of the cargo channel image to a standard planar view, and perform region segmentation on the standard planar view according to the preset cargo channel physical coordinates to obtain several independent cargo channel sub-images. Step 3, construct and locate the candidate regions of the goods target: In each channel sub-image, the bounding boxes of the goods target are extracted using a region proposal network based on feature pyramids. By calculating the response values of feature maps at different scales, all suspected goods individuals from the outermost end to the innermost end of the channel are initially identified. The non-maximum suppression algorithm is used to remove redundant candidate boxes with an overlap rate higher than the preset overlap threshold, and the located goods target image is obtained. Step 4, perform product feature extraction and category recognition: input the located product target image into the deep residual network model, extract the preset dimension feature vector containing color distribution, texture structure and geometric contour, and compare the feature vector with the preset product feature library by cosine similarity to determine the specific product category information and recognition confidence of each target. When the confidence is higher than the preset confidence threshold, it is judged as successful recognition; if it is lower than or equal to the threshold, the recognition fails. Step 5: Calculate the inventory of goods in the vending machine and update the inventory status: Based on the position of the identified goods in the depth direction of the vending machine and the preset vending machine capacity parameters, calculate the real-time remaining quantity of goods in each vending machine, and perform logical verification in conjunction with the vending machine's historical transaction records. Finally, synchronize the generated inventory data package to the cloud management database through an encrypted transmission protocol.
[0008] Preferably, in step 1, the vision sensor uses multiple sets of wide-angle industrial cameras with preset resolutions, and the size of the photosensitive element is a predetermined size with a specific single pixel size to ensure that clear edge features of the product can still be obtained in low-light environments.
[0009] Preferably, the visual sensors are arranged as follows: multiple sets of cameras are set above each shelf of the vending machine, wherein the main optical axis of the first set of cameras forms a preset angle with the shelf plane and is used to photograph the brand logo on the top of the product, and the main optical axis of the second set of cameras is perpendicular to the shelf plane and is used to detect the gaps between products to help determine the stacking depth.
[0010] Preferably, the triggering mechanism in step 1 includes active triggering and passive triggering. Active triggering is to perform a full data acquisition once at a predetermined time interval, and passive triggering is to immediately start an image acquisition process when a specific door opening and closing behavior is detected by a magnetic induction switch deployed on the vending machine door frame.
[0011] Preferably, in step 2, the image contrast adjustment adopts an adaptive local histogram equalization algorithm, which divides the image into sub-regions of a preset size, calculates the cumulative distribution function for each sub-region, and uses interpolation to eliminate the block effect at the boundaries of the sub-regions, thereby suppressing overexposure of reflective areas while preserving the details of the product label.
[0012] Preferably, the matrix parameters of the perspective transformation operator are obtained through a multi-point calibration method. By pasting calibration stickers with preset reflectivity at multiple corner points of the cargo channel, a mapping relationship between the physical coordinate system of the cargo channel and the image pixel coordinate system is established, and the reprojection error is controlled within a preset pixel error range.
[0013] Preferably, in step 3, the region proposal network uses anchor frames with various preset aspect ratios to cover the product height range within a preset pixel range, so as to adapt to different sizes of bottled water, canned beverages and bagged snacks.
[0014] Preferably, the positioning process incorporates a depth prediction branch. By analyzing the changes in the projected size of the product in the image, the actual physical distance between the product and the camera is calculated. For partially occluded targets caused by stacking, the bounding box regression algorithm is used to complete the product outline. The area deviation after completion is less than a preset deviation threshold.
[0015] The localization process also integrates a rotation-aware bounding box prediction branch. For the nonlinear deformation of pixel contours caused by the tilting or displacement of cylindrical goods in the channel, the region proposal network adds an additional rotation factor parameter to the output of four basic coordinate offset values to fit the yaw angle of the goods in the channel plane. Through bounding box prediction with rotation parameters, the system can make the four sides of the bounding box adaptively fit the actual edge of the goods, reducing false detections caused by the overlap of adjacent goods in closely arranged scenarios.
[0016] Preferably, in step 4, the deep residual network model adopts a network structure with a preset number of layers, and embeds a squeezing and incentive attention module after each residual block to enhance the response to key brand identity features by learning the correlation between channels.
[0017] Preferably, the product feature library contains visual templates of no less than a predetermined number of common retail products. Each product stores multiple sets of standard feature vectors under different angles and lighting conditions. The feature vectors are updated at a preset period and incremental training is performed by automatically crawling data.
[0018] Preferably, during the category identification process, if multiple categories of goods are detected mixed together in the same aisle, the system will mark the aisle as abnormal and trigger an alarm via the voice prompt module, while simultaneously generating a special record of misplaced goods in the inventory report.
[0019] Preferably, in step 5, the calculation logic for the real-time remaining quantity of goods is as follows: first, obtain the visual inventory based on the image recognition results, then obtain the number of orders sold since the last update, and if the difference between the visual inventory and the theoretical inventory after order deduction is greater than a preset value, then trigger the second image verification process.
[0020] Preferably, the inventory data package is formatted using structured data serialization and includes vending machine number, channel index, product barcode, current quantity, last update timestamp, and confidence score of the recognition result.
[0021] Preferably, the vending machine inventory management method based on image recognition also includes dynamic monitoring of the brightness of the vending machine's internal environment. When the photosensitive sensor detects that the brightness of the internal environment is lower than a preset brightness threshold, an array of light-emitting diodes with preset power and preset color temperature is automatically turned on, and the supplementary lighting time continues until the image acquisition ends.
[0022] Preferably, the identification process also includes visual assistance in judging the shelf life of the product. The production date string on the outer packaging of the product is identified by text recognition technology, the time information is extracted and compared with the current system time. If the remaining shelf life percentage is less than a preset percentage threshold, an expiration reminder is sent to the management backend.
[0023] Preferably, the cloud management database has an inventory prediction model that analyzes sales data within a preset historical time period to predict the demand for goods within a preset future time period and automatically generates a replenishment list to be sent to the mobile terminal of maintenance personnel.
[0024] Preferably, the computational tasks of the image recognition-based vending machine inventory management method are allocated as follows: the image preprocessing and target localization in steps 2 to 3 are completed in the embedded computing unit on the vending machine, which has a preset computing power; and the high-precision recognition and data synchronization in steps 4 to 5 are completed on the cloud server, so as to balance the system's response speed and recognition accuracy.
[0025] Compared with the prior art, the beneficial technical effects of the present invention are as follows: This invention effectively overcomes the recognition challenges of traditional inventory management technologies when faced with multi-layered stacking of goods, partial occlusion, and insufficient light in the depths of the aisles by coordinating the arrangement of multi-view visual sensors and deeply integrating them with deep convolutional neural networks. By employing adaptive local histogram equalization technology, the system can still extract clear visual features of goods in strong reflective or low-light environments, achieving an extremely high overall accuracy in goods recognition and significantly reducing inventory statistical errors caused by visual blind spots.
[0026] This invention introduces a strategy that combines a passive data collection mechanism triggered by door opening behavior with a fixed-period active monitoring mechanism. This multimodal triggering scheme ensures that after each transaction is completed, the inventory data can be parsed and synchronized within a predetermined very short time, which greatly improves the real-time performance of inventory monitoring, provides timely data support for the precise scheduling of the supply chain, and effectively avoids the risk of stockouts.
[0027] This invention has strong versatility and can be quickly deployed in existing vending machines of various specifications and structures. The simplified system structure and low-cost deployment eliminate the need for complex and easily damaged mechanical lifting devices. It relies entirely on fixed-installation high-resolution cameras and efficient software algorithms to achieve all-round spatial coverage. The software-based design significantly reduces the hardware manufacturing cost and subsequent mechanical maintenance cost of a single vending machine.
[0028] This invention represents a leap from simple inventory counting to in-depth business analysis. The system can automatically warn of near-expiry products and provide accurate replenishment suggestions, significantly reducing the workload of manual inventory counting. At the same time, by monitoring the mixed placement of goods in aisles in real time, it ensures the standardization of shelf display, improves the consumer's purchasing experience and the output efficiency of vending machines, and builds a closed-loop digital management system for the modern retail industry. Attached Figure Description
[0029] Figure 1 This is a schematic diagram of the overall technical solution architecture of a vending machine inventory management method based on image recognition provided by the present invention; Figure 2 This is a schematic diagram of the core principle framework of commodity feature extraction and category recognition based on deep residual networks in this invention; Figure 3 This is a flowchart illustrating the logical process of inventory statistics update in this invention, which combines multi-view image acquisition with cloud data synchronization. Detailed Implementation
[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] Example 1 In the image recognition-based vending machine inventory management method provided by this invention, the entire system's operational logic is built upon deep collaboration between the vending machine's local embedded computing unit and a cloud-based management database. This method achieves fully automated, real-time monitoring of the vending machine's internal inventory levels through multi-dimensional image acquisition, high-precision preprocessing, and deep learning-driven target localization and category recognition. (Refer to...) Figure 1 The specific steps are as follows: For step 1, raw image data of the vending machine's internal aisles is collected. In this embodiment, the deployment of the vision sensors adopts a multi-level, multi-view redundant design. Multiple sets of wide-angle industrial cameras with preset resolutions are independently configured on the top layer of the vending machine cabinet and above each shelf. These cameras use CMOS sensors with a size of 1 / 2.3 inches or larger, and the single pixel size is maintained at 1.55 micrometers or larger to ensure a very high signal-to-noise ratio even in low-light environments.
[0032] Specifically, the visual sensors are arranged in two complementary functional arrays: the first set of cameras forms a preset angle of 30 to 60 degrees with the shelf plane, its core function being to capture brand logos, brand markings, and specific shape features on the top of the merchandise; the second set of cameras has its main optical axis strictly perpendicular to the shelf plane, looking downwards to detect gaps between merchandise. This dual-view layout, through geometric calculations, can effectively assist in determining the stacking situation in the depth direction of the aisle. The installation angles of all cameras are pre-calibrated at the factory, and the intrinsic and extrinsic parameter matrices of each camera are obtained using a calibration board.
[0033] Regarding the data acquisition triggering mechanism, the system integrates dual-mode logic of active and passive triggering. The active triggering logic is set to perform a full data acquisition every 3600 seconds to perform system self-checks and periodic inventory calibration. The passive triggering logic relies on a magnetic induction switch deployed on the vending machine door frame. When the Hall sensor detects that the door lock status changes from open to closed, the trigger signal is immediately transmitted to the image acquisition module through the interrupt control circuit. To eliminate the impact of machine vibration caused by the moment the door closes on imaging, the system delays for 500 milliseconds after detecting the door closing action before starting the synchronous exposure program.
[0034] For step 2, multi-dimensional preprocessing is performed on the raw image data. The raw image data first enters the embedded computing unit on the vending machine for streaming processing. To suppress thermal noise or impulse noise generated by the CMOS sensor during imaging, the system calls the median filtering algorithm. The median filter uses a 5×5 sliding window to traverse the image pixel matrix, replacing the center pixel value of the window with the median value of the neighboring pixels after sorting, thereby filtering out noise while preserving the edge contrast of the product to the maximum extent.
[0035] Subsequently, the system executes an adaptive local histogram equalization algorithm to adjust image contrast. This algorithm divides an original image into 8×8 or 16×16 sub-regions and independently calculates the cumulative distribution function of gray levels for each sub-region. To prevent excessive amplification of noise in areas with extremely low contrast, the system introduces a contrast limiting factor. During processing, bilinear interpolation is used to smooth the boundaries between sub-regions, thereby eliminating blockiness and ensuring that details of product labels in shadowed areas deep within the aisles are clearly visible, while suppressing overexposure caused by specular reflections from outer product packaging.
[0036] The perspective transformation operator plays a crucial role in geometric correction during the preprocessing stage. Due to the physical offset of the camera's mounting position, the original image often exhibits trapezoidal distortion. The system employs a multi-point calibration method, affixing calibration stickers with preset reflectivities to the four corners of the cargo channel. By establishing a mapping relationship between the physical coordinate system of the cargo channel and the image pixel coordinate system, a transformation matrix is constructed. The mathematical description of the perspective transformation is as follows: in, These are the original pixel coordinates. The coordinates are the corrected standard plan view coordinates. , Responsible for translation transformations, , Responsible for perspective projection, Set to 1, , , , It is responsible for the linear transformation of the image; through this transformation, the system maps the non-frontal view of the cargo lane image into a frontal projection view; according to the preset physical coordinates of each cargo lane, the panoramic image is divided into several independent cargo lane sub-images using the slicing operator, and each sub-image corresponds to a specific SKU cargo lane.
[0037] For step 3, candidate regions for product targets are constructed and located. In each product channel sub-image, the system invokes a feature pyramid-based region proposal network (FPN). The FPN fuses shallow high-resolution features with deep, rich semantic features through top-down paths and lateral connections, generating feature maps with five levels, P2 to P6. At each feature layer, the system deploys anchor boxes with various preset aspect ratios, such as 1:1, 1:2, and 2:1. These anchor boxes cover a scale range from 32 pixels to 512 pixels, enabling precise matching of product targets of different sizes, such as bottled beverages and canned snacks.
[0038] The localization process also incorporates a depth prediction branch. By inversely calculating the scaling ratio of the product's projected size in the image, the actual physical distance between the product and the camera is estimated. For partially occluded targets deep within the product aisles due to overlap, the system uses a bounding box regression algorithm for contour completion. When the visible area of a detected target is less than 60%, the algorithm predicts and completes its full logical bounding box based on the standard aspect ratio model for that product category, with the area deviation after completion strictly controlled within 10%. To eliminate duplicate detections, the system employs a non-maximum suppression (NMS) algorithm. By calculating the intersection-over-union (IoU) ratio between candidate boxes, redundant boxes with an overlap rate higher than 0.45 are eliminated, ensuring that each product corresponds to a unique localization identifier.
[0039] Before performing depth prediction, the system first completes the camera intrinsic parameter calibration, controls the reprojection error to within 0.1 pixels, and builds a physical size library for the products, including the actual height, actual width, and standard aspect ratio of each SKU.
[0040] Depth estimation employs a pinhole camera model. For fully visible products, the visible area is greater than or equal to 60%. The depth is calculated using the formula "physical distance equals normalized focal length multiplied by actual product height divided by detection box pixel height," and a depth confidence score is output. Redundancy verification is performed using the detection box width, and the weighted average of the height and width estimates is taken as the final depth. For partially occluded targets deep within the aisle due to overlap, where the visible area is less than 60%, the system first identifies the occlusion type: front occlusion, horizontal occlusion, or border occlusion. Then, contour completion is performed based on the standard aspect ratio model for that product category: the width or height of the visible portion is extracted, and the complete bounding box size is calculated using the standard aspect ratio. After completion, quality verification is performed to ensure the area deviation is strictly controlled within 10%, and the depth is recalculated using the completed dimensions. If the completion quality is substandard, the product is marked as "depth unreliable" and downgraded. To improve the spatiotemporal consistency of depth estimation, the system also integrates multi-view depth information for cross-validation: the depth estimate from the top-view camera is transformed to the world coordinate system and then reprojected onto the side-view camera plane. The reprojection error is calculated, and when the reprojection error is less than 5 pixels, it is considered consistent, and the average depth of the two views is output. Otherwise, a second acquisition is triggered, and a Kalman filter is introduced to perform temporal smoothing of the depth values between consecutive frames. When the depth change exceeds ±15 cm in three consecutive frames, inventory recalculation is triggered. To eliminate duplicate detection, the system adopts a depth-aware non-maximum suppression algorithm. The algorithm dynamically adjusts the cross-union ratio (CUR) suppression threshold based on the depth difference between two candidate boxes. The greater the depth difference, the higher the allowed CUR threshold. This threshold can be adjusted from 0.45 to a maximum of 0.75. By calculating the CUR between candidate boxes, redundant boxes with an overlap rate higher than the adaptive threshold are eliminated. Candidate boxes with similar depths and a depth difference of less than 10 cm, and a CUR higher than 0.45, are suppressed. Candidate boxes with significant depth differences and a depth difference greater than or equal to 15 cm are retained even if the CUR value is high, thus ensuring that different product individuals in consecutive positions are not incorrectly merged. When the depth confidence score is lower than 0.3, the detection box height is less than 20 pixels, or the multi-view consistency test fails, the system triggers a degradation processing strategy: first, it prioritizes the prior assignment based on the channel, and calculates the depth according to the recognition order index and the preset channel spacing; second, it uses adjacent frame depth interpolation; and third, it uses multi-view median voting. If all of these fail, it is pushed to the manual review queue. After the above processing, each retained candidate box is assigned a unique logical identifier to ensure that each product individual corresponds to a unique positioning identifier.
[0041] refer to Figure 2In step 4, product feature extraction and category identification are performed. The located product target image is scaled to a preset 224×224 pixels and input into a deep residual network model. This network model adopts a deep architecture with 50 or 101 layers, and a squeezing and incentive attention module is embedded after each residual block. The SE module compresses feature channels through a global average pooling layer, and then learns the weight correlation between channels through a fully connected layer, automatically enhancing the response values of key brand identifier features such as specific color blocks and brand fonts, and suppressing background interference.
[0042] A set of 512-dimensional feature vectors extracted from the fully connected layer of the network represents the highly abstract semantic features of the product. The system compares these vectors in real time with a pre-built product feature library stored in the cloud. The feature library contains visual templates of no fewer than 5,000 common retail products, and each product stores standard vectors under 12 different angles and 3 lighting conditions. The comparison process uses a cosine similarity algorithm, calculated as follows: in, To identify confidence levels, The dimension of the feature vector. This is a 512-dimensional feature vector extracted from the image of the product to be identified. For vectors In the Values in the dimension This refers to the standard feature vector of a specific product stored in a cloud-based feature library. For vectors In the The system calculates similarity values (i.e., recognition confidence) higher than 0.85, indicating successful recognition; otherwise, recognition fails. If multiple feature vectors of different product categories are found within the same product aisle during the recognition process, and all have high confidence levels, it indicates an abnormal mixing of products within the aisle. In this case, the vending machine's local voice prompt module will immediately broadcast an alarm and generate a special record in the background inventory report, marking the index number of the product aisle. In addition, this step integrates OCR text recognition technology, specifically designed to capture the production date string on the product packaging. By extracting the date information and logically comparing it with the current system time, if the remaining shelf life days are less than 20%, the system will automatically send an expiration warning message to the maintenance personnel.
[0043] refer to Figure 3For step 5, the system calculates the inventory level in each vending machine and updates the inventory status. Based on the projection sequence of the identified product targets along the depth direction of the vending machine, and combined with the preset total length parameter of the vending machine, the system calculates the real-time visual inventory level for each vending machine. To ensure absolute data accuracy, the system introduces a logical verification mechanism. The management unit accesses the vending machine's historical transaction records in real time to obtain the number of orders sold since the last inventory update. The system calculates the verification residual using the following formula: Let the inventory identified by vision be... The stock at the time of the last update was During this period, the number of orders sold was The verification residual is ;like If the value is 0, then the inventory status is directly confirmed; if If the value is greater than the preset value of 1, it is determined that there may be a visual blind spot or abnormal handling. The system will trigger a second image verification process and dispatch cameras at different angles to retake the image.
[0044] The final generated inventory data package is encapsulated using structured data serialization. The package contains the vending machine's globally unique identifier (UUID), the product channel index number, the product's international barcode, the current inventory balance, the last update timestamp, and the confidence score distribution of the identification results. This data package is synchronized to the cloud management database via TLS 1.2 encrypted transmission protocol. The cloud database has an inventory prediction model based on a Long Short-Term Memory (LSTM) network. By analyzing the sales curve of the past 30 days, it predicts the consumption of goods in the next 48 hours and automatically generates the optimal replenishment list.
[0045] In this embodiment, the allocation of computing tasks follows the principle of edge-cloud collaboration. Tasks with extremely high real-time requirements and massive data volumes, such as image preprocessing and target localization, are completed on the vending machine's local embedded GPU unit with a computing power of no less than 4 TOPS. Meanwhile, high-precision category recognition and complex inventory logic verification, which involve comparisons of a large feature library, are completed on a cluster of high-performance servers in the cloud. This allocation method significantly reduces the power consumption and data transmission bandwidth pressure on the front-end hardware while ensuring recognition accuracy.
[0046] Furthermore, this method integrates dynamic monitoring of the ambient brightness inside the vending machine. The system reads the values from the photosensitive sensors deployed in the corners of the shelves in real time. When the ambient illuminance is detected to be below 50 lux, the control circuit automatically turns on the array of LED supplementary lights. The color temperature of the supplementary lights is set between 5500K and 6500K, with a color rendering index of no less than 90, to simulate a natural light environment and ensure the accuracy of color reproduction of the product packaging. The supplementary lighting process is synchronized with image acquisition at the microsecond level, and the lights are immediately turned off after acquisition to save energy.
[0047] Example 2 Based on Example 1, this example further refines the image recognition-based vending machine inventory management method in engineering, especially in terms of product positioning accuracy and multimodal data verification in complex occlusion environments.
[0048] In this embodiment, a more robust bounding box regression optimization strategy is introduced for the localization process in step 3. During the sub-image processing of the cargo channel, since bottled goods typically have a cylindrical structure, their pixel contours in the orthographic projection view will undergo nonlinear deformation when the goods are slightly tilted or displaced within the channel. Therefore, the region proposal network in this embodiment not only outputs four basic coordinate offset values but also adds an additional rotation factor parameter to fit the yaw angle of the goods within the channel. Through this bounding box prediction with a rotation parameter, the system can more accurately isolate the edges of closely packed goods, reducing false detections caused by overlapping adjacent goods.
[0049] Among them, the boundary prediction with rotation parameters elaborates on the principle of improving the edge stripping accuracy of goods in closely arranged scenarios from three dimensions: geometric deformation, feature alignment, and overlap differentiation.
[0050] The geometric deformation caused by tilting cylindrical products poses a significant challenge in the actual operation of vending machines. Bottled products often tilt or shift slightly during customer handling. This tilt manifests physically as a deflection of the product around its vertical axis, i.e., a change in the yaw angle. When a cylindrical product tilts, its pixel outline in the frontal projection view undergoes non-linear deformation. Specifically, the projection area, which should be rectangular, no longer has straight edges perpendicular to the horizontal plane, but instead appears as sloping lines. Simultaneously, the projection of the circular bottle cap at the top changes from a perfect circle to an ellipse, with the major axis of the ellipse aligned with the tilt direction. This non-linear deformation prevents traditional axis-aligned bounding boxes from closely fitting the actual outline of the product. The bounding box may contain a large number of background pixels or pixels from adjacent products, thus interfering with the accuracy of feature extraction.
[0051] To overcome the aforementioned problems, this embodiment adds a rotation factor parameter to the output layer of the region proposal network, in addition to the traditional four coordinate bias values: the x-coordinate of the bounding box center point, the y-coordinate of the center point, the width, and the height. The rotation factor parameter is an angle value used to describe the yaw angle of the product within the channel plane, i.e., the angle between the product's centerline and the channel depth direction. The rotation factor parameter, together with the four basic coordinate bias values, constitutes a bounding box with a rotation angle. The five-tuple expression of this bounding box is: x-coordinate of the center point, y-coordinate of the center point, width, height, and rotation angle. During the training process of the region proposal network, annotators need to annotate not only the bounding rectangle of the product but also the product's true yaw angle. By learning from a large number of samples with angle annotations, the network gradually masters the ability to regress the rotation angle from image pixel features, such as edge gradient direction and texture principal direction.
[0052] In this embodiment, edge stripping accuracy is improved by rotating the bounding box. The bounding box with rotation parameters can adaptively rotate according to the actual tilt angle of the product, so that the four sides of the bounding box always remain parallel to the actual edge of the product. For tilted cylindrical products, the left and right sides of the rotated bounding box can accurately fit the two generatrices of the cylindrical surface of the product, and the top and bottom sides can accurately fit the top of the bottle cap and the bottom of the bottle. This precise fitting ensures that most of the pixels within the bounding box belong to the target product itself, and background pixels and interfering pixels of adjacent products are effectively excluded, thereby improving the purity of subsequent feature extraction. In closely arranged channels, there is often partial overlap between adjacent products. When two side-by-side products are tilted to different degrees, their respective projection areas on the image will intersect each other. When using traditional axis-aligned bounding boxes, the bounding boxes of two products will overlap significantly, making it difficult for non-maximum suppression algorithms (NMS) to distinguish between the same product and two different products. However, by using rotated bounding boxes, each product's bounding box extends along its respective tilt direction, significantly reducing the overlap area and the cross-union ratio (CUI). This allows NMS to clearly distinguish between two independent products, reducing false detections caused by overlapping adjacent products. Before inputting the image region within the bounding box into the deep residual network for feature extraction, the system performs rotation correction on the image region based on the rotation factor parameter, rotating the tilted product image to a standard pose, i.e., the product's central axis is parallel to the image's vertical axis. This spatial alignment operation eliminates the need for the subsequent feature extraction network to learn rotation invariance, allowing the entire network capacity to focus on distinguishing the visual features of the products themselves. Brand logos, colors, textures, etc., improve the accuracy of category recognition. In scenarios where goods are closely arranged and pressed together in aisles, the physical gap between adjacent goods is extremely small, or even non-existent. Traditional methods rely on the horizontal spacing between detection boxes to distinguish different goods. When goods are tilted, the edges of adjacent goods will interlock, the horizontal spacing will disappear, and the detection boxes will merge into a large connected region. This embodiment introduces a rotation factor parameter, which allows the bounding box of each goods to extend along its own tilt direction. Even if goods are physically pressed together, their rotated bounding boxes have different extension directions in the image space, thus being effectively separated in the feature space. By analyzing the difference in rotation angle, the system can determine that these are two independent goods in different postures in space, rather than two fragments of one goods, thereby ensuring that each individual goods are correctly identified and counted.
[0053] In the feature extraction stage of step 4, this embodiment further employs a multi-scale feature fusion strategy. While extracting global semantic features, the deep residual network opens a parallel shallow texture branch. This branch specifically targets minute feature points on product packaging, such as anti-counterfeiting marks on the bottle cap and unique anti-slip textures. The system dynamically weights and concatenates the shallow texture features with the deep category features. This method provides a more discriminative basis for distinguishing products of the same brand but with different capacities and flavors but highly similar packaging, such as 500ml and 600ml bottles of the same brand of purified water.
[0054] To address the data verification logic in step 5, this embodiment adds an auxiliary verification step based on weight sensors. High-precision strain gauge weight sensors are pre-embedded at the bottom of each product aisle in the vending machine. After the image recognition system completes a full inventory count, the obtained inventory quantity is compared with the real-time total weight transmitted from each product aisle. Let the preset standard weight for a single item be... The number of identified products is The total weight sensor returned a value of The system determines whether it exists. The logical relationship, in which This is a preset tolerance. If there is a significant mismatch between the two, the system will prioritize trusting the reduction amount from the weight sensor and trigger the vision system to take multiple consecutive shots of the aisle, using the pixel change patterns over time to capture items in the back row that may be completely obscured by the goods in front.
[0055] At the hardware collaboration level, the visual sensor in this embodiment has the function of autonomously adjusting the focal length. When step 5 triggers the second image verification, the computing unit controls the voice coil motor (VCM) to quickly fine-tune the focus. Three rapid snapshots are taken at the front, middle, and rear of the cargo channel, and then synthesized. This depth synthesis technology can produce a panoramic depth image with extremely high clarity from the front to the back, greatly improving the edge extraction accuracy of goods deep in the cargo channel.
[0056] For cloud-based database management, this embodiment constructs a more complex product knowledge graph. In addition to basic image features, it integrates non-visual data such as promotional status of each SKU, weather correlation, and surrounding pedestrian traffic. The inventory forecasting model utilizes this multimodal data and employs an ensemble learning algorithm to more accurately predict inventory depletion rates during holidays or specific marketing campaigns. For example, when the system receives a high-temperature warning from the weather interface, it automatically raises the lower inventory limit alarm threshold for cold drinks, prompting maintenance personnel to increase replenishment frequency in advance.
[0057] Regarding communication security, the encrypted transmission protocol in this embodiment adds a device fingerprint authentication step. Before uploading each image frame, the vending machine's local computing unit uses its built-in Hardware Security Module (HSM) to generate a digital signature containing a unique hardware serial number and the current timestamp. The cloud server only decrypts the image data packet after verifying the signature's validity, thus preventing man-in-the-middle attacks and malicious data injection, ensuring the security and privacy of the company's inventory data.
[0058] Furthermore, the inventory management system in this embodiment also possesses self-learning capabilities. For product images with a recognition confidence level bordering on a threshold, such as 0.75 to 0.85, the system automatically marks them as "samples awaiting review" and asynchronously transmits them to the incremental training pool in the cloud. After manual online labeling, these new samples are used to fine-tune the deep residual network. This incremental learning mechanism enables the system to continuously and autonomously improve its recognition accuracy as new SKUs are added to the shelves and as ambient lighting conditions change seasonally.
[0059] Example 3 In this embodiment, the image recognition-based vending machine inventory management method provided by the present invention is deployed in a large-scale integrated automatic vending machine scenario with a high-density product aisle arrangement. Because the internal space of such vending machines is extremely limited and the product categories are extremely complex, this embodiment specifically enhances the calibration accuracy of step 1 and the image segmentation logic of step 2.
[0060] In step 1, to achieve complete coverage of the narrow, elongated vending machine aisle exceeding 1.5 meters in depth, the installation angle of the vision sensors underwent precise spatial geometric calculations. Multiple wide-angle cameras, with a field of view exceeding 120 degrees, were installed diagonally along the inner wall of the vending machine. This diagonal perspective allows the system to capture gaps between products from the side, solving the problem of complete overlap between products when shooting from the front. The calibration sticker uses a special reflective material that produces a high-brightness response in the infrared band. In extremely poor lighting conditions, the system switches to infrared mode for auxiliary calibration, ensuring that the reprojection error is always controlled within one pixel.
[0061] For image segmentation in step 2, this embodiment introduces a semantic segmentation operator based on a fully convolutional neural network (FCN) to replace simple geometric slicing. FCN can classify pixels in the panoramic image point by point, identifying non-product areas such as shelf borders, aisle rails, and label slots, and masking these interfering areas at the pixel level. Through this dynamic segmentation, even if the vending machine experiences slight deformation of the shelf due to physical impact, the system can adjust the boundaries of the aisle sub-images in real time based on the identified rail positions, ensuring that the segmented sub-images always accurately encompass the effective product storage area.
[0062] In step 3, for the localization of small-sized items such as chocolate bars and chewing gum, the Region Proposal Network (RPN) employs an enhanced receptive field (RFB) module. RFB simulates the centrifugal characteristics of the human visual system through dilated convolutions with different dilation rates, significantly improving the feature extraction capability for extremely small targets. After the NMS algorithm is executed, the system also adds a topological constraint-based filtering operator. This operator utilizes prior knowledge of the physical width of the aisle to perform lateral alignment verification on the identified candidate boxes. If the total width of the identified boxes on the same cross-section exceeds the physical width of the aisle, the system automatically reduces the confidence of overlapping boxes, effectively solving the ghosting interference caused by the high reflectivity of the product packaging.
[0063] In step 4, the feature comparison stage employs a local feature association strategy in this embodiment. For product series with highly similar packaging designs, the system no longer relies solely on a single global feature vector, but instead extracts specific local areas of interest (ROIs) on the packaging, such as the capacity label area and the barcode printing area. The system uses a Spatial Transformation Network (STN) to correct these local areas to a standard orientation before performing a second high-precision feature extraction. A weight decay mechanism is introduced during the comparison process, giving higher decision weights to local features containing key distinguishing information, ensuring the uniqueness of the identification results.
[0064] For inventory data synchronization in step 5, this embodiment introduces an intermediate caching mechanism for edge computing nodes. When the vending machine is in an environment with unstable network signals, such as an underground parking lot, the generated inventory data packets are temporarily stored in local non-volatile memory and marked with priority tags. Once the network connection is detected to be restored, the system uses a breakpoint resume protocol to prioritize the synchronization of high-priority inventory warning data. Simultaneously, the data packet format employs differential encoding technology, transmitting only the channel data that has changed since the last successful synchronization. This compression strategy reduces the amount of data synchronized in a single operation by approximately 85%, significantly improving system response speed in weak network environments.
[0065] Regarding anomaly handling, this embodiment employs a specific logical definition for the "empty aisle" state. When the visual recognition system detects no product targets in a particular aisle with extremely high confidence, the system does not immediately mark it as having zero inventory. Instead, it retrieves historical sales speed data for that aisle. If the product is a best-selling SKU and its inventory is depleted within a very short time, the system triggers the video playback function of the associated camera, automatically capturing 20-second video clips before and after the inventory is depleted and uploading them to the management backend. This mechanism helps operators quickly determine whether the inventory anomaly is due to normal sell-out or aisle congestion or mechanical malfunction.
[0066] Furthermore, this embodiment also includes a defogging preprocessing step to address fog interference inside the vending machine. In low-temperature refrigerated environments, condensation often occurs on camera lenses or product packaging due to temperature differences, leading to a significant decrease in image contrast. In step 2, the system integrates a defogging algorithm based on dark chromatic a priori theory. By estimating the global atmospheric light intensity and transmittance distribution, color compensation and contrast restoration are performed on the affected pixels. The real-time operation of this algorithm ensures that the inventory management system maintains stable recognition accuracy even in extremely humid refrigerated environments, eliminating the need for frequent manual lens wiping.
[0067] In this embodiment, the cloud-based management database further integrates the trajectory scheduling logic for maintenance personnel. Once the inventory prediction model generates a replenishment list, the system automatically generates the optimal inspection route based on the maintenance personnel's current location, traffic conditions, and the urgency of stockouts at each station, using a heuristic path optimization algorithm. The replenishment list is pushed to the maintenance personnel's handheld device as a dynamic H5 page, containing a real-world comparison image of each vending machine, guiding personnel to quickly and accurately replenish stock. This closed-loop intelligent management significantly reduces the vending machine's stockout duration.
[0068] Example 4 Within the technical framework of the above embodiments, this embodiment focuses on the in-depth application of an image recognition-based vending machine inventory management method in intelligent monitoring of product shelf life and automated pricing strategies.
[0069] In step 4, this embodiment significantly enhances the ability to extract detailed text from product packaging. Besides identifying product categories, the system utilizes a specially trained deep text detection network, CRAFT, to locate the printed areas for production dates and expiration dates in the product image. Considering the diversity of printing formats from different manufacturers, such as "YYYY / MM / DD" and "MFG / EXP," the system incorporates a date parsing engine based on Natural Language Processing (NLP). This engine can convert various non-standard date strings into standard timestamps.
[0070] Subsequently, the system executes a shelf-life risk assessment. For each identified item, the system calculates its remaining lifespan. If the shelf life of the item at the front of a vending machine is more than halfway past its expiration date, the system automatically marks that item's status as "priority promotion" in the cloud management database. At this point, the electronic price tag or LCD screen at the front of the vending machine receives instructions from the cloud via API and adjusts the price of the item in real time, implementing a tiered discount strategy. This dynamic pricing mechanism based on visual recognition effectively reduces the spoilage rate of expired goods and improves asset turnover.
[0071] For image acquisition in step 1, this embodiment introduces a multi-frame fusion super-resolution reconstruction technique. When image acquisition is triggered, the camera continuously captures 5 to 8 low-resolution images with sub-pixel offsets at an extremely high frequency. In the vending machine's local computing unit, these multi-frame information are aligned and fused using a motion compensation algorithm to reconstruct a high-resolution, high-detail composite image. This technique compensates for edge chromatic aberration and resolution loss caused by inexpensive wide-angle lenses, allowing even the smallest characters deep within the product aisle, such as the small print of the expiration date, to be accurately resolved by the OCR algorithm.
[0072] In step 5, inventory statistics, this embodiment adds a "product display compliance" scoring dimension. The system not only counts quantities but also analyzes the center point coordinates and rotation angle of the product targets to assess whether the products are neatly arranged within the aisles and whether there are any instances of inverted or tilted packaging. The compliance score is synchronized to the backend via an encrypted protocol. For sites with consistently low scores, the system automatically generates a "display maintenance" task order, reminding maintenance personnel to manually tidy up the displays during the next restocking, thus ensuring a good user experience.
[0073] To further reduce cloud storage pressure, this embodiment performs "visual feature desensitization and summary extraction" in step 5. Before uploading data, the vending machine converts the original high-definition product images into extremely small feature fingerprints. Unless a recognition conflict or abnormal alarm is detected, the cloud only stores this summary information and the final inventory value. This approach meets auditing requirements while reducing long-term cloud storage costs by more than 70%, and also complies with relevant data security management regulations, protecting consumer privacy.
[0074] For multi-camera collaboration, this embodiment implements a "dynamic field-of-view weight allocation" mechanism. During inventory verification, if the first camera group assesses a high occlusion rate for a particular aisle, the computational unit automatically increases the weight of the second camera group's vertical viewpoint in the feature fusion of that aisle. The system establishes a dynamic probabilistic graphical model (PGM) to synthesize the observation results from multiple cameras and outputs a globally optimal inventory estimation distribution. This multi-sensor fusion strategy makes the system highly adaptable to changes in the way goods are stacked within the aisles. Even in extreme abnormal arrangements such as tilted or overturned goods, it can still arrive at correct inventory conclusions through multi-angle verification.
[0075] This embodiment also explores the automatic initialization capability of this method in large-scale distributed deployment. Once a new vending machine is installed, the system automatically initiates a "spatial self-calibration" program. The camera captures background images when the aisles are empty, and a deep learning model automatically identifies the number of shelf layers, the number of aisles per layer, and the spacing between the paddles. The system automatically compares the identified physical topology with preset configuration parameters, achieving a plug-and-play deployment effect. This automatic initialization capability significantly reduces the technical barrier to on-site installation, enabling the method to be rapidly replicated in large-scale retail networks.
[0076] Finally, the cloud management database in this embodiment integrates an image simulator based on a generative adversarial network. This simulator can simulate virtual samples under various extreme lighting and viewing angles based on existing product feature vectors, and feed these samples back to the recognition model for adversarial training. This closed-loop self-evolution capability enables the system to adapt its model in a very short time when faced with newly introduced specially packaged products, such as transparent packaging or highly reflective metal packaging, maintaining the leading edge and robustness of the system's core recognition algorithm.
[0077] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An image recognition-based vending machine inventory management method, characterized by, Includes the following steps: Step 1: Using multiple visual sensors deployed on the top layer of the vending machine cabinet and between each shelf layer, multi-view original images of all product aisle areas are simultaneously acquired at preset time intervals or when triggered by door lock status. Step 2: Perform multi-dimensional preprocessing on the original image data. Use the median filtering algorithm to remove impulse noise from the original image and use adaptive local histogram equalization to compensate for uneven ambient light. Then, use the perspective transformation operator to correct the non-frontal view of the cargo channel image to a standard planar view. And use the preset cargo channel physical coordinates to segment the standard planar view into regions to obtain independent cargo channel sub-images. Step 3: In each cargo channel sub-image, the bounding boxes of the target goods are extracted using a region proposal network based on feature pyramids. By calculating the response values of feature maps at different scales, all suspected individual goods from the outermost end to the innermost end of the cargo channel are initially identified. Then, the redundant candidate boxes with an overlap rate higher than the preset overlap threshold are removed using a non-maximum suppression algorithm to obtain the localized target goods image. Step 4: Perform product feature extraction and category recognition. Input the located product target image into the deep residual network model, extract the feature vector of the preset dimension, and compare the feature vector with the preset product feature library by cosine similarity to determine the specific product category information and recognition confidence of each target. When the recognition confidence is higher than the preset confidence threshold, it is judged as successful recognition; if it is lower than or equal to the threshold, the recognition fails. Step 5: Based on the arrangement of the identified product targets in the depth direction of the vending channel and the preset vending channel capacity parameters, calculate the real-time remaining quantity of products in each vending channel, and perform logical verification in conjunction with the vending machine's historical transaction records. Finally, synchronize the generated inventory data package to the cloud management database through an encrypted transmission protocol.
2. The image recognition-based vending machine inventory management method of claim 1, characterized by: In step 1, the installation angle of the vision sensor is pre-calibrated to cover the entire space in the depth direction of the aisle. The vision sensor is arranged in a dual-function array structure: multiple sets of industrial cameras are set above each shelf of the vending machine; the main optical axis of the first set of cameras forms a preset angle with the shelf plane to capture the brand logo, logo and shape features on the top of the goods. The main optical axis of the second set of cameras is perpendicular to the shelf plane and is used to detect physical gaps between adjacent products. The vision sensor uses hardware with a preset size photosensitive element and a preset single pixel size, and the intrinsic and extrinsic parameter matrices of each camera are obtained through a calibration board at the factory.
3. The image recognition-based vending machine inventory management method of claim 1, wherein: The triggering mechanism in step 1 consists of both active and passive triggering logic: the active triggering logic is set to perform a full image acquisition once every preset time interval; the passive triggering logic relies on the magnetic induction switch deployed on the vending machine door frame. When the Hall sensor detects a change in the door lock state, the trigger signal is transmitted to the image acquisition module through the interrupt control line; the system starts the synchronous exposure program within a preset delay time after detecting the door closing action to eliminate the impact of physical vibration generated at the moment the door closes on the image quality.
4. The image recognition-based vending machine inventory management method of claim 1, wherein: In step 2, the adjustment of image contrast is achieved through an adaptive local histogram equalization algorithm: the original image is divided into multiple sub-regions of a preset size, the cumulative distribution function of gray level distribution is calculated independently for each sub-region, and a limiting contrast factor is introduced to prevent local noise amplification; during the processing, bilinear interpolation is called to smooth the boundary between sub-regions, eliminate block effects, enhance the details of product labels in the shadow areas deep in the aisle, and suppress the mirror reflection of the packaging of the front-end products.
5. The image recognition-based vending machine inventory management method of claim 1, wherein: In step 2, the transformation matrix of the perspective transformation operator is determined by the multi-point calibration method: calibration stickers with preset reflectivity are pasted at multiple corner points in the physical space of the cargo channel to establish a mapping relationship between the physical coordinate system of the cargo channel and the image pixel coordinate system; the perspective transformation operator converts the original pixel coordinates into the corrected standard planar view coordinates according to the mapping relationship, and restores the tilted cargo channel image from the non-frontal viewpoint to the orthographic projection view. After geometric correction, the panoramic image is segmented into corresponding channel sub-images according to the preset channel index coordinates by a slicing operator.
6. The image recognition-based vending machine inventory management method of claim 1, wherein: In step 3, the feature pyramid fuses shallow high-resolution features with deep semantic features through a top-down path and horizontal connections, generating feature maps of multiple levels. Multiple anchor frames with preset aspect ratios are deployed on each feature layer, with the scale range of the anchor frames covering a preset pixel interval to match product targets of different physical sizes. The localization process integrates a depth prediction branch, which inversely calculates the actual physical distance of the product relative to the visual sensor by analyzing the scaling ratio of the product's projection size in the image. The localization process also integrates a rotation-aware bounding box prediction branch. For the nonlinear deformation of pixel contours caused by the tilting or displacement of cylindrical goods in the channel, the region proposal network adds an additional rotation factor parameter to the output of four basic coordinate offset values to fit the yaw angle of the goods in the channel plane. Through bounding box prediction with rotation parameters, the system can make the four sides of the bounding box adaptively fit the actual edge of the goods, reducing false detections caused by the overlap of adjacent goods in closely arranged scenarios.
7. The image recognition-based vending machine inventory management method of claim 1, wherein: In step 3, for partially obscured targets caused by overlapping goods in the cargo channel, the bounding box regression algorithm is used to complete the outline: when the visible area of a single target is detected to be lower than a preset ratio, the algorithm predicts its complete logical boundary according to the physical model of the standard aspect ratio of the product category; the non-maximum suppression algorithm calculates the intersection-union ratio between candidate boxes and eliminates redundant positioning results with an overlap rate higher than a preset threshold to ensure that each individual product corresponds to a unique logical identifier in the cargo channel.
8. The image recognition-based vending machine inventory management method of claim 1, wherein: In step 4, the preset dimension feature vector includes color distribution, texture structure and geometric contour; the deep residual network model includes a network architecture of preset depth, and a squeezing and excitation attention module is embedded after each residual block; the attention module compresses feature channels through a global average pooling layer and uses the fully connected layer to learn the weight correlation between channels to enhance the model's response value to key brand identity features. The product feature library stores multiple sets of standard feature vectors under different angles and lighting conditions. The comparison process uses the cosine similarity algorithm to calculate the similarity value between the extracted feature vectors and the preset feature templates.
9. The image recognition-based vending machine inventory management method of claim 1, wherein: Step 4 also involves commodity anomaly monitoring and information analysis: when multiple product category feature vectors exist in the same cargo channel and the recognition confidence is higher than the preset threshold, it is determined to be a cargo channel mixed storage anomaly, triggering the voice prompt module to broadcast an alarm and mark the cargo channel index in the inventory report; at the same time, text recognition technology is used to locate the production date area of the product's outer packaging, the date string is parsed and converted into a timestamp, and when the remaining shelf life percentage is less than the preset near-expiration threshold, an early warning message is sent to the management backend.
10. The image recognition-based vending machine inventory management method of claim 1, wherein: In step 5, the logical verification of the remaining quantity of goods in real time is achieved by calculating the verification residual: obtaining the visual inventory value, the last updated inventory value, and the number of orders sold during the period; when the verification residual is greater than the preset tolerance value, the second image verification process is triggered, scheduling visual sensors at different installation angles to repeatedly take pictures and make corrections based on the time series pixel change pattern; the calculation tasks are allocated according to the edge-cloud collaboration principle, wherein image preprocessing and target positioning are executed on the vending machine's local computing unit with preset computing power, and category recognition and inventory logical verification are executed on the cloud server.