An unmanned cabinet commodity state intelligent identification method based on multi-source perception data

CN122821670APending Publication Date: 2026-09-25MACDEXI VISION TECHNOLOGY NANJING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610933763.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]然而,现有技术方案在定位识别精度、追踪连续性以及异常状态判别能力上存在显著缺陷

Benefits of technology

1.通过截断图像流将目标局部图像输入特征网络输出视觉特征与遮挡比例,将遮挡比例转换输出置信度系数。在此基础上,进一步将实际重量转换输出重量特征,并将重量特征、视觉特征、置信度系数输入跨模态模型输出基础标识,建立了一种多维度物理特征的自适应调节机制,能够客观量化当前视野空间的受阻程度,并据此平滑调节不同感知维度之间的数据采信分布;当纯视觉环境的数据获取受限时,依循置信度系数将特征研判的侧重转移至重力等其他物理属性,从而使得设备在应对各类复杂物理遮挡状态时,依然能够持续获取连续、稳定且有效的数据支撑,提升了整体数据采集与多模态融合过程的物理抗干扰能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821670A_ABST
    Figure CN122821670A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of inventory management, in particular to a kind of unmanned cabinet commodity state intelligent identification method based on multi-source perception data.It includes receiving the actual weight, reference weight, global image in cabinet, operating part trajectory and storage rack space grid when weight mutation is intercepted target image, extracts visual feature and occlusion confidence coefficient, and inputs cross-modal model fusion output basic identification with weight feature;Action trajectory, grid and actual weight are input into sliding alignment window, extract misalignment feature and match with basic identification, output storage rack mapping relationship;Compare misalignment weight identification and basic identification, respond to difference signal to extract attenuation waveform to verify material density, combined with mapping relationship output inventory update instruction.The present application makes visual appearance, weight mutation and acoustic damping realize the close logical interlocking in time series and three-dimensional space, enhances the complementary and cross-verification of global perception data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of inventory management technology, specifically to a method for intelligent identification of the status of goods in unmanned vending machines based on multi-source sensing data. Background Technology

[0002] In closed, unmanned distribution and micro-warehouse environments, real-time inventory tracking and accurate ledger management of large-scale, dispersed materials is a significant technological application in the field of smart supply chain and inventory management. Current technologies generally employ a combination of basic visual imaging, gravity sensing, and internet-based material verification methods. This involves monitoring the optical image characteristics or overall load changes within the cabinet; when external physical quantities reach specific thresholds, the system automatically extracts and compares the changed material information. This is currently the mainstream technology for achieving automated inventory counting and data synchronization in unmanned cabinets.

[0003] However, existing technical solutions have significant shortcomings in terms of positioning accuracy, tracking continuity, and the ability to identify abnormal states. First, the features used for determining the status of materials are static and have coarse accuracy. Existing solutions are prone to range deviations and misjudgments when faced with severe visual obstruction caused by densely packed materials in cabinets, or when different materials have similar basic weights and heights. Second, they usually only perform "snapshot-style" position determination at the moment a specific action is triggered, which cannot achieve continuous tracking in spatial sequence. They cannot accurately capture the complex trajectory of materials being temporarily moved and misplaced to other shelves during the interaction process, nor can they achieve automatic correction of physical grid mapping relationships, resulting in serious confusion in the actual layout. Furthermore, existing verification methods mostly rely on superficial recognition of appearance outlines and basic weights. They often cannot make accurate judgments when faced with abnormal situations such as undamaged appearances but depleted core materials inside, and cannot cope with complex environmental obstructions and abnormal interaction trajectories, resulting in data that lacks validity and deep security.

[0004] To address this, a method for intelligent identification of the status of goods in unmanned vending machines based on multi-source sensing data is proposed. Summary of the Invention

[0005] The purpose of this invention is to provide an intelligent identification method for the status of goods in unmanned vending machines based on multi-source sensing data, so as to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for intelligent identification of the status of goods in unmanned vending machines based on multi-source sensing data, comprising: Receives data from the unmanned cabinet's sensor terminal, including the actual weight, reference weight, global image inside the cabinet, trajectory of the operating parts, and storage rack space grid. When the difference between the actual weight and the reference weight is greater than the difference threshold, extract the sudden change coordinates of the storage rack; truncate the global image inside the cabinet according to the sudden change coordinates of the storage rack and output the target local image; input the target local image into the feature network to output visual features and occlusion ratio, convert the occlusion ratio to output the confidence coefficient; convert the actual weight to output the weight feature; input the weight feature, visual feature, and confidence coefficient into the cross-modal model to output the basic identifier; Input the trajectory of the operating part, the spatial grid of the storage rack, and the actual weight into the sliding alignment window to output the misaligned target features and the weight of the misaligned storage rack; match the misaligned target features with the basic identifier to output the storage rack mapping relationship; Input the weight of the misaligned storage rack into the classification model to output a weight label, and compare the basic label with the weight label to output a difference signal; in response to the difference signal, receive the historical attenuation waveform sampled by the unmanned cabinet sensor terminal, input it into the material seismic wave feature library to output the density result; combine the storage rack mapping relationship, density result, and basic label to output an inventory update instruction.

[0007] Preferably, the specific acquisition process of actual weight, reference weight, global image inside the cabinet, trajectory of operating parts, and storage rack space grid includes: periodically polling the storage rack load value through the gravity sensing matrix included in the unmanned cabinet sensing terminal and storing it as the reference weight; acquiring the current load value through the gravity sensing matrix in interactive mode and outputting the actual weight; acquiring a continuous video frame sequence through the visual acquisition matrix included in the unmanned cabinet sensing terminal and outputting the global image inside the cabinet; acquiring acoustic waveforms through the microphone array included in the unmanned cabinet sensing terminal and storing them in the audio buffer queue; parsing the action node coordinate sequence in the global image inside the cabinet through an image skeleton point extraction algorithm and outputting the trajectory of the operating parts; retrieving the three-dimensional shelf physical coordinate topology model, extracting grid parameters, and outputting the storage rack space grid.

[0008] Preferably, the specific generation process of the target local image includes: calculating the numerical difference between the actual weight and the reference weight; locating the gravity sensing matrix target sensing unit that generates the numerical difference in response to an interruption condition where the numerical difference exceeds the difference threshold; extracting the device number of the target sensing unit; inputting the device number into a grid coordinate mapping table for querying, extracting the three-dimensional position coordinates and outputting them as the storage rack mutation coordinates; inputting the storage rack mutation coordinates into a spatial perspective mapping matrix and outputting two-dimensional pixel bounding box parameters; performing pixel array slicing operation on the global image inside the cabinet using the two-dimensional pixel bounding box parameters and outputting an initial truncated image stream; extracting the static background pixel template of the storage rack spatial grid mapping; performing pixel contrast subtraction operation on the initial truncated image stream and the static background pixel template and outputting a preliminary foreground pixel set; removing hand pixels from the preliminary foreground pixel set according to the hand mask mapped by the trajectory of the operation part and outputting a differentiated foreground pixel set; reconstructing the differentiated foreground pixel set into an independent pixel tensor and outputting the target local image.

[0009] Preferably, the specific generation process of the confidence coefficient and visual features includes: inputting the target local image into the convolutional feature extraction layer included in the feature network to perform multi-scale feature convolution operation, and outputting the visual features; inputting the target local image into the instance semantic segmentation layer included in the feature network to output the hand object mask and the target object mask; extracting the number of overlapping pixels between the area of ​​the hand object mask and the area of ​​the target object mask; extracting the quotient of the number of overlapping pixels divided by the total number of pixels in the target object mask, and outputting the occlusion ratio; inputting the occlusion ratio into a nonlinear decay function to extract the inverse weight value; inputting the inverse weight value into a normalization processing layer to perform a range limitation operation, and outputting the confidence coefficient.

[0010] Preferably, the specific generation process of the basic identifier includes: inputting the actual weight into the scalar feature embedding layer of the cross-modal model and outputting the weight feature; assigning the confidence coefficient as a penalty factor; inputting the visual feature, the weight feature, and the penalty factor into the feature adjustment layer, using the penalty factor to perform a channel dimension numerical attenuation operation on the visual feature and outputting a corrected visual feature, and using the complementary value of the penalty factor to perform a splicing dimension numerical enhancement operation on the weight feature and outputting a corrected weight feature; inputting the corrected visual feature and the corrected weight feature into the perceptron fusion layer to perform a feature-level splicing operation and outputting a joint feature vector; inputting the joint feature vector into the fully connected classification layer to perform a classification node distance calculation operation, extracting the classification node labels that match the spatial distance and outputting them as the basic identifier.

[0011] Preferably, the specific generation process of the storage rack mapping relationship includes: converting the trajectory of the operating part and the actual weight into a trajectory time series and a weight change time series, respectively; inputting the trajectory time series and weight change time series into the dynamic time warping algorithm layer contained in the sliding alignment window to generate a spatiotemporal distance matrix; extracting singular nodes that deviate from the matching path in the spatiotemporal distance matrix and outputting action conflict coordinates; mapping the action conflict coordinates to the storage rack spatial grid and outputting the misalignment occurrence grid node; extracting the local weight jump value of the misalignment occurrence grid node and outputting it as the weight of the misaligned storage rack; and extracting the misalignment occurrence grid node matching based on the trajectory of the operating part. The trajectory interaction features are output as the misaligned target features; the basic identifier is input into the attribute database of the Internet platform for querying, and the associated standard three-dimensional shape size features and standard packaging outline deformation features are output; the misaligned target features are parsed, and the interactive volume features and interactive outline deformation features are output; the volume difference value between the interactive volume features and the standard three-dimensional shape size features is calculated; the deformation deviation value between the interactive outline deformation features and the standard packaging outline deformation features is calculated; a two-dimensional cross deviation matrix is ​​constructed based on the volume difference value and the deformation deviation value; the two-dimensional cross deviation matrix is ​​substituted into the preset position determination formula to perform the attribution determination calculation, and the storage rack mapping relationship is output.

[0012] Preferably, the specific generation process of the difference signal includes: inputting the weight of the misaligned storage rack into the discretization and quantization mapping layer included in the classification model to perform interval classification operation, and outputting the interval segment to which the weight belongs; inputting the interval segment to which the weight belongs into the association lookup table included in the classification model for querying, and outputting the weight identifier; performing interval assignment logic determination operation on the weight identifier and the basic identifier; and generating a hardware interrupt pulse level and outputting it as the difference signal in response to the identifier matching failure result output by the interval assignment logic determination.

[0013] Preferably, the specific generation process of the inventory update instruction includes: using the difference signal as an extraction trigger condition, extracting the acoustic waveform at the moment of weight change from the audio buffer queue, and outputting the historical attenuation waveform; inputting the historical attenuation waveform into a fast Fourier transform to perform a time-frequency domain conversion operation, and outputting a frequency-domain damping coefficient sequence; inputting the frequency-domain damping coefficient sequence into the material seismic feature library to perform an Euclidean distance comparison operation, selecting the material category that matches the Euclidean distance and outputting the density result; using the density result to overwrite the original material attribute corresponding to the basic identifier, and outputting the verified identifier; binding the verified identifier to the grid physical coordinates indicated by the storage rack mapping relationship, and outputting the inventory update instruction to the Internet platform.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By truncating the image stream, the target local image is input into a feature network to output visual features and occlusion ratio. The occlusion ratio is then converted into a confidence coefficient. Based on this, the actual weight is further converted into a weight feature. The weight feature, visual feature, and confidence coefficient are then input into a cross-modal model to output a basic identifier. This establishes a multi-dimensional physical feature adaptive adjustment mechanism that can objectively quantify the degree of obstruction in the current visual space and smoothly adjust the data acceptance distribution between different perceptual dimensions. When data acquisition in a purely visual environment is limited, the focus of feature analysis is shifted to other physical attributes such as gravity, based on the confidence coefficient. This allows the device to continuously acquire stable and effective data support even when dealing with various complex physical occlusion states, improving the overall physical anti-interference capability of the data acquisition and multi-modal fusion process.

[0015] 2. By sliding the alignment window to output the misaligned target features and misaligned storage rack weight, based on the trajectory of the operating part, the spatial grid of the storage rack, and the actual weight input, and matching the mapping relationship between the misaligned target features and the basic identifier output storage rack, a physical logic for trajectory tracking and node verification based on the spatiotemporal coordination dimension was constructed. In the cross-mapping dimension of the time flow sequence and the three-dimensional spatial grid, the corresponding relationship between local weight jump nodes and external action trajectories was found. With the help of the spatiotemporal feature sliding alignment continuously executed within the continuous operation cycle, the flow and displacement path of the target object in the spatial coordinate system can be objectively restored. This helps the device to maintain the mapping relationship of environmental nodes robustly when facing the interference of disordered physical interaction or irregular action trajectory, and ensures the synchronization accuracy of multi-source sensors under complex spatiotemporal conditions.

[0016] 3. By expanding the detection dimension of physical properties in a specific data comparison stage, in response to the attenuation waveform sampled by the unmanned cabinet sensor terminal in the response to the difference signal, the attenuation waveform is then input into the material vibration feature library to output the density result. Finally, the storage rack mapping relationship, density result, and basic identification are combined to output the inventory update instruction. By using the sensor terminal to capture the acoustic damping law at the moment of physical interaction of the item, the comparison and verification are carried out at the level of the intrinsic density and material physical properties of the target material. This adds hardware detection capability at the level of acoustic damping to the overall perception link, so that the recognition process breaks through the limitation of relying only on static weight and appearance pixels. It ensures that even when encountering extreme situations where the appearance is intact but the internal core material structure has changed, the device can still output highly objective deep physical identification results.

[0017] 4. By deeply integrating confidence-based cross-modal weight adjustment, sliding window-based spatiotemporal misalignment tracking, and attenuation waveform-based material density verification mechanism, a highly coupled multi-source feature global verification closed loop is constructed, enabling tight logical interlocking of visual representation, weight mutations, and acoustic damping in time series and three-dimensional space. When shallow visual features are limited by environmental occlusion, gravity attributes are adaptively compensated through dynamic weights; when disordered jumps occur at multiple levels, motion trajectories intervene for spatiotemporal alignment; and when logical discrepancies arise in the comparison of basic physical quantities, transient sound waves, as ontological attributes, trigger the underlying decision. This cross-modal and cross-dimensional holistic linkage processing enhances the complementary advantages and cross-verification of global perception data, constructing a robust physical state defense line from a global perspective, resulting in a final output judgment with high state authenticity and overall effectiveness. Attached Figure Description

[0018] Figure 1 This is a flowchart of a method for intelligent identification of the status of goods in unmanned lockers based on multi-source sensing data, as proposed in an embodiment of this invention application. Figure 2 This is a flowchart of the cross-modal dynamic weight allocation and fusion process proposed in an embodiment of this invention application; Figure 3 This is a flowchart of spatiotemporal sliding alignment and acoustic anti-counterfeiting verification proposed in an embodiment of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Please see Figures 1-3 The present invention provides an intelligent identification method for the status of goods in unmanned vending machines based on multi-source sensing data, the specific steps of which are as follows: Receives data from the unmanned cabinet's sensor terminal, including the actual weight, reference weight, global image inside the cabinet, trajectory of the operating parts, and storage rack space grid. When the difference between the actual weight and the reference weight is greater than the difference threshold, extract the sudden change coordinates of the storage rack; truncate the global image inside the cabinet according to the sudden change coordinates of the storage rack and output the target local image; input the target local image into the feature network to output visual features and occlusion ratio, convert the occlusion ratio to output the confidence coefficient; convert the actual weight to output the weight feature; input the weight feature, visual feature, and confidence coefficient into the cross-modal model to output the basic identifier; Input the trajectory of the operating part, the spatial grid of the storage rack, and the actual weight into the sliding alignment window to output the misaligned target features and the weight of the misaligned storage rack; match the misaligned target features with the basic identifier to output the storage rack mapping relationship; Input the weight of the misaligned storage rack into the classification model to output a weight label, and compare the basic label with the weight label to output a difference signal; in response to the difference signal, receive the historical attenuation waveform sampled by the unmanned cabinet sensor terminal, input it into the material seismic wave feature library to output the density result; combine the storage rack mapping relationship, density result, and basic label to output an inventory update instruction.

[0021] The technical solution of the present invention will be further described in detail below with reference to specific embodiments.

[0022] Example 1 This application discloses an intelligent identification method for the status of goods in unmanned vending machines based on multi-source sensing data. (See attached document.) Figure 1 The specific steps proposed in this invention include: S1, receiving the actual weight, reference weight, global image inside the cabinet, trajectory of the operating part, and spatial grid of the storage rack collected by the unmanned cabinet sensing terminal; S2, when the difference between the actual weight and the reference weight is greater than the difference threshold, extracting the abrupt change coordinates of the storage rack; truncating the global image inside the cabinet according to the abrupt change coordinates of the storage rack and outputting the target local image; inputting the target local image into a feature network to output visual features and occlusion ratio, converting the occlusion ratio to output a confidence coefficient; converting the actual weight to output weight features; inputting the weight features, visual features, and confidence coefficient into a cross-modal model to output a basic identifier; S3, inputting the trajectory of the operating part, spatial grid of the storage rack, and actual weight into a sliding alignment window to output misaligned target features and misaligned storage rack weight; matching the misaligned target features and basic identifiers to output the storage rack mapping relationship; S4, inputting the misaligned storage rack weight into a classification model to output a weight identifier, comparing the basic identifier and the weight identifier to output a difference signal; responding to the difference signal, receiving the historical attenuation waveform sampled by the unmanned cabinet sensing terminal, inputting it into a material seismic wave feature library to output a density result; combining the storage rack mapping relationship, density result, and basic identifiers to output an inventory update instruction.

[0023] Furthermore, the system receives the actual weight, reference weight, global image of the cabinet interior, trajectory of the operating part, and storage rack space grid collected by the unmanned cabinet's sensor terminal; corresponding to step S1 above; the specific implementation process includes: The unmanned cabinet sensor terminal uses a gravity sensing matrix to periodically poll the load value of the storage rack and stores it as the baseline weight. In interactive mode, the gravity sensing matrix collects the current load value and outputs the actual weight. The unmanned cabinet sensor terminal uses a visual acquisition matrix to acquire a continuous video frame sequence and outputs a global image of the cabinet interior. The unmanned cabinet sensor terminal uses a microphone array to collect acoustic waveforms and store them in an audio buffer queue. An image skeleton point extraction algorithm is used to analyze the motion node coordinate sequence in the global image of the cabinet interior and outputs the trajectory of the operating part. A three-dimensional shelf physical coordinate topology model is retrieved, and mesh parameters are extracted to output the storage rack spatial mesh.

[0024] Specifically, the acquisition process for actual weight, baseline weight, global image inside the cabinet, trajectory of operating parts, and storage rack space grid is as follows: The unmanned cabinet's sensing terminal periodically polls the storage rack's load values ​​using a gravity sensing matrix. This gravity sensing matrix consists of an array of resistance strain gauge miniature weighing units deployed at the bottom of the rack. In the resting state, the polling sampling frequency is set to 10 Hz. This resting sampling frequency was selected based on experimental results balancing the maximum bandwidth of the gravity sensor's data bus with standby power consumption. When the cabinet door is locked and closed in the resting state, the polled values ​​are stored in non-volatile memory as a baseline weight.

[0025] When the cabinet door opens and the user enters the interactive state, the current load value is collected through the gravity sensing matrix. To capture the high-frequency dynamic load fluctuations at the moment the user's hand touches the product, the sampling frequency for the interactive state is set to 50 Hz based on the Nyquist sampling theorem. To reduce the impact of mechanical vibration noise generated by the user's hand touching the shelf, the current load value is smoothed using a Kalman filter algorithm to output the actual weight.

[0026] The perception module acquires continuous video frame sequences through a visual acquisition matrix included in the unmanned vending machine's sensing terminal. This visual acquisition matrix employs a binocular stereo camera architecture, installed on the top edge and side walls inside the vending machine, providing a field of view for environmental perception. The camera's frame rate is 30 frames per second, and its resolution is 1920×1080 pixels. The camera's frame rate and resolution are selected based on a combination of the edge computing unit's video decoding throughput and the pixel density required to achieve the minimum recognizable characters on product packaging. Real-time three-channel image tensors are output as a global image of the vending machine's interior.

[0027] On the parallel channel, acoustic waveforms are acquired in real time via a microphone array included in the unmanned vending machine's sensing terminal. The microphone sampling rate is selected as follows: to cover the high-frequency acoustic characteristic bands generated by common product drops or touches, the sampling rate is set to 44.1 kHz, referencing Nyquist frequency coverage requirements. The acquired digital audio data stream is continuously stored in an audio buffer queue. This audio buffer queue adopts a circular buffer structure, permanently storing the acoustic waveforms of the past 5.0 seconds. The retention period for this audio buffer queue is selected by combining the maximum processing delay time from the user triggering an anomaly to completing cross-modal calculations, along with the required physical cycle margin for the complete attenuation of the acoustic waveform. The queue is updated using a first-in, first-out (FIFO) principle for backtracking and retrieving historical attenuation waveforms.

[0028] For image data, the computing node analyzes the motion node coordinate sequence in the global image inside the cabinet using a pre-set image skeleton point extraction algorithm, extracting the two-dimensional pixel coordinates and depth coordinates of key nodes on the user's arm. The continuous time-series coordinates are then stitched together to output the trajectory of the operated part.

[0029] The system synchronously retrieves a pre-burned 3D shelf physical coordinate topology model from the firmware. This model divides the cabinet space into several standard-volume 3D grids, such as 5-centimeter cubes. The parameters of the standard-volume 3D grids are selected based on the circumscribed length, width, and height dimensions of the smallest volume item (such as chewing gum or gummy candy) allowed to be accommodated in the unmanned cabinet design. Finally, the boundary parameters, center point physical coordinates, and corresponding shelf level numbers of the grids are extracted, and the storage shelf space grid is output as the physical coordinate mapping reference.

[0030] By continuously storing waveforms into an audio buffer queue using a microphone array, a hardware backtracking foundation for acoustic transient data is established, avoiding sampling lag when abnormal triggers occur and ensuring the temporal continuity of physical capture.

[0031] Furthermore, when the difference between the actual weight and the reference weight exceeds a threshold, the abrupt change coordinates of the storage rack are extracted; based on the abrupt change coordinates, the global image inside the cabinet is truncated to output a local target image; the local target image is input into a feature network to output visual features and occlusion ratio, and the occlusion ratio is converted to output a confidence coefficient; the actual weight is converted to output a weight feature; the weight feature, visual feature, and confidence coefficient are input into a cross-modal model to output a basic identifier; this corresponds to step S2 above; see [link to relevant documentation]. Figure 2 The specific implementation process includes: Calculate the numerical difference between the actual weight and the reference weight; in response to an interruption condition where the numerical difference exceeds the difference threshold, locate the gravity sensing matrix target sensing unit that generates the numerical difference; extract the device number of the target sensing unit; input the device number into a grid coordinate mapping table for querying, extract the three-dimensional position coordinates and output them as the storage rack mutation coordinates; input the storage rack mutation coordinates into a spatial perspective mapping matrix and output two-dimensional pixel bounding box parameters; use the two-dimensional pixel bounding box parameters to perform pixel array slicing operation on the global image inside the cabinet and output an initial truncated image stream; extract the static background pixel template of the storage rack spatial grid mapping; perform pixel contrast subtraction operation on the initial truncated image stream and the static background pixel template to output a preliminary foreground pixel set; based on the hand mask mapped by the trajectory of the operation part, remove hand pixels from the preliminary foreground pixel set and output a differentiated foreground pixel set; reconstruct the differentiated foreground pixel set into an independent pixel tensor and output the target local image.

[0032] Specifically, the process of generating the target local image is as follows: First, the actual weight and the reference weight are read, both in grams. The numerical difference between the actual weight and the reference weight is then calculated. This digital signal of the difference is input into a digital comparator for judgment and calculation.

[0033] The digital comparator has a preset difference threshold. The threshold is selected by collecting the mass of the lightest item allowed to be sold in the vending machine, measuring the noise fluctuation of the gravity sensor under environmental vibration, and setting it using a safety margin principle of superimposing the mean noise level with three standard deviations; for example, it is set to 10.0 grams. When the absolute value of the difference between the actual weight and the reference weight (in grams) is greater than 10.0 grams, the digital comparator outputs a pulse signal.

[0034] In response to an interruption condition where the numerical difference exceeds a threshold, the target sensing unit of the gravity sensing matrix that generated the numerical difference is located. The device physical address number of this target sensing unit is extracted. The device number is entered into a grid coordinate mapping table for lookup to extract the sensor's three-dimensional position coordinates on the shelf, and then output as the storage shelf transition coordinates.

[0035] After obtaining the mutable coordinates, input them into a pre-calibrated spatial perspective mapping matrix to output two-dimensional pixel bounding box parameters.

[0036] In the calculation of the homogeneous pixel coordinates of the center of the 2D image plane in spatial perspective mapping, to ensure the correct dimension of the final output pixel, a pre-calibrated camera intrinsic parameter matrix is ​​first extracted, with a dimension of pixels. This intrinsic parameter matrix is ​​then multiplied by an extrinsic parameter matrix consisting of dimensionless rotation parameters and translation parameters with a dimension of millimeters to form a perspective mapping matrix. Next, this perspective mapping matrix is ​​multiplied by the extracted 3D homogeneous vector of the storage rack abrupt coordinates, with a dimension of millimeters. Finally, the resulting product vector is divided by a scaling factor representing the depth distance of the object from the camera's optical center, also with a dimension of millimeters. After the above dimensional reduction, the millimeters in the numerator and denominator cancel each other out, and the first two terms of the final output vector are the center 2D pixel coordinates of the truncated image, with a dimension of pixels.

[0037] The selection method for the camera's intrinsic parameter matrix parameters is as follows: using the standard Zhang Zhengyou checkerboard calibration method, calibration images from different perspectives are collected before leaving the factory, and nonlinear optimization algorithms are used to calculate and extract the parameters.

[0038] Using the obtained center pixel coordinates as a reference, the system expands outward by a preset pixel offset parameter, outputting a two-dimensional pixel bounding box parameter. The pixel offset is selected by reading the pixel width and height of the largest item in the unmanned cabinet's projection area on the camera calibration plane, thus determining the expansion boundary range. For example, it is selected to expand outward by 150 pixels in each direction (top, bottom, left, and right). Using the two-dimensional pixel bounding box parameter, a pixel array slicing operation is performed on the global image inside the cabinet, outputting an initial truncated image stream.

[0039] Based on the storage rack mutation coordinates, extract the static background pixel template mapped in the storage rack spatial grid. Input the initial truncated image stream and the static background pixel template into the image difference processor.

[0040] The initial foreground pixel set is generated by calculating the absolute value of the difference between the grayscale intensity values ​​at the same coordinates in the initial truncated image stream and the static background pixel template using an image difference processor. Both quantities involved in the calculation are dimensionless grayscale levels from 0 to 255. If the absolute difference is strictly greater than a preset grayscale threshold, the pixel value at that coordinate is retained as the original pixel value of the truncated image; otherwise, if the absolute difference is not greater than the grayscale threshold, the pixel value at that coordinate is forcibly assigned the number 0.

[0041] The grayscale threshold is selected as follows: under different ambient light intensities, the background noise pixel fluctuation difference of the empty shelf is collected, and the maximum difference range between the peak and trough is taken as the limit against ambient light interference, which is set to 25 for example. This operation outputs a preliminary set of foreground pixels.

[0042] Subsequently, the trajectory of the manipulated part is invoked, and a hand mask is generated in its projection area. Based on the hand mask mapped from the trajectory, pixels belonging to the hand region are removed from the initial foreground pixel set using bitwise NAND logic, outputting a differentiated foreground pixel set. This differentiated foreground pixel set is then reconstructed into an independent pixel tensor of fixed resolution size using an interpolation algorithm, outputting the target local image. The resolution parameters for the independent pixel tensor reconstruction are selected by consistent alignment based on the standard receptive field dimensions required by the input layer of the subsequent deep residual feature network model, for example, setting both the length and width to 224 pixels.

[0043] By eliminating irrelevant pixels through background subtraction and hand masking, physical interference from environmental background and user limbs is effectively eliminated, improving the purity of visual framing of the target local image and avoiding subsequent semantic pollution.

[0044] The target local image is input into the convolutional feature extraction layer of the feature network to perform multi-scale feature convolution operations, and the visual features are output. The target local image is input into the instance semantic segmentation layer of the feature network to output a hand object mask and a target object mask. The number of overlapping pixels between the areas of the hand object mask and the target object mask is extracted. The quotient of the number of overlapping pixels divided by the total number of pixels in the target object mask is extracted, and the occlusion ratio is output. The occlusion ratio is input into a nonlinear decay function to extract inverse weight values. The inverse weight values ​​are input into a normalization processing layer for range limitation operations, and the confidence coefficient is output.

[0045] Specifically, the generation process of the confidence coefficient and visual features is as follows: The execution instructions input the local image of the target into the convolutional feature extraction layer of the feature network, and perform multi-scale feature convolution and pooling operations.

[0046] After convolution extraction, a one-dimensional visual feature vector is output. The feature dimension parameter is selected based on the matching principle of multimodal data information entropy to ensure that the dimension can cover the spatial representation requirements of complex packaging textures, and is usually set to 2048 dimensions.

[0047] The target local image is input into the instance semantic segmentation layer of the feature network, and the output is a hand object mask and a target object mask, both of which are binarized matrices.

[0048] In this embodiment, the initial intersection boundary between the hand object mask and the target object mask is extracted. Pixel gradient change rates are collected along the normal direction. Misjudged diffuse reflection pixels are eliminated using an adaptive dual-threshold method, and an optimized hand object mask is output. Specifically, after obtaining the hand object mask and the target object mask, the set of binarized overlapping boundary pixels of the two is extracted. A micro-transition zone is formed by extending 5 pixels to each side along the normal direction from this boundary set as the center. The grayscale values ​​of the target local image within this micro-transition zone are convolved using the Sobel operator, and the magnitude of the first derivative in two-dimensional space is extracted as the pixel gradient change rate. A dual-threshold range for skin color diffuse reflection interference is preset, with the upper threshold set to the grayscale difference value of 45 and the lower threshold set to 15. The gradient change rate of each pixel within the transition zone is compared: if the gradient change rate of a pixel is less than 15, it is determined to be a high-reflectivity misjudgment area caused by the hand being close to the product, and its binarized state in the hand object mask is forcibly set to zero; if the gradient change rate is greater than 45, it is determined to be a real physical occlusion boundary, and its state remains unchanged; for pixels between 15 and 45, the Euclidean distance between them and the initial boundary is calculated, and those with a distance less than 3 pixels are retained, otherwise they are set to zero. After gradient filtering, an optimized hand object mask is generated. The above dual threshold interval selection method is as follows: extract the pixel gradient distribution data of the pre-collected halo and the real physical boundary, and take the statistical empirical boundary that maximizes the separation of the two types of features to set the upper and lower limits respectively; the distance threshold selection method is as follows: calculate the typical lateral diffusion range of the halo based on the imaging magnification ratio of the camera at the normal interaction distance inside the cabinet. The above process eliminates the visual halo interference caused by the palm being close to the highly reflective packaging, and improves the numerical accuracy of the occlusion ratio calculation.

[0049] Subsequently, the number of overlapping pixels between the areas of the hand object mask and the target object mask is extracted. Specifically, the calculation of the occlusion ratio involves multiplying the dimensionless values ​​of each coordinate point in the hand object mask matrix by the corresponding dimensionless values ​​in the target object mask matrix. All these products are then globally summed to obtain the numerator representing the number of overlapping pixels, with the dimension of pixels. Simultaneously, all pixels with a value of 1 in the target object mask matrix are independently summed to form the baseline denominator, also with the dimension of pixels. Finally, the numerator is divided by the denominator, causing the pixel dimensions to cancel each other out. The resulting dimensionless quotient is then extracted and output as the occlusion ratio.

[0050] The calculated occlusion ratio is input into a nonlinear attenuation function to extract the inverse weight value. The specific calculation method for extracting the inverse weight value is as follows: First, the square of the dimensionless occlusion ratio value is calculated, and it is multiplied by a preset dimensionless attenuation factor. Then, a negative sign is added before the product result as the exponent of the natural logarithm base, and then a power operation is performed to obtain the dimensionless inverse weight value.

[0051] The attenuation factor is selected by pre-establishing a large number of test sets with different occlusion ratios and a function scatter plot fitting curve of visual recognition accuracy, and setting the corresponding curvature parameter at the steepest position of the slope of the decrease in visual feature confidence. For example, it is set to 4.0.

[0052] Next, a range-limited calculation operation is performed on the inverse weight values ​​to output the confidence coefficient. The specific calculation steps are as follows: calculate the difference between the preset dimensionless confidence upper limit value and the confidence lower limit value, multiply this difference by the input inverse weight value of this layer, and finally add the confidence lower limit value to the product result, thereby outputting the dimensionless confidence coefficient that is always constrained within the safe range.

[0053] The confidence lower and upper limits are selected as follows: the lower limit is set based on the expected probability of the model when randomly guessing without any visual features, and is set to 0.1 in the example; the upper limit is set based on the confidence of the feature network under ideal baseline prediction conditions without any occlusion or reflection, and is set to 1.0 in the example.

[0054] By calculating the intersection ratio between the hand and the target mask, the confidence coefficient was output, which quantitatively assessed the degree of visual obstruction and provided an objective physical metric for subsequent cross-modal weight allocation.

[0055] The actual weight is input into the scalar feature embedding layer of the cross-modal model, and the weight feature is output. The confidence coefficient is assigned and output as a penalty factor. The visual feature, the weight feature, and the penalty factor are input into the feature adjustment layer. The penalty factor is used to perform a channel dimension numerical decay operation on the visual feature to output a corrected visual feature. The complementary value of the penalty factor is used to perform a splicing dimension numerical enhancement operation on the weight feature to output a corrected weight feature. The corrected visual feature and the corrected weight feature are input into the perceptron fusion layer to perform a feature-level splicing operation and output a joint feature vector. The joint feature vector is input into the fully connected classification layer to perform a classification node distance calculation operation and extract the classification node labels that match the spatial distance to output as the basic identifier.

[0056] Specifically, the generation process of the basic identifier is as follows: The actual weight is input into the scalar feature embedding layer. The embedding layer uses a multi-level fully connected structure to perform data dimensionality upscaling and outputs a weight feature vector. The selection method for the number of neurons in each level of the scalar feature embedding layer is as follows: adopting a doubling proportional expansion principle, expanding from a single input scalar to 64, then to 128, and finally reaching 256 output dimensions. This selection aims to ensure that the scalar data can uniformly expand the feature representation space during the dimensionality upscaling process, thereby matching the high-dimensional visual features.

[0057] The dimensionless confidence coefficients output are input into the penalty factor mapping unit of the cross-modal model, directly converting the coefficients into penalty factors. This unit does not contain trainable network weights; it only performs an identity mapping operation. Using the penalty factor, a numerical decay operation is performed on the visual features along the channel dimension. For the calculation of corrected visual features, a channel traversal operation is performed, directly multiplying each numerical element of the original visual feature vector along the channel dimension by the transformed penalty factor. This operation can reduce the weights of visual features when occlusion is severe.

[0058] Subsequently, the bias complement value is calculated using a penalty factor. The specific calculation process is as follows: the penalty factor is multiplied by a preset scaling factor (set to 0.5), and then the product is subtracted by a constant 1. The difference obtained is used as the bias complement value to ensure that the lower limit of the complement value is always not lower than 0.5. Then, the bias complement value is used to perform a scalar multiplication scaling operation on the channel dimension of the weight feature to output the corrected weight feature.

[0059] The 2048-dimensional corrected visual features and 256-dimensional corrected weight features are directly concatenated at the channel dimension to form an initial joint feature vector of 2304 dimensions. This initial joint feature vector is then input into the perceptron fusion layer, where linear dimensionality reduction and nonlinear activation operations are performed through the hidden layer to output a 512-dimensional reduced joint feature vector. The hidden layer dimensionality reduction parameters are selected by evaluating the information redundancy in the feature vector using principal component analysis, and determining the final length by extracting the number of principal components with a cumulative feature contribution rate of over 95%. An example length of 512 dimensions is used.

[0060] The joint feature vector is input into the fully connected classification layer of the cross-modal model to calculate the distance between classification nodes. The specific steps for calculating the distance between classification nodes are as follows: In the computational space, for each input joint feature vector and the standard central feature vector of each product category pre-stored in the classification layer, the numerical difference in the corresponding feature dimension is calculated, and the dimension of this difference is the feature space distance. Next, each calculated difference is squared. Then, the squared results of all feature dimensions are summed. Finally, the square root of the sum is taken to obtain a scalar value representing the multidimensional similarity, and the dimension of this output is restored to a unified feature space distance unit. The standard central feature vector of each product category pre-stored in the classification layer is obtained as follows: During the cross-modal model training convergence phase, all valid samples belonging to the same product category in the training set are input into the network for forward propagation. The dimensionality-reduced joint feature vectors corresponding to these samples are extracted, and their feature mean values ​​in 512 dimensions are calculated to generate the mean central vector of the category, which is then stored.

[0061] After traversing and calculating the spatial distance of all categories, the label of the category node with the smallest distance value is extracted and output as the base identifier.

[0062] By applying a penalty factor to visual performance attenuation and enhancing weight features, a smooth transition in heterogeneous data acceptance is achieved, which contributes to the physical robustness and adaptability of feature fusion when vision is limited.

[0063] Furthermore, the operation location trajectory, storage rack space grid, and actual weight are input into the sliding alignment window to output the misaligned target features and the weight of the misaligned storage rack; the misaligned target features are matched with the basic identifiers to output the storage rack mapping relationship; corresponding to the above S3 step; the specific implementation process includes: The trajectory of the operating part and the actual weight are converted into a trajectory time series and a weight change time series, respectively. The trajectory time series and weight change time series are input into the dynamic time warping algorithm layer contained in the sliding alignment window to generate a spatiotemporal distance matrix. Singular nodes deviating from the matching path in the spatiotemporal distance matrix are extracted, and action conflict coordinates are output. These action conflict coordinates are mapped to the storage rack space grid, and misalignment occurrence grid nodes are output. Local weight jump values ​​of the misalignment occurrence grid nodes are extracted and output as the weight of the misaligned storage rack. Based on the trajectory of the operating part, the trajectory interaction features matched by the misalignment occurrence grid nodes are extracted and output as… The misaligned target features are described; the basic identifier is input into the attribute database of the Internet platform for querying, and the associated standard three-dimensional shape size features and standard packaging outline deformation features are output; the misaligned target features are parsed, and interactive volume features and interactive outline deformation features are output; the volume difference value between the interactive volume features and the standard three-dimensional shape size features is calculated; the deformation deviation value between the interactive outline deformation features and the standard packaging outline deformation features is calculated; a two-dimensional cross deviation matrix is ​​constructed based on the volume difference value and the deformation deviation value; the two-dimensional cross deviation matrix is ​​substituted into the preset position determination formula to perform the attribution determination calculation, and the storage rack mapping relationship is output.

[0064] Specifically, the process for generating the storage rack mapping relationship is as follows: The movement characteristics of each node in the trajectory of the operating part are arranged by time, and the Euclidean space displacement distance of the hand center node between adjacent time frames is calculated to form a one-dimensional trajectory time series representing the motion velocity scalar. Simultaneously, the load fluctuation values ​​in the continuously collected actual weight are filtered for baseline and then converted into a one-dimensional weight change time series.

[0065] The trajectory time series and weight change time series are input into the dynamic time warping algorithm layer contained in the sliding alignment window to generate a spatiotemporal distance matrix. A maximum time offset constraint parameter is introduced into the algorithm layer. This maximum time offset constraint parameter is selected by statistically analyzing the asynchronous time difference between a large number of user actions of picking up and placing goods and misplacing goods during the trial operation phase, and setting a delay range index that covers the time distribution characteristics of the main operating habits to constrain the search bandwidth.

[0066] The calculation process for generating the spatiotemporal distance matrix employs a dynamic programming strategy. To ensure consistent underlying dimensions and prevent computational failure, the trajectory time series and weight change time series are standardized fractionally before comparison calculations, strictly transforming them into dimensionless distribution data with a mean of zero and a variance of one. The specific steps for calculating the cumulative cost of nodes corresponding to any row and column coordinates in the matrix are as follows: First, calculate the absolute difference between the current scanned dimensionless trajectory feature point value and the current scanned dimensionless weight change feature point value. Next, extract the cumulative cost of the node's left-hand adjacent node, the upper-hand adjacent node, and the upper-left diagonal adjacent node in the matrix. Compare and filter these three extracted costs, finding the smallest one. Finally, add the calculated absolute difference to the selected smallest value to obtain the final cumulative cost of the node. This process is executed iteratively, ultimately outputting the complete spatiotemporal distance matrix.

[0067] In this embodiment, a strip-shaped global constraint mechanism is introduced, which sets a threshold value for the main diagonal based on the sensor sampling frequency to limit the search path. Out-of-bounds nodes are masked during cost calculation to avoid unreasonable time distortion matching, resulting in an optimized spatiotemporal distance matrix. Specifically, when inputting dimensionless trajectory feature points and dimensionless weight change feature points into the dynamic time warping algorithm layer for calculation, a Sakoe-Chiba global constraint band mechanism is introduced to prevent excessive time distortion that could force unrelated actions and weight jumps to be matched. In the specific calculation, the lowest sampling frequency (30 Hz) between the camera's 30 Hz sampling rate and the gravity sensor's 50 Hz sampling rate is extracted and multiplied by a preset maximum user action delay constant (set to 1.5 in this embodiment) to calculate the maximum allowable offset bandwidth parameter of the matrix's main diagonal (rounded down to 45 time frame nodes after calculation). Before calculating the cumulative cost of the spatiotemporal distance matrix row by row and column by column, it is first determined whether the absolute difference between the row index and column index of the current scan node is greater than this bandwidth parameter. If the difference is strictly greater than 45, the node is determined to belong to a non-physical matching zone caused by extreme time distortion. Its absolute difference is then forced to the system's maximum floating-point value, causing it to be naturally eliminated in the subsequent minimum cumulative cost selection of adjacent nodes. If the difference is less than or equal to 45, the original difference calculation and cumulative summation continue. By limiting the search corridor width, the final output is a spatiotemporal distance matrix that excludes false matching paths. This process reduces the algorithm's global search computational overhead and prevents misjudgments caused by extreme time distortion due to user pauses.

[0068] Singular nodes deviating from the matching path are extracted based on the spatiotemporal distance matrix, and their timestamps are output as action conflict coordinates. These action conflict coordinates are mapped to the storage rack space grid to locate the grid index of abnormal load fluctuations, and the grid node where the misalignment occurred is output.

[0069] The local weight jump value of the misalignment grid node is extracted and output as the weight of the misaligned storage rack. Simultaneously, based on the trajectory backtracking of the operation part, the image matching the misalignment grid node is extracted to extract the product's outline information. Specifically, using the binocular disparity map of the visual acquisition matrix and combining it with the previously truncated pixel bounding box, the 3D point cloud of the target product is reconstructed. The directed bounding box of this point cloud is extracted, and its length, width, and height outline dimensions are used as interactive volume features. At the same time, the discrete curvature extremum sequence of the product's outer contour in the 2D target local image is extracted as interactive contour deformation features, outputting the misaligned target features.

[0070] The basic identifier is input into a pre-set attribute database, and the associated standard three-dimensional shape dimension features and standard packaging outline deformation features are output. Specifically, the standard three-dimensional shape dimension features are the length, width, and height triplets of the product category in millimeters; the standard packaging outline deformation features are the allowable curvature variation parameters of the product category's packaging. After calculating the standard volume using the length, width, and height triplets, the volume difference between the interactive volume feature and the standard volume is calculated. The absolute difference between the curvature extremum of the interactive outline deformation features and the allowable curvature feature parameter in the standard packaging outline deformation features is also calculated, and the output is the deformation deviation value. These are combined to construct a two-dimensional cross-deviation matrix. The construction rules of the attribute database are as follows: during initialization, the length, width, and height parameters from the product specification are retrieved as the standard three-dimensional shape dimension features; simultaneously, the standard undamaged packaging is photographed from multiple angles to obtain the baseline outline, and its corresponding maximum allowable outline curvature variation threshold parameter is calibrated through mechanical extrusion deformation resistance testing, and then entered after being bound to the product ID.

[0071] The two-dimensional cross-deviation matrix is ​​substituted into the preset position determination formula to perform the attribution determination calculation. The calculation method of the attribution determination value is a linear weighted evaluation: to ensure dimensional consistency, the volume difference value is preprocessed into a dimensionless ratio of the actual absolute difference to the standard volume, while the deformation deviation value itself adopts a dimensionless structural similarity index. During the calculation, the preset dimensionless volume difference weight parameter is multiplied by the dimensionless volume difference value, and the preset dimensionless deformation weight parameter is multiplied by the dimensionless deformation deviation value. Then, the two independent product results are summed. Subsequently, a conditional judgment is made on the final summation result. If the dimensionless summation result is less than or equal to the preset dimensionless safety limit threshold, the attribution logic is determined to be true and a true value is output; otherwise, a false value is output.

[0072] The selection method for volume difference weight and deformation weight is as follows: A random forest decision algorithm is used to analyze the feature importance of the historical misplaced product dataset. Since volume features are relatively less affected by lighting and perspective distortion, a higher normalized importance value is assigned to the volume difference weight, and the remaining weight is assigned to the deformation weight, resulting in a weight sum of 1. The selection method for the safety threshold is as follows: Multiple random misplacement tests are conducted on the same product in an experimental environment. The distribution set of interactive volume and deformation deviations is collected, and the maximum comprehensive deviation value satisfying the set confidence interval is taken as the threshold limit.

[0073] Based on the judgment result, output the storage rack mapping relationship that confirms the destination of the goods.

[0074] By extracting singular nodes using the spatiotemporal distance matrix and performing deviation determination, the physical alignment of the motion trajectory with the gravity jump was achieved, thus restoring the true displacement path under abnormal interaction.

[0075] Further, the weight of the misaligned storage rack is input into the classification model to output a weight identifier, and the difference signal is output by comparing the basic identifier and the weight identifier; in response to the difference signal, the historical attenuation waveform sampled by the unmanned cabinet sensor terminal is received, input into the material seismic wave feature library, and the density result is output; combined with the storage rack mapping relationship, density result, and basic identifier, an inventory update instruction is output; corresponding to step S4 above; see [link to relevant documentation]. Figure 3 The specific implementation process includes: The weight of the misaligned storage rack is input into the discretization and quantization mapping layer of the classification model to perform interval classification operation and output the interval to which the weight belongs; the interval to which the weight belongs is input into the association lookup table of the classification model for query and output the weight identifier; the weight identifier and the basic identifier are subjected to interval assignment logic determination operation; in response to the identifier matching failure result output by the interval assignment logic determination, a hardware interrupt pulse level is generated and output as the difference signal.

[0076] Specifically, the generation process of the difference signal includes: The weight of the misaligned storage rack is input into the discretization and quantization mapping layer of the classification model. The specific calculation process for performing interval classification to obtain the lower limit of the interval to which the weight belongs is as follows: The collected weight of the misaligned storage rack, which is a floating-point number and has the dimension of grams, is directly divided by a preset quantization step size, which also has the dimension of grams. For the dimensionless quotient obtained from this division operation, the decimal part is rounded down to retain the integer part. Then, the dimensionless integer value after rounding is multiplied again by the original quantization step size with the dimension of grams to obtain the lower limit of the discrete interval to which the weight belongs, and the dimension of the output result is restored to grams.

[0077] The quantization step size parameter is selected by collecting the weight distribution map of all specifications of goods sold in the unmanned vending machine to extract the minimum weight difference span. This is combined with the maximum linear drift error amplitude of the gravity sensing unit under continuous working conditions. These two indicators are combined to prevent the weight value of a single item from crossing different quantization levels; in this example, it is set to 5.0 grams. Based on this value, the quantization mapping layer queries the corresponding discrete category in a pre-configured association lookup table. The construction process of the association lookup table is as follows: During the device initialization phase, the standard weight of all goods for sale in the database is read, interval calculations are performed using the same quantization step size, and the category labels of goods belonging to the same weight interval are binary-encoded and concatenated to store the mapping value of that interval. When the query matches a corresponding interval, the mapping value of that interval is output as the weight identifier. If the same interval contains multiple discrete categories, the weight identifier is represented by a composite binary string containing multiple category feature bits to facilitate subsequent bit-by-bit comparison with the basic identifier.

[0078] The extracted weight identifier and the base identifier are input together into a multi-level logic comparator to perform interval attribution logic determination. This multi-level logic comparator internally includes a semantic interval mapping subunit and a hardware XOR logic circuit unit: first, the mapping subunit retrieves the standard weight interval corresponding to the base identifier; then, the XOR logic circuit unit performs a strict bitwise XOR comparison between the binary string of the real-time weight interval identifier and the binary string of the standard interval identifier. If any bit is not 0, a conflict is determined, and the identifier matching failure result is output.

[0079] In response to a failed match, a hardware interrupt pulse level is generated as a difference signal output. The duration parameter of the hardware interrupt pulse level is selected by combining the minimum pulse width capture and hold requirements of the external interrupt controller of the main microprocessor with an additional delay margin to resist voltage jitter, for example, set to 100 milliseconds.

[0080] By performing bitwise state comparisons using an XOR logic comparator, logical discrepancies between multi-source fusion results and basic weight attributes are captured, providing hardware triggering conditions for deep physical verification.

[0081] Using the difference signal as the extraction trigger condition, the acoustic waveform at the moment of weight change is extracted from the audio buffer queue, and the historical attenuation waveform is output. The historical attenuation waveform is input into a fast Fourier transform to perform a time-frequency domain transformation operation, and the frequency domain damping coefficient sequence is output. The frequency domain damping coefficient sequence is input into the material seismic feature library to perform an Euclidean distance comparison operation, and the material category with Euclidean distance matching is selected and output as the density result. The density result is used to overwrite the original material attribute corresponding to the basic identifier, and the verified identifier is output. The verified identifier is bound to the grid physical coordinates indicated by the storage rack mapping relationship, and the inventory update instruction is output to the Internet platform.

[0082] Specifically, the process for generating an inventory update instruction is as follows: In response to the trigger command of the difference signal, the main control module extracts the acoustic waveform at the moment of weight change from the audio buffer queue and outputs the historical attenuation waveform.

[0083] In this embodiment, a pure ambient audio segment before the cut-off window is retrieved as a noise template. Frequency domain spectral subtraction is performed on the acoustic waveform, and a signal-to-noise ratio constraint parameter is set to dynamically calculate the over-subtraction factor to filter out background noise, outputting a clean acoustic waveform. Specifically, after extracting the acoustic waveform, an audio segment of 1.0 second in length, between 1.5 and 0.5 seconds before the weight abrupt change, is extracted from the audio buffer queue as a steady-state noise template. The acoustic waveform and the steady-state noise template are both windowed using Hamming windows, with a frame length of 256 sampling points and a frame shift of 128 sampling points. A short-time Fourier transform is performed on the windowed data to extract the power spectral density matrix of the acoustic waveform and the average power spectral density matrix of the steady-state noise template. In each frequency band, the power spectral density of the acoustic waveform is subtracted from the average power spectral density of the steady-state noise multiplied by the over-subtraction factor. The over-attenuation factor is dynamically calculated based on the signal-to-noise ratio (SNR): when the SNR of the frame is greater than 20 dB, the over-attenuation factor is set to 1.0; when the SNR is less than 5 dB, the over-attenuation factor is set to 3.0, and linear interpolation is performed in other intervals. If the power spectral density after subtraction is negative, it is forcibly replaced with 1% of the power spectral density of the original acoustic waveform in that frequency band. The processed power spectral density matrix is ​​combined with the original phase information and reconstructed into a one-dimensional time-domain audio through short-time inverse Fourier transform, and the output is the historical attenuation waveform. The upper and lower limits of the SNR are selected according to the mean value of the impact sound signal that is not masked and the critical value that is completely submerged in the environmental noise floor test, respectively; the over-attenuation factor is selected according to the acoustic denoising experience curve, taking into account both the physical energy retention rate under high SNR and the maximum denoising factor under low SNR; the fallback replacement ratio is selected as the minimum numerical boundary preset to prevent abnormal zero values ​​from occurring in the subsequent logarithmic attenuation rate calculation. The above process effectively suppressed the pollution of sound wave characteristics by background noise in the shopping mall environment and mechanical refrigeration resonance, ensuring the purity of the damping coefficient.

[0084] The selection method for the time window parameters before and after capturing the historical attenuation waveform is as follows: by physically measuring the diffusion and dissipation period of the sound wave energy generated when common packaging materials fall onto the test shelf, the lower limit of the time window can be set to include the high-frequency establishment stage before the impact, and the upper limit can be set to cover the tail oscillation dissipation stage.

[0085] The historical attenuation waveform is input into the Fast Fourier Transform (FFT) to perform a time-frequency domain transformation. The selection method for the frequency band sampling dimension parameter is as follows: based on the time decimation operation architecture constraints of the FFT core algorithm, and considering the frequency band resolution requirements for distinguishing high-frequency physical damping characteristics, a comprehensive setting is made.

[0086] Subsequently, a frequency-domain damping coefficient sequence characterizing the attenuation properties of the sound wave is output. The extraction and calculation method of the frequency-domain damping coefficient sequence at a specific frequency point is as follows: First, the dimensionless logarithmic attenuation rate index of the acoustic waveform after conversion at that frequency point is extracted; then, the resonant angular frequency value corresponding to that frequency is calculated, with the dimension of radians per second, and multiplied by the acoustic recording observation time length value with the dimension of seconds to obtain the product result representing the cumulative phase, with the dimensionless dimension of radians; finally, the previously obtained signal logarithmic attenuation rate is divided by the radian product result, and the resulting pure dimensionless quotient value is recorded as the damping coefficient representing the physical dissipation characteristics of the material.

[0087] The frequency domain damping coefficient sequence is input into the material seismic feature library. The detailed calculation steps for performing Euclidean distance comparison in the feature library are as follows: For the extracted actual dimensionless frequency domain damping coefficient sequence vector and the dimensionless damping coefficient sequence vector stored for a certain material template in the feature library, a traversal comparison is performed one by one across all the divided frequency band dimensions. That is, the dimensionless feature difference is obtained by subtracting the damping coefficient of the corresponding template frequency point from the damping coefficient of the actual detection frequency point; then, the difference is multiplied by itself and squared; then, the squared results of all sampled frequency band dimensions are added together; finally, the square root mathematical operation is performed on the summed sum, and the final dimensionless scalar value obtained by this operation is used as the comparison distance of the acoustic features.

[0088] The process of establishing the material vibration feature library is as follows: During the initialization phase of the unmanned cabinet system, for various preset standard packaging materials (such as glass, PET plastic, aluminum cans, cardboard boxes, etc.), a robotic arm performs drop tests at standardized heights on each storage shelf within the cabinet; acoustic samples generated by each collision are collected, and the average damping coefficient at each frequency point is extracted through the same short-time Fourier transform and attenuation rate calculation process as in the real-time sensing phase. These are then combined and spliced ​​into a standard dimensionless damping coefficient sequence vector, and this vector is bound to the corresponding material label to form the material vibration feature library. To eliminate the system error introduced by the differences in acoustic transmission characteristics of different shelf panels (such as metal / glass bottom panels), the material vibration feature library adopts a layered independent library construction mode; when performing subsequent Euclidean distance comparison operations, the damping templates that belong to the same shelf layer as the coordinates of the conflict with the preceding action are limited to comparison.

[0089] After traversing all known material templates in the feature library, select the corresponding material category with the smallest comparison distance and output its label as the density result.

[0090] Density results serve as the acoustic physical basis for identifying the item. Using these density results, the original material properties attached to the base identifier are checked for their status. When the physical material properties do not match, the category label corresponding to the base identifier is overwritten, the replacement status is marked, and the verified identifier is output.

[0091] This verified identifier is then bound to the grid physical coordinates indicated by the storage rack mapping relationship confirmed at the third level prior to the verification. Based on this, the coordinate position is recorded, the product tag is associated, the database status is updated, and an inventory update command containing a complete timeline and verification results is output to the control terminal.

[0092] By extracting frequency domain damping from the acoustic waveform captured from the cache and comparing the material density at the acoustic resonance level, the objectivity of the physical identification results under extreme interference such as equal weight and undamaged appearance is ensured.

[0093] The underlying architecture details of the network model and the complete training and solidification process involved in this embodiment are as follows: The core model involved in this method is constructed by combining multiple network branches. Specifically, the feature network uses a deep residual network (such as ResNet-50) as the backbone network to perform multi-scale feature convolution to extract visual features, and connects a feature pyramid network based on the Mask R-CNN architecture in parallel as an instance semantic segmentation layer to extract masks; the cross-modal model is mainly composed of multi-layer perceptron stacks, including scalar feature embedding layers for weight upscaling, fusion dimensionality reduction units based on feature channel concatenation, and fully connected layers to achieve alignment and classification dimensionality reduction of heterogeneous features.

[0094] In the model training preparation phase, a multimodal aligned dataset containing image / text / weight pairs was pre-constructed, covering different occlusion patterns of at least 200 categories of daily-sold goods. The data collection process encompassed various ambient lighting conditions (such as normal, low light, and partial reflection) in real-world application scenarios within unmanned vending machines, different product sizes, and diverse hand interaction occlusion states (light, heavy, and complete coverage). Each set of samples was simultaneously bound to visual image data and corresponding weight load jump values. Simultaneously, pixel-level semantic masks were manually calibrated for product outlines and hand regions in the images, and the true identity category (One-hot encoding) and standard physical attributes of the products were set as hard labels.

[0095] The constructed multimodal alignment dataset is randomly shuffled and divided into training, validation, and test sets according to a preset ratio (e.g., 8:1:1). The training set occupies the largest proportion and is dedicated to the gradient calculation and iterative update of the underlying weight matrix and bias parameters during backpropagation. The validation set does not participate in weight updates; it is only used to monitor the model's convergence status after each training epoch, for dynamic hyperparameter tuning, and to trigger an early stopping mechanism to prevent overfitting on the training set. The test set is strictly isolated and is only used after the model has been fully solidified to conduct a final objective evaluation of its generalization ability and overall accuracy on unknown data.

[0096] During the training phase, image tensors and their corresponding weight scalars from the training set are input in batches into the unfixed model to perform forward propagation.

[0097] A global joint loss function is constructed for different network output branches: for the basic output labels (classification task), the cross-entropy loss function is used to calculate the deviation between the predicted class and the true label; for the output instance object mask (segmentation task), a combination of binary cross-entropy and Dice loss functions is used to evaluate the fitting accuracy of the occlusion ratio edge; for continuous feature mappings such as weight, the mean squared error loss function is used. The global joint loss function is a weighted sum of the above three loss functions. In this embodiment, the weight of the cross-entropy loss function used for basic label classification is set to 0.6, the weight of the Dice loss function used for occlusion mask prediction is set to 0.3, and the weight of the mean squared error loss function used for weight feature alignment is set to 0.1 to ensure that the network prioritizes recognition accuracy.

[0098] Subsequently, the gradient of the joint loss function with respect to the parameters of each layer was calculated using the backpropagation algorithm. For training parameter settings, the Adam optimizer was selected, with an initial learning rate of 0.001, and a cosine annealing decay strategy was used to ensure smooth convergence in later stages; the batch size was set to 32, and the weight decay coefficient was set to 10. -4The maximum number of training iterations is set to 200. An early stopping mechanism is triggered when the joint loss function on the validation set stops decreasing for 15 consecutive iterations. Finally, the network weights of the nodes that achieve the highest accuracy on the validation set are extracted, solidified, and deployed in the edge computing module of the unmanned vending machine's sensing terminal.

[0099] This invention provides an intelligent identification method for the status of goods in unmanned vending machines based on multi-source sensing data. It extracts confidence coefficients including visual and occlusion ratios, combines weight features with a cross-modal model to output basic identifiers, and extracts misalignment features under spatiotemporal alignment of trajectory and weight. Finally, it verifies acoustic density using transient attenuation waveforms based on response difference signals, establishing a global verification closed loop that links multi-dimensional physical features. This achieves tight logical interlocking between visual appearance, weight abrupt changes, and acoustic damping. When encountering complex physical occlusion, it maintains sensing stability through adaptive weight adjustment; when irregular picking and placing occurs, it locks the true displacement through spatiotemporal sliding alignment; and when facing extreme states where the appearance is intact but the core material has changed, it directly triggers the underlying decision through transient acoustic wave physical properties. This strengthens the objective complementarity of multi-source sensing data, establishes a robust physical state defense line, and ensures the overall effectiveness of the final state determination result.

[0100] Example 2 This embodiment uses the example of an intelligent identification method for the status of goods in an unmanned vending machine based on multi-source sensing data, applied to a situation where the arrangement of goods in the vending machine is disrupted, to elaborate on the execution process of each step of perception and identification.

[0101] During the data acquisition phase, the gravity sensing matrix at the bottom continuously collects actual load data at a frequency of 50 Hz, while the binocular stereo camera at the top acquires internal global images at a rate of 30 frames per second. The microphone array simultaneously records acoustic waveforms and stores them in a circular buffer queue. Meanwhile, a skeletal point algorithm extracts the hand movement trajectory of the interacting user in real time and uses a 5-centimeter-side vertical shelving topology network as the physical mapping benchmark.

[0102] When the user picks up a 500-gram glass bottle of juice from the shelf, the difference between the actual load and the baseline weight instantly exceeds the set 10-gram threshold. The digital comparator immediately outputs a pulse signal, which the processing center uses to extract the shelf's sudden change coordinates. This is then converted into a two-dimensional pixel bounding box using a perspective mapping matrix, and the global image is sliced ​​and cropped to output a local image of the target. At this moment, the user's hand largely covers the juice bottle. After processing by the semantic segmentation layer, it is calculated that the item is occluded by as much as 75%. After negative exponential operation using a nonlinear decay function, the visual confidence coefficient drops significantly. The cross-modal network uses the calculated complementary values ​​to correspondingly increase the weight of the gravity feature, initially inferring that the basic identifier of the item is a glass bottle of juice.

[0103] Subsequently, instead of removing the item, the staff randomly placed it into an empty area on another floor, disrupting the original display of the goods. To track this misplacement, the computation module synchronously input the extracted hand trajectory and weight change time series into a sliding alignment window. A dynamic time warping algorithm was used to eliminate the time difference between the hand movement and the item's fall, generating a spatiotemporal distance matrix and identifying singular nodes deviating from the normal path, accurately locating the grid node where the misplacement occurred. A weighted calculation combining volume differences and deformation deviation values ​​confirmed the new physical attribution of the randomly placed juice.

[0104] To ensure absolute accuracy in state tracking, the classification model converts the weight collected from the misaligned grid into binary weight identifiers, which are then synchronously input into the underlying XOR logic comparator along with the original identifiers. Due to slight stress deformation in the base plate of the vacant area, a boundary offset occurred during weighing, triggering a difference signal. The main control unit immediately traced back the audio queue, capturing the attenuation waveform at the moment the juice touched the shelf. A fast Fourier transform was used to extract the frequency domain damping coefficient sequence, and after comparison with the material feature library, it was confirmed that the sound wave dissipation characteristics perfectly matched the thick glass material. Finally, the control logic eliminated the interference from underlying physical errors, re-bound the juice labels to the new, disordered coordinates, and output an inventory update command.

[0105] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for intelligent identification of the status of goods in unmanned vending machines based on multi-source sensing data, characterized in that, include: Receives data from the unmanned cabinet's sensor terminal, including the actual weight, reference weight, global image inside the cabinet, trajectory of the operating parts, and storage rack space grid. When the difference between the actual weight and the reference weight is greater than the difference threshold, extract the sudden change coordinates of the storage rack; truncate the global image inside the cabinet according to the sudden change coordinates of the storage rack and output the target local image; input the target local image into the feature network to output visual features and occlusion ratio, convert the occlusion ratio to output the confidence coefficient; convert the actual weight to output the weight feature; input the weight feature, visual feature, and confidence coefficient into the cross-modal model to output the basic identifier; Input the trajectory of the operating part, the spatial grid of the storage rack, and the actual weight into the sliding alignment window to output the misaligned target features and the weight of the misaligned storage rack; match the misaligned target features with the basic identifier to output the storage rack mapping relationship; Input the weight of the misaligned storage rack into the classification model to output a weight label, and compare the basic label with the weight label to output a difference signal; in response to the difference signal, receive the historical attenuation waveform sampled by the unmanned cabinet sensor terminal, input it into the material seismic wave feature library to output the density result; combine the storage rack mapping relationship, density result, and basic label to output an inventory update instruction.

2. The intelligent identification method for the status of goods in unmanned vending machines based on multi-source sensing data according to claim 1, characterized in that, The specific acquisition process for actual weight, reference weight, global image inside the cabinet, trajectory of operating parts, and storage rack space grid includes: periodically polling the storage rack load value using the gravity sensing matrix included in the unmanned cabinet sensing terminal and storing it as the reference weight; acquiring the current load value using the gravity sensing matrix in interactive mode and outputting the actual weight; acquiring a continuous video frame sequence using the visual acquisition matrix included in the unmanned cabinet sensing terminal and outputting the global image inside the cabinet; acquiring acoustic waveforms using the microphone array included in the unmanned cabinet sensing terminal and storing them in the audio buffer queue; parsing the action node coordinate sequence in the global image inside the cabinet using an image skeleton point extraction algorithm and outputting the trajectory of the operating parts; retrieving the three-dimensional shelf physical coordinate topology model, extracting grid parameters, and outputting the storage rack space grid.

3. The intelligent identification method for the status of goods in unmanned vending machines based on multi-source sensing data according to claim 1, characterized in that, The specific generation process of the target local image includes: calculating the numerical difference between the actual weight and the reference weight; locating the gravity sensing matrix target sensing unit that generates the numerical difference in response to an interruption condition where the numerical difference exceeds the difference threshold; extracting the device number of the target sensing unit; inputting the device number into a grid coordinate mapping table for querying, extracting the three-dimensional position coordinates and outputting them as the storage rack mutation coordinates; inputting the storage rack mutation coordinates into a spatial perspective mapping matrix and outputting two-dimensional pixel bounding box parameters; performing pixel array slicing operation on the global image inside the cabinet using the two-dimensional pixel bounding box parameters and outputting an initial truncated image stream; extracting the static background pixel template of the storage rack spatial grid mapping; performing pixel contrast subtraction operation on the initial truncated image stream and the static background pixel template and outputting a preliminary foreground pixel set; removing hand pixels from the preliminary foreground pixel set based on the hand mask mapped by the trajectory of the operation part and outputting a differentiated foreground pixel set; reconstructing the differentiated foreground pixel set into an independent pixel tensor and outputting the target local image.

4. The intelligent identification method for the status of goods in unmanned vending machines based on multi-source sensing data according to claim 1, characterized in that, The specific generation process of the confidence coefficient and visual features includes: inputting the target local image into the convolutional feature extraction layer of the feature network to perform multi-scale feature convolution operation, and outputting the visual features; inputting the target local image into the instance semantic segmentation layer of the feature network to output the hand object mask and the target object mask; extracting the number of overlapping pixels between the area of ​​the hand object mask and the area of ​​the target object mask; extracting the quotient of the number of overlapping pixels divided by the total number of pixels in the target object mask, and outputting the occlusion ratio; inputting the occlusion ratio into a nonlinear decay function to extract the inverse weight value; inputting the inverse weight value into a normalization processing layer to perform a range limitation operation, and outputting the confidence coefficient.

5. The intelligent identification method for the status of goods in unmanned vending machines based on multi-source sensing data according to claim 1, characterized in that, The specific generation process of the basic identifier includes: inputting the actual weight into the scalar feature embedding layer of the cross-modal model and outputting the weight feature; assigning the confidence coefficient as a penalty factor; inputting the visual feature, the weight feature, and the penalty factor into the feature adjustment layer, using the penalty factor to perform channel dimension numerical attenuation operation on the visual feature and outputting a corrected visual feature, and using the complementary value of the penalty factor to perform splicing dimension numerical enhancement operation on the weight feature and outputting a corrected weight feature; inputting the corrected visual feature and the corrected weight feature into the perceptron fusion layer to perform feature-level splicing operation and outputting a joint feature vector; inputting the joint feature vector into the fully connected classification layer to perform classification node distance calculation operation, extracting the classification node labels that match the spatial distance and outputting them as the basic identifier.

6. The intelligent identification method for the status of goods in unmanned vending machines based on multi-source sensing data according to claim 1, characterized in that, The specific generation process of the storage rack mapping relationship includes: converting the trajectory of the operating part and the actual weight into a trajectory time series and a weight change time series, respectively; inputting the trajectory time series and weight change time series into the dynamic time warping algorithm layer contained in the sliding alignment window to generate a spatiotemporal distance matrix; extracting singular nodes that deviate from the matching path in the spatiotemporal distance matrix and outputting action conflict coordinates; mapping the action conflict coordinates to the storage rack space grid and outputting the misalignment occurrence grid node; extracting the local weight jump value of the misalignment occurrence grid node and outputting it as the weight of the misaligned storage rack; and extracting the track matching the misalignment occurrence grid node based on the operating part trajectory. The interaction features are analyzed, and the output is the misaligned target feature. The basic identifier is input into the attribute database of the Internet platform for querying, and the associated standard three-dimensional shape size features and standard packaging outline deformation features are output. The misaligned target features are analyzed, and the interactive volume features and interactive outline deformation features are output. The volume difference value between the interactive volume features and the standard three-dimensional shape size features is calculated. The deformation deviation value between the interactive outline deformation features and the standard packaging outline deformation features is calculated. A two-dimensional cross deviation matrix is ​​constructed based on the volume difference value and the deformation deviation value. The two-dimensional cross deviation matrix is ​​substituted into the preset position determination formula to perform the attribution determination calculation, and the storage rack mapping relationship is output.

7. The intelligent identification method for the status of goods in unmanned vending machines based on multi-source sensing data according to claim 1, characterized in that, The specific generation process of the difference signal includes: inputting the weight of the misaligned storage rack into the discretization and quantization mapping layer included in the classification model to perform interval classification operation, and outputting the interval segment to which the weight belongs; inputting the interval segment to which the weight belongs into the association lookup table included in the classification model for querying, and outputting the weight identifier; performing interval assignment logic determination operation on the weight identifier and the basic identifier; and generating a hardware interrupt pulse level and outputting it as the difference signal in response to the identifier matching failure result output by the interval assignment logic determination.

8. The intelligent identification method for the status of goods in unmanned vending machines based on multi-source sensing data according to claim 1, characterized in that, The specific generation process of the inventory update instruction includes: using the difference signal as the extraction trigger condition, extracting the acoustic waveform at the moment of weight change from the audio buffer queue, and outputting the historical attenuation waveform; inputting the historical attenuation waveform into a fast Fourier transform to perform a time-frequency domain conversion operation, and outputting a frequency domain damping coefficient sequence; inputting the frequency domain damping coefficient sequence into the material seismic feature library to perform an Euclidean distance comparison operation, selecting the material category that matches the Euclidean distance and outputting the density result; using the density result to overwrite the original material attribute corresponding to the basic identifier, and outputting the verified identifier; binding the verified identifier to the grid physical coordinates indicated by the storage rack mapping relationship, and outputting the inventory update instruction to the Internet platform.