Vending Cabinet Product Recognition Algorithm Based on Artificial Intelligence Model

Through the sales container product recognition algorithm based on artificial intelligence models, interactive graphs are constructed and graph neural networks are used to predict consumer purchasing intentions, which solves the problem of real-time optimization of product display in the existing technology, and achieves high-precision purchasing intention prediction and product display optimization.

CN119904792BActive Publication Date: 2025-06-03ZHEJIANG HI CONVENIENCE NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510372893.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-03
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The prior art cannot predict consumer purchasing intentions in real time when optimizing the display of vending container products, and does not fully utilize consumer behavior characteristics for dynamic optimization.

Method used

Using a sales container product recognition algorithm based on artificial intelligence models, we use the product area and consumer hand area to construct interactive graphs and use graph neural network to capture global space-time dependence, and output the potential purchase intention index of each product to achieve optimization of product display.

Benefits of technology

Based on single-camera video data, it realizes that high-precision prediction of consumer purchasing intentions, dynamically optimizes product display, and improves user purchasing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904792B_ABST
    Figure CN119904792B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of vending cabinet commodity recognition, and specifically relates to a vending cabinet commodity recognition algorithm based on an artificial intelligence model. Based on single-camera video data, after continuous frame-by-frame feature tracking, a consumer interaction graph is constructed, and then local feature aggregation is performed using a graph neural network. Subsequently, by capturing global spatio-temporal dependencies, the potential purchase intention index of each commodity is predicted. The interaction graph is used to structurally describe consumer behavior, and the detailed features of each interaction are retained. Multilevel information fusion is achieved through graph convolution and Transformer. Under the premise of relying only on single visual data, high-precision behavior modeling and purchase intention prediction are realized, so as to iteratively optimize the commodity layout, preferentially allocate high-purchase-intention commodities to more reasonable display positions, realize the display configuration of vending cabinet commodities, and also improve the user purchase experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vending cabinet product recognition, and particularly to a vending cabinet product recognition algorithm based on an artificial intelligence model. Background Art

[0002] In the modern retail environment, the display method of vending cabinet products has an important impact on consumers' purchase decisions. How to optimize the product placement strategy to maximize the sales conversion rate has become an important research direction in the intelligent retail system. Currently, intelligent retail mainly relies on technologies such as product recognition, consumer behavior analysis, and inventory management, but there are still many challenges in practical applications, especially in accurately obtaining the purchase intention index of consumers and adjusting the product display plan based on this, where there are obvious deficiencies.

[0003] Currently, the optimization of product display mainly relies on empirical rules and sales data analysis. For example, vending cabinets usually adjust the shelf display according to the historical sales data of best-selling and slow-selling products, combined with marketing strategies (such as promotional activities, bundle sales). This method has the following problems:

[0004] Lag: The optimization method based on sales data can only analyze the product sales after the fact, and cannot predict consumers' purchase intentions before their decisions, thus lacking the ability to adjust in real time;

[0005] Failure to consider consumer behavior factors: The current optimization scheme is mainly based on sales volume data and does not fully utilize the behavioral characteristics of consumers during the shopping process (such as touch interaction, etc.) to dynamically optimize the display. Summary of the Invention

[0006] In view of the above-mentioned shortcomings of the prior art, the present invention provides a vending cabinet product recognition algorithm based on an artificial intelligence model, which can effectively solve the problem in the prior art that it is impossible to dynamically optimize the vending cabinet product display in combination with the behavioral characteristics of consumers.

[0007] To achieve the above objectives, the present invention is realized through the following technical solutions:

[0008] The present invention provides a vending cabinet product recognition algorithm based on an artificial intelligence model, including the following steps:

[0009] Detect the product area and the consumer's hand area in the product image;

[0010] Construct the interaction between the consumer and the product into a graph structure to form an interaction graph, including:

[0011] Calculate the intersection over union, the curvature of the hand movement trajectory, and the acceleration respectively in the candidate interaction events, where:

[0012] Calculate the curvature of the hand movement trajectory using the fitted curve and its derivative;

[0013] Introduce the Bayesian estimation method, combine the acceleration of the previous frame and the observed data of the speed change of the current frame, establish the prior distribution and the observation probability, realize the maximum a posteriori estimation of the acceleration, and based on the Kalman filtering method, dynamically adjust the process noise covariance and the observation noise covariance according to the prediction error to obtain the updated acceleration;

[0014] Determine the comprehensive score based on the intersection over union, the curvature of the hand movement trajectory, and the updated acceleration, and judge the interaction event of the commodity occurring in the image frame;

[0015] Take the interaction event as a node in the graph, construct edges using the time interval and the spatial distance, and form a graph structure through the nodes and edges to realize the construction of the interaction graph;

[0016] Input the interaction graph into the graph neural network for node feature aggregation, output the potential purchase intention index of each commodity by capturing the global spatio-temporal dependence, and use this purchase intention index to optimize the commodity display.

[0017] Further, when obtaining the consumer's hand region, determine the optical flow matching points by tracking the feature points, as follows:

[0018] Select the set of feature points in the image frame , It is expressed that the coordinate of the -th feature point in the image frame is , It represents the total number of feature points selected in the image frame;

[0019] Solve the motion displacement of each feature point in the next frame ;

[0020] Assume that the motion of the feature point remains consistent within its local neighborhood window , and minimize the error within the local neighborhood window:

[0021]

[0022] Among them, represents the neighborhood window centered on the feature point, , respectively represent the displacements of the feature point in the and directions;

[0023] By finding the optimal motion displacement, obtain the optical flow matching points and form matching point pairs;

[0024] Input the feature point matching set;

[0025] Calculate the candidate transformation matrix;

[0026] Calculate the transformation error of all optical flow matching points, and re-determine the optical flow matching points in combination with the transformation error.

[0027] Further, the method for determining the curvature of the hand motion trajectory is as follows:

[0028] Suppose the sequence of hand center positions in consecutive multiple frames is ;

[0029] denotes the coordinates of the hand center in the image frame at time , denotes the starting time, denotes the number of consecutive frames collected;

[0030] Perform robust curve fitting:

[0031]

[0032] where, denotes the fitted curve, denotes the coordinate function of the fitted curve, denotes the second derivative of the fitted curve, denotes the smoothing parameter;

[0033] Calculate the curvature of the hand motion trajectory:

[0034]

[0035] where, , respectively denote , with respect to derivatives, , respectively denote , second derivatives.

[0036] Further, the method for determining the maximum a posteriori estimation of acceleration is as follows:

[0037] Determine the prior estimation of the current frame, denotes the acceleration estimated in the previous frame, denotes the process noise covariance;

[0038] Observation probability :

[0039]

[0040] where, denotes the observation noise covariance, denotes the time interval between frames, represents a normal distribution;

[0041] Calculate the posterior acceleration distribution , represents the normalization constant;

[0042] Determine the maximum a posteriori estimate:

[0043]

[0044] where, represents the acceleration value obtained by the maximum a posteriori estimate at time , , respectively represent and the inverse values of.

[0045] Furthermore, the method for obtaining the updated acceleration is:

[0046] Determine the predicted acceleration of the current frame , represents the acceleration value calculated by the maximum a posteriori estimate (MAP) at time ;

[0047] Update the covariance: , represents the acceleration estimation error covariance predicted to time ,

[0048] Calculate the Kalman gain ;

[0049] Calculate the updated hand acceleration , represents the acceleration observation value obtained by observation of the current frame.

[0050] After the interaction event is generated, the method for forming an interaction graph is:

[0051] Each interaction event generates an interaction event node ;

[0052] If the interaction event node and the interaction event node respectively correspond to interaction events within adjacent frames, then construct an edge , and construct an interaction graph based on the interaction event nodes and edges.

[0053] The interaction event node , which includes:

[0054] Product category, product location, hand touch duration, hand movement speed.

[0055] Furthermore, the features of the edge include:

[0056] Node time interval , 、 respectively represent the occurrence times of interaction event nodes 、 .

[0057] Spatial distance , 、 respectively represent the coordinates of interaction event node , 、 respectively represent the coordinates of interaction event node .

[0058] Furthermore, the method for inputting the interaction graph into the graph neural network for node feature aggregation is as follows:

[0059] Let the nodes participating in feature aggregation in the interaction graph be , and the initial feature be ;

[0060] The update formula of graph convolution is:

[0061]

[0062] Among them, represents the feature of the node at the th layer, represents the set of neighbor nodes of node , and respectively represent the weight matrix and bias at the th layer, represents the activation function.

[0063] Furthermore, the method for outputting the potential purchase intention index of goods by capturing global spatio-temporal dependencies is as follows:

[0064] Calculate the purchase intention index for each node :

[0065]

[0066] Among them, represents the purchase intention index of the th node, represents in the th feature vector of the node, 、 respectively represent the weight matrix and bias of the fully connected layer.

[0067] The technical solution provided by the present invention has the following beneficial effects compared with the known prior art:

[0068] Based on single-camera video data, construct a consumer interaction graph through continuous inter-frame feature tracking, then use a graph neural network for local feature aggregation, and then capture global spatio-temporal dependencies to predict the potential purchase intention index of each commodity. Use the interaction graph to structurally describe consumer behavior and retain the detailed features of each interaction (such as commodity category, location, hand touch duration, and movement speed). Achieve multi-level information fusion through graph convolution and Transformer, and realize high-precision behavior modeling and purchase intention prediction on the premise of relying only on single visual data, so as to iteratively optimize the commodity layout, give priority to allocating high-purchase-intention commodities to more reasonable display positions, realize the display configuration of commodities in the vending cabinet, and also improve the user's purchase experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0070] Figure 1 It is a schematic diagram of the overall method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0072] With the rapid development of intelligent retail and unmanned vending technologies, automated vending cabinets have become an important part of the retail field. Traditional vending cabinet management mainly relies on the following methods:

[0073] Single visual recognition:

[0074] Static images are captured using a single camera or multiple cameras, and convolutional neural networks (CNNs) are employed for object detection and recognition. This method can achieve good results under ideal lighting and unobstructed conditions, but in practical applications, it often faces problems such as uneven lighting, occlusion, and limited viewing angles, resulting in a decrease in the recognition rate.

[0075] Multi-modal sensor fusion

[0076] Some systems compensate for the deficiencies of single sensors by fusing data from multiple sensors such as vision, weight, and RFID. However, the core drawbacks of this solution are high technical complexity, increased costs, and the possibility of information redundancy or noise interference during the fusion of multi-sensor data, which reduces the overall robustness and maintainability of the system.

[0077] In unmanned vending cabinets, product recognition is not only for inventory counting, but more importantly, by obtaining real-time interaction data between consumers and products, the following goals can be achieved:

[0078] Consumer behavior analysis:

[0079] By recording every product contact and operation of consumers in front of the vending cabinet, the purchase path, hesitation behavior, preferences, etc. of consumers can be analyzed in detail, providing data support for subsequent personalized recommendations, promotional strategies, and product display optimization.

[0080] Display optimization and marketing strategy adjustment:

[0081] Based on consumer behavior data (such as potential purchase intention, interaction frequency, stay duration, etc.), dynamically adjust the display position plan of products in the vending cabinet to achieve reasonable display of vending cabinet products, ensure the sale of normal goods, and also improve the user experience of purchasing products (by reasonably displaying products with high purchase intention, which is conducive to users quickly finding and purchasing products).

[0082] However, in practical applications, the existing technologies mainly face the following challenges:

[0083] Complexity of the environment and operations:

[0084] Factors such as internal lighting, occlusion, and multi-angle viewing angles in the vending cabinet affect the accuracy of visual recognition; the diversity and continuity of hand movements during the operation of consumers also pose difficulties for behavior modeling.

[0085] In response to the above problems, the solution provided by the present invention is:

[0086] Robust tracking of consumers' hand movements to improve robustness under occlusion and lighting change conditions; by defining interaction events and constructing spatio-temporal interaction graphs, each contact between a consumer's hand and a product is converted into a graph node, and edges are used to record spatio-temporal correlation information. This graph structure is intuitive and has stronger interpretability, and can effectively reflect consumers' complex decision-making paths; combined with real-time PWI data, the simulated annealing algorithm is used to guide product display adjustment, realizing dynamic optimization of product exposure and simultaneously improving the convenient experience of users purchasing products in vending cabinets.

[0087] The present invention will be further described below in conjunction with embodiments.

[0088] Embodiment 1 (refer to Figure 1 ): A vending cabinet product recognition algorithm based on an artificial intelligence model, including the following steps:

[0089] Step 1: Detect all product areas and consumers' hand areas from the product images captured by a single camera, and perform inter-frame tracking to ensure the continuity and stability of the detection results, specifically as follows:

[0090] Product area detection:

[0091] Obtain an image frame at time in the product image, and determine the bounding box and class information by a pre-trained object detection model (such as YOLOv5 or EfficientDet):

[0092] represents the bounding box (coordinates) of the -th product, represents the class label of the -th product (such as cola, potato chips, etc.), represents the number of products detected in the image frame .

[0093] Consumers' hand area detection:

[0094] Based on the object detection algorithm, detect the consumers' hand area in the image frame to obtain the hand bounding box , which will not be elaborated here;

[0095] Track the feature points in the consumers' hand area, adopt the Lucas-Kanade optical flow algorithm, and eliminate abnormal matches through RANSAC to robustly track the hand movement, ensuring that stable two-dimensional trajectory information can still be obtained in the presence of partial occlusion and lighting changes, providing accurate spatio-temporal position information for subsequent construction of the interaction graph, specifically as follows:

[0096] Select a set of feature points in the image frame , denote the th feature point in the image frame (the feature points are extracted using Shi-Tomasi corner detection or Harris corner detection, and corner detection will select points with significant gradient changes in the image, such as the edges and corners of goods), and the coordinates are , denote the total number of feature points selected in the image frame;

[0097] Use the Lucas-Kanade algorithm to solve the motion displacement of each feature point in the next frame , , respectively denote the displacements of the feature point in the and directions. Thus, it is determined whether a good is taken away or put in;

[0098] Assume that the motion of the feature point remains consistent within its local neighborhood window . Then, the Lucas-Kanade method minimizes the error within this neighborhood window:

[0099]

[0100] wherein, denotes the neighborhood window centered on the feature point, usually 3×3 or 5×5. By finding an optimal motion displacement , the pixel difference within the local neighborhood window is minimized, that is, the best matching position of the feature point in the next frame is found, that is, the position of the feature point in the next frame is , that is, the optical flow matching point is obtained, and then a pair of matching points is formed.

[0101] It should be noted that the motion displacement calculated by Lucas-Kanade may contain incorrect optical flow matching points, because:

[0102] Occlusion and noise: Some feature points are occluded between two frames, resulting in invalid optical flow matching points calculated;

[0103] Non-rigid deformation: If the image contains non-rigid objects (such as human bodies, fluids, etc.), then the assumed local consistency may not hold, resulting in incorrect matching.

[0104] Illumination change: Since the Lucas-Kanade method assumes photometric constancy (i.e., the pixel values of the same point remain unchanged), but in reality, there may be brightness changes, affecting the calculation of optical flow matching points;

[0105] Therefore, based on RANSAC to eliminate incorrect optical flow matching points and ensure the accuracy of subsequent transformation matrix estimation, we have:

[0106] Input feature point matching set ;

[0107] Calculate candidate transformation matrices, including:

[0108] Calculate the homography matrix : ;

[0109] Calculate the fundamental matrix : ;

[0110] Calculate the transformation error of all optical flow matching points:

[0111] , from which the optical flow matching points with transformation errors exceeding the threshold can be eliminated to ensure the detection accuracy.

[0112] Step 2: Construct each interaction between the consumer and the product (such as the hand approaching or touching the product) into a graph structure and record the spatio-temporal information to form an interaction graph reflecting the consumer's behavior path, including:

[0113] In the image frame For each candidate interaction event (a candidate where there is an overlap between the consumer's hand region and the product region in the image frame), calculate the following metrics:

[0114] Calculate the intersection over union of the two rectangular regions according to the bounding box of the product and the bounding box of the hand ;

[0115] where represents the intersection area of the two rectangles, represents the union area of the two rectangles;

[0116] Curvature of the hand movement trajectory (describing the degree of curvature of the hand movement trajectory, which can reflect the subtle movement changes of the consumer when picking up or browsing):

[0117] Suppose the sequence of hand center positions in consecutive frames is ;

[0118] represents the coordinates of the hand center in the image frame at time , represents the starting time, represents the number of consecutive frames collected;

[0119] Perform robust curve fitting:

[0120]

[0121] Among them, represents the fitting curve, , represents the coordinate function of the fitting curve, describing the hand movement trajectory obtained by fitting, represents the second derivative of the fitting curve, representing the acceleration information of the curve, that is, the change in the degree of curve bending, used to measure the smoothness of the curve, represents the smoothing parameter;

[0122] Approximate calculation of local curvature using the second derivative:

[0123]

[0124] Among them, , respectively represent , the derivatives with respect to , reflecting the instantaneous velocity components of the curve at time , , respectively represent , the second derivatives of, describing the acceleration or bending change of the curve, represents the curvature of the hand movement trajectory. It should be noted that the curvature calculation uses the analytical expression of the fitting curve and its derivative to avoid the influence of noise in discrete differences, so as to obtain a more stable and continuous curvature of the hand movement trajectory to accurately describe the geometric characteristics of the hand movement trajectory.

[0125] The acceleration of the hand (reflecting the change in hand movement speed, helping to distinguish rapid saccades from actual grasping behaviors):

[0126]

[0127] Among them, represents the hand movement speed in the image frame (current frame).

[0128] It should be noted that considering that the calculation of hand acceleration only depends on the difference between adjacent frames and is easily affected by noise, motion blur, and changes in dynamic interaction patterns, resulting in unstable acceleration. Therefore, a Bayesian estimation method is introduced. By combining the acceleration of the previous frame and the observed data of the current frame speed change, a prior distribution and an observation probability are established, so as to realize the maximum a posteriori estimation of acceleration; at the same time, in order to further smooth and optimize this estimation, an adaptive Kalman filtering method is adopted to dynamically adjust the process noise covariance and the observation noise covariance according to the prediction error to adaptively update the acceleration​​ The estimation is as follows:

[0129] At time The estimated acceleration is used as the prior estimate for the current frame , denotes the acceleration estimated for the previous frame, denotes the process noise covariance, which describes the uncertainty of the acceleration estimate;

[0130] Observation probability , which represents the acceleration distribution deduced based on the current velocity :

[0131]

[0132] where denotes the observation noise covariance, which describes the error of the velocity measurement, denotes the time interval between frames, denotes the normal distribution;

[0133] The posterior acceleration distribution is calculated by combining the prior distribution and the observation probability , denotes the normalization constant, i.e., the probability of observing the current velocity under all possible acceleration values, ensuring that the integral of the posterior acceleration distribution is 1;

[0134] The solution of the maximum a posteriori estimate is:

[0135]

[0136] where denotes the acceleration value obtained by the maximum a posteriori estimate at time , i.e., the optimal acceleration estimate obtained by combining the prior distribution and the observation probability, , respectively denote and inverse values;

[0137] The predicted acceleration for the current frame , denotes the acceleration value calculated by the maximum a posteriori estimate (MAP) at time ;

[0138] Update the covariance: , denotes the covariance of the acceleration estimation error predicted to time ,

[0139] Calculate the Kalman gain , Determines the weighting ratio of the predicted value and the observed value in the update stage, reflecting the relative relationship between the prediction uncertainty and the observation noise;

[0140] Calculate the updated hand acceleration , represents the acceleration observation value obtained through observation in the current frame.

[0141] Through the above, it is possible to maintain the smoothness and accuracy of the acceleration calculation in the face of environmental noise, motion blur, and different interaction modes (such as grasping, pushing, or sliding), providing more stable and reliable basic data for subsequent consumer behavior analysis and merchandise display optimization.

[0142] Furthermore, , , are weighted and summed to obtain a comprehensive score , if exceeds the threshold, it is determined that an interaction event of the commodity has occurred in the image frame;

[0143] Based on the intersection over union (measuring the overlap degree between the consumer's hand area and the commodity area), the curvature of the hand movement trajectory, and the acceleration characteristics, capture more subtle behavior changes of consumers during the interaction, such as:

[0144] When the consumer quickly glances at the commodity, although there may be a certain overlap, the curvature of its movement trajectory and the acceleration change are large, and the comprehensive score will be low;

[0145] When the consumer seriously picks up or carefully browses, the overlap is high and the movement is relatively stable, so the score is higher, and thus it can more accurately reflect whether an interaction event of the commodity has occurred in the image frame.

[0146] Regarding each interaction event as a node in the graph, using the time interval and spatial distance to construct edges, and forming a graph structure through nodes and edges to form an interaction graph, there is:

[0147] Each interaction event generates an interaction event node , representing an interaction event, which includes:

[0148] Commodity category, commodity location, hand touch duration (interaction duration within consecutive multiple frames), hand movement speed;

[0149] If the interaction event node and the interaction event node correspond to the interaction events in adjacent frames respectively, then construct an edge , and the characteristics of the edge include:

[0150] Node time interval , , respectively represent the occurrence times of interaction event nodes 、 ;

[0151] Spatial distance , 、 respectively represent the coordinates of interaction event node , 、 respectively represent the coordinates of interaction event node . It should be noted that the interaction graph structures the multiple interaction behaviors of consumers in front of the vending cabinet, retaining both the detailed features of each interaction event and recording the spatio-temporal relationships between behaviors.

[0152] Furthermore, after obtaining the interaction graph as described above, the interaction graph can be input into a graph neural network for node feature aggregation, and then combined with the Transformer module to capture global spatio-temporal dependencies. Finally, the potential purchase intention index PWI of each commodity is output, and this index is used to guide the optimization of commodity display to achieve the management of vending cabinet commodities, specifically including:[[]]

[0153] Update the features of each node in the interaction graph using graph convolution to achieve the aggregation of local information, then there is:[[]]

[0154] Let the nodes participating in feature aggregation in the interaction graph be , and the initial feature be ;

[0155] The update formula of graph convolution is:[[]]

[0156]

[0157] Among them, represents the feature of the node at the th layer, represents the set of neighbor nodes of node , and respectively represent the weight matrix and bias at the th layer, represents the activation function, such as Relu;

[0158] It is worth noting that the graph convolution operation can aggregate the information of neighbor nodes, so that the feature of each node not only contains its own information, but also integrates the context of its adjacent interaction events for subsequent commodity management.

[0159] Input the node feature sequence updated by graph convolution into the Transformer module, and capture the long-distance dependencies and temporal dynamics between nodes through the self-attention mechanism, then there is:[[]]

[0160] Arrange the node features into a sequence , represents the total number of nodes in the interaction graph, that is, all interaction events;

[0161] Self-attention calculation:

[0162]

[0163] Among them, forms a node feature matrix, is the dimension of each node feature, represents the dimension of the key vector , that is, the length of each column vector of , 、 、 respectively represent learnable weight matrices with dimensions of ;

[0164] Calculate the similarity matrix :

[0165]

[0166] Calculate the similarity scores between all query vectors and all key vectors , that is, the attention score matrix, and thus obtain the fused node features of the entire graph, represents the Transformer module, which processes the node features through the self-attention mechanism to capture long-range dependencies and global temporal information. After being processed by the Transformer, the features of each node are fused with global context information, enabling the model to analyze interaction behaviors more accurately.

[0167] Therefore, use a fully connected layer to calculate the purchase intention index of each product:

[0168] Calculate the purchase intention index for each node (representing a product interaction):

[0169]

[0170] Among them, represents the purchase intention index of the th node, represents the th feature vector of the node in , respectively represent the weight matrix and bias of the fully connected layer, represents the activation function, and the sigmoid function can be adopted to normalize the output value to the range of [0, 1]. It should be noted that the purchase intention index reflects the potential purchase willingness of consumers for a certain commodity. A higher value indicates a higher degree of attention and purchase tendency for this commodity in consumer behavior data.

[0171] Finally, based on the calculated geometric simulated annealing algorithm, the optimal display scheme of the commodity in the vending cabinet can be determined, that is, the commodity display and recommendation strategies of the vending cabinet are guided by the optimal solution, including:

[0172] Define the commodity set of the vending cabinet , which represents all the commodities to be displayed in the vending cabinet, represents the total number of commodities;

[0173] Define the display position set , which represents the fixed display positions in the vending cabinet (such as different heights and left - right positions on the cabinet surface), and the number of positions is the same as the number of commodities;

[0174] Define the layout solution : The layout solution is a permutation vector or mapping, indicating which position each commodity is assigned to , which can be represented by , where represents the commodity is placed at the position ;

[0175] Construct the objective function , where the goal is to maximize the overall effect;

[0176] Among them, , represents the set of commodities with higher PWI values, represents the threshold, , respectively represent the corresponding weight coefficients, used to balance the influence of the scores of the two parts, represents the exposure score of the commodity at the position . The exposure score is usually related to the conspicuousness of the position. For example, the upper - middle area of the vending cabinet has a higher score, represents the strategy score of the commodity , reflecting factors such as inventory and profit;

[0177] By simulating the physical annealing process, the system temperature is gradually reduced, and some inferior solutions are accepted to jump out of the local optimum, thus approaching the global optimum solution. Then we have:

[0178] Initial layout solution ;

[0179] Set a relatively high temperature value ;

[0180] Define the temperature attenuation coefficient : After each iteration, the temperature is updated to , multiplying the current temperature by the attenuation coefficient to obtain a new temperature value, indicating that the temperature gradually decreases exponentially, thereby gradually reducing the probability of accepting inferior solutions and finally making the algorithm converge to an approximate global optimum solution;

[0181] Number of iterations : The number of iterations at each temperature;

[0182] Stopping criterion: The current temperature drops to a certain low threshold or reaches the maximum number of iterations;

[0183] Make a slight perturbation to the current layout solution to generate a new solution. For example, randomly select two items and exchange their positions, that is, exchange and to obtain a new neighborhood solution ;

[0184] Calculate the objective function values of the current layout solution and the neighborhood solution ; and ;

[0185] , and similarly obtain ;

[0186] Calculate the energy difference ;

[0187] If is less than or equal to 0, it means accepting the neighborhood solution ;

[0188] If is greater than 0, it means accepting the neighborhood solution with a probability , and the probability ;

[0189] Generate a random number , in the range 0 - 1. If , then accept , otherwise keep the original solution. Therefore, through PW calculation, the simulated annealing algorithm finally finds an optimal placement plan to maximize the exposure of tall products. According to , a reasonable display distribution of products is carried out, which is conducive to subsequent users quickly finding and purchasing products in the vending cabinet, improving the user's purchase experience.

[0190] The above embodiments are only used to illustrate the technical solutions of the present invention, not to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. The vending machine commodity recognition algorithm based on artificial intelligence model is characterized by: The steps include: Detect the product area and the consumer's hand area in the product image; The graph structure is constructed based on the interaction between consumers and products to form an interaction graph, including: The intersection-and-union ratio, curvature of the hand motion trajectory, and acceleration are calculated in the candidate interaction events, where: The curvature of the hand motion trajectory is calculated using the fitting curve and its derivative; The Bayesian estimation method is introduced to combine the acceleration of the previous frame and the observed data of the speed change of the current frame to establish the prior distribution and observation probability, and realize the maximum a posteriori estimation of acceleration. Based on the Mann filtering method, the process noise covariance and the observation noise covariance are dynamically adjusted according to the prediction error to obtain the updated acceleration. Determine the comprehensive score based on the intersection-and-union ratio, the curvature of the hand motion trajectory, and the updated acceleration to determine the interaction events of the products in the image frame; Treat interaction events as nodes in the graph, use time intervals and spatial distances to construct edges, and construct a graph structure through nodes and edges to achieve the construction of an interaction graph. The interaction graph is input into the graph neural network for node feature aggregation. By capturing the global spatiotemporal dependency, the potential purchase intention index of each product is output, and the purchase intention index is used to optimize the product display. The method of inputting the interaction graph into the graph neural network for node feature aggregation is: Suppose the nodes participating in feature aggregation in the interaction graph are The initial features are ; The update formula of graph convolution is: in, Indicates Layer Node Features, Representation Node The set of neighbor nodes of and Respectively represent The weight matrices and biases of the layers, represents the activation function; The method of capturing the global spatiotemporal dependency and outputting the potential purchase intention index of a product is as follows: For each node Calculate the purchase intention index: in, Indicates The purchase intention index of nodes, express The The feature vector of a node, Represent the weight matrix and bias of the fully connected layer respectively; Fusion node features of the entire graph represents the Transformer module, Represents the node feature matrix. The node features are processed through the self-attention mechanism to capture long-distance dependencies and global timing information. After being processed by Transformer, the features of each node are integrated with the global context information.

2. The vending machine commodity recognition algorithm based on artificial intelligence model according to claim 1 is characterized in that: When the consumer's hand area is obtained, the feature points are tracked to determine the optical flow matching points, as follows: Select a set of feature points in the image frame Indicates the first The coordinates of the feature points are Indicates the total number of feature points selected in the image frame; Solve the motion displacement of each feature point in the next frame ; Assume that the motion of the feature point is in its local neighborhood window Keep consistent within the local neighborhood window and minimize the error: in, represents the neighborhood window centered on the feature point, Respectively represent the feature points in and Direction displacement, Indicates time image frame; By finding the optimal motion displacement, the optical flow matching points are obtained to form matching point pairs; Input feature point matching set; Calculate candidate transformation matrices; Calculate the transformation errors of all optical flow matching points, and re-determine the optical flow matching points based on the transformation errors.

3. The vending machine commodity identification algorithm based on artificial intelligence model according to claim 2 is characterized in that: The method for determining the curvature of the hand motion trajectory is: Assume that the sequence of hand center positions in multiple consecutive frames is ; Indicates at time The coordinates of the hand center in the image frame, Indicates the starting time, Indicates the number of consecutive frames collected; Perform robust curve fitting: in, represents the fitting curve, represents the coordinate function of the fitting curve, represents the second-order derivative of the fitted curve, represents the smoothing parameter; Calculate the curvature of the hand motion trajectory: in, Respectively about The derivative of Respectively The second derivative of .

4. The vending machine commodity identification algorithm based on artificial intelligence model according to claim 3 is characterized in that: The maximum a posteriori estimate of the acceleration is determined by: Determine the prior estimate of the current frame represents the estimated acceleration of the previous frame, represents the process noise covariance; Observation probability : in, represents the observation noise covariance, represents the time interval between frames, represents normal distribution; Compute the posterior acceleration distribution represents the normalization constant; Determine the maximum a posteriori estimate: in, Indicates at time The acceleration value obtained by maximum a posteriori estimation at Respectively and The inverse value of Indicates the hand movement speed.

5. The vending machine commodity identification algorithm based on artificial intelligence model according to claim 4 is characterized in that: The method to get the updated acceleration is: Determine the predicted acceleration for the current frame Indicates at time The acceleration value calculated by maximum a posteriori estimation; Update the covariance: Indicates the predicted time The acceleration estimation error covariance is Calculate Kalman gain ; Calculate updated hand acceleration Represents the acceleration observation value obtained through observation in the current frame.

6. The vending machine commodity identification algorithm based on artificial intelligence model according to claim 5 is characterized in that: After the interaction event is generated, the method of forming the interaction graph is: Each interaction event generates an interaction event node ; If the interaction event node Interaction event nodes Corresponding to the interaction events in adjacent frames, the edge is constructed , construct an interaction graph based on interaction event nodes and edges.

7. The vending machine commodity identification algorithm based on artificial intelligence model according to claim 6 is characterized in that: The interaction event node , which contains: Product category, product location, hand contact duration, and hand movement speed.

8. The vending machine commodity identification algorithm based on artificial intelligence model according to claim 6 is characterized in that: The characteristics of the edge include: Node Time Interval Represents the interaction event nodes The moment of appearance; Spatial distance Represents the interaction event nodes The coordinates of Represents the interaction event nodes The coordinates of .

Citation Information

Patent Citations

  • Offline customer behavior identification method and system, storage medium and terminal

    CN113326816A

  • Other Solution Automation & Interface Analysis Implementations

    US20230044564A1