Automatic selling method and system based on machine vision
By using binocular cameras, blind deconvolution recovery algorithms and YOLO v5 models in the vending system for product identification, and combining with a finite state machine to manage payment status, the problem of low recognition accuracy and frequent mechanical failures under specific lighting conditions is solved, and an efficient, safe and flexible vending process is achieved.
Patent Information
- Application Number
- CN202510090747.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
The existing automatic sales methods are difficult to accurately identify products under specific lighting conditions, with low recognition accuracy and frequent mechanical failures, resulting in a decline in user experience, difficult to meet safety needs, and high maintenance costs.
Multiple images of product interactions between users are collected through binocular cameras, preprocessed and fused, and product identification is used using blind deconvolution recovery algorithm and YOLO v5 model, and payment status is managed through a finite state machine to realize the automatic sales process.
It improves the robustness and accuracy of image recognition, improves the efficiency and user experience of automatic sales, reduces labor costs, and enhances the security and operational flexibility of the system.
Smart Images

Figure CN120014615A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to an automatic vending method and system based on machine vision. Background Art
[0002] Machine vision refers to the process of using computer vision technology to allow machines to perceive, process and understand visual information through cameras or sensors. Automatic vending refers to the process of automatically completing the selection, identification, settlement and delivery of goods by the system without human intervention. Automatic vending based on machine vision refers to the use of visual devices such as cameras to capture images, process these images through algorithms, identify goods and user behavior, and complete the process of automatic settlement and product delivery.
[0003] By using machine vision and advanced image processing and artificial intelligence technologies, we can achieve an unmanned, self-service sales process, which is of great significance to product layout and user services, thereby promoting the development of the retail industry towards intelligence, improving customer satisfaction and shopping experience, and providing data support for marketing strategies and business decisions.
[0004] Although existing vending methods provide many conveniences, they still have some defects. First, machine vision systems are sometimes not very adaptable to different ambient light and may have difficulty accurately identifying goods under specific lighting conditions. In addition, for goods with similar appearances but different varieties, the risk of identification errors increases. Mechanical failures of vending machines, such as stuck coins and goods, can also affect the user experience. At the same time, vending machines have a high demand for anti-theft measures and require more security and monitoring technologies to prevent theft. Although vending machines can reduce labor costs, the initial investment and maintenance costs are high, especially for advanced models using high-end technology. These factors limit the popularity and efficiency of vending systems. In short, existing vending methods have difficulty accurately identifying goods under specific lighting conditions, resulting in low recognition accuracy, and frequent mechanical failures, resulting in a decline in user experience, difficulty in meeting security requirements, and high maintenance costs. Summary of the invention
[0005] In order to solve the technical problems in the prior art that the existing automatic vending methods are difficult to accurately identify commodities under specific lighting conditions, resulting in low recognition accuracy, frequent mechanical failures, resulting in a decline in user experience, difficulty in meeting safety requirements, and high maintenance costs, the present invention provides an automatic vending method and system based on machine vision.
[0006] The technical solution provided by the embodiment of the present invention is as follows:
[0007] First aspect
[0008] An automatic vending method based on machine vision provided by an embodiment of the present invention includes:
[0009] S1: Collect multiple images of products and user interactions through a binocular camera;
[0010] S2: Preprocess each product image to generate a fused image;
[0011] S3: restore the fused image through a blind deconvolution restoration algorithm to determine the restored image of the product image;
[0012] S4: Build a product image recognition model based on YOLO v5;
[0013] S5: Input the restored image into the product image recognition model and output the category of the product taken by the user;
[0014] S6: Generate a payment QR code based on the product category;
[0015] S7: Obtain the user's payment status for the payment QR code, and obtain the character string corresponding to the payment status;
[0016] S8: Determine the maximum execution time of the string through a finite state machine, and use the maximum execution time as the state maintenance time of the vending machine;
[0017] S9: If the state maintenance time is exceeded, return to step S6; otherwise, complete the sale of the current product and reset the vending machine.
[0018] Second aspect
[0019] An automatic vending system based on machine vision provided by an embodiment of the present invention includes:
[0020] processor;
[0021] A memory having computer-readable instructions stored therein, wherein when the computer-readable instructions are executed by the processor, the automatic vending method based on machine vision as described in the first aspect is implemented.
[0022] The third aspect
[0023] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the automatic vending method based on machine vision as described in the first aspect is implemented.
[0024] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0025] In an embodiment of the present invention, multiple images of the interaction between the commodity and the user are collected by a binocular camera, and each commodity image is preprocessed to generate a fused image, and then the fused image is restored by a blind deconvolution restoration algorithm to determine the restored image of the commodity image, thereby improving the robustness of image recognition. Afterwards, a commodity image recognition model is constructed based on YOLO v5, and the restored image is input into the commodity image recognition model, the category of the commodity taken by the user is output, and a payment QR code is generated. The string corresponding to the payment QR code state is further obtained, and the maximum execution time of the string is determined by a finite state machine, and the maximum execution time is used as the state maintenance time of the vending machine. If the state maintenance time is exceeded, the payment QR code is regenerated, otherwise, the sale of the current commodity is completed and the vending machine is reset. The present invention combines machine vision for automatic vending, which can accurately identify products and user behavior, greatly improve sales efficiency, reduce labor costs, and enhance user experience. At the same time, the invention can adapt to a variety of scenarios, such as unmanned supermarkets, vending machines, etc., thereby improving operational flexibility and enabling operators to optimize inventory management, product layout and user services through data analysis, thereby promoting the development of the retail industry towards intelligence. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0027] Figure 1 A schematic diagram of a flow chart of an automatic vending method based on machine vision provided by an embodiment of the present invention;
[0028] Figure 2 A schematic diagram of the structure of an automatic vending system based on machine vision provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0029] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0030] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0031] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0032] Reference Manual Attached Figure 1 , showing a flow chart of an automatic vending method based on machine vision provided by an embodiment of the present invention.
[0033] The embodiment of the present invention provides an automatic vending method based on machine vision, which can be implemented by an automatic vending device based on machine vision, and the automatic vending device based on machine vision can be a terminal or a server. The processing flow of the automatic vending method based on machine vision can include the following steps:
[0034] S1: Collect multiple images of products and user interactions through a binocular camera.
[0035] Among them, the binocular camera is a device that simulates human binocular parallax through two cameras. It can collect depth information and three-dimensional structure of the scene. The interactive multiple images refer to multiple pictures collected by the binocular camera from different angles, which contain depth and surface information, and provide multi-dimensional data support for subsequent processing.
[0036] It should be noted that the acquisition of multiple images of products and user interactions by binocular cameras lays the foundation for the entire automatic vending process. Compared with monocular cameras, binocular cameras can capture more information, including the three-dimensional features and depth data of products, which helps to identify products in complex environments. Multi-image acquisition can reduce the recognition errors that may be caused by single-view images and provide rich input data for subsequent image fusion and restoration.
[0037] S2: Preprocess each product image to generate a fused image.
[0038] Among them, preprocessing is the initial processing and processing of the original product image, including corner extraction, feature matching, geometric alignment, color migration and other steps, the purpose is to improve the image quality and consistency, to facilitate subsequent analysis and recognition. Fusion image refers to the integration of information of multiple product images through a variety of algorithms to generate a high-quality image with comprehensive features.
[0039] It should be noted that preprocessing improves the quality and consistency of product images, laying a solid foundation for subsequent image recognition.
[0040] In a possible implementation, S2 specifically includes:
[0041] S201: Calculate the brightness change of each product image pixel through the Harris corner feature matching algorithm:
[0042]
[0043] Among them, E(u,v) represents the change of pixel brightness in the horizontal and vertical directions within the window, u represents the horizontal movement of the image, v represents the vertical movement of the image, M represents the structure tensor, and I x Represents the horizontal gradient of the image, I y Represents the gradient of the image in the vertical direction.
[0044] It should be noted that the Harris corner feature matching algorithm is an algorithm that calculates corner points based on the change in pixel brightness in an image, and is mainly used to detect key feature points in an image.
[0045] S202: Calculate the corner point response value according to the brightness change:
[0046] R = det(M) - k (trace(M)) 2
[0047] Among them, R represents the corner point response value, det represents the determinant, k represents the empirical constant, and trace represents the trace of the matrix.
[0048] Specifically, the corner point response value refers to an indicator value used to evaluate whether a point in an image is a corner point. The larger the value, the more likely it is a corner point.
[0049] S203: Perform feature matching on the corner points extracted from the product images to obtain matching feature points between the images.
[0050] S204: Generate a triangular mesh through Delaunay triangulation according to the matched feature points.
[0051] Among them, Delaunay triangulation is an algorithm that divides a planar point set into non-overlapping triangular meshes and is often used for geometric alignment in image processing.
[0052] S205: Perform affine transformation on the triangular mesh to complete the geometric alignment of the image:
[0053]
[0054] Among them, Ps.x represents the horizontal coordinate of the point x in the target image after the transformation, Ps.y represents the horizontal coordinate of the point y in the target image after the transformation, 1 represents the third dimension of the homogeneous coordinates, a11 and a12 represent the control of the scaling and rotation in the x direction, a21 and a22 represent the control of the scaling and rotation in the y direction, a13 and a23 represent the translation in the x and y directions, Px represents the point x in the product image, and Py represents the point y in the product image.
[0055] S206: Calculate the affine transformation matrix:
[0056]
[0057] Among them, A represents the affine transformation matrix, P s1 .x and P s1 .y respectively represent the coordinates x and y of the first vertex of the transformed triangle, P s2 .x and P s2 .y respectively represent the coordinates x and y of the second vertex of the transformed triangle, P s3 .x and P s3 .y represents the third vertex coordinate x and coordinate y of the transformed triangle, P1.x and P1.y represent the first vertex coordinate x and coordinate y of the triangle before the transformation, P2.x and P2.y represent the first vertex coordinate x and coordinate y of the triangle before the transformation, P3.x and P3.y represent the first vertex coordinate x and coordinate y of the triangle before the transformation, and -1 represents the inverse of the matrix.
[0058] S207: Calculate the RGB three-channel histogram of the product image through color migration.
[0059] Among them, color migration is a technology based on color distribution adjustment, which is used to adjust the color characteristics of one image to be similar to another image.
[0060] S208: Determine the cumulative histogram probability of each gray level:
[0061]
[0062] Among them, HistproR-A represents the histogram probability of the red channel in image A, R represents the red channel, j=0,…,m, m represents the total number of gray levels, HistR-A represents the histogram of the red channel in image A, w1 represents the width of the image, and h1 represents the height of the image.
[0063] S209: Establishing a grayscale mapping table through histogram matching to generate an image after color migration.
[0064] S210: Generate a fused image through Alpha fusion.
[0065] Among them, Alpha fusion is a method of fusing multiple images in a weighted manner, which can smoothly combine image information to generate the final fused image.
[0066] It should be noted that by extracting and matching image feature points through the Harris corner feature matching algorithm, the key parts of the product image can be accurately identified and aligned, effectively reducing the recognition errors caused by angle and position differences. Delaunay triangulation and affine transformation further realize the geometric alignment of the image to ensure the consistency of the product image. Color migration technology reduces the interference of different lighting conditions on image recognition by adjusting the image color distribution. Alpha fusion integrates multi-view data to generate high-quality fused images, which improves the robustness and accuracy of subsequent processing. The processing effect of image data is enhanced, so that the system can still operate stably in complex environments.
[0067] S3: Restore the fused image through a blind deconvolution restoration algorithm to determine a restored image of the product image.
[0068] Among them, the blind deconvolution restoration algorithm is an algorithm used to restore a clear image from a blurred or degraded image. It does not require knowing the exact cause or parameters that cause the image degradation in advance. The restored image refers to a clear and optimized image obtained from the fused image through image processing technology, especially the blind deconvolution algorithm.
[0069] It should be noted that by applying the blind deconvolution restoration algorithm to process the fused image, the image quality is significantly improved, especially in terms of eliminating noise and enhancing image details. The advantage is that it does not require prior knowledge of the specific cause of image degradation, such as the type or degree of blur, and can still effectively restore the image.
[0070] In a possible implementation, S3 specifically includes:
[0071] S301: Restore the fused image using a blind deconvolution restoration algorithm:
[0072] ||h×fg|| 2 =E[∫(h×fg) 2 dx]=E(∫n 2 dx)=σ 2 E(e)
[0073] Among them, ‖‖ represents the second norm, h represents the point spread function, f represents the restored image, g represents the fused image, E represents the expected value, e represents the random variable, n represents the noise, σ 2 represents the variance of the noise.
[0074] S302: Convert the solution process of blind deconvolution recovery into a Lagrangian optimization problem:
[0075] minL(f,h)=min[||h×fg|| 2 +α1r(f)+α2r(h)]
[0076] Among them, min means taking the minimum value, L represents the Lagrangian function, α1 represents the weight coefficient of the regularization term of the restored image, r(f) represents the regularization term of the restored image, α2 represents the weight coefficient of the regularization term of the point spread function, and r(h) represents the regularization term of the point spread function.
[0077] S303: Solve the Lagrangian optimization problem to determine a restored image of the product image.
[0078] Among them, the Lagrangian optimization problem is a mathematical optimization problem. The Lagrangian multiplier method is used to transform the constrained optimization problem into an unconstrained optimization problem to find the optimal solution. The Lagrangian optimization problem is used to optimize the model for restoring image quality and balance data fitting and regularization constraints.
[0079] It should be noted that by solving the Lagrangian optimization problem, the blurred area in the fused image is further optimized, and finally a product image with higher definition (restored image) is generated. This process optimizes the quality of the product image while ensuring the accuracy and robustness of subsequent image recognition.
[0080] S4: Build a product image recognition model based on YOLO v5.
[0081] Among them, YOLO (You Only Look Once) is a real-time target detection algorithm, and v5 is one of its versions. It is known for its efficient and fast performance and is suitable for small target detection tasks. The product image recognition model refers to an image classification and target detection model built based on deep learning, which is specifically used to identify product categories or attributes.
[0082] It should be noted that by building a product image recognition model based on YOLO v5, the system's product detection and classification capabilities have been significantly improved. With its efficient single-stage target detection method, YOLO v5 enables the system to strike a balance between real-time performance and accuracy, meeting the requirements of automatic vending for rapid response.
[0083] S5: Input the restored image into the product image recognition model, and output the category of the product taken by the user.
[0084] Among them, the product category is the result output by the system after analyzing the restored image through the image recognition model, which indicates the specific classification of the product taken by the user. By inputting the restored image into the product image recognition model in YOLO format, efficient and accurate recognition of the product category is achieved.
[0085] In a possible implementation, S5 specifically includes:
[0086] S501: Input the restored image into the product image recognition model in YOLO format.
[0087] S502: Initialize the anchor box size through the K-means clustering algorithm.
[0088] It should be noted that K-means clustering automatically determines the most suitable anchor box size based on the distribution of the real bounding box of the target in the training data, so that the anchor box fits the data characteristics better and improves the detection performance of the model. Reasonable initialization of the anchor box reduces the deviation between the target area and the anchor box, making the target detection model (such as YOLO v5) converge faster during training, while improving the accuracy of detection, especially the recognition effect of targets of different sizes.
[0089] S503: According to the size of the anchor box, the target area of the restored image is freely sampled through a deformable convolution layer.
[0090] S504: Determine the target area features through the ECA-Net attention mechanism.
[0091] S505: Calculate the loss function of the product image recognition model.
[0092] S506: Using a gradient descent algorithm, adjust the hyperparameters of the product image recognition model until the loss function value is less than a preset loss function value, and output a prediction result.
[0093] S506 specifically includes:
[0094] S5061: Initialize hyperparameters.
[0095] S5062: Calculate the loss function value of the product image recognition model based on the hyperparameters.
[0096] S5063: Determine the gradient based on the loss function value:
[0097]
[0098] Among them, g t represents the gradient vector of the tth iteration, B represents the batch size, represents the gradient operator, x i represents the target area of the i-th input restored image, θ t Represents the parameter vector at the tth iteration.
[0099] S5064: Calculate the gradient norm based on the gradient:
[0100]
[0101] Among them, g norm represents the gradient vector g t The Euclidean norm of , d represents the dimension of the gradient vector, gt [i] represents the gradient vector g in the tth iteration t The i-th component of .
[0102] S5065: Update the historical gradient norm according to the gradient norm and calculate the scaling factor:
[0103] e t =γe t-1 +(1-γ)g norm
[0104]
[0105] Among them, e t represents the smoothed value of the gradient norm of the tth iteration, γ represents the smoothing coefficient, e t-1 It represents the smoothed value of the gradient norm in the t-1th iteration, s represents the gradient scaling factor, ε represents a constant, and when the gradient vector is greater than the smoothed value, the gradient is scaled. When the gradient scaling factor is equal to 1, the gradient is not adjusted.
[0106] S5066: Update momentum and model hyperparameters:
[0107] m t =βm t-1 +(1-β)s·g t
[0108] θ t =θ t -αm t
[0109] Among them, m t represents the momentum term of the t-th iteration, β represents the momentum coefficient, and α represents the learning rate.
[0110] S5067: Determine whether the loss function value is less than the preset loss function value. If so, output the prediction result; otherwise, go to step S5062.
[0111] It should be noted that by combining the gradient descent algorithm, the optimization process is made smoother, the convergence speed is accelerated, and the convergence problem caused by excessive or small gradients is avoided, so that the model parameters that meet the loss conditions can be found more efficiently, the convergence is accelerated, and the classification accuracy and generalization ability of the model are improved.
[0112] S507: Perform non-maximum suppression on each prediction result to remove overlapping areas.
[0113] S508: Determine the category of the product based on the target area of the product in the prediction result and by using the confidence factor.
[0114] It should be noted that by optimizing the size of the anchor box and using deformable convolution, the system can more flexibly capture the features of the target area, thereby improving the recognition performance. Combined with the ECA-Net attention mechanism, the model can better focus on the key feature area and further improve the detection accuracy. The optimization of the loss function and the use of the gradient descent algorithm ensure that the model converges quickly and reduces errors. In addition, non-maximum suppression removes redundant detection boxes, improves the simplicity and accuracy of the recognition results, and through the calculation of the confidence factor, the system can effectively filter low-confidence results to ensure that the output prediction results are more reliable. Overall, it provides efficient and stable product recognition capabilities for the automatic vending system.
[0115] In a possible implementation manner, S503 specifically includes:
[0116] The target area of the restored image is freely sampled according to the following formula:
[0117]
[0118] Δp n =1,2,…,N
[0119] Among them, y(p0) represents the value of the output feature map at position p0, p n represents the fixed offset position of the convolution kernel, w represents the weight, △p n represents the learnable offset, △m n represents the learnable amplitude, and x represents the target area of the input restored image.
[0120] In a possible implementation, the loss function is specifically:
[0121]
[0122] Among them, Loss CIoU represents the CIoU loss function, IoU, α, v all represent intermediate variables, ρ represents the Euclidean distance, b represents the coordinates of the center point of the prediction box, and b gt represents the coordinates of the center point of the real box, c represents the diagonal distance of the minimum enclosing rectangle, w represents the width of the predicted box, and w gt represents the width of the real box, h represents the height of the predicted box, and h gt represents the real box height, and arctan represents the tangent function.
[0123] In a possible implementation manner, S504 specifically includes:
[0124] S5041: Perform global average pooling on the output feature map.
[0125] S5042: Based on the feature map after global average pooling, determine the local cross-channel interaction through a one-dimensional convolutional layer:
[0126]
[0127] Among them, k represents the size of the convolution kernel, represents the adaptive function, C represents the number of channels, γ represents the adjustment parameter, b represents the bias value, lb(C) represents the logarithmic value of channel C, |.| odd It means to round the calculated result to the nearest odd number.
[0128] S5043: Generate channel weights through the sigmoid function to obtain target area features with channel attention.
[0129] It should be noted that global average pooling is used to extract global information from feature maps, reduce interference in spatial dimensions, and focus on important feature relationships between channels. Then, one-dimensional convolution is used to achieve cross-channel interaction, dynamically adjust the weights of different channels, and adaptively optimize the depth and accuracy of feature extraction. The size of the convolution kernel is dynamically determined by the function to enhance adaptability, and the stability of the calculation is ensured by odd number constraints. Finally, the channel weights generated by the sigmoid function can effectively enhance the important features of the target area and suppress irrelevant feature information.
[0130] In a possible implementation manner, S508 specifically includes:
[0131] S5081: Set the confidence threshold of the product image recognition model:
[0132] CF=0.5+d
[0133] Where CF represents the global confidence factor and d represents the positive value of fine-tuning the global confidence factor.
[0134] S5082: Calculate the local confidence factor according to the confidence threshold:
[0135]
[0136] Among them, CF n represents the local confidence factor of the nth target region, and N represents the number of target regions.
[0137] S5083: Calculate the comprehensive confidence factor based on the global confidence factor and the local confidence factor:
[0138] CF k =sum({CF|A=A k},{CF1,…CF n ,…,CF N |An =A k}),k=1,…K,
[0139] Among them, CF k represents the comprehensive confidence factor of the kth class, A represents the classification result, and A k represents the kth classification category, A n =A k It means that the nth target area is classified as the kth category, and K represents the total number of classification categories.
[0140] S5084: Determine the category of the product based on the maximum comprehensive confidence factor:
[0141] CF max =max(CF1,…,CF K )
[0142] Among them, CF max Represents the maximum certainty factor.
[0143] It should be noted that by setting the global confidence threshold, the system can perform unified basic control over the overall recognition accuracy and adapt to different scenarios by fine-tuning parameters. Secondly, the calculation of the local confidence factor focuses on the fine-grained information of the target area, making the confidence assessment of each target area more accurate. Then, through the calculation of the comprehensive confidence factor, the system integrates the global and local confidences to form a more comprehensive assessment of the product category, avoiding the deviation that may be caused by a single confidence factor. Finally, based on the maximum comprehensive confidence factor, the product category with the highest confidence is selected as the final output, ensuring the reliability and robustness of the classification results.
[0144] S6: Generate a payment QR code based on the product category.
[0145] Among them, the payment QR code is a QR code generated based on the product category, which usually contains the payment amount, product information and transaction link for users to scan and complete the payment.
[0146] It should be noted that by automatically generating a payment QR code based on the identified product category, the payment process for users to purchase goods is greatly simplified. Compared with the traditional manual settlement or keyboard price input method, this method of automatically generating a QR code is faster and more efficient, while avoiding errors in human operation. The payment QR code can be directly connected to the backend system to achieve real-time deduction and inventory management.
[0147] S7: Obtain the user's payment status for the payment QR code, and obtain the character string corresponding to the payment status.
[0148] The payment status is the result status after the user operates the payment QR code, usually including information such as payment success, payment failure or payment timeout. The string is the payment status encoded by the system into a text form (string), for example, "success" means payment success, and "fail" means payment failure.
[0149] It should be noted that by obtaining the payment status of the user's payment QR code and converting it into a string for processing, the payment process is automatically monitored and managed. This method can determine in real time whether the payment is successful, thereby triggering the corresponding vending machine behavior, reducing human intervention, and improving the intelligence level of the system. In addition, the payment status is encoded in the form of a string, which facilitates the system to quickly parse and classify various payment results.
[0150] S8: Determine the maximum execution time of the string through a finite state machine, and use the maximum execution time as the state maintenance time of the vending machine.
[0151] Among them, the finite state machine is a mathematical model used to represent a finite set of states in a system and the transition rules between their states. In a vending machine, the finite state machine can manage different payment states (such as success, failure, or timeout) and the corresponding system behaviors. The maximum execution time of the string indicates the maximum time for the vending machine to perform an operation based on the payment status string, thereby limiting the response time of the system under abnormal circumstances and avoiding long-term stagnation of the system. The state maintenance time refers to the maintenance time of the current state of the vending machine (such as waiting for payment). After this time, a rollback or reinitialization operation will be triggered.
[0152] It should be noted that by introducing a finite state machine to dynamically manage the payment status and maximum execution time, the operating efficiency and robustness of the vending machine have been significantly improved. The finite state machine allows the system to flexibly switch behaviors according to the payment status, such as entering the commodity release state after payment is completed, or returning to the initial state after a timeout, thereby avoiding long-term freezes or waiting problems. This dynamic time management mechanism can effectively deal with payment anomalies or delays, and ensure the stable operation of the system in complex scenarios.
[0153] In a possible implementation manner, S8 specifically includes:
[0154] By calculating the time complexity of the vending machine, the maximum execution time of the string is dynamically determined, and the maximum execution time is used as the state maintenance time of the vending machine:
[0155]
[0156] in, represents the maximum execution time of the finite state machine M when processing an input string of length n with a complexity parameter of c, Σ n represents the set of all possible input strings of length n, It represents the execution time of the finite state machine M when processing an input string of length z with complexity parameter c.
[0157] S9: If the state maintenance time is exceeded, return to step S6; otherwise, complete the sale of the current product and reset the vending machine.
[0158] It should be noted that by setting the status maintenance time, the system can automatically return for reprocessing in the event of a payment timeout or abnormal situation to avoid freezes or dead loops, while ensuring the smoothness and stability of the sales process. The system automatically resets after the sale is completed, improving operational efficiency and user experience, and enhancing the fault tolerance of the vending machine.
[0159] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0160] In an embodiment of the present invention, multiple images of the interaction between the commodity and the user are collected by a binocular camera, and each commodity image is preprocessed to generate a fused image, and then the fused image is restored by a blind deconvolution restoration algorithm to determine the restored image of the commodity image, thereby improving the robustness of image recognition. Afterwards, a commodity image recognition model is constructed based on YOLO v5, and the restored image is input into the commodity image recognition model, the category of the commodity taken by the user is output, and a payment QR code is generated. The string corresponding to the payment QR code state is further obtained, and the maximum execution time of the string is determined by a finite state machine, and the maximum execution time is used as the state maintenance time of the vending machine. If the state maintenance time is exceeded, the payment QR code is regenerated, otherwise, the sale of the current commodity is completed and the vending machine is reset. The present invention combines machine vision for automatic vending, which can accurately identify products and user behavior, greatly improve sales efficiency, reduce labor costs, and enhance user experience. At the same time, the invention can adapt to a variety of scenarios, such as unmanned supermarkets, vending machines, etc., thereby improving operational flexibility and enabling operators to optimize inventory management, product layout and user services through data analysis, thereby promoting the development of the retail industry towards intelligence.
[0161] Reference Manual Attached Figure 2 , showing a structural schematic diagram of an automatic vending system based on machine vision provided by the present invention.
[0162] The present invention further provides an automatic vending system 20 based on machine vision, which is applied to the automatic vending method based on machine vision, and comprises:
[0163] Processor 201.
[0164] The memory 202 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 201 , the automatic vending method based on machine vision as in the method embodiment is implemented.
[0165] The automatic vending system 20 based on machine vision provided by the present invention can execute the automatic vending method based on machine vision described above and achieve the same or similar technical effects. To avoid repetition, the present invention will not elaborate on them.
[0166] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0167] In an embodiment of the present invention, multiple images of the interaction between the commodity and the user are collected by a binocular camera, and each commodity image is preprocessed to generate a fused image, and then the fused image is restored by a blind deconvolution restoration algorithm to determine the restored image of the commodity image, thereby improving the robustness of image recognition. Afterwards, a commodity image recognition model is constructed based on YOLO v5, and the restored image is input into the commodity image recognition model, the category of the commodity taken by the user is output, and a payment QR code is generated. The string corresponding to the payment QR code state is further obtained, and the maximum execution time of the string is determined by a finite state machine, and the maximum execution time is used as the state maintenance time of the vending machine. If the state maintenance time is exceeded, the payment QR code is regenerated, otherwise, the sale of the current commodity is completed and the vending machine is reset. The present invention combines machine vision for automatic vending, which can accurately identify products and user behavior, greatly improve sales efficiency, reduce labor costs, and enhance user experience. At the same time, the invention can adapt to a variety of scenarios, such as unmanned supermarkets, vending machines, etc., thereby improving operational flexibility and enabling operators to optimize inventory management, product layout and user services through data analysis, thereby promoting the development of the retail industry towards intelligence.
[0168] It should be understood that the processor in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0169] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).
[0170] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When a computer instruction or computer program is loaded or executed on a computer, a process or function according to an embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0171] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.
[0172] In the present invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.
[0173] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0174] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0175] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0176] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0177] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0178] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0179] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.
[0180] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the automatic vending method based on machine vision as described in the method embodiment is implemented.
[0181] A computer-readable storage medium provided by the present invention can implement the steps and effects of the automatic vending method based on machine vision in the above method embodiment. To avoid repetition, the present invention will not go into details.
[0182] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0183] In an embodiment of the present invention, multiple images of the interaction between the commodity and the user are collected by a binocular camera, and each commodity image is preprocessed to generate a fused image, and then the fused image is restored by a blind deconvolution restoration algorithm to determine the restored image of the commodity image, thereby improving the robustness of image recognition. Afterwards, a commodity image recognition model is constructed based on YOLO v5, and the restored image is input into the commodity image recognition model, the category of the commodity taken by the user is output, and a payment QR code is generated. The string corresponding to the payment QR code state is further obtained, and the maximum execution time of the string is determined by a finite state machine, and the maximum execution time is used as the state maintenance time of the vending machine. If the state maintenance time is exceeded, the payment QR code is regenerated, otherwise, the sale of the current commodity is completed and the vending machine is reset. The present invention combines machine vision for automatic vending, which can accurately identify products and user behavior, greatly improve sales efficiency, reduce labor costs, and enhance user experience. At the same time, the invention can adapt to a variety of scenarios, such as unmanned supermarkets, vending machines, etc., thereby improving operational flexibility and enabling operators to optimize inventory management, product layout and user services through data analysis, thereby promoting the development of the retail industry towards intelligence.
[0184] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
[0185] There are a few points to note:
[0186] (1) The drawings of the embodiments of the present invention only relate to the structures related to the embodiments of the present invention, and other structures may refer to the general design.
[0187] (2) For the sake of clarity, in the drawings used to describe the embodiments of the present invention, the thickness of the layers or regions is exaggerated or reduced, that is, these drawings are not drawn according to the actual scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element may be "directly" "on" or "under" the other element or there may be intermediate elements.
[0188] (3) In the absence of conflict, the embodiments of the present invention and the features therein may be combined with each other to obtain new embodiments.
[0189] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. An automatic vending method based on machine vision, characterized in that: include: S1: Collect multiple images of products and user interactions through a binocular camera; S2: Preprocess each product image to generate a fused image; S3: Restoring the fused image by a blind deconvolution restoration algorithm to determine a restored image of the product image; S4: Build a product image recognition model based on YOLO v5; S5: inputting the restored image into the commodity image recognition model, and outputting the category of the commodity taken by the user; S6: Generate a payment QR code according to the commodity category; S7: Obtain the user's payment status for the payment QR code, and obtain a character string corresponding to the payment status; S8: Determine the maximum execution time of the character string through a finite state machine, and use the maximum execution time as the state maintenance time of the vending machine; S9: When the state maintenance time is exceeded, return to step S6; otherwise, complete the sale of the current product and reset the vending machine.
2. The automatic vending method based on machine vision according to claim 1, characterized in that: The S2 specifically includes: S201: Calculate the brightness change of each product image pixel through the Harris corner feature matching algorithm: Among them, E(u,v) represents the change of pixel brightness in the horizontal and vertical directions within the window, u represents the horizontal movement of the image, v represents the vertical movement of the image, M represents the structure tensor, and I x Represents the horizontal gradient of the image, I y Represents the gradient of the image in the vertical direction; S202: Calculate the corner point response value according to the brightness change: R=det(M)-k·(trace(M)) 2 Among them, R represents the corner point response value, det represents the determinant, k represents the empirical constant, and trace represents the trace of the matrix; S203: performing feature matching on the corner points extracted from the product images to obtain matching feature points between the images; S204: generating a triangular mesh through Delaunay triangulation according to the matching feature points; S205: Perform affine transformation on the triangular mesh to complete geometric alignment of the image: Where Ps.x represents the horizontal coordinate of point x in the target image after transformation, Ps.y represents the horizontal coordinate of point y in the target image after transformation, 1 represents the third dimension of homogeneous coordinates, a11 and a12 represent the control of scaling and rotation in the x direction, a21 and a22 represent the control of scaling and rotation in the y direction, a13 and a23 represent the translation in the x and y directions, Px represents point x in the product image, and Py represents point y in the product image; S206: Calculate the affine transformation matrix: Where A represents the affine transformation matrix, P s1 .x and P s1 .y respectively represent the coordinates x and y of the first vertex of the transformed triangle, P s2 .x and P s2 .y respectively represent the coordinates x and y of the second vertex of the transformed triangle, P s3 .x and P s3 .y represents the coordinate x and coordinate y of the third vertex of the transformed triangle, P1.x and P1.y represent the coordinate x and coordinate y of the first vertex of the triangle before the transformation, P2.x and P2.y represent the coordinate x and coordinate y of the first vertex of the triangle before the transformation, P3.x and P3.y represent the coordinate x and coordinate y of the first vertex of the triangle before the transformation, and -1 represents the inverse of the matrix; S207: Calculating the RGB three-channel histogram of the product image through color migration; S208: Determine the cumulative histogram probability of each gray level: Among them, HistproR - A represents the histogram probability of the red channel in image A, R represents the red channel, j=0,…,m, m represents the total number of gray levels, HistR-A represents the histogram of the red channel in image A, w1 represents the width of the image, and h1 represents the height of the image; S209: Establishing a grayscale mapping table through histogram matching to generate an image after color migration; S210: Generate a fused image through Alpha fusion.
3. The automatic vending method based on machine vision according to claim 1, characterized in that: The S3 specifically includes: S301: Restoring the fused image by a blind deconvolution restoration algorithm: ||h×f-g|| 2 =E[∫(h×f-g) 2 dx]=E(∫n 2 dx)=σ 2 E(e) Among them, || || represents the second norm, h represents the point spread function, f represents the restored image, g represents the fused image, E represents the expected value, e represents the random variable, n represents the noise, σ 2 represents the variance of the noise; S302: Convert the solution process of blind deconvolution recovery into a Lagrangian optimization problem: minL(f,h)=min[||h×f-g|| 2 +α1r(f)+α2r(h)] Wherein, min means taking the minimum value, L means the Lagrangian function, α1 means the weight coefficient of the regularization term of the restored image, r(f) means the regularization term of the restored image, α2 means the weight coefficient of the regularization term of the point spread function, and r(h) means the regularization term of the point spread function; S303: Solve the Lagrangian optimization problem to determine a restored image of the product image.
4. The automatic vending method based on machine vision according to claim 1, characterized in that: The S5 specifically includes: S501: Inputting the restored image into the product image recognition model in YOLO format; S502: Initialize the anchor box size through K-means clustering algorithm; S503: freely sampling the target area of the restored image through a deformable convolution layer according to the size of the anchor frame; S504: Determine the target area features through the ECA-Net attention mechanism; S505: Calculating the loss function of the commodity image recognition model; S506: adjusting the hyperparameters of the product image recognition model by a gradient descent algorithm until the loss function value is less than a preset loss function value, and outputting a prediction result; S507: Perform non-maximum suppression on each prediction result to remove overlapping areas; S508: Determine the category of the product based on the target area of the product in the prediction result and by using the confidence factor.
5. The automatic vending method based on machine vision according to claim 4, characterized in that: The S503 is specifically as follows: The target area of the restored image is freely sampled according to the following formula: Δp n =1,2,…,N Among them, y(p0) represents the value of the output feature map at position p0, p n represents the fixed offset position of the convolution kernel, w represents the weight, △p n represents the learnable offset, △m n represents the learnable amplitude, and x represents the target area of the input restored image.
6. The automatic vending method based on machine vision according to claim 4, characterized in that: The loss function is specifically: Among them, Loss CIoU represents the CIoU loss function, IoU, α, v all represent intermediate variables, ρ represents the Euclidean distance, b represents the coordinates of the center point of the prediction box, and b gt represents the coordinates of the center point of the real box, c represents the diagonal distance of the minimum enclosing rectangle, w represents the width of the predicted box, and w gt represents the width of the real box, h represents the height of the predicted box, and h gt represents the real box height, and arctan represents the tangent function.
7. The automatic vending method based on machine vision according to claim 4, characterized in that: The S504 specifically includes: S5041: Performing global average pooling on the output feature map; S5042: Based on the feature map after global average pooling, determine the local cross-channel interaction through a one-dimensional convolutional layer: Among them, k represents the size of the convolution kernel, represents the adaptive function, C represents the number of channels, γ represents the adjustment parameter, b represents the bias value, lb(C) represents the logarithmic value of channel C, |.| odd It means taking the nearest odd number for the calculated result; S5043: Generate channel weights through the sigmoid function to obtain target area features with channel attention.
8. The automatic vending method based on machine vision according to claim 4, characterized in that: The S508 specifically includes: S5081: Setting the confidence threshold of the product image recognition model: CF=0.5+d Where CF represents the global confidence factor, and d represents the positive value of the fine-tuned global confidence factor; S5082: Calculate a local confidence factor according to the confidence threshold: Among them, CF n represents the local confidence factor of the nth target region, and N represents the number of target regions; S5083: Calculate a comprehensive confidence factor according to the global confidence factor and the local confidence factor: CF k =sum({CF|A=A k },{CF1,…CF n ,…,CF N |A n =A k }),k=1,…K, Among them, CF k represents the comprehensive confidence factor of the kth class, A represents the classification result, and A k represents the kth classification category, A n =A k Indicates that the nth target area is classified as the kth category, and K represents the total number of classification categories; S5084: Determine the category of the product according to the maximum comprehensive confidence factor: CF max =max(CF1,…,CF K ) Among them, CF max Represents the maximum certainty factor.
9. The automatic vending method based on machine vision according to claim 4, characterized in that: The S8 is specifically: By calculating the time complexity of the vending machine, the maximum execution time of the string is dynamically determined, and the maximum execution time is used as the state maintenance time of the vending machine: in, represents the maximum execution time of the finite state machine M when processing an input string of length n with a complexity parameter of c, Σ n represents the set of all possible input strings of length n, It represents the execution time of the finite state machine M when processing an input string of length z with complexity parameter c.
10. An automatic vending system based on machine vision, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the automatic vending method based on machine vision as claimed in any one of claims 1 to 9 is implemented.